Credit risk assessment method and system based on knowledge graph, and storage medium

Through the credit risk assessment method based on the knowledge graph, the problems of multi-source heterogeneous data fusion and dynamic risk modeling are solved, and efficient assessment and accurate identification of credit risks are achieved, which is especially suitable for the risk identification of complex mutual insurance chains for small and medium-sized enterprises.

CN120563231APending Publication Date: 2025-08-29IND & COMMERCIAL BANK OF CHINA CO LTD ZHENGZHOU BRANCH
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510836164.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing credit risk assessment methods have problems such as low fusion efficiency of multi-source heterogeneous data, insufficient dynamic relationship modeling, uncertainty in risk state expression and high computational complexity, resulting in insufficient evaluation accuracy and limited risk identification capabilities.

Method used

Using a knowledge graph-based method, semantic standardization is performed through ontology mapping rules, dynamic maps are constructed by self-organizing algorithms, graph convolution risk propagation operators are used to model states, and risk aggregation is performed in combination with adaptive attention mechanisms to realize multi-dimensional feature extraction and evaluation.

Benefits of technology

It has achieved effective integration of multi-source heterogeneous data, strong dynamic risk modeling capabilities, can accurately express the uncertainty and complexity of credit risks, improves the comprehensiveness and accuracy of risk identification, and is suitable for the evaluation of complex credit scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563231A_ABST
    Figure CN120563231A_ABST
Patent Text Reader

Abstract

The invention discloses a credit risk assessment method and system based on a knowledge graph and a storage medium, and the method comprises the steps: carrying out the standardization processing of multi-source credit data through an ontology mapping rule, and obtaining an RDF triple data set; constructing a dynamic knowledge graph containing a guarantee chain, a fund flow direction and a risk factor by adopting a self-organizing algorithm; carrying out modeling through a graph convolution risk propagation operator to obtain a risk state vector; performing time sequence embedding extraction to obtain a five-dimensional comprehensive risk feature vector; and through adaptive attention mechanism aggregation processing, a credit risk assessment result and a confidence interval are obtained. The technical problems of multi-source heterogeneous data fusion, dynamic relation modeling, risk state quantification and multi-dimensional feature aggregation in credit risk assessment based on the knowledge graph are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a credit risk assessment method, system and storage medium based on knowledge graph. Background Art

[0002] Existing credit risk assessment methods primarily rely on traditional scorecard models and machine learning algorithms, building risk assessment models by collecting structured data such as borrowers' basic information, financial data, and credit records. These methods employ algorithms such as logistic regression, decision trees, and random forests to predict borrowers' default probabilities and make credit decisions based on these predictions. Traditional assessment methods are somewhat effective in processing standardized data and historical statistical patterns, and are widely used in financial institutions.

[0003] However, existing technologies suffer from significant shortcomings: severe data silos and the difficulty in effectively integrating heterogeneous data from multiple sources, resulting in incomplete risk assessment information; static assessment models are unable to adapt to the dynamically changing market environment and lack a real-time update mechanism; single-dimensional feature extraction ignores the complex network of borrowers' relationships and fails to capture hidden risk transmission paths; models are poorly interpretable, and the decision-making process lacks transparency; and computational complexity is high, creating a conflict between real-time requirements and computational efficiency. These technical limitations result in insufficient assessment accuracy and limited risk identification capabilities for complex credit scenarios.

[0004] Based on the shortcomings of existing technologies, further analysis revealed deeper technical problems: the semantic standardization processing of multi-source heterogeneous credit data lacks a unified ontology framework, resulting in inefficient data fusion; the dynamic graph construction of entity relationships lacks an adaptive learning mechanism and cannot automatically discover complex association patterns; risk state modeling lacks the multi-state combination capability of quantum superposition states, making it difficult to express risk uncertainty; time series embedding extraction lacks a bidirectional encoding mechanism and cannot simultaneously capture historical evolution and future trends; risk aggregation processing lacks an adaptive attention mechanism and cannot dynamically adjust the importance weights of different risk dimensions. Summary of the Invention

[0005] This application provides a knowledge graph-based credit risk assessment method, system and storage medium, which are used to solve the technical problems of multi-source heterogeneous data fusion, dynamic relationship modeling, risk status quantification and multi-dimensional feature aggregation in knowledge graph-based credit risk assessment.

[0006] In the first aspect, the present application provides a credit risk assessment method based on a knowledge graph, which includes: semantically standardizing multi-source heterogeneous credit data through ontology mapping rules to obtain a credit subject triple data set in RDF format; dynamically constructing entity relationships using a self-organizing algorithm based on the credit subject triple data set to obtain a multi-level dynamic knowledge graph containing a guarantee chain network, a capital flow graph, and a risk factor graph; performing risk state modeling on the multi-level dynamic knowledge graph through a graph convolution risk propagation operator to obtain a credit subject risk state vector; performing time series embedding and extraction processing on the guarantee chain network, the capital flow graph, and the risk factor graph based on the credit subject risk state vector to obtain a comprehensive risk feature vector of the credit history dimension, the social relationship dimension, the financial status dimension, the behavior pattern dimension, and the external environment dimension; performing risk aggregation processing on the comprehensive risk feature vector through an adaptive attention mechanism to obtain a credit risk assessment result and a confidence interval.

[0007] Optionally, the semantic standardization processing of multi-source heterogeneous credit data is performed using ontology mapping rules to obtain a credit subject triple dataset in RDF format, including: Establish a credit domain ontology concept system, define the semantic structure of entity classes, relationship classes, and attribute classes, and obtain an ontology library; Perform entity recognition processing on the credit data, transaction data, behavior data and related party data based on the ontology library to obtain an entity identification set; Perform entity disambiguation processing on the entity identifier set using an edit distance algorithm to obtain an entity mapping table; Performing relationship extraction and attribute extraction processing on multi-source data according to the entity mapping table to obtain relationship attribute pairs; The relationship attribute pairs are converted into triple format and constructed to obtain a credit subject triple dataset in RDF format.

[0008] Optionally, the self-organizing algorithm is used to construct a dynamic graph of entity relationships based on the credit subject triple data set to obtain a multi-level dynamic knowledge graph including a guarantee chain network, a capital flow graph, and a risk factor graph, including: Inputting the credit subject triple data set into the input layer of the self-organizing neural network for feature vector encoding processing to obtain entity embedding vectors and relationship embedding vectors; Based on the entity embedding vector and the relationship embedding vector, the Euclidean distance is calculated through each neuron in the competition layer to perform optimal matching processing to obtain the winning neuron position; According to the position of the winning neuron, a neighborhood function is used to perform a learning and updating process on the weights of the surrounding neurons to obtain an adaptive weight matrix; The adaptive weight matrix is ​​classified and mapped according to entity type and relationship type through the output layer to obtain a guarantee chain network, a capital flow map and a risk factor map; The guarantee chain network, capital flow map and risk factor map are hierarchically combined to obtain a multi-level dynamic knowledge map including the guarantee chain network, capital flow map and risk factor map.

[0009] Optionally, the multi-level dynamic knowledge graph is subjected to risk state modeling processing through a graph convolution risk propagation operator to obtain a credit subject risk state vector, including: Convolution kernels of different sizes are designed for the guarantee chain network, capital flow map, and risk factor map to perform subgraph feature extraction processing to obtain guarantee relationship features, capital flow features, and risk factor features; The guarantee relationship characteristics, capital flow characteristics and risk factor characteristics are vectorized through multi-dimensional feature coding to obtain a subgraph embedding vector; Based on the subgraph embedding vector, a quantum superposition state modeling method is used to perform multi-state combination processing to obtain a composite risk state set; Performing historical risk adjustment processing based on the composite risk state set by time decay weight calculation to obtain a time series risk weight matrix; Inputting the temporal risk weight matrix into the risk propagation operator to perform multi-hop neighbor influence calculation processing to obtain a risk propagation intensity matrix; The risk propagation intensity matrix is ​​subjected to vectorized projection and feature fusion processing to obtain a credit subject risk state vector.

[0010] Optionally, the step of inputting the temporal risk weight matrix into a risk propagation operator to perform multi-hop neighbor influence calculation processing to obtain a risk propagation intensity matrix includes: Inputting the temporal risk weight matrix into the risk propagation operator to perform one-hop direct neighbor node identification processing to obtain a set of directly related nodes; Performing two-hop indirect neighbor node expansion processing based on the directly associated node set through a graph traversal algorithm to obtain a secondary associated node set; Continue to perform graph traversal expansion according to the second-level associated node set to perform three-hop neighbor node discovery processing to obtain a third-level associated node set; Calculating the product of the path weight between each node in the directly associated node set, the secondary associated node set, and the tertiary associated node set and the target node to perform influence quantification processing to obtain a node influence value; The node influence value is weighted by the attenuation coefficient according to the hop distance to obtain the distance-weighted influence; The distance-weighted influences are matrix-arranged and normalized to obtain a risk propagation intensity matrix.

[0011] Optionally, the guarantee chain network, capital flow map and risk factor map are subjected to time-series embedding and extraction processing based on the credit subject risk status vector to obtain a comprehensive risk feature vector in the credit history dimension, social relationship dimension, financial status dimension, behavior pattern dimension and external environment dimension, including: Inputting the credit subject risk state vector into a bidirectional temporal encoder for forward historical trajectory encoding processing to obtain a historical risk evolution vector; Based on the historical risk evolution vector, a future risk trend encoding process is performed using a backward prediction trend encoder to obtain a predicted risk trend vector; Extracting time series relationships from the guarantee chain network based on the historical risk evolution vector and the predicted risk trend vector to obtain credit history dimension features and social relationship dimension features; Performing transaction behavior pattern analysis on the fund flow map using the predicted risk trend vector to obtain financial status dimension features and behavior pattern dimension features; Based on the historical risk evolution vector, the risk factor map is subjected to macro-environmental factor extraction processing to obtain external environment dimension characteristics; Vector splicing and feature fusion processing are performed on the credit history dimension features, social relationship dimension features, financial status dimension features, behavioral pattern dimension features and external environment dimension features to obtain a comprehensive risk feature vector.

[0012] Optionally, the step of performing risk aggregation processing on the comprehensive risk feature vector through an adaptive attention mechanism to obtain a credit risk assessment result and a confidence interval includes: Inputting the comprehensive risk feature vector into a multi-head attention network to generate a query vector, a key vector, and a value vector to obtain an attention calculation matrix; Performing attention weight normalization processing based on the attention calculation matrix through the softmax function to obtain a standardized attention weight; Performing weighted fusion processing on different risk dimensions according to the standardized attention weights to obtain an aggregated risk feature vector; Inputting the aggregated risk feature vector into a Bayesian neural network for uncertainty quantification to obtain a risk assessment mean and variance; Confidence interval calculation and risk level classification are performed based on the risk assessment mean and variance to obtain a credit risk assessment result and a confidence interval.

[0013] In a second aspect, the present application provides a credit risk assessment system based on a knowledge graph, the credit risk assessment system based on a knowledge graph comprising: The processing module is used to perform semantic standardization on multi-source heterogeneous credit data through ontology mapping rules to obtain a credit subject triple dataset in RDF format; A construction module is used to construct a dynamic graph of entity relationships using a self-organizing algorithm based on the credit subject triple data set to obtain a multi-level dynamic knowledge graph including a guarantee chain network, a capital flow graph, and a risk factor graph; A modeling module, configured to perform risk state modeling processing on the multi-level dynamic knowledge graph through a graph convolution risk propagation operator to obtain a credit subject risk state vector; An extraction module is used to perform time-series embedding and extraction processing on the guarantee chain network, capital flow map, and risk factor map based on the credit subject risk status vector to obtain a comprehensive risk feature vector in the credit history dimension, social relationship dimension, financial status dimension, behavioral pattern dimension, and external environment dimension; The aggregation module is used to perform risk aggregation processing on the comprehensive risk feature vector through an adaptive attention mechanism to obtain a credit risk assessment result and a confidence interval.

[0014] In a third aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned knowledge graph-based credit risk assessment method.

[0015] Positive and beneficial effects: The technical solution provided in this application solves the core technical difficulties of multi-source heterogeneous data processing and dynamic risk modeling in the existing technology by adopting a credit risk assessment method based on knowledge graph. The semantic standardization of multi-source heterogeneous credit data is carried out through ontology mapping rules, and a unified RDF format triple data set is established, which eliminates the problems of data silos and format inconsistency in traditional methods and realizes the effective integration of credit data, transaction data, behavioral data and related party data. The self-organizing algorithm constructs a dynamic graph of entity relationships and can automatically discover complex guarantee chain networks, capital flow maps and risk factor maps. Compared with traditional static modeling methods, it has the ability of adaptive learning and dynamic updating, which significantly improves the comprehensiveness and accuracy of risk identification. The graph convolution risk propagation operator performs risk state modeling and realizes multi-state combination through quantum superposition state modeling, which overcomes the limitation that traditional methods can only handle deterministic risk states and can more accurately express the uncertainty and complexity of credit risk. The time series embedding extraction processing adopts a bidirectional encoding mechanism to simultaneously capture historical risk evolution and future trend predictions, solving the deficiency of existing technologies in the lack of time dimension analysis and realizing comprehensive feature extraction in five dimensions: credit history, social relations, financial status, behavioral patterns and external environment.

[0016] In the specific application area of ​​credit risk assessment, the core algorithmic features of this invention significantly contribute to the effectiveness of the solution. When processing complex guarantee relationship networks, the self-organizing algorithm can automatically identify hidden risk propagation paths. Compared with traditional rule-driven relationship identification methods, it has stronger pattern discovery capabilities and adaptability, making it particularly suitable for identifying risks in complex mutual guarantee chains in small and medium-sized enterprise credit. The graph convolutional risk propagation operator, combined with quantum superposition state modeling, demonstrates unique advantages in processing systemic risk propagation, capable of simultaneously considering the probabilistic combination of multiple risk states, providing financial institutions with more accurate risk quantification results. The adaptive attention mechanism dynamically adjusts the weights of various risk dimensions in risk aggregation processing based on the characteristics of different borrowers, avoiding the limitations of traditional fixed-weight models. This is particularly true when assessing risk across different industries and enterprises of different sizes, automatically adapting to industry and enterprise characteristics. The Bayesian neural network quantifies uncertainty, outputting not only risk assessment results but also confidence intervals, providing more comprehensive information support for financial institutions' risk decision-making. This can effectively quantify forecast uncertainty and reduce decision-making risks, especially when facing emerging industries or rapidly changing economic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 This is a schematic diagram of an embodiment of a credit risk assessment method based on a knowledge graph in an embodiment of the present application; Figure 2 This is a schematic diagram of an embodiment of a credit risk assessment system based on a knowledge graph in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The embodiments of the present application provide a method, system and storage medium for credit risk assessment based on a knowledge graph. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.

[0020] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of the credit risk assessment method based on the knowledge graph includes: Step S101: semantically standardize the multi-source heterogeneous credit data using ontology mapping rules to obtain a credit subject triple dataset in RDF format; Step S102: Based on the credit subject triple data set, a self-organizing algorithm is used to construct a dynamic graph of entity relationships to obtain a multi-level dynamic knowledge graph including a guarantee chain network, a capital flow graph, and a risk factor graph; Step S103: Perform risk state modeling on the multi-level dynamic knowledge graph using a graph convolution risk propagation operator to obtain a credit subject risk state vector; Step S104: Perform time-series embedding and extraction processing on the guarantee chain network, capital flow map, and risk factor map based on the credit subject risk status vector to obtain a comprehensive risk feature vector in the credit history dimension, social relationship dimension, financial status dimension, behavioral pattern dimension, and external environment dimension; Step S105: Perform risk aggregation processing on the comprehensive risk feature vector through the adaptive attention mechanism to obtain the credit risk assessment result and confidence interval.

[0021] It is understandable that the execution entity of this application can be a credit risk assessment system based on a knowledge graph, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution entity as an example.

[0022] Specifically, a conceptual ontology system for the credit domain was established, defining semantic structures such as borrower entity classes, institution entity classes, and guarantee relationship classes to form an ontology library. Entity recognition was then performed on credit data, transaction data, behavioral data, and related party data. An edit distance algorithm was used to eliminate naming differences between entities in different data sources. Finally, the processed relationship attribute pairs were converted into a subject-predicate-object triple format. The edit distance algorithm measures similarity by calculating the minimum number of edit operations between two strings. For example, if a borrower is recorded as "Zhang San" in the financial system and "Zhang Mou San" in the credit system, the algorithm recognizes them as the same entity and uniformly identifies them.

[0023] A self-organizing algorithm is used to construct a multi-level dynamic knowledge graph. Triple data is fed into the input layer of the self-organizing neural network for feature vector encoding. Each neuron in the competitive layer calculates the Euclidean distance between the input vector and its own weight vector. The neuron with the smallest distance wins. A neighborhood function then updates the weights of the winning neuron and its neighbors. The output layer performs classification and mapping based on entity type and relationship type. The self-organizing neural network is an unsupervised learning algorithm that automatically discovers clustering patterns in data through a competitive learning mechanism. When dealing with borrower guarantee relationships, the algorithm clusters entities with similar guarantee patterns in the same area, forming three subgraph structures: a guarantee chain network, a capital flow graph, and a risk factor graph.

[0024] Risk states are modeled using a graph convolutional risk propagation operator. Convolutional kernels of different sizes are designed for each of the three subgraphs to extract features. The extracted guarantee relationship features, capital flow features, and risk factor features are multi-dimensionally encoded to form subgraph embedding vectors. Quantum superposition modeling is used to probabilistically combine multiple risk states. The impact of historical risk is adjusted using time-decaying weights. The risk propagation operator calculates the influence of multi-hop neighboring nodes. The graph convolutional network updates node representations by aggregating node neighbor information. When a borrower has a guarantee relationship, its risk state is affected by the risk state of the guaranteed party. The propagation operator calculates the influence strength of one-hop direct guarantee relationships, two-hop indirect relationships, and three-hop and longer relationships.

[0025] Time series embedding and extraction are performed, and the credit entity's risk state vector is input into a bidirectional time series encoder. The forward encoder processes historical risk evolution trajectories, while the backward encoder predicts future risk trends. Based on the encoding results, the guarantee chain network is extracted from time series relationships to obtain features in the credit history and social relationship dimensions. Transaction behavior patterns are analyzed on the capital flow graph to obtain features in the financial status and behavior pattern dimensions. Macro-environmental factors are extracted from the risk factor graph to obtain features in the external environment dimension. The bidirectional time series encoder combines historical information with future predictions. When analyzing a corporate borrower, the forward encoder processes historical data such as past repayment records and changes in financial statements, while the backward encoder predicts future risk trends based on industry trends and policy changes.

[0026] To achieve risk aggregation, the comprehensive risk feature vector is input into a multi-head attention network to generate a query vector, a key vector, and a value vector. The attention weights are normalized using the softmax function, and different risk dimensions are weighted and fused to form an aggregated risk feature vector. A Bayesian neural network quantifies uncertainty to calculate the mean and variance of the risk assessment. Finally, confidence intervals are calculated and risk levels are assigned. The multi-head attention mechanism allows the model to simultaneously focus on risk information across different dimensions. When a borrower's social relationships indicate high risk while their financial status indicates low risk, the attention mechanism assigns corresponding weights based on similar outcomes in historical data. The Bayesian neural network expresses predictions through probability distributions rather than deterministic numerical values, providing an uncertainty measure for risk assessment.

[0027] In a specific embodiment, the process of executing step S101 may specifically include the following steps: Establish a credit domain ontology concept system, define the semantic structure of entity classes, relationship classes, and attribute classes, and obtain an ontology library; Perform entity recognition processing on credit data, transaction data, behavior data and related party data based on the ontology library to obtain an entity identification set; Perform entity disambiguation on the entity identification set through the edit distance algorithm to obtain an entity mapping table; Perform relationship extraction and attribute extraction on multi-source data according to the entity mapping table to obtain relationship-attribute pairs; Convert the relationship-attribute pairs into a triple format for construction processing to obtain a credit subject triple dataset in RDF format.

[0028] Specifically, starting from establishing an ontology concept system in the credit field, first define entity classes including core concepts such as individual borrowers, financial institutions, and guarantee companies. Relationship classes cover connection methods such as guarantee relationships, lending relationships, and equity relationships. Attribute classes describe characteristic information such as credit ratings, financial indicators, and risk levels. The ontology library is constructed using the RDF schema, and each concept has a unique URI identifier and detailed semantic descriptions. When defining an individual borrower in the entity class, it includes basic attributes such as ID number, name, and contact information. The financial institution class includes attributes such as institution code, business scope, and registered capital. When defining a guarantee relationship in the relationship class, key elements such as the guarantor, the guaranteed party, the guarantee amount, and the guarantee period are clarified. The attribute class sets specific numerical ranges and value constraints for each entity and relationship. Based on the constructed ontology library, when performing entity recognition on multi-source data, the borrower's name in the credit investigation data needs to be matched with the individual borrower class in the ontology library, the account information in the transaction data needs to be corresponded with the financial institution class, the operation records in the behavior data need to be associated with the lending relationship class, and the enterprise information in the associated party data needs to be mapped with the institutional entity class. The entity recognition process uses a method that combines string matching and semantic similarity calculation. When "XX Technology Co., Ltd." appears in the credit investigation system and is recorded as "XX Technology Limited Liability Company" in the industrial and commercial data, the algorithm will calculate the semantic similarity of the two strings and determine them as the same entity. Each successfully recognized entity will obtain a unique entity identifier and be added to the entity identification set.

[0029] When performing entity disambiguation using the edit distance algorithm, calculate the minimum number of character operations required between two strings to measure similarity. The operations include three types: insertion, deletion, and replacement. When processing the borrower names Zhang Zhiqiang and Zhang Zhijian, the algorithm finds that only one operation of replacing "qiang" with "jian" is required, and the edit distance is 1. Combine context information such as the first few digits of the ID number and the contact phone number to assist in judging whether they are the same person. The algorithm sets a threshold parameter. When the edit distance is less than the preset threshold and the matching degree of other attributes exceeds the specified value, it is determined as the same entity and a corresponding relationship is established in the entity mapping table. The mapping table records the original entity identifier, the unified entity identifier, and the confidence score.

[0030] Relationship extraction and attribute extraction processes uniformly process multi-source data based on entity mapping tables. They extract information such as loan records, repayment history, and overdue payments from credit data, extract information such as capital flow, transaction frequency, and counterparties from transaction data, identify attributes such as login time, operating habits, and device information from behavioral data, and discover complex relationships such as equity, guarantee, and supply chain relationships from related party data. During attribute extraction, numeric attributes such as loan amount and credit score are directly extracted and converted to different data types. Text attributes such as industry classification and risk level require standardized mapping. Time attributes such as account opening date and expiration date require a unified format and time interval calculation. Boolean attributes such as whether the payment is overdue or whether there is a guarantee require binary processing.

[0031] The triple format construction process converts relationship attribute pairs into the standard RDF format of subject-predicate-object. The subject is typically the borrower or institutional entity, the predicate represents the relationship type or attribute name, and the object represents the associated entity or attribute value. During the construction process, each triple is assigned a unique identifier and annotated with the data source and update time. When borrower A provides a guarantee for enterprise B, the generated triple is A-guarantee-B. When A's credit rating is BBB, the generated triple is A-credit rating-BBB. The RDF format uses Uniform Resource Identifiers (URIs) to represent entities and relationships, ensuring global uniqueness and interoperability of data while supporting reasoning and query operations.

[0032] In a specific embodiment, the process of executing step S102 may specifically include the following steps: Input the credit subject triple data set into the input layer of the self-organizing neural network for feature vector encoding processing to obtain entity embedding vectors and relationship embedding vectors; Based on the entity embedding vector and the relationship embedding vector, the Euclidean distance is calculated through each neuron in the competition layer to perform optimal matching and obtain the winning neuron position; According to the position of the winning neuron, the neighborhood function is used to learn and update the weights of the surrounding neurons to obtain an adaptive weight matrix; The adaptive weight matrix is ​​classified and mapped according to entity type and relationship type through the output layer to obtain the guarantee chain network, capital flow map and risk factor map; The guarantee chain network, capital flow map and risk factor map are hierarchically combined to obtain a multi-level dynamic knowledge map containing the guarantee chain network, capital flow map and risk factor map.

[0033] Specifically, the data processing process begins by feeding a dataset of credit subject triples into the input layer of a self-organizing neural network. The input layer receives information about the subject entity, predicate relationship, and object entity of each triple, converting the textual entities and relationships into numerical feature vectors. The feature vector encoding process utilizes word embedding technology, decomposing each entity name into a character sequence and mapping it into a high-dimensional vector space. The entity embedding vector reflects the semantic characteristics of the entity, while the relationship embedding vector captures the type characteristics of the relationship. During the encoding process, the borrower entity "a certain manufacturing enterprise" is converted into a vector representation containing semantic information such as enterprise attributes, industry information, and scale characteristics. The guarantee relationship "providing a guarantee" is encoded into a vector reflecting relationship characteristics such as guarantee strength, term, and amount. The dimensions of each vector are set to a fixed length to ensure consistency in subsequent calculations.

[0034] When each neuron in the competitive layer calculates the Euclidean distance for optimal matching, it stores a weight vector. When a new entity embedding vector and relationship embedding vector are input, all neurons in the competitive layer simultaneously calculate the Euclidean distance between their weight vectors and the input vectors. Euclidean distance calculation involves taking the square root of the sum of the squared differences in each dimension of the vector. The smaller the distance, the more similar the neuron is to the input vector. The competitive learning mechanism ensures that only one neuron wins at a time. That is, the neuron with the smallest distance becomes the winning neuron, and the position of the winning neuron corresponds to a specific coordinate in the two-dimensional topology of the neural network. When processing a guarantee relationship triple, if the neuron in the fifth row and third column of the competitive layer has the smallest distance to the embedding vector of the triple, then this neuron becomes the winning neuron, and its position coordinates are recorded.

[0035] The neighborhood function performs a learning and updating process on the weights of surrounding neurons. Based on the location of the winning neuron, the neighborhood range that needs to be updated is determined. The neighborhood function usually takes the form of a Gaussian function. The closer the neuron is to the winning neuron, the larger the update amplitude, while the farther the neuron is, the smaller the update amplitude. The learning and update process adjusts the weight vectors of the winning neuron and the neurons in its neighborhood toward the input vector. The adjustment amplitude is controlled by the learning rate parameter, which gradually decays as the training progresses. When the winning neuron is located in the fifth row and third column, its neighborhood includes eight neighboring neurons. The weight vector of each neighboring neuron is updated to varying degrees based on its distance from the winning neuron, forming a local learning effect. The adaptive weight matrix records the weight update results of all neurons. The rows and columns of the matrix correspond to the topology of the neural network, and each position stores the weight vector of the corresponding neuron.

[0036] The output layer performs classification mapping based on entity type and relationship type, categorizing and archiving neurons in the adaptive weight matrix according to their learned feature patterns. Neurons with similar weight patterns are grouped into the same category. Classification mapping categorizes neurons into three main categories based on the characteristics of credit operations: guarantee relationship neurons primarily respond to triples related to the guarantee chain, capital flow neurons primarily respond to triples related to transactions and capital transfers, and risk factor neurons primarily respond to triples related to risk indicators and external environmental factors. The mapping process determines the classification results by analyzing the characteristics of each neuron's weight vector and the type of triples it primarily responds to during training. The weight vectors of guarantee relationship neurons have large values ​​in dimensions such as guarantee amount and guarantee period, while capital flow neurons respond strongly to dimensions such as transaction frequency and capital amount. Risk factor neurons are highly sensitive to dimensions such as risk rating and industry indicators.

[0037] Hierarchical combination processing constructs the three types of neurons and their corresponding triplet data into three subgraph structures: a guarantee chain network, a capital flow graph, and a risk factor graph. The entities and relationships within each subgraph form a specific network topology. The guarantee chain network connects entities with guarantee relationships as the core, forming a complex guarantee dependency network. The capital flow graph constructs the capital flow paths between entities based on capital transfer relationships. The risk factor graph constructs a risk propagation network with risk-related attributes and external factors as nodes. The combination processing maintains the entity correspondence between the three subgraphs. When the same borrower entity appears as a guarantor in the guarantee chain network, as a fund recipient in the capital flow graph, and as a high-risk entity in the risk factor graph, the algorithm establishes entity mapping relationships across the subgraphs, forming a unified multi-level dynamic knowledge graph structure.

[0038] In a specific embodiment, the process of executing step S103 may specifically include the following steps: Convolution kernels of different sizes are designed for the guarantee chain network, capital flow map, and risk factor map to perform subgraph feature extraction processing, thereby obtaining guarantee relationship features, capital flow features, and risk factor features. The guarantee relationship characteristics, capital flow characteristics and risk factor characteristics are vectorized through multi-dimensional feature encoding to obtain the subgraph embedding vector; Based on the subgraph embedding vector, quantum superposition state modeling is used to perform multi-state combination processing to obtain a set of composite risk states; According to the composite risk state set, historical risk adjustment is performed through time decay weight calculation to obtain the time series risk weight matrix; The time series risk weight matrix is ​​input into the risk propagation operator to calculate the multi-hop neighbor influence degree and obtain the risk propagation intensity matrix; The risk propagation intensity matrix is ​​subjected to vectorized projection and feature fusion processing to obtain the credit subject risk state vector.

[0039] Specifically, data processing begins by designing convolutional kernels of different sizes for each of the three subgraphs. The guarantee chain network employs a three-by-three convolutional kernel to specifically capture the local connectivity patterns of guarantee relationships. The capital flow graph employs a five-by-five convolutional kernel to extract mid-range correlation features of capital flows. The risk factor graph employs a seven-by-seven convolutional kernel to capture the global impact of risk factors. Convolutional kernels are core components of graph convolutional networks, similar to filters in image processing. They extract features by sliding across the graph structure and performing convolution operations with local subgraphs. When processing guarantee relationships, the convolutional kernels of the guarantee chain network consider both the guarantor and the guaranteed party, as well as their directly connected entities, to form a guarantee relationship feature vector. This vector contains key information such as the strength of the guarantee amount, the duration of the guarantee, and the creditworthiness of the guarantor. The convolutional kernels of the capital flow graph have a wider coverage, capturing the paths and patterns of capital flow between multiple entities. This creates a capital flow feature vector that captures characteristics such as the stability of the capital source, the concentration of capital flows, and the frequency of transactions. The convolution kernel range of the risk factor map is the largest, and it can analyze the comprehensive impact of external risk factors such as macroeconomic indicators, industry policy changes, and market fluctuations on the borrower, and form a risk factor feature vector.

[0040] Multi-dimensional feature encoding and vector representation processing convert the three types of feature vectors into a unified numerical representation format. The encoding process utilizes a multi-layer neural network structure, and each feature vector undergoes nonlinear transformation and dimensional mapping to obtain a fixed-length embedding representation. The guarantee relationship features are encoded into a dense vector containing information about the guarantee network topology and guarantee strength. The capital flow features are encoded as a vector representation reflecting capital flow patterns and transaction behavior. The risk factor features are encoded as a vector form that captures external environmental risks and systemic risks. During the encoding process, the numerical range and distribution of the original features are standardized to ensure that different types of features have equal weight in subsequent calculations. The three encoded vectors are spliced ​​into a comprehensive subgraph embedding vector according to preset combination rules.

[0041] Quantum superposition modeling for multi-state combination processing, based on the concept of superposition in quantum computing, represents each borrower's risk state as a linear combination of multiple underlying risk states. These underlying risk states include three discrete states: low risk, medium risk, and high risk. Each state corresponds to a specific probability amplitude. The borrower's actual risk state is a probability-weighted combination of these three underlying states. During the combination process, the subgraph embedding vector is decomposed into components corresponding to different underlying states, with the weight of each component reflecting the probability of the borrower being in the corresponding risk state. Quantum superposition allows borrowers to simultaneously possess multiple risk characteristics, rather than simply being categorized into a single risk level. This representation better reflects the uncertainty and complexity of credit risk. The composite risk state set contains quantum superposition representations of all borrowers, with each state containing both probability amplitude and phase information.

[0042] Time-decay weight calculations perform historical risk adjustment, dynamically adjusting the impact of risk events based on their time of occurrence. Recent risk events receive higher weights, while risk events with a longer history have their weights gradually decay. The decay calculation utilizes an exponential decay function, with the rate of decay controlled by a decay constant, which is set based on the duration of the impact of different risk events. The impact of a guarantee default decays slowly, the impact of abnormal capital flows decays moderately, and the impact of short-term market fluctuations decays rapidly. During the weight calculation process, each risk event is assigned a timestamp. The algorithm calculates a decay weight based on the difference between the current time and the time of the event, then applies the decay weight to the corresponding risk status component. The time series risk weight matrix records the changes in risk weights for all borrowers at different points in time. The rows of the matrix correspond to the borrowers, and the columns correspond to the time series. Each element records the risk weight value of a specific borrower at a specific time.

[0043] The multi-hop neighbor influence calculation process uses a risk propagation operator to analyze the risk propagation path and intensity within the knowledge graph. A one-hop neighbor is an entity directly connected to the target borrower, a two-hop neighbor is connected to the target borrower through an intermediate entity, and a three-hop neighbor requires two intermediate entities to reach the target borrower. The influence calculation considers the attenuation effect of path length on risk propagation; the more distant the neighbor, the less influence it has on the target borrower. During the calculation, the algorithm traverses all possible paths from the target borrower, calculates the product of the weights on each path as the propagation intensity of that path, and then takes the weighted sum of all path intensities to obtain the total influence. The risk propagation intensity matrix records the risk propagation intensity between each pair of entities. The matrix's symmetry reflects the bidirectional nature of risk propagation, with the diagonal elements representing the risk intensity of the entity itself.

[0044] Vectorized projection and feature fusion convert the risk propagation intensity matrix into a vector format suitable for machine learning algorithms. The projection process uses matrix decomposition techniques to compress the high-dimensional sparse matrix into a low-dimensional dense vector. Each borrower corresponds to a row in the matrix, which records the risk propagation intensity of that borrower to all other entities. Dimensionality reduction techniques such as principal component analysis or singular value decomposition are used to compress this row of data into a fixed-length feature vector. Feature fusion combines feature vectors from different propagation hops to form a comprehensive risk status vector. The fusion weights are set based on the importance of different hops, with the highest weight given to one-hop neighbors and the weights decreasing for two-hop and three-hop neighbors.

[0045] In a specific embodiment, the process of inputting the temporal risk weight matrix into the risk propagation operator to calculate the multi-hop neighbor influence may specifically include the following steps: Input the time series risk weight matrix into the risk propagation operator to perform one-hop direct neighbor node identification processing to obtain the set of directly related nodes; Based on the directly associated node set, a two-hop indirect neighbor node expansion process is performed through a graph traversal algorithm to obtain a secondary associated node set; Continue to perform graph traversal expansion according to the second-level associated node set to perform three-hop neighbor node discovery processing to obtain the third-level associated node set; For each node in the directly associated node set, the second-level associated node set, and the third-level associated node set, the path weight product between the node and the target node is calculated and the influence is quantified to obtain the node influence value; The node influence value is weighted by the attenuation coefficient according to the hop distance to obtain the distance-weighted influence; The distance-weighted influence is matrixed and normalized to obtain the risk transmission intensity matrix.

[0046] Specifically, the risk propagation operator begins the process of identifying one-hop direct neighbor nodes by inputting the temporal risk weight matrix into a risk propagation operator. This operator is a computational component specifically designed to analyze the path of risk propagation within a knowledge graph. The operator reads the row data corresponding to each borrower in the temporal risk weight matrix and identifies all neighbor nodes directly connected to the target node. Direct neighbor nodes are entities in the knowledge graph that are directly connected to the target borrower via an edge. These include the guaranteed party in a guarantee relationship, the source or recipient of funds in a fund flow, and external factors that directly influence risk factors. During the identification process, the operator traverses all outgoing and incoming edges of the target node and adds the other end of the edge to the set of directly connected nodes. Each directly connected node retains its connection weight and relationship type information with the target node. The set of directly connected nodes records key attributes such as node identity, connection weight, and relationship type for all one-hop neighbors. This information is derived from the edge weights and node attributes in the previously constructed multi-level dynamic knowledge graph.

[0047] The graph traversal algorithm performs two-hop indirect neighbor node expansion, expanding one layer outward for each node in the directly connected node set. Graph traversal is a standard algorithm in graph theory used to access all nodes in a graph, encompassing two basic strategies: depth-first search and breadth-first search. The expansion process employs a breadth-first search strategy, starting from each node in the directly connected node set and accessing all directly connected neighbor nodes. These newly discovered nodes constitute the second-level connected node set. Two-hop indirect neighbors are entities that require an intermediate node to reach from the target node. For example, if target borrower A guarantees enterprise B, and enterprise B has a business relationship with supplier C, then supplier C is A's two-hop neighbor. During the expansion process, the algorithm checks whether each newly discovered node already appears in the previous set to avoid duplicate counting. The algorithm also records the complete path from the target node to the two-hop neighbor, including the intermediate nodes and the weights of the connecting edges.

[0048] The three-hop neighbor node discovery process continues by performing a graph traversal expansion based on the set of second-level associated nodes. Starting from each node in the set, the algorithm expands outward one more layer to discover all neighboring nodes within three hops of the target node. Three-hop neighbors are entities that require two intermediate nodes to reach from the target node. These long-distance relationships often reflect the propagation path of systemic risk in credit risk assessment. During the discovery process, the algorithm maintains path integrity, recording all possible paths from the target node to three-hop neighbors. Each path consists of three edges and two intermediate nodes. Because three-hop expansion generates a large number of nodes, the algorithm imposes a visit depth limit and an upper limit on the number of nodes to prevent excessive computational complexity. At the same time, it prioritizes paths with higher weights and filters out long-distance connections with weaker impact.

[0049] Path weight product quantification is used to calculate the influence of each node on the target node within a set of connected nodes at three levels. This influence quantification is based on the product of all edge weights along the path. Starting from the target node, the path weight product is calculated for each edge along the connecting path. For a one-hop neighbor, the influence is equal to the connecting edge weight. For a two-hop neighbor, the influence is equal to the edge weight from the target node to the intermediate node multiplied by the edge weight from the intermediate node to the two-hop neighbor. For a three-hop neighbor, the influence is equal to the product of the three edge weights along the path. When multiple paths connect the same pair of nodes, the algorithm calculates the sum of the weighted products of all paths as the total influence of the node. During the product calculation, edge weights reflect the strength of the connection. Guarantees with large collateral amounts are given higher weights, capital flows with frequent transactions are given higher weights, and risk factors with wide impact are given higher weights. This product operation accurately reflects the attenuation effect of risk as it propagates along a specific path.

[0050] The weighted attenuation coefficient calculation adjusts the node influence value based on the number of hops. The farther away the node, the less influence it has on the target node. The attenuation coefficient reflects this distance-attenuation effect. The attenuation coefficient is set to different values ​​based on the number of hops: 1 for one-hop neighbors, 0.6 for two-hop neighbors, and 0.36 for three-hop neighbors. The attenuation ratio reflects the distance-loss pattern of risk propagation. The weighted calculation multiplies each node's raw influence value by the corresponding attenuation coefficient to obtain a distance-weighted influence value. This value more accurately reflects the actual impact of neighboring nodes at different distances on the target node's risk status. The attenuation coefficient is based on the actual patterns of credit risk propagation: direct connections have the strongest influence, while indirect connections rapidly decay with increasing distance. The influence of connections beyond three hops is generally negligible.

[0051] Matrix permutation and normalization convert the distance-weighted influence into a standardized matrix format. The rows of the matrix correspond to all entities in the knowledge graph, and the columns also correspond to all entities. Each element records the risk propagation intensity of the corresponding row entity to the column entity. The permutation process constructs the matrix according to the order of the unique identifiers of the entities in the knowledge graph to ensure the consistency and repeatability of the matrix structure. The normalization process numerically scales each row of the matrix so that the sum of the elements in each row is equal to one. In this way, the risk propagation intensity distribution of each entity forms a probability distribution, which is convenient for subsequent probability calculation and risk aggregation. The normalization formula divides each element by the sum of the elements in the corresponding row, eliminating the deviation caused by the absolute numerical differences in the influence of different entities and highlighting the distribution pattern of relative influence intensity.

[0052] In a specific embodiment, the process of executing step S104 may specifically include the following steps: Input the credit subject risk state vector into the bidirectional temporal encoder for forward historical trajectory encoding to obtain the historical risk evolution vector; Based on the historical risk evolution vector, the future risk trend is encoded through the backward prediction trend encoder to obtain the predicted risk trend vector; Based on the historical risk evolution vector and the predicted risk trend vector, the guarantee chain network is processed for temporal relationship extraction to obtain the credit history dimension characteristics and social relationship dimension characteristics; Use the predicted risk trend vector to analyze the transaction behavior pattern of the capital flow map to obtain the dimension characteristics of financial status and the dimension characteristics of behavior pattern; Based on the historical risk evolution vector, the macro-environmental factors of the risk factor map are extracted and processed to obtain the external environment dimension characteristics; The credit history dimension features, social relationship dimension features, financial status dimension features, behavioral pattern dimension features and external environment dimension features are processed by vector splicing and feature fusion to obtain a comprehensive risk feature vector.

[0053] Specifically, the forward historical trajectory encoding process begins by inputting the credit entity's risk state vector into a bidirectional temporal encoder. The bidirectional temporal encoder is a neural network structure capable of processing both historical sequences and future predictions. It consists of two components: a forward encoder and a backward encoder. The forward historical trajectory encoding process arranges the credit entity's risk state vector in chronological order, starting from the earliest historical moment and progressing towards the current moment. The encoder's hidden state is updated at each time step, accumulating information from all previous time steps. During the encoding process, the encoder reads the risk state vector at each time point and uses a gating mechanism to determine which historical information to retain and which outdated information to discard. The gating mechanism consists of three components: an input gate, a forget gate, and an output gate, which control the input of new information, the forgetting of old information, and the output of the current state, respectively. The historical risk evolution vector contains the complete risk trajectory of the borrower from the beginning to the current moment. This vector encodes the upward and downward trends in risk levels, fluctuation patterns, and the impact of key risk events.

[0054] The backward prediction trend coding process models future risk trends based on historical risk evolution vectors using a backward prediction trend encoder. The backward encoder uses a time direction opposite to the forward encoder, starting from the current moment and moving forward into the future. The prediction trend coding process combines historical risk evolution patterns with information about the current market environment, predicting future risk development directions by learning from the cyclical patterns and trend characteristics in historical data. During processing, the encoder identifies seasonal patterns in risk changes, cyclical fluctuations, and the impact of external environmental changes on risk trends, encoding these complex time dependencies into fixed-length vector representations. The predicted risk trend vector reflects the expected direction of change in the borrower's risk level over a period of time, including the probability of increased risk, the duration of risk stability, and the timing of potential risk peaks.

[0055] Temporal relationship extraction processing conducts in-depth analysis of the guarantee chain network based on historical risk evolution vectors and predicted risk trend vectors, extracting guarantee relationship features related to time changes. Temporal relationship extraction is a feature extraction technology specifically designed for dynamic network structures, capable of identifying patterns in how nodes and edges in the network change over time. During processing, the algorithm analyzes changes in the topological structure of the guarantee chain network in different historical periods, identifying new guarantee relationships, expired guarantee relationships, and dynamic adjustments to guarantee strength. Credit history dimension features are derived by analyzing the historical role changes of the borrower in the guarantee chain, including changes in the frequency of acting as a guarantor, the evolution of the guaranteed party's credit status, and the temporal distribution pattern of the guarantee amount. Social relationship dimension features are derived by analyzing the stability of the relationship between the borrower and other entities in the guarantee network, including the identification of long-term guarantee partners, changes in the concentration of guarantee relationships, and the expansion or contraction trend of the guarantee network.

[0056] Transaction behavior pattern analysis and processing applies predicted risk trend vectors to deep mining of the capital flow map, analyzing how the borrower's capital flow behavior will change under future trends. Transaction behavior pattern analysis is a graph-based behavior recognition technology that identifies abnormal behavior patterns by analyzing the paths, frequency, and scale of capital flows between nodes. The processing combines predicted risk trend information to evaluate the borrower's capital management behavior under different risk scenarios and identify the regularity, suddenness, and abnormality of capital flows. Financial status dimension features are derived by analyzing the inflow and outflow balance relationship in the capital flow map, including the stability of capital sources, the rationality of capital use, and the cyclical changes in cash flow. Behavioral pattern dimension features are derived by analyzing transaction time distribution, counterparty selection, and changes in transaction scale, reflecting the borrower's operating behavior characteristics and risk appetite.

[0057] The macro-environmental factor extraction process systematically analyzes the risk factor map based on historical risk evolution vectors to extract risk characteristics related to changes in the external environment. Macro-environmental factor extraction is an analytical technique that links micro-individual risks with macro-environmental changes. It predicts the impact of environmental factors by identifying the correlation patterns between historical risk evolution and external environmental changes. During the processing process, the algorithm analyzes the correlation between the borrower's historical risk changes and macro factors such as industry cycles, economic policies, and market fluctuations, identifying the external factors that most significantly affect the risk status of the borrower. The external environment dimension characteristics include indicators such as industry risk exposure, policy sensitivity, market volatility responsiveness, and system and storage medium risk correlation. These characteristics reflect the extent and manner in which the borrower is affected by changes in the external environment.

[0058] Vector concatenation and feature fusion combine the five-dimensional feature vectors in a predetermined order to form a unified comprehensive risk feature vector. Vector concatenation connects multiple independent vectors into a longer vector in the order of their dimensions. The concatenation order follows the credit history dimension, social relationship dimension, financial status dimension, behavioral pattern dimension, and external environment dimension. Feature fusion, based on concatenation, analyzes correlations and adjusts weights between dimensions. It identifies redundant and complementary information by calculating correlation coefficients between features from different dimensions. The fusion process utilizes a combination of weighted averaging and attention mechanisms, assigning higher weights to important features and downgrading redundant features to ensure that the comprehensive risk feature vector comprehensively and efficiently represents the borrower's multi-dimensional risk profile.

[0059] In a specific embodiment, the process of executing step S105 may specifically include the following steps: The comprehensive risk feature vector is input into the multi-head attention network to generate query vector, key vector and value vector to obtain the attention calculation matrix; Based on the attention calculation matrix, the attention weight is normalized by the softmax function to obtain the standardized attention weight; Perform weighted fusion processing on different risk dimensions according to the standardized attention weights to obtain the aggregated risk feature vector; The aggregated risk feature vector is input into the Bayesian neural network for uncertainty quantification to obtain the risk assessment mean and variance; Confidence interval calculation and risk level classification are performed based on the risk assessment mean and variance to obtain the credit risk assessment results and confidence intervals.

[0060] Specifically, the query vector, key vector, and value vector generation process begins by inputting the comprehensive risk feature vector into a multi-head attention network. The multi-head attention network is a neural network architecture capable of processing multiple attention subspaces in parallel and consists of multiple independent attention heads. The query vector, key vector, and value vector generation process transforms the input comprehensive risk feature vector into a query vector, a key vector, and a value vector, respectively, using three different linear transformation matrices. The dimensions of each vector are divided according to the number of attention heads. The query vector represents the risk dimension information currently under consideration, the key vector represents the risk feature pattern available for matching, and the value vector contains the actual risk feature value. During the generation process, the credit history dimension of the comprehensive risk feature vector is generated into a corresponding query vector using the query transformation matrix, the social relationship dimension is generated into a key vector using the key transformation matrix, and the financial status, behavioral pattern, and external environment dimensions are generated into value vectors using the value transformation matrix. The attention calculation matrix is ​​generated by calculating the dot product of the query vector and the key vector. Each element in the matrix represents the correlation strength between the corresponding query and key. Dimension combinations with higher correlations receive higher attention weights in subsequent processing.

[0061] Attention weight normalization is based on the numerical normalization of the attention matrix using the softmax function. The softmax function is a mathematical function that converts any real vector into a probability distribution, ensuring that the sum of all weights equals one. During normalization, the softmax function first applies an exponential operation to each element in the attention matrix, then divides each exponential value by the sum of all exponential values ​​to obtain a normalized probability distribution. The mathematical representation of the function uses the relative comparison mechanism of exponential operation. Elements with larger values ​​retain their relative advantage after normalization, but all elements are compressed to the range of zero to one and sum to one. Normalized attention weights reflect the distribution of importance between different risk dimensions. When the credit history dimension is highly correlated with the current risk assessment, the corresponding attention weight will be significantly higher than other dimensions. When multiple dimensions are equally influential, the weight distribution will be relatively uniform. The normalization process eliminates absolute numerical differences in the original relevance scores and highlights the distribution pattern of relative importance.

[0062] Weighted fusion calculates linear combinations of different risk dimensions based on standardized attention weights. The fusion process multiplies the feature vector of each risk dimension by its corresponding attention weight, and then sums all weighted vectors element-wise. Weighted fusion is a feature combination technique based on importance weights, where dimensions of higher importance have a greater weight in the final result, while dimensions of lower importance have a relatively smaller impact. During the fusion process, each element of the feature vector of the credit history dimension is multiplied by its corresponding attention weight. The feature vectors of the social relationship dimension, financial status dimension, behavioral pattern dimension, and external environment dimension are also multiplied by the same weights. The five weighted vectors are then summed according to their corresponding positions. The aggregated risk feature vector is the output of the fusion process. This vector integrates information from all risk dimensions while highlighting the most critical risk features based on the attention weights. Each element of the vector reflects the weighted contribution of features from multiple dimensions.

[0063] Uncertainty quantification involves inputting the aggregated risk feature vector into a Bayesian neural network for probability distribution modeling. A Bayesian neural network is a machine learning model that incorporates probability theory into neural networks. The network's weights and biases are modeled as probability distributions rather than deterministic numerical values. During the quantification process, the network generates multiple predictions through multiple forward propagations. Each propagation samples different weight values ​​from the probability distribution of the weights to form a distribution of predictions. The advantage of Bayesian neural networks is that they can simultaneously output both the prediction results and the uncertainty of the predictions. Uncertainty is categorized into epistemic uncertainty and stochastic uncertainty. Epistemic uncertainty arises from the model's inadequate understanding of the data, while stochastic uncertainty arises from inherent noise in the data. The risk assessment mean is calculated by averaging multiple predictions and reflects the model's best estimate of the borrower's risk level. The risk assessment variance is calculated by calculating the variance of multiple predictions and reflects the uncertainty of the model's predictions.

[0064] Confidence interval calculation and risk classification are based on statistical inference based on the risk assessment mean and variance. A confidence interval is a statistical method used to estimate the range of a population parameter, indicating the probability that the true risk value falls within that interval. The calculation assumes a normal distribution. The lower bound of the confidence interval is equal to the risk assessment mean minus the standard deviation corresponding to the confidence level multiplied by the risk assessment standard deviation. The upper bound is equal to the risk assessment mean plus the same value. The confidence level is typically set at 95%, corresponding to a standard deviation multiple of approximately 1.96. When the risk assessment variance is small, the confidence interval is narrow, indicating greater predictive certainty. When the variance is large, the confidence interval is wide, indicating greater predictive uncertainty. Risk classification is determined based on the risk assessment mean and confidence interval. Low risk corresponds to situations where the risk assessment mean is low and the upper bound of the confidence interval is still within a safe range; medium risk corresponds to situations where the risk assessment mean is moderate; and high risk corresponds to situations where the risk assessment mean is high or the lower bound of the confidence interval exceeds the risk threshold.

[0065] The above describes the credit risk assessment method based on the knowledge graph in the embodiment of the present application. The following describes the credit risk assessment system based on the knowledge graph in the embodiment of the present application. Figure 2 In the embodiments of the present application, an embodiment of a credit risk assessment system based on a knowledge graph includes: The processing module is used to perform semantic standardization on multi-source heterogeneous credit data through ontology mapping rules to obtain a credit subject triple dataset in RDF format; A construction module is used to construct a dynamic graph of entity relationships using a self-organizing algorithm based on the credit subject triple data set to obtain a multi-level dynamic knowledge graph including a guarantee chain network, a capital flow graph, and a risk factor graph; A modeling module, configured to perform risk state modeling processing on the multi-level dynamic knowledge graph through a graph convolution risk propagation operator to obtain a credit subject risk state vector; An extraction module is used to perform time-series embedding and extraction processing on the guarantee chain network, capital flow map, and risk factor map based on the credit subject risk status vector to obtain a comprehensive risk feature vector in the credit history dimension, social relationship dimension, financial status dimension, behavioral pattern dimension, and external environment dimension; The aggregation module is used to perform risk aggregation processing on the comprehensive risk feature vector through an adaptive attention mechanism to obtain a credit risk assessment result and a confidence interval.

[0066] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the knowledge graph-based credit risk assessment method.

[0067] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0068] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a knowledge graph-based credit risk assessment device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0069] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. The credit risk assessment method based on knowledge graph is characterized by: The method comprises: The semantic standardization of multi-source heterogeneous credit data is carried out through ontology mapping rules to obtain a credit subject triple dataset in RDF format; Based on the credit subject triple data set, a self-organizing algorithm is used to construct a dynamic graph of entity relationships to obtain a multi-level dynamic knowledge graph including a guarantee chain network, a capital flow graph, and a risk factor graph; Performing risk state modeling on the multi-level dynamic knowledge graph through a graph convolution risk propagation operator to obtain a credit subject risk state vector; Based on the credit subject risk status vector, the guarantee chain network, capital flow map and risk factor map are subjected to time series embedding and extraction processing to obtain a comprehensive risk feature vector in the credit history dimension, social relationship dimension, financial status dimension, behavior pattern dimension and external environment dimension; The comprehensive risk feature vector is subjected to risk aggregation processing through an adaptive attention mechanism to obtain a credit risk assessment result and a confidence interval.

2. The credit risk assessment method based on knowledge graph according to claim 1, characterized in that: The semantic standardization of multi-source heterogeneous credit data is performed through ontology mapping rules to obtain a credit subject triple dataset in RDF format, including: Establish a credit domain ontology concept system, define the semantic structure of entity classes, relationship classes, and attribute classes, and obtain an ontology library; Perform entity recognition processing on the credit data, transaction data, behavior data and related party data based on the ontology library to obtain an entity identification set; Perform entity disambiguation processing on the entity identifier set using an edit distance algorithm to obtain an entity mapping table; Performing relationship extraction and attribute extraction processing on multi-source data according to the entity mapping table to obtain relationship attribute pairs; The relationship attribute pairs are converted into triple format and constructed to obtain a credit subject triple dataset in RDF format.

3. The credit risk assessment method based on knowledge graph according to claim 1, characterized in that: The self-organizing algorithm is used to construct a dynamic graph of entity relationships based on the credit subject triple data set to obtain a multi-level dynamic knowledge graph including a guarantee chain network, a capital flow graph, and a risk factor graph, including: Inputting the credit subject triple data set into the input layer of the self-organizing neural network for feature vector encoding processing to obtain entity embedding vectors and relationship embedding vectors; Based on the entity embedding vector and the relationship embedding vector, the Euclidean distance is calculated through each neuron in the competition layer to perform optimal matching processing to obtain the winning neuron position; According to the position of the winning neuron, a neighborhood function is used to perform a learning and updating process on the weights of the surrounding neurons to obtain an adaptive weight matrix; The adaptive weight matrix is ​​classified and mapped according to entity type and relationship type through the output layer to obtain a guarantee chain network, a capital flow map and a risk factor map; The guarantee chain network, capital flow map and risk factor map are hierarchically combined to obtain a multi-level dynamic knowledge map including the guarantee chain network, capital flow map and risk factor map.

4. The credit risk assessment method based on knowledge graph according to claim 1, characterized in that: The multi-level dynamic knowledge graph is subjected to risk state modeling processing through a graph convolution risk propagation operator to obtain a credit subject risk state vector, including: Convolution kernels of different sizes are designed for the guarantee chain network, capital flow map, and risk factor map to perform subgraph feature extraction processing to obtain guarantee relationship features, capital flow features, and risk factor features; The guarantee relationship characteristics, capital flow characteristics and risk factor characteristics are vectorized through multi-dimensional feature coding to obtain a subgraph embedding vector; Based on the subgraph embedding vector, a quantum superposition state modeling method is used to perform multi-state combination processing to obtain a composite risk state set; Performing historical risk adjustment processing based on the composite risk state set by time decay weight calculation to obtain a time series risk weight matrix; Inputting the temporal risk weight matrix into the risk propagation operator to perform multi-hop neighbor influence calculation processing to obtain a risk propagation intensity matrix; The risk propagation intensity matrix is ​​subjected to vectorized projection and feature fusion processing to obtain a credit subject risk state vector.

5. The credit risk assessment method based on knowledge graph according to claim 4 is characterized in that: The temporal risk weight matrix is ​​input into the risk propagation operator to perform multi-hop neighbor influence calculation processing to obtain a risk propagation intensity matrix, including: Inputting the temporal risk weight matrix into the risk propagation operator to perform one-hop direct neighbor node identification processing to obtain a set of directly related nodes; Performing two-hop indirect neighbor node expansion processing based on the directly associated node set through a graph traversal algorithm to obtain a secondary associated node set; Continue to perform graph traversal expansion according to the second-level associated node set to perform three-hop neighbor node discovery processing to obtain a third-level associated node set; Calculating the product of the path weight between each node in the directly associated node set, the secondary associated node set, and the tertiary associated node set and the target node to perform influence quantification processing to obtain a node influence value; The node influence value is weighted by the attenuation coefficient according to the hop distance to obtain the distance-weighted influence; The distance-weighted influences are matrix-arranged and normalized to obtain a risk propagation intensity matrix.

6. The credit risk assessment method based on knowledge graph according to claim 1, characterized in that: The guarantee chain network, capital flow map and risk factor map are subjected to time series embedding and extraction processing based on the credit subject risk status vector to obtain a comprehensive risk feature vector in the credit history dimension, social relationship dimension, financial status dimension, behavior pattern dimension and external environment dimension, including: Inputting the credit subject risk state vector into a bidirectional temporal encoder for forward historical trajectory encoding processing to obtain a historical risk evolution vector; Based on the historical risk evolution vector, a future risk trend encoding process is performed using a backward prediction trend encoder to obtain a predicted risk trend vector; Extracting time series relationships from the guarantee chain network based on the historical risk evolution vector and the predicted risk trend vector to obtain credit history dimension features and social relationship dimension features; Performing transaction behavior pattern analysis on the fund flow map using the predicted risk trend vector to obtain financial status dimension features and behavior pattern dimension features; Based on the historical risk evolution vector, the risk factor map is subjected to macro-environmental factor extraction processing to obtain external environment dimension characteristics; Vector splicing and feature fusion processing are performed on the credit history dimension features, social relationship dimension features, financial status dimension features, behavioral pattern dimension features and external environment dimension features to obtain a comprehensive risk feature vector.

7. The credit risk assessment method based on knowledge graph according to claim 1, characterized in that: The comprehensive risk feature vector is subjected to risk aggregation processing through an adaptive attention mechanism to obtain a credit risk assessment result and a confidence interval, including: Inputting the comprehensive risk feature vector into a multi-head attention network to generate a query vector, a key vector, and a value vector to obtain an attention calculation matrix; Performing attention weight normalization processing based on the attention calculation matrix through the softmax function to obtain a standardized attention weight; Performing weighted fusion processing on different risk dimensions according to the standardized attention weights to obtain an aggregated risk feature vector; Inputting the aggregated risk feature vector into a Bayesian neural network for uncertainty quantification to obtain a risk assessment mean and variance; Confidence interval calculation and risk level classification are performed based on the risk assessment mean and variance to obtain a credit risk assessment result and a confidence interval.

8. A credit risk assessment system based on knowledge graph, characterized in that: For implementing the knowledge graph-based credit risk assessment method according to any one of claims 1 to 7, the knowledge graph-based credit risk assessment system comprises: The processing module is used to perform semantic standardization on multi-source heterogeneous credit data through ontology mapping rules to obtain a credit subject triple dataset in RDF format; A construction module is used to construct a dynamic graph of entity relationships using a self-organizing algorithm based on the credit subject triple data set to obtain a multi-level dynamic knowledge graph including a guarantee chain network, a capital flow graph, and a risk factor graph; A modeling module, configured to perform risk state modeling processing on the multi-level dynamic knowledge graph through a graph convolution risk propagation operator to obtain a credit subject risk state vector; An extraction module is used to perform time-series embedding and extraction processing on the guarantee chain network, capital flow map, and risk factor map based on the credit subject risk status vector to obtain a comprehensive risk feature vector in the credit history dimension, social relationship dimension, financial status dimension, behavioral pattern dimension, and external environment dimension; The aggregation module is used to perform risk aggregation processing on the comprehensive risk feature vector through an adaptive attention mechanism to obtain a credit risk assessment result and a confidence interval.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is enabled to perform the credit risk assessment method based on a knowledge graph according to any one of claims 1 to 7.

Citation Information

Cited By

  • Credit score simulation method and device based on large model

    CN121168616A

  • Credit scoring simulation method and device based on large model

    CN121168616B

  • Knowledge graph modeling-based energy consumption risk index system construction method, system and equipment and medium

    CN121169106A

  • Credit assessment method and system based on multi-modal data

    CN121213229A

  • Component risk prediction method and system fused with time decay factor, and storage medium

    CN121304309A