A cross-business credit risk assessment method and system based on a graph neural network and a bank joint knowledge graph

CN122390861BActive Publication Date: 2026-08-21NANJING DAYAN DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610865907.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-08-21
Estimated Expiration
2046-06-16

AI Technical Summary

Technical Problem

[0007]鉴于现有银行信用风险评估技术存在多业务数据融合不足、跨业务关联关系刻画不完整、风险传导路径难以识别以及路径风险贡献缺乏动态校准的问题,提出了本发明

Benefits of technology

[0013]与现有技术相比,本发明有益效果为:通过采集银行综合业务管理系统、对公业务系统和交易风险评估系统中的多业务客户数据,构建多源异构原始数据集,使授信、交易、担保、股权及风控数据形成统一的数据基础,避免单一业务数据导致的风险识别片面问题;通过以客户实体和业务事件为图节点、以跨业务关联关系为图边并配置时间衰减权重,构建跨业务联合知识图谱,使客户间担保传导、股权穿透和资金往来关系能够按照关联类型和时间有效性进行表达,提高风险关系建模的完整性;通过预设担保链、股权穿透和资金往来三类业务元路径,并计算节点级注意力系数和语义级注意力系数,得到客户融合嵌入向量,使不同邻居节点和不同业务路径对客户风险的贡献能够被差异化度量,提高客户风险表征的准确性;通过对候选风险传导路径进行路径遮蔽对比,计算遮蔽前后的风险表示差异,并结合时间衰减权重生成路径风险确认系数,使风险传导路径能够经过影响程度校验,减少虚假关联、弱关联或过期关联对风险评分的干扰;通过将客户融合嵌入向量输入图神经网络模型,并按照路径风险确认系数进行跨业务风险聚合,得到跨业务风险特征向量,计算客户信用风险评分并生成正常级、关注级和预警级的分级预警结果,使风险评估结果兼具评分判断、路径解释和预警处置依据,从而提高跨业务信用风险识别的准确性、时效性和可解释性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122390861B_ABST
    Figure CN122390861B_ABST
Patent Text Reader

Abstract

The application discloses a cross-business credit risk assessment method and system based on a graph neural network and a bank joint knowledge graph, relates to the technical field of financial business data processing and bank credit risk assessment, and comprises the following steps: collecting multi-business customer data, and constructing a multi-source heterogeneous original data set; constructing a cross-business joint knowledge graph based on a customer entity, a business event and a cross-business association relationship, and configuring a time decay weight; extracting a customer fusion embedding vector and a candidate risk transmission path set based on a guarantee chain, equity penetration and a fund flow meta path; generating a path risk confirmation coefficient through path masking comparison, and performing cross-business risk aggregation based on the path risk confirmation coefficient to obtain a customer credit risk score and a grading early warning result. The application improves the accuracy, interpretability, early warning timeliness and the like of cross-business credit risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of financial business data processing and bank credit risk assessment technology, and in particular to a cross-business credit risk assessment method and system based on graph neural networks and bank joint knowledge graphs. Background Technology

[0002] With the development of digital operations and integrated credit management in commercial banks, customer credit risk assessment has gradually evolved from traditional financial indicator scoring and rule-based threshold judgments to intelligent methods such as knowledge graphs, machine learning, and graph neural networks. Bank customers form complex relationships across various business areas, including credit, guarantees, transactions, equity, and fund transfers. Risks are often not isolated events but are implicitly transmitted along relationship networks such as guarantee chains, equity control chains, and funding chains. While existing credit risk analysis methods can improve scoring efficiency by utilizing basic customer information, transaction data, or partial relationship graphs, they still have shortcomings in multi-business data fusion, cross-business risk transmission path identification, risk path contribution confirmation, and time decay impact modeling. This can easily lead to models providing only single-point risk scores, failing to explain the source of risk and its cross-business diffusion mechanisms.

[0003] CN115471323A discloses a method and apparatus for analyzing credit risk of bank customers. It primarily uses encrypted credit rating prediction model parameters, a time-recurrent neural network computation unit, and a secure multi-party computation method to obtain encrypted personal risk levels while protecting user privacy. Furthermore, it enables a banking business server to construct a customer credit knowledge graph based on the personal risk level and unencrypted user personal data. While this solution focuses on calculating customer risk levels under privacy protection, thus improving data security, its credit knowledge graph is mainly built around customer personal data and risk levels. It does not form a unified cross-business joint knowledge graph for multiple business scenarios such as comprehensive banking business, corporate banking business, and transaction risk assessment. It also does not perform layered modeling of typical business meta-paths such as guarantee chains, equity penetration, and fund transfers. Therefore, it is difficult to characterize the transmission direction and contribution differences of risk between different business relationships.

[0004] CN116012140A discloses a credit scoring method and apparatus for corporate bank clients. It obtains a set of corporate clients and their information, determines the industry relationships between these clients, constructs a relationship graph based on these relationships, and then uses this graph to train a graph neural network architecture for credit scoring. While this approach can improve the efficiency and accuracy of credit scoring for corporate clients by leveraging industry relationships, its relationship graph primarily reflects industry connections between clients, with relatively simple semantics at the edges. It does not further incorporate cross-business relationships such as guarantee relationships, equity penetration relationships, and fund transfer relationships, nor does it adjust the validity of historical relationships using time decay weights. Furthermore, it fails to confirm the actual impact of candidate risk transmission paths on risk representation through path masking comparison, resulting in insufficient calibration and explanatory capabilities for cross-business risk transmission chains.

[0005] Therefore, existing bank credit risk assessment technologies suffer from insufficient integration of multi-business data, inadequate identification of cross-business risk transmission paths, and difficulty in calibrating relationship timeliness and path risk contribution. This invention provides a cross-business credit risk assessment method and system based on graph neural networks and a joint bank knowledge graph. It incorporates customer entities, business events, and cross-business relationships into the graph structure and combines time decay weights, business element path attention, path masking comparison, and path risk confirmation coefficients to identify, confirm, and calibrate risk transmission paths. This addresses the lack of transmission explanation and dynamic calibration in credit risk scoring across bank business scenarios. Summary of the Invention

[0006] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract and title of the invention. Such simplifications or omissions shall not be used to limit the scope of the present invention.

[0007] In view of the problems of insufficient integration of multi-business data, incomplete characterization of cross-business relationships, difficulty in identifying risk transmission paths, and lack of dynamic calibration of path risk contribution in existing bank credit risk assessment technologies, this invention is proposed.

[0008] Therefore, the problem to be solved by this invention is how to construct a joint knowledge graph that can represent customer entities, business events and cross-business relationships based on multi-source heterogeneous data such as comprehensive banking business, corporate business and transaction risk assessment, and use graph neural networks to identify, confirm and aggregate cross-business risk transmission paths such as guarantee chains, equity penetration and fund transfers, so as to improve the accuracy, interpretability and timeliness of customer credit risk assessment.

[0009] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, embodiments of the present invention provide a cross-business credit risk assessment method based on graph neural networks and a joint bank knowledge graph, comprising, Collect customer data from multiple services, extract customer entity sets, business event sets, and cross-business relationship sets, and construct a multi-source heterogeneous raw dataset; A cross-business joint knowledge graph is constructed by using the set of customer entities and the set of business events as graph nodes, the set of cross-business relationships as graph edges, and configuring time decay weights for the graph edges. For cross-business joint knowledge graphs, three types of business meta-paths are preset: guarantee chain, equity penetration and fund transfer. Node-level attention coefficient and semantic-level attention coefficient are calculated to obtain customer fusion embedding vector and candidate risk transmission path set. The candidate risk transmission path set is compared by path masking, the difference in risk representation before and after path masking is calculated, and the path risk confirmation coefficient is generated by combining time decay weight. The path risk confirmation coefficient is used as the risk transmission calibration parameter. The customer is fused into the vector input graph neural network model, and cross-business risk is aggregated according to the path risk confirmation coefficient to obtain cross-business risk feature vector. The customer credit risk score is calculated and a graded early warning result is generated.

[0010] Secondly, embodiments of the present invention provide a cross-business credit risk assessment system based on graph neural networks and a joint bank knowledge graph, comprising: The multi-source data construction module is used to collect customer data from multiple businesses, extract customer entity sets, business event sets, and cross-business relationship sets, and construct multi-source heterogeneous raw datasets. The joint graph construction module is used to construct a cross-business joint knowledge graph using customer entity sets and business event sets as graph nodes, cross-business relationship sets as graph edges, and time decay weights configured for graph edges. The risk path identification module is used to pre-set three types of business meta-paths for cross-business joint knowledge graphs: guarantee chain, equity penetration and fund transfer. It calculates node-level attention coefficients and semantic-level attention coefficients to obtain customer fusion embedding vectors and candidate risk transmission path sets. The path risk calibration module is used to compare the path masking of the candidate risk transmission path set, calculate the difference in risk representation before and after path masking, and generate the path risk confirmation coefficient by combining the time decay weight. The path risk confirmation coefficient is used as the risk transmission calibration parameter. The risk scoring and early warning module is used to integrate customers into a vector input graph neural network model, aggregate cross-business risks according to the path risk confirmation coefficient, obtain cross-business risk feature vectors, calculate customer credit risk scores, and generate graded early warning results.

[0011] Thirdly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any step of the above-described cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs.

[0012] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the above-described cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs.

[0013] Compared with existing technologies, the beneficial effects of this invention are as follows: By collecting multi-business customer data from the bank's integrated business management system, corporate business system, and transaction risk assessment system, a multi-source heterogeneous original dataset is constructed, enabling credit granting, transaction, guarantee, equity, and risk control data to form a unified data foundation, avoiding the one-sided risk identification problem caused by single business data; by constructing a cross-business joint knowledge graph with customer entities and business events as graph nodes, cross-business relationships as graph edges, and configuring time decay weights, the relationship between customer guarantee transmission, equity penetration, and fund transfers can be expressed according to the association type and time validity, improving the completeness of risk relationship modeling; by pre-setting three types of business meta-paths—guarantee chain, equity penetration, and fund transfer—and calculating node-level attention coefficients and semantic-level attention coefficients, a customer fusion embedding vector is obtained, enabling... The contributions of neighboring nodes and different business paths to customer risk can be measured differentially, improving the accuracy of customer risk characterization. By comparing the path masking of candidate risk transmission paths, the difference in risk representation before and after masking is calculated, and a path risk confirmation coefficient is generated by combining time decay weights. This allows the risk transmission path to pass the impact degree verification, reducing the interference of false associations, weak associations, or expired associations on risk scoring. By embedding customer data into a vector input graph neural network model and aggregating cross-business risks according to the path risk confirmation coefficient, a cross-business risk feature vector is obtained. This vector is used to calculate customer credit risk scores and generate graded early warning results of normal, attention, and warning levels. This makes the risk assessment results have the basis for scoring judgment, path interpretation, and early warning handling, thereby improving the accuracy, timeliness, and interpretability of cross-business credit risk identification. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 The flowchart shows a cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs. Figure 2 This is a structural diagram of a cross-business credit risk assessment system based on graph neural networks and a bank joint knowledge graph. Detailed Implementation

[0015] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0016] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort should fall within the scope of protection of this invention.

[0017] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0018] As mentioned in the background section, while existing technologies can score credit risk based on customer information, industry relationships, or privacy-preserving calculation methods, they typically focus on risk assessment under a single business dimension or relationship type, failing to reflect the complex relationships among bank customers across multiple business scenarios such as credit granting, guarantees, transactions, equity control, and fund transfers. Furthermore, existing methods do not adequately consider the timeliness of historical business relationships, lack a mechanism for dynamically correcting risk transmission relationships using time decay weights, and lack technical means to verify the actual contribution of candidate risk transmission paths through path masking comparisons. This results in credit risk scoring results that tend to remain at the static prediction level, making it difficult to clearly identify the source of risk, the transmission chain, and key impact paths. To address these issues, this invention provides a cross-business credit risk assessment method based on graph neural networks and a joint bank knowledge graph.

[0019] Reference Figures 1-2 , Figure 1 This is a flowchart illustrating a cross-business credit risk assessment method based on graph neural networks and a joint bank knowledge graph, according to an embodiment of the present invention. Figure 1 As shown, a cross-business credit risk assessment method based on graph neural networks and a joint bank knowledge graph includes: S1: Collect customer data from multiple services, extract customer entity sets, business event sets, and cross-business relationship sets, and construct multi-source heterogeneous raw datasets; S2: Construct a cross-business joint knowledge graph using the customer entity set and business event set as graph nodes and the cross-business relationship set as graph edges, and configure time decay weights for the graph edges; S3: For cross-business joint knowledge graphs, three types of business meta-paths are preset: guarantee chain, equity penetration and fund transfer. Node-level attention coefficient and semantic-level attention coefficient are calculated to obtain customer fusion embedding vector and candidate risk transmission path set. S4: Compare the path masking of the candidate risk transmission path set, calculate the difference in risk representation before and after path masking, and generate the path risk confirmation coefficient by combining the time decay weight. Use the path risk confirmation coefficient as the risk transmission calibration parameter. S5: The customer is fused into the vector input graph neural network model, and cross-business risk is aggregated according to the path risk confirmation coefficient to obtain cross-business risk feature vector, calculate customer credit risk score and generate graded early warning results.

[0020] In this embodiment of the application, step S1 includes: S1.1: Initiate data interface calls to the bank's integrated business management system, corporate business system, and transaction risk assessment system respectively, and obtain multi-business customer data according to the preset collection cycle.

[0021] Specifically, multi-business customer data includes basic customer information, credit records, transaction history, guarantee relationships, and equity relationship data. Basic customer information includes customer ID, registered capital, years of operation, and legal representative identification. Credit records include credit limit, credit period, and number of overdue payments. Transaction history includes transaction timestamps, transaction amounts, and counterparty IDs. Guarantee relationships include guarantor ID, guaranteed party ID, and guarantee amount. Equity relationship data includes shareholder customer ID, invested customer ID, shareholding ratio, voting rights ratio, actual control relationship identifier, and equity registration timestamp. This equity relationship data originates from the customer equity relationship table or group customer association table in the corporate banking system. When the same equity relationship simultaneously involves shareholding ratio and voting rights ratio, the voting rights ratio is used as the basis for determining the equity penetration boundary.

[0022] S1.2: Perform missing value detection on multi-business customer data, mark records in multi-business customer data whose field missing rate exceeds the preset missing threshold as low-quality records and remove them, and construct source tag fields for the retained multi-business customer data according to the source identifiers of the bank's integrated business management system, corporate business system and transaction risk assessment system, respectively, to obtain multi-business customer data with source tags.

[0023] It should be noted that the preset missing threshold is uniformly configured by the data quality management module of the bank's big data analysis and early warning platform, with a default value of 30%. The basis for determining this default value is as follows: In banking business data practice, when the missing rate of key fields (such as registered capital, credit line, and transaction amount) exceeds 30%, the contribution of the record to risk modeling decreases significantly, and forcibly retaining it will introduce a large imputation error; while a missing rate of less than 30% can be repaired by mean imputation or median imputation without affecting the calculation quality of subsequent feature vectors; the above threshold can be dynamically adjusted by the data quality management module within the range of 10% to 50% according to the actual data integrity level of each bank.

[0024] S1.3: Using the customer ID and legal entity identifier in the multi-business customer data with source tags as the primary key, the entity alignment rules are used to merge customer records with the same name but different codes in the bank's integrated business management system, corporate business system and transaction risk assessment system to obtain a deduplicated set of customer entities.

[0025] Furthermore, the entity alignment rule, based on the combined hash value of the prefix organization code of the customer number and the legal entity identifier, uniquely merges duplicate customer records across systems. Each customer entity in the customer entity set carries a unified global customer identifier. The extraction rule for the prefix organization code is as follows: the first non-numeric separator (including hyphens, underscores, or forward slashes) in the customer number string is used as the truncation point, and the substring before the truncation point is taken as the prefix organization code; if there are no non-numeric separators in the customer number, the first six characters of the customer number are taken as the prefix organization code. These extraction rules are uniformly configured by the data quality management module of the bank's big data analysis and early warning platform and can be dynamically adjusted according to the numbering specifications of each business system.

[0026] S1.4: Using credit records and transaction flows in multi-business customer data with source tags as the data source, extract credit granting events and overdue default events from the credit records according to triples, and extract fund transfer events and abnormal transaction events from the transaction flows. Summarize the above four types of events into a business event set.

[0027] Preferably, the triplet format is event type-occurrence time-associated customer entity; the determination of abnormal transaction events is based on a preset multiple threshold that the amount of a single transaction exceeds the average transaction amount of the corresponding customer in the customer entity set over the past 180 days, where the preset multiple threshold is 5 times. The basis for determining this default value is: referring to the definition of abnormal transactions in the banking industry's abnormal transaction monitoring practice, when the amount of a single transaction reaches 5 times or more of the customer's historical average, the degree of deviation has entered the tail of the distribution in a statistical sense (corresponding to about 4σ beyond the normal distribution), and has significant value in identifying abnormal transactions; if the multiple is too low (such as 2 to 3 times), a large number of normal fluctuations will be misidentified, and if the multiple is too high (such as more than 10 times), the false negative rate will increase significantly. This threshold can be adjusted by the risk control parameter module within the range of 3 to 10 times.

[0028] It should be noted that the average transaction amount over the past 180 days is calculated as follows: using the current transaction timestamp as the cutoff time, tracing back 180 days, the arithmetic average of all transaction amounts for the client entity within the aforementioned time range is calculated; if there are fewer than 30 transaction records within the tracing period, the average is calculated using all available transaction records; if there are no historical transaction records, no abnormal transaction event determination is performed for the client entity, and no abnormal transaction events are generated for all transaction flows of the client entity within the current collection period.

[0029] S1.5: Using the guarantee relationships in the multi-business customer data with source tags as the data source, extract the global customer identifier of the guarantor, the global customer identifier of the guaranteed party, and the guarantee amount to construct a guarantee association triplet; simultaneously, for each guarantee relationship, obtain the registered capital corresponding to the guaranteed party from the customer entity set and calculate the proportion of the guarantee amount to the registered capital of the guaranteed party; if the proportion exceeds 50%, generate a controlling stake identifier and set it to true, otherwise set it to false; if the guaranteed party is an external entity, its registered capital field takes the default value, and the effective proportion of the guarantee amount to the registered capital cannot be calculated. In this case, the controlling stake identifier is forcibly set to false, and the directed weighted edge corresponding to this guarantee relationship will only be included in the guarantee chain edge subset in subsequent steps, not in the equity penetration edge subset; store the controlling stake identifier as an additional field of this guarantee relationship.

[0030] Furthermore, by mapping the counterparty number in the transaction flow to the global customer identifier in the customer entity set, the fund transfer association triplet is extracted; the guarantee association triplet (carrying the holding identifier bit) and the fund transfer association triplet are merged to obtain the cross-business association set, where each association in the cross-business association set carries the source system identifier and transaction timestamp to support the calculation of time decay weight in subsequent steps.

[0031] It should be noted that when extracting the fund transfer triplet, if the counterparty number in the transaction log cannot be found in the customer entity set, the counterparty is regarded as an external entity, a virtual global customer identifier is temporarily generated for the external entity, and the virtual global customer identifier is included in the customer entity set. At the same time, the registered capital, years of operation and legal person identity in its customer basic information are all set to default values ​​so as to preserve the integrity of fund transfer relationships in subsequent graph construction.

[0032] S1.6: The customer entity set, business event set, and cross-business relationship set are associated and indexed according to a unified global customer identifier, and encapsulated into a multi-source heterogeneous original dataset in a structured storage format.

[0033] It should be noted that the multi-source heterogeneous original dataset uses a unified global customer identifier as the main index, and mounts the corresponding customer basic information, credit records, transaction flow, guarantee relationship and cross-business relationship set, which can be directly called for the construction of cross-business joint knowledge graph.

[0034] In this embodiment of the application, step S2 includes: S2.1: Read the customer entity set and business event set from the multi-source heterogeneous original dataset. Use the unified global customer identifier of each customer entity in the customer entity set as the customer class node identifier and the triple of each business event in the business event set as the event class node identifier. Initialize the customer class node feature vector and the event class node feature vector respectively.

[0035] Specifically, the feature vector of customer-type nodes is obtained by concatenating the registered capital, years of operation, and legal person identification from the customer's basic information after normalization. The feature vector of event-type nodes is obtained by concatenating the one-hot encoding of the event type, the time difference encoding of the transaction timestamp, and the logarithmic transformation value of the related transaction amount. The time difference encoding is calculated as follows: taking the system time of the current collection period as the reference time point, the time interval (in days) between the transaction timestamp and the reference time point is calculated, and the time interval is normalized and used as the time difference encoding value. The reference time point is consistent with the current system time used to calculate the time decay weight in step S2.3.

[0036] S2.2: Read the set of cross-business relationships from the multi-source heterogeneous original dataset, and divide the set of cross-business relationships into a set of guarantee relationship edges and a set of fund transfer relationship edges according to the relationship type, specifically: Read the cross-business relationship set from the multi-source heterogeneous original dataset. Each relationship in the cross-business relationship set contains at least the following fields: source node identifier, target node identifier, relationship type label, guarantee amount or transaction amount, transaction timestamp, and source system identifier. The relationship type label includes guarantee type and fund transfer type. Iterate through each relationship in the cross-business relationship set and extract the relationship type label for each relationship; if the relationship type label is "guarantee", then add the current relationship to the temporary list of guarantee relationship edges; if the relationship type label is "fund transfer", then add the current relationship to the temporary list of fund transfer relationship edges. For each relationship in the temporary list of guaranteed related edges, the source node identifier in the relationship is used as the global customer identifier of the guarantor, the target node identifier in the relationship is used as the global customer identifier of the guaranteed party, and the guarantee amount is used as the initial edge weight to form a guarantee related edge triple (global customer identifier of the guarantor, global customer identifier of the guaranteed party, guarantee amount); the guarantee related edge triple is encapsulated together with the corresponding transaction timestamp and source system identifier to obtain the guarantee related edge set; For each relationship in the temporary list of fund transfer relationships, the source node identifier in the relationship is used as the global customer identifier of the fund remitter, the target node identifier in the relationship is used as the global customer identifier of the fund remitter, and the single transaction amount is used as the initial edge weight to form a fund transfer relationship triple (global customer identifier of fund remitter, global customer identifier of fund remitter, transaction amount); the fund transfer relationship triple is then encapsulated together with the corresponding transaction timestamp and source system identifier to obtain the fund transfer relationship edge set.

[0037] It should be noted that each graph edge in the guarantee-related edge set and the fund transfer-related edge set retains the original transaction timestamp, which is used to calculate the time interval; each graph edge in the guarantee-related edge set and the fund transfer-related edge set retains the source system identifier to support subsequent steps in tracing the original business system of the relationship; when constructing the cross-business joint knowledge graph in step S2, the association type label of the directed weighted edge adopts a two-level classification system: the first-level classification includes two types of labels, guarantee type and fund transfer type, which are directly generated by step S2.2 and stored in the graph edge structure; the second-level classification further subdivides the guarantee type edge into guarantee chain type and equity penetration type according to the holding identifier bit in step S3.1, and finally forms three edge subsets of guarantee chain type, equity penetration type and fund transfer type; in the two-level classification system, the first-level label is persistently stored with the graph, and the second-level label is dynamically generated when step S3.1 is executed and is not written back to the graph storage structure.

[0038] S2.3: For each graph edge in the set of guarantee-related edges and the set of fund transfer-related edges, calculate the time decay weight according to the time decay function, using the difference between the current system time and the transaction timestamp as the time interval.

[0039] Preferably, the specific formula for the time decay function is as follows: ; in, Let be the time decay weight of the k-th type edge between node i and node j at the current sampling time. The initial weight coefficients are for the k-th edge type. The decay rate coefficient corresponding to the k-th edge type. For time intervals.

[0040] Preferably, the decay rate coefficient of the guarantee-related edge set is 0.005, and the decay rate coefficient of the fund transfer-related edge set is 0.01, to reflect the business characteristic that the timeliness of the guarantee relationship is weaker than that of the fund transfer relationship; for graph edges whose time interval exceeds the preset time window threshold, the time decay weight is forcibly set to zero, and the corresponding graph edge is marked as an invalid edge and removed from the subsequent graph construction; the initial weight coefficient corresponding to the guarantee-related edge set is 1, and the initial weight coefficient corresponding to the fund transfer-related edge set is 1; the initial weight coefficient can be adjusted by the risk control parameter module of the bank's big data analysis and early warning platform according to the business scenario, and the above values ​​are only the default configuration.

[0041] It should be noted that the preset time window threshold is uniformly configured by the risk control parameter module of the bank's big data analysis and early warning platform. The default value of the preset time window threshold corresponding to the guarantee association edge set is 730 days, and the default value of the preset time window threshold corresponding to the fund transfer association edge set is 180 days. The above values ​​are only the default configurations and can be dynamically adjusted by the risk control parameter module according to the business scenario.

[0042] S2.4: The valid graph edges from the guarantee-related edge set and the fund-transaction-related edge set after time decay weighting are input into the graph construction module along with the feature vectors of customer-type nodes and event-type nodes. The node set and the directed weighted edge set are stored in the form of an adjacency list of the graph, forming a heterogeneous directed weighted graph structure, specifically: The pre-deployed graph construction module is invoked. This graph construction module is a graph data structuring component pre-initialized by the graph storage engine of the bank's big data analysis and early warning platform before the start of step S2. This module is used to receive node feature vectors and graph edge information and convert them into the adjacency list storage format supported by the graph database. The graph construction module does not involve graph neural network computation, but is only responsible for the physical storage and index construction of the graph structure. The graph construction module extracts all graph edges that are not marked as invalid edges from the set of guarantee-related edges and the set of fund-related edges after the time decay weight is calculated, and records them as the set of valid graph edges. For each valid graph edge in the set of valid graph edges, its source node identifier, target node identifier, directed direction, time decay weight, association type label and source system identifier are retained. A node set is created by the graph construction module. The node set includes a subset of customer-class nodes and a subset of event-class nodes. Each node in the customer-class node subset is uniquely identified by a unified global customer identifier in the customer entity set and is attached with a customer-class node feature vector. Each node in the event-class node subset is uniquely identified by an event triple (event type-occurrence time-associated customer entity) in the business event set and is attached with an event-class node feature vector. A set of directed weighted edges is created by the graph construction module. Each directed weighted edge in the set is generated based on the corresponding record in the set of valid graph edges. Specifically, the generation method is as follows: the source node identifier in the valid graph edge is used as the starting node of the directed weighted edge, the target node identifier in the valid graph edge is used as the ending node of the directed weighted edge, the time decay weight in the valid graph edge is used as the edge weight of the directed weighted edge, the association type label in the valid graph edge is used as the edge type label of the directed weighted edge, and the source system identifier in the valid graph edge is used as the source label of the directed weighted edge. For directed weighted edges with the association type label of guarantee, the corresponding holding identifier bit needs to be read from the guarantee relationship record of the multi-source heterogeneous original dataset, and the holding identifier bit is stored as an additional field in the directed weighted edge for direct use in the three-class division of the edge in the subsequent step S3.1, without having to re-retrieve the multi-source heterogeneous original dataset. The graph construction module stores the set of nodes and the set of directed weighted edges in the form of an adjacency list of the graph. The adjacency list is a key-value storage structure in the graph database, where the key is the unique identifier of each node in the node set, and the value is a list of all directed weighted edges originating from that node. Each directed weighted edge is recorded in the list as the target node identifier, edge weight, edge type label, and source marker. The graph construction module associates and assembles the node set and the directed weighted edge set through an adjacency list to form a heterogeneous directed weighted graph structure. The heterogeneous directed weighted graph structure contains two types of heterogeneous nodes (customer nodes and event nodes) and weighted directed edges. The direction of the edge represents the direction of the guarantee relationship or the flow of funds (from the guarantor to the guaranteed party, from the fund sender to the fund receiver), and the weight of the edge is a time decay weight.

[0043] It should be noted that the node set includes two types of heterogeneous nodes: client nodes and event nodes. Each edge in the directed weighted edge set carries an association type label, a time decay weight, and a source system identifier.

[0044] S2.5: Perform connectivity verification on heterogeneous directed weighted graph structures.

[0045] S2.5.1: Read the node set and the directed weighted edge set from the heterogeneous directed weighted graph structure. The node set contains a subset of client-class nodes and a subset of event-class nodes. Each directed weighted edge in the directed weighted edge set records the start node identifier and the end node identifier. S2.5.2: Traverse the nodes in the node set, starting from the current node, and search the directed weighted edge set for a directed weighted edge that starts or ends at the current node. If at least one directed weighted edge is found that starts or ends at the current node, then the current node is marked as a connected node. If no directed weighted edge is found that starts or ends with the current node, the current node is marked as an isolated candidate node. S2.5.3: Summarize all marked isolated candidate nodes to obtain an isolated node candidate set, which includes customer-class isolated candidate nodes and event-class isolated candidate nodes; S2.5.4: For isolated candidate nodes in the isolated node candidate set, obtain the unique identifier of the isolated candidate node; if the isolated candidate node is a customer type node, its unique identifier is the unified global customer identifier; if the isolated candidate node is an event type node, its unique identifier is the event triple (event type-occurrence time-associated customer entity) in the business event set. S2.5.5: For isolated candidate nodes, re-retrieve the multi-source heterogeneous original dataset based on their unique identifier: using the unified global customer identifier (if it is a customer-class node) of the isolated candidate node or the unified global customer identifier (if it is an event-class node) corresponding to the associated customer entity in the event triple as the search key, query whether there are any missing association records in the cross-business association relationship set of the multi-source heterogeneous original dataset that match the source node identifier or target node identifier. The missing association records refer to the association relationships that were not correctly classified into the above two sets due to abnormal data reading, format parsing errors, or temporary data loss when dividing the guarantee association edge set and the fund transaction association edge set. S2.5.6: If, after re-searching in step S2.5.5, there is a missing association record in the cross-business association set that matches the current isolated candidate node, then the following steps are executed: S2.5.6.1: Extract missing related records from the multi-source heterogeneous original dataset, and supplement them with corresponding guarantee related edges or fund transfer related edges according to the method of generating the guarantee related edge set or fund transfer related edge set in step S2.2; S2.5.6.2: Calculate the time decay weights corresponding to the supplementary generated associated edges; S2.5.6.3: The supplementary associated edges calculated with time decay weights are included in the set of valid graph edges, and the graph construction module is triggered to update the heterogeneous directed weighted graph structure. The current isolated candidate node is connected to the supplementary target node through the supplementary associated edges. After the update is completed, the current node is removed from the isolated node candidate set and re-marked as a connected node. S2.5.7: If, after re-searching in step S2.5.5, there is no missing association record matching the current isolated candidate node in the cross-business association set, then the isolated candidate node is finally confirmed as an isolated node, and the following differentiated processing operations are performed: S2.5.7.1: For the finally confirmed isolated node, if the node type of the isolated node is a customer node, then retain this customer node in the heterogeneous directed weighted graph structure, and change the node type label of this customer node to an unrelated customer node. S2.5.7.2: If the node type of the isolated node is an event node, then delete the event node from the node set and ignore the business event information corresponding to the event node in subsequent steps. If the event node is an isolated node, it means that the business event is not connected with any customer entity through guarantee relationship or fund transaction relationship, and its contribution to credit risk assessment is regarded as zero. Therefore, it is deleted to reduce graph structure redundancy. S2.5.7.3: For customer-class nodes marked as unrelated customer nodes, retain the customer-class node feature vector attached to them, and add a flag field to the metadata of the node: Semantic aggregation disabled flag, with a value of true. This flag is used to indicate that unrelated customer nodes are prohibited from participating in the aggregation operation of semantic attention coefficients in subsequent semantic attention coefficient calculations. S2.5.7.4: Simultaneously, the participation rights of unrelated customer nodes in the calculation of node-level attention coefficients are retained: when calculating node-level attention coefficients subsequently, unrelated customer nodes will not participate in the neighbor aggregation operation as direct neighbors of any other customer class nodes. That is, in step S3.4.1, if the node type of a neighbor node is labeled as an unrelated customer node, then the neighbor node is excluded from the set of direct neighbors and will not participate in the calculation of node-level attention coefficients and subsequent weighted aggregation; the customer fusion embedding vector of an unrelated customer node is directly replaced by its customer class node feature vector, and neighbor aggregation and semantic fusion are not performed. S2.5.8: After performing steps S2.5.5 to S2.5.7 sequentially on all nodes in the candidate set of isolated nodes, a heterogeneous directed weighted graph structure is obtained after connectivity verification and processing of all isolated nodes.

[0046] S2.6: The heterogeneous directed weighted graph structure after connectivity verification, together with the feature vectors of customer-type nodes, feature vectors of event-type nodes, time decay weights, association type labels and source system identifiers, are encapsulated and stored in a unified graph data format to obtain a cross-business joint knowledge graph. It should be noted that the cross-business joint knowledge graph uses a unified global customer identifier as the main index, and attaches the corresponding customer-type node feature vector, event-type node feature vector, and time decay weight, which can be directly called for meta-path pre-setting and attention coefficient calculation.

[0047] In this embodiment of the application, step S3 includes: S3.1: Read the feature vectors of customer-type nodes, feature vectors of event-type nodes, time decay weights, association type labels, and source system identifiers from the cross-business joint knowledge graph. According to the association type labels carried by the directed weighted edges, divide the directed weighted edges in the cross-business joint knowledge graph into three subsets of directed weighted edges: guarantee chain edges, equity penetration edges, and fund transfer edges.

[0048] S3.1.1: Read all directed weighted edges from the cross-business joint knowledge graph and traverse the association type label of each directed weighted edge; S3.1.2: If the association type label of the current directed weighted edge is a guarantee type, then the holding identifier is directly read from the additional field of the directed weighted edge. The holding identifier has been extracted from the guarantee relationship record of the multi-source heterogeneous original dataset in step S2.4 and stored as an additional field in the directed weighted edge structure of the graph. The holding identifier has been generated in step S1.5 according to the ratio of the guarantee amount to the registered capital of the guaranteed party and stored as an additional field in the graph along with the directed weighted edge. It is read directly here and will not be calculated again. S3.1.3: For directed weighted edges whose association type label value is in the guarantee category, if the controlling stake identifier is true, the edge is assigned to the equity penetration category edge subset; if the controlling stake identifier is false, the edge is assigned to the guarantee chain category edge subset. S3.1.4: If the association type label of the current directed weighted edge is of the fund transaction class, then the edge is assigned to the fund transaction class edge subset; Based on the above division, three mutually exclusive directed weighted edge subsets are obtained: the guarantee chain edge subset, the equity penetration edge subset, and the fund transfer edge subset. The edges in the guarantee chain edge subset represent ordinary guarantee relationships without a controlling relationship, the edges in the equity penetration edge subset represent strong control relationships with a controlling relationship, and the edges in the fund transfer edge subset represent fund transfer relationships.

[0049] S3.2: Based on the edge subsets of guarantee chains, equity penetration, and fund transfers, three types of business meta-paths are pre-defined in the cross-business joint knowledge graph. These three types of pre-defined business meta-paths include guarantee chain meta-paths, equity penetration meta-paths, and fund transfer meta-paths, specifically: The guarantee chain meta-path is defined as: a directed connection sequence from the starting customer node to the next customer node via one or more directed edges in a subset of guarantee chain edges; the guarantee chain meta-path is used to characterize the guarantee transmission path without a controlling relationship. The equity penetration meta-path is defined as follows: the starting customer node penetrates down through one or more directed edges in the equity penetration edge subset to the subsequent controlled customer node in a directed connection sequence, with the path length not exceeding the maximum path length constraint; the equity penetration meta-path is used to represent the control relationship at each level in the actual controller or control chain, and the path terminates when the path length reaches the maximum path length constraint, or when the current node has no outgoing edges in the equity penetration edge subset; The meta-path of fund transfers is defined as: a directed connection sequence from the starting customer node to the counterparty customer node via one or more directed edges from a subset of fund transfer edges. The meta-path of fund transfers is used to characterize the path of fund flow.

[0050] It should be noted that the maximum path length of the three types of business meta-paths is uniformly configured by the risk control parameter module of the bank's big data analysis and early warning platform; the default maximum path length of the guarantee chain meta-path is four hops, the maximum path length of the equity penetration meta-path is six hops, and the maximum path length of the fund transfer meta-path is three hops; the maximum path length is used to constrain the search range of path enumeration in subsequent steps to avoid combinatorial explosion.

[0051] S3.3: For each customer-class node in the cross-business joint knowledge graph, enumerate all directed connection sequences that satisfy the corresponding maximum path length constraint along the guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path. Record the customer-class node sequence and the corresponding time decay weight sequence that each directed connection sequence passes through as candidate path instances to obtain a set of candidate risk transmission paths.

[0052] S3.3.1: Obtain the unified global customer identifier of all customer class nodes in the cross-business joint knowledge graph, and denote it as the starting node set; S3.3.2: For each starting customer node in the starting node set, according to the definitions of the guarantee chain meta-path, equity penetration meta-path and fund transfer meta-path, the depth-first search algorithm is used to enumerate paths on the directed weighted edges. During the search process, starting from the current node, it is only allowed to move along the direction of the directed edges in the corresponding edge subset, and the path length (number of edges) does not exceed the maximum path length preset for the meta-path. S3.3.3: For each enumerated directed connection sequence, record the unified global customer identifier of the customer-class nodes that the sequence passes through in sequence to form a customer-class node sequence; at the same time, record the time decay weight corresponding to each directed edge in the sequence in the same order to form a time decay weight sequence. S3.3.4: Store the customer class node sequence in association with the time decay weight sequence, and add the following fields: the business meta-path type label (guarantee chain meta-path, equity penetration meta-path or fund transfer meta-path), and the unified global customer identifier of the starting customer class node of the path. The above combination constitutes a candidate path instance. S3.3.5: Summarize all candidate path instances obtained by enumerating all starting customer class nodes to obtain a candidate risk transmission path set. Each candidate path instance in the candidate risk transmission path set contains a customer class node sequence, a time decay weight sequence, a business element path type label, and a unified global customer identifier of the starting customer class node, which can be directly called for path occlusion comparison.

[0053] S3.4: For each customer-class node in the cross-business joint knowledge graph, calculate the node-level attention coefficient between this customer-class node and its direct neighbor nodes under the guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path, respectively. Specifically: S3.4.1: Obtain the feature vector of each customer-type node in the cross-business joint knowledge graph, and the set of direct neighbor nodes of this customer-type node under each business meta-path. Direct neighbor nodes are defined as: there exists a directed edge from the current customer-type node to the neighbor node, and the directed edge belongs to the edge subset of the corresponding meta-path (i.e., the edge subset of guarantee chain, equity penetration, or fund transfer). The type of neighbor node can be a customer-type node or an event-type node. It should be noted that when obtaining the set of direct neighbor nodes under each business element path, if the node type of a neighbor node is marked as an unrelated customer node (i.e. the semantic aggregation disable flag set in step S2.5.7 is true), then the neighbor node will be excluded from the set of direct neighbor nodes of the current target customer class node and will not participate in the calculation of the node-level attention coefficient. S3.4.2: For any target customer node and its direct neighbor node under a certain business element path, concatenate the customer node feature vector of the target customer node with the feature vector of the neighbor node; where, if the neighbor node is a customer node, the customer node feature vector is taken; if the neighbor node is an event node, the event node feature vector is taken. S3.4.3: The concatenated vector is input into a shared linear transformation layer that is independently initialized for the current business meta-path type to obtain the original attention score. The shared linear transformation layer contains a trainable weight matrix and bias vector, which are independently initialized for the guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path, respectively. The shared linear transformation layer is randomly initialized by the model initialization module during the model training phase and is jointly trained with the graph neural network parameters using historical default labels as supervision signals. After training, the parameters are fixed. During the inference execution phase of this method, the weight matrix and bias vector of the shared linear transformation layer remain unchanged, and the fixed parameters after training are directly loaded to calculate the original attention score. S3.4.4: After processing the original attention score with the LeakyReLU activation function, the node-level attention coefficient is calculated using the softmax function under the same business metapath type, with all direct neighbor nodes of the target customer class node as the normalization range.

[0054] It should be noted that steps S3.4 and S3.3 are independent of each other. Both take the cross-business joint knowledge graph as input and can be executed in parallel if computing resources allow. The calculation of node-level attention coefficients only depends on the feature vectors of customer-class nodes and event-class nodes, and does not include time decay weights.

[0055] S3.5: Based on node-level attention coefficients, the feature vectors of all direct neighbor nodes of each customer-class node in the cross-business joint knowledge graph are weighted and aggregated respectively under the guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path. This yields the path semantic embedding vectors of this customer-class node under the guarantee chain meta-path, the equity penetration meta-path, and the fund transfer meta-path, specifically: S3.5.1: For each target customer class node and each of its direct neighbor nodes under a certain business meta-path, obtain the node-level attention coefficient and calculate and store the time decay weight of the directed weighted edge with the cross-business joint knowledge graph. S3.5.2: Multiply the node-level attention coefficient by the time decay weight to obtain the aggregate weight, where the aggregate weight simultaneously integrates the semantic relevance between nodes and the timeliness of the association relationship; S3.5.3: Perform a weighted summation on the feature vectors of the direct neighbor nodes, and process the result of the weighted summation through a non-linear activation function to obtain the path semantic embedding vector under the business meta path, wherein the dimension of the path semantic embedding vector is the same as the dimension of the feature vector of the customer class node. S3.5.4: For each customer-type node, execute S3.5.1 to S3.5.3 for the guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path respectively, to obtain three path semantic embedding vectors, namely the path semantic embedding vector under the guarantee chain meta-path, the path semantic embedding vector under the equity penetration meta-path, and the path semantic embedding vector under the fund transfer meta-path.

[0056] S3.6: For the path semantic embedding vectors under the guarantee chain meta-path, the equity penetration meta-path, and the fund transfer meta-path, calculate the semantic-level attention coefficients corresponding to the three types of business meta-paths, specifically: S3.6.1: During the model initialization phase, a trainable semantic query vector is created, wherein the dimension of the semantic query vector is the same as the dimension of the path semantic embedding vector. This semantic query vector is jointly updated with the parameters of the shared linear transformation layer and the graph neural network during the model training process. The semantic query vector is jointly trained with the parameters of the shared linear transformation layer and the graph neural network during the model training phase, and is fixed after training. It is directly loaded and used during the inference execution phase. S3.6.2: For each customer-type node, calculate the dot product of the path semantic embedding vector under the guarantee chain meta-path, the path semantic embedding vector under the equity penetration meta-path, and the path semantic embedding vector under the fund transfer meta-path with the semantic query vector to obtain the original semantic score of each business meta-path. S3.6.3: The original semantic scores of each business metapath are normalized using the softmax function within the scope of the customer-class node itself (i.e., between the three metapaths of the guarantee chain metapath, equity penetration metapath, and fund transfer metapath) to obtain semantic-level attention coefficients. The semantic-level attention coefficients reflect the relative importance of the guarantee chain metapath, equity penetration metapath, and fund transfer metapath to the credit risk assessment of the customer-class node.

[0057] It should be noted that when the set of direct neighbor nodes of a certain customer class node under a certain business meta-path is empty, the corresponding path semantic embedding vector is set to a zero vector, and a validity mask with a value of zero is set for the business meta-path. Before performing softmax normalization, the validity mask of each business meta-path is checked first. The original semantic score corresponding to the business meta-path with a validity mask value of zero is forced to be negative infinity, so that the semantic level attention coefficient output of the business meta-path after softmax normalization is zero. Business meta-paths with a validity mask value of one participate in the softmax normalization calculation normally. If the validity masks of all three types of business meta-paths of a certain customer class node are all zero, then all semantic level attention coefficients of this customer class node are set to zero, and its customer fusion embedding vector is directly composed of the feature vector of the customer class node, without performing semantic level weighted fusion.

[0058] Preferably, for customer-type nodes labeled as unrelated customer nodes, the path semantic embedding vectors under the guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path calculated in step S3.5 are not included in the semantic-level attention coefficient calculation in step S3.6. When calculating the semantic-level attention coefficient, the semantic-level attention coefficients corresponding to each business meta-path of this type of node are directly set to zero, and its customer fusion embedding vector is composed only of the customer-type node feature vector itself, without semantic-level weighted fusion.

[0059] S3.7: Using semantic-level attention coefficients as weights, the path semantic embedding vectors under the guarantee chain meta-path, the equity penetration meta-path, and the fund transfer meta-path are weighted and summed to obtain the customer fusion embedding vector for each customer-type node, specifically: For each customer-class node, if the node is not an unrelated customer node, its path semantic embedding vector and semantic-level attention coefficient are obtained; the path semantic embedding vector under the guarantee chain meta-path is multiplied by the semantic-level attention coefficient corresponding to the guarantee chain meta-path, the path semantic embedding vector under the equity penetration meta-path is multiplied by the semantic-level attention coefficient corresponding to the equity penetration meta-path, and the path semantic embedding vector under the fund transfer meta-path is multiplied by the semantic-level attention coefficient corresponding to the fund transfer meta-path. Then, the above three products are added together to obtain the customer fusion embedding vector of the customer-class node; if the node is an unrelated customer node, its customer-class node feature vector is directly used as the customer fusion embedding vector.

[0060] In this embodiment of the application, step S4 includes: S4.1: Read all candidate path instances from the candidate risk transmission path set. Based on the business element path type tag carried by each candidate path instance, divide the candidate risk transmission path set into a guarantee chain candidate path subset, an equity penetration candidate path subset, and a fund transfer candidate path subset, specifically: S4.1.1: Traverse each candidate path instance in the candidate risk transmission path set and read the business element path type label of the candidate path instance; S4.1.2: If the value of the business meta-path type label is a guarantee chain meta-path, then the candidate path instance is assigned to the guarantee chain candidate path subset; if the value of the business meta-path type label is an equity penetration meta-path, then the candidate path instance is assigned to the equity penetration candidate path subset; if the value of the business meta-path type label is a fund transfer meta-path, then the candidate path instance is assigned to the fund transfer candidate path subset. By storing the above three subsets separately, we obtain the candidate path subsets for guarantee chains, equity penetration candidate paths, and fund transfer candidate paths.

[0061] S4.2: From the customer fusion embedding vector, according to the unified global customer identifier of the starting customer class node of each candidate path instance, read the customer fusion embedding vector of the corresponding customer class node, and use the customer fusion embedding vector as the benchmark risk representation vector for path occlusion comparison.

[0062] S4.3: For each candidate path instance in the subsets of guarantee chain candidate paths, equity penetration candidate paths, and fund transfer candidate paths, perform path masking operations one by one, specifically as follows: S4.3.1: For the current candidate path instance to be occluded, read its customer class node sequence, where the customer class node sequence contains a starting customer class node, one or more intermediate customer class nodes, and an ending customer class node; S4.3.2: Perform temporary masking on the cross-business joint knowledge graph: Temporarily set the feature vectors of all intermediate customer nodes and the terminal customer node in the customer node sequence (excluding the starting customer node) to zero. Since step S4.4.2 will re-execute the complete calculation process of S3.4 to S3.7 for the starting customer node, the generation of the path semantic embedding vector depends on the feature vectors of neighboring nodes. Therefore, masking the feature vectors of customer nodes will directly affect the result of the weighted aggregation of neighboring nodes during the recalculation process, which is reflected in the masking risk representation vector. For the customer fusion embedding vector that has been calculated in step S3.7, no direct modification is performed in the path masking comparison stage. Instead, the equivalent customer fusion embedding vector after masking, i.e., the masking risk representation vector, is obtained by re-executing the calculation process of S3.4 to S3.7. S4.3.3: At the same time, temporarily set each time decay weight in the time decay weight sequence corresponding to the current candidate path instance to zero, that is, cover all directed weighted edges covered by the current candidate path instance. S4.3.4: The path occlusion operation is performed independently for each candidate path instance. Only one candidate path instance is occluded at a time, while the node feature vectors and time decay weights corresponding to the remaining candidate path instances remain unchanged. It should be noted that after completing the masking operation for a candidate path instance, before performing the masking operation for the next candidate path instance, it is necessary to restore the node feature vectors and edge weights that were temporarily set to zero in the cross-business joint knowledge graph. The operation of restoring the cross-business joint knowledge graph is as follows: before the masking operation for each candidate path instance begins, a deep copy of the node feature vector copy and edge weight copy of the current graph is made to a temporary cache. After masking and recalculation are completed, the original node feature vectors and edge weights are restored from the cache before processing the next instance. The operation of restoring the cross-business joint knowledge graph does not rely on external transaction mechanisms, adopts memory-level copying, has efficient recovery performance under the normal graph size, and does not affect the overall processing efficiency.

[0063] S4.4: Taking the masked cross-business joint knowledge graph as input, and following the calculation methods of node-level attention coefficients and semantic-level attention coefficients, the weighted aggregation of neighbor nodes and the semantic-level weighted fusion of the three types of business meta-paths are re-executed on the initial customer class node to obtain the masking risk representation vector of the candidate path instance, specifically: S4.4.1: After completing the masking operation on the current candidate path instance, obtain the masked cross-business joint knowledge graph; in the masked graph, the feature vectors of the non-starting nodes traversed by the current candidate path instance are all zero vectors, and the weights of the corresponding edges are zero. S4.4.2: Taking the masked cross-business joint knowledge graph as input, for the starting customer class node of the current candidate path instance, re-execute the calculation process from S3.4 to S3.7, including: recalculating the node-level attention coefficients of the starting customer class node under each business meta-path (note that the feature vectors of neighboring nodes may become zero vectors due to masking at this time), recalculating the path semantic embedding vector and semantic-level attention coefficients, and finally re-weighting and summing to obtain the masked customer fusion embedding vector; S4.4.3: The recalculated customer fusion embedding vector of the starting customer class node is denoted as the occlusion risk representation vector of the current candidate path instance.

[0064] It should be noted that during the calculation of the occlusion risk representation vector, the weight matrix of the shared linear transformation layer and the trainable semantic query vector in step S3 remain unchanged and are not retrained, so as to eliminate the interference of model parameter changes on the calculation results of the risk representation difference before and after occlusion.

[0065] S4.5: For each candidate path instance, calculate the risk representation difference between the baseline risk representation vector and the occlusion risk representation vector, specifically: S4.5.1: Obtain the baseline risk representation vector and the occlusion risk representation vector of the current candidate path instance, both of which are vectors with the same dimension; S4.5.2: Calculate the Euclidean distance between the baseline risk representation vector and the occlusion risk representation vector, and use this Euclidean distance as the original risk representation difference value. The larger the Euclidean distance value, the more significant the contribution of the current candidate path instance to the cross-business risk characteristics of the starting customer class node. S4.5.3: Perform maximum-minimum normalization on the original risk representation difference values ​​of all candidate path instances under the same starting customer class node: Subtract the minimum value of all original risk representation difference values ​​under the starting node from each original risk representation difference value, and then divide by the difference between the maximum and minimum values ​​to obtain the normalized risk representation difference value. The normalized risk representation difference value ranges from zero to one, in order to eliminate the impact of the difference in the dimensions of the benchmark risk representation vector between different customer class nodes on the calculation of subsequent path risk confirmation coefficients.

[0066] Preferably, if the maximum value of all original risk representation differences under the same starting customer class node is equal to the minimum value, then the normalized risk representation differences of all candidate path instances under that starting customer class node are set to zero.

[0067] S4.6: Combining the time decay weight sequence, perform timeliness correction on the normalized risk representation difference value of each candidate path instance to generate a path risk confirmation coefficient, specifically: S4.6.1: Obtain the time decay weight sequence of the current candidate path instance, where the length of the time decay weight sequence is the same as the number of edges in the customer class node sequence, and each element in the sequence is the time decay weight of the corresponding directed weighted edge; S4.6.2: Calculate the arithmetic mean of all time decay weights in the time decay weight sequence, denoted as the path timeliness factor; S4.6.3: Multiply the normalized risk representation difference value of the current candidate path instance by the path timeliness factor to obtain the path risk confirmation coefficient; S4.6.4: For candidate path instances whose path risk confirmation coefficient is lower than the preset path validity threshold, the candidate path instance is marked as an inefficient path and removed from the candidate risk transmission path set. The preset path validity threshold is uniformly configured by the risk control parameter module of the bank's big data analysis and early warning platform, and the default value is 0.1.

[0068] It should be noted that when there are node pairs connected by nodes that have been marked as invalid edges in the customer class node sequence of a candidate path instance, the path risk confirmation coefficient of the candidate path instance is forcibly set to zero, and it is directly marked as an inefficient path and eliminated, and no further calculation is performed; this fallback judgment is used to ensure the correctness of the path risk confirmation coefficient calculation results under the above abnormal scenarios.

[0069] S4.7: The path risk confirmation coefficient corresponding to each candidate path instance in the candidate risk transmission path set after inefficient path elimination is grouped according to the unified global customer identifier of the starting customer class node, and the path risk confirmation coefficients of all valid candidate path instances under each customer class node are stored in an ordered list. The path risk confirmation coefficients as a whole are used as risk transmission calibration parameters, specifically: S4.7.1: Obtain the remaining candidate path instances after eliminating inefficient paths. Each candidate path instance carries a unified global customer identifier of the starting customer class node, a business element path type label, a customer class node sequence, and a calculated path risk confirmation coefficient. S4.7.2: Group all valid candidate path instances according to the unified global customer identifier, and arrange the path risk confirmation coefficients of all valid candidate path instances under the customer class node name in order of calculation or path length to form an ordered list; S4.7.3: Use the unified global customer identifier of each customer-class node as the primary key, and attach the business element path type label of the corresponding valid candidate path instance, the customer-class node sequence and the ordered list of path risk confirmation coefficients to jointly constitute the risk transmission calibration parameters.

[0070] In this embodiment of the application, step S5 includes: S5.1: Read the customer fusion embedding vectors corresponding to all customer class nodes from the customer fusion embedding vectors. Read the business meta-path type labels, customer class node sequences, and path risk confirmation coefficients of all valid candidate path instances corresponding to all customer class nodes from the risk transmission calibration parameters. Associate the customer fusion embedding vectors with the risk transmission calibration parameters according to the unified global customer identifier to obtain the fusion input data pair for each customer class node, specifically: S5.1.1: The graph neural network model has been initialized in the model training module of the bank's big data analysis and early warning platform. The graph neural network model includes a message passing layer, a cross-business fusion linear transformation layer, and a credit risk scoring layer. The parameters of each layer are pre-trained using historical default labels as supervision signals, and the model parameters are fixed after training. Before the model inference is executed, the trainable parameters in the graph neural network model and the attention mechanism are jointly trained. The training process is as follows: Training data construction: A training sample set containing customer credit outcomes is extracted from historical data of the bank's big data analysis and early warning platform. Each training sample corresponds to a customer class node, and its label is a historical default label, with a value of 0 or 1. Where 1 indicates that the customer defaulted within the preset observation period (including loan delinquency exceeding 90 days, guarantee compensation, or non-performing asset identification), and 0 indicates that no default occurred. The default observation period length is 12 months, which is configured by the risk control parameter module. The training set, validation set, and test set are randomly divided in a ratio of 70%, 15%, and 15%, respectively. Loss function: The binary cross-entropy loss function is used, and the formula is as follows: ; in, For the historical default label of the i-th training sample, The customer credit risk score output by the model; Optimizer and hyperparameters: The Adam optimizer is used, with an initial learning rate of 0.001 and a weight decay coefficient of 1e-5. The training batch size is 64, the maximum number of training epochs is 200, and an early stopping strategy is adopted, stopping training when the validation set loss does not decrease for 10 consecutive epochs. Trainable parameter initialization: The weight matrix of the shared linear transformation layer is initialized uniformly by Xavier, and the bias vector is initialized to zero; the semantic query vector is randomly initialized using a Gaussian distribution with a mean of 0 and a standard deviation of 0.01; the parameters of the fully connected layers in the message passing layer, cross-business fusion linear transformation layer, and credit risk scoring layer of the graph neural network are all initialized using He normal distribution. Training process: In each round, training samples are loaded in batches, forward propagation is used to calculate customer credit risk scores, loss functions are calculated, and backpropagation is used to update all trainable parameters; after each round, loss and AUC are calculated on the validation set, and the model parameters with the minimum loss on the validation set are saved as the graph neural network model.

[0071] In an optional embodiment, historical corporate customer data from a bank is used as the source, taking data from the past 3 years. Default labels are defined as loans overdue for more than 90 days or guarantee repayments occurring. Each training sample includes a customer ID, credit record, guarantee relationship, transaction history, and corresponding default label. The training process is as follows: each epoch iterates through the training set in batches of 64, calculates the score using forward propagation, performs backpropagation of cross-entropy loss, and updates parameters using the Adam optimizer. After each epoch, the AUC is calculated on the validation set. If the validation set loss does not decrease for 10 consecutive epochs, training stops, and the model parameters with the minimum validation set loss are restored. After training, the weights of the shared linear transformation layer, semantic query vectors, and parameters of each layer of the graph neural network are exported as binary files for loading during the inference phase.

[0072] S5.1.2: Using the unified global customer identifier as the association key, the customer fusion embedding vector and the risk transmission calibration parameters are left-joined; for each customer class node, the customer fusion embedding vector (a vector) and the path risk confirmation coefficients of the valid candidate path instances corresponding to that node in the risk transmission calibration parameters are obtained in an ordered list.

[0073] S5.2: The message passing layer of the vector input graph neural network model embeds the customer fusion of each customer class node in the fused input data pair into the message passing layer. Under the adjacency structure constraint of the cross-business joint knowledge graph, the neighbor node message passing is executed according to the three types of business meta-paths: guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path. Specifically: S5.2.1: For target customer-type nodes, obtain all valid candidate path instances under the node name from the risk transmission calibration parameters. Each valid candidate path instance contains a business element path type label, a customer-type node sequence, and a path risk confirmation coefficient. S5.2.2: For each valid candidate path instance, extract all neighboring customer class nodes (including intermediate and terminal nodes) from the customer class node sequence, excluding the starting customer class node; obtain the customer fusion embedding vector of these neighboring customer class nodes (read from the fusion input data pair in S5.1 using the unified global customer identifier). S5.2.3: Divide the path risk confirmation coefficient corresponding to the current candidate path instance by the total number of neighboring customer nodes contained in the candidate path instance to obtain the single node aggregation weight; use the single node aggregation weight to perform a weighted summation of the customer fusion embedding vectors of the above neighboring customer nodes to obtain the path aggregation message vector of the current candidate path instance; the above normalization process eliminates the influence of path length differences on the aggregation result, so that the magnitude of the path aggregation message vector is determined only by the path risk confirmation coefficient and is independent of the number of neighboring nodes in the path.

[0074] S5.3: For each customer-type node, the path aggregation message vector of all valid candidate path instances under the three types of business meta-paths—guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path—is summed within the domain according to the business meta-path type to obtain the guarantee chain domain aggregation vector, equity penetration domain aggregation vector, and fund transfer domain aggregation vector, specifically: S5.3.1: Initialize three zero vectors with the same dimensions as the customer fusion embedding vector, denoted as the guarantee chain domain cumulative vector, equity penetration domain cumulative vector, and fund transfer domain cumulative vector, respectively; S5.3.2: Traverse all valid candidate path instances under the target customer class node name. For each candidate path instance, obtain its business element path type label and path aggregation message vector. S5.3.3: If the business meta path type label is a guarantee chain meta path, then the path aggregation message vector is accumulated to the guarantee chain domain accumulation vector; if it is an equity penetration meta path, then it is accumulated to the equity penetration domain accumulation vector; if it is a fund transfer meta path, then it is accumulated to the fund transfer domain accumulation vector. S5.3.4: After traversal, the cumulative vector of the guarantee chain domain is used as the aggregate vector of the guarantee chain domain, the cumulative vector of the equity penetration domain is used as the aggregate vector of the equity penetration domain, and the cumulative vector of the fund transfer domain is used as the aggregate vector of the fund transfer domain.

[0075] It should be noted that when the number of valid candidate path instances for a certain customer class node under a specific business element path type is zero, the corresponding domain aggregation vector remains a zero vector to maintain the dimensional consistency of the three types of domain aggregation vectors.

[0076] S5.4: The customer fusion embedding vector, guarantee chain domain aggregation vector, equity penetration domain aggregation vector, and fund transfer domain aggregation vector of the target customer node are concatenated column-wise, and then dimensionality-reduced through the cross-business fusion linear transformation layer of the graph neural network model to obtain the cross-business risk feature vector of each customer node, specifically: S5.4.1: The four vectors of the target customer class node, namely the customer fusion embedding vector, the guarantee chain domain aggregation vector, the equity penetration domain aggregation vector, and the fund transfer domain aggregation vector, are concatenated in sequence to form a concatenated vector with a dimension four times that of the customer fusion embedding vector. S5.4.2: Input the concatenated vector into the cross-service fusion linear transformation layer, which consists of a fully connected layer and a batch normalization layer connected in series; the output dimension of the fully connected layer is half the dimension of the customer fusion embedded vector, and the batch normalization layer normalizes the output of the fully connected layer. S5.4.3: The cross-business fusion linear transformation layer outputs the cross-business risk feature vector of this customer-type node. The cross-business risk feature vector contains the cross-business fusion features of the target customer-type node itself and the calibration aggregation information from the three risk transmission paths of guarantee chain, equity penetration and fund transfer.

[0077] S5.5: Input the cross-business risk feature vector into the credit risk scoring layer to obtain the customer credit risk score for each customer class node, specifically: S5.5.1: The credit risk scoring layer consists of a fully connected layer and a sigmoid activation function connected in series. The fully connected layer maps cross-business risk feature vectors to scalar output values, and the sigmoid activation function normalizes the scalar output values ​​to the range of zero to one. S5.5.2: Input the cross-business risk feature vector into the credit risk scoring layer to calculate the normalized scalar value. This scalar value is the customer credit risk score. The closer the customer credit risk score is to one, the higher the cross-business credit default probability of the corresponding customer class node.

[0078] S5.6: For each customer category node, customer credit risk is scored and classified according to a preset risk threshold range, generating a graded early warning result, specifically as follows: The system retrieves preset risk threshold ranges configured uniformly by the risk control parameter module of the bank's big data analysis and early warning platform. These preset risk threshold ranges are divided into three levels: normal level, attention level, and early warning level. When a customer's credit risk score is less than the first preset threshold, this customer type node is determined to be normal and normal level operation is performed: the unified global customer identifier, customer credit risk score and normal level label of this customer type node are recorded to the normal level customer list, without triggering any warning action and without generating risk transmission path record. When a customer's credit risk score is greater than or equal to the first preset threshold and less than the second preset threshold, this customer-type node is determined to be at the attention level, and attention level operations are performed: the unified global customer identifier, customer credit risk score, and attention level label of this customer-type node are recorded to the attention level customer list. At the same time, the effective candidate path instance with the highest path risk confirmation coefficient under this customer-type node is extracted from the risk transmission calibration parameters. The customer-type node sequence of this effective candidate path instance is used as the attention transmission path. The attention transmission path, customer credit risk score, and attention level label are encapsulated together into an attention prompt record. The attention prompt record is for manual review by bank risk control personnel and does not trigger proactive early warning push. When a customer's credit risk score is greater than or equal to the second preset threshold, this customer type node is determined to be at the warning level, and warning level operations are performed: the unified global customer identifier, customer credit risk score, and warning level label of this customer type node are recorded in the warning level customer list. At the same time, the effective candidate path instance with the highest path risk confirmation coefficient under this customer type node is extracted from the risk transmission calibration parameters. The customer type node sequence of this effective candidate path instance is used as the risk transmission path. The risk transmission path, customer credit risk score, and warning level label are encapsulated together into a warning detail record, which is used by the bank's big data analysis and warning platform to trigger real-time warning push.

[0079] Preferably, the first preset threshold is 0.4 and the second preset threshold is 0.7. Both the first and second preset thresholds are dynamically adjusted by the risk control parameter module of the bank's big data analysis and early warning platform. The above values ​​are only the default configurations. The parameter adjustment methods are: modifying the configuration file and restarting the scoring service, or dynamically updating the parameter copy in memory through the interface, without affecting the model structure.

[0080] S5.7: Record the customer credit risk scores and tiered early warning results for all customer-related nodes, as well as the transmission path records corresponding to customer-related nodes classified as "watch list" and "early warning," using a unified global customer identifier as the primary index. Encapsulate this data into a structured data package according to the data interface specifications of the bank's big data analysis and early warning platform, and write it to the bank's big data analysis and early warning platform via the data interface call. Specifically: S5.7.1: Traverse all customer class nodes. For each customer class node, assemble a write record with the unified global customer identifier as the primary key. The write record contains the following fields: unified global customer identifier, customer credit risk score, and level label of the graded early warning result (normal level, attention level, or early warning level). S5.7.2: For customer-type nodes whose graded early warning results are at the attention level, add two additional fields to the written record: the sequence of customer-type nodes in the attention transmission path and the path risk confirmation coefficient corresponding to the attention transmission path; S5.7.3: For customer-type nodes whose graded early warning result is at the early warning level, add two additional fields to the written record: the customer-type node sequence of the risk transmission path and the path risk confirmation coefficient corresponding to the risk transmission path; S5.7.4: All write records are structured and encapsulated according to the data interface specifications of the bank's big data analysis and early warning platform (e.g., using JSON format or Protocol Buffers format) to form a structured write data packet; S5.7.5: Call the data writing interface of the bank's big data analysis and early warning platform to send the structured data packet to the platform for storage, querying, and visualization of risk transmission paths.

[0081] In summary, this invention constructs a multi-source heterogeneous original dataset by collecting customer data from the bank's integrated business management system, corporate banking system, and transaction risk assessment system. This unifies credit granting, transaction, guarantee, equity, and risk control data, avoiding the problem of one-sided risk identification caused by single-business data. By using customer entities and business events as graph nodes and cross-business relationships as graph edges, and configuring time decay weights, a cross-business joint knowledge graph is constructed. This allows the relationships between customers regarding guarantee transmission, equity penetration, and fund transfers to be expressed according to association type and time validity, improving the completeness of risk relationship modeling. Furthermore, by pre-setting three types of business meta-paths—guarantee chains, equity penetration, and fund transfers—and calculating node-level and semantic-level attention coefficients, a customer fusion embedding vector is obtained, enabling different neighboring nodes to... The contribution of different business paths to customer risk can be measured differentially, improving the accuracy of customer risk characterization. By comparing the path masking of candidate risk transmission paths, the difference in risk representation before and after masking is calculated, and a path risk confirmation coefficient is generated by combining time decay weights. This allows the risk transmission path to pass the degree of influence verification, reducing the interference of false associations, weak associations, or expired associations on risk scoring. By embedding customer data into a vector input graph neural network model and aggregating cross-business risks according to the path risk confirmation coefficient, a cross-business risk feature vector is obtained. This vector is used to calculate customer credit risk scores and generate graded early warning results of normal, attention, and warning levels. This makes the risk assessment results have the basis for scoring judgment, path interpretation, and early warning handling, thereby improving the accuracy, timeliness, and interpretability of cross-business credit risk identification.

[0082] Based on the teachings of the above embodiments, other aspects of the present invention also propose a cross-business credit risk assessment system based on graph neural networks and a joint bank knowledge graph, such as... Figure 2 As shown, it includes: The multi-source data construction module is used to collect multi-business customer data from the bank's integrated business management system, corporate business system and transaction risk assessment system, extract customer entity sets, business event sets and cross-business relationship sets, and construct multi-source heterogeneous raw datasets. The joint graph construction module is used to construct a cross-business joint knowledge graph using customer entity sets and business event sets as graph nodes, cross-business relationship sets as graph edges, and time decay weights configured for graph edges. The risk path identification module is used to pre-set three types of business meta-paths for cross-business joint knowledge graphs: guarantee chain, equity penetration and fund transfer. It calculates node-level attention coefficients and semantic-level attention coefficients to obtain customer fusion embedding vectors and candidate risk transmission path sets. The path risk calibration module is used to compare the path masking of the candidate risk transmission path set, calculate the difference in risk representation before and after path masking, and generate the path risk confirmation coefficient by combining the time decay weight. The path risk confirmation coefficient is used as the risk transmission calibration parameter. The risk scoring and early warning module is used to integrate customers into a vector input graph neural network model, aggregate cross-business risks according to the path risk confirmation coefficient, obtain cross-business risk feature vectors, calculate customer credit risk scores, and generate graded early warning results.

[0083] This embodiment also provides a computer device applicable to the cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as proposed in the above embodiment.

[0084] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0085] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as proposed in the above embodiments.

[0086] The storage medium proposed in this embodiment and the data storage method proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0087] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A cross-business credit risk assessment method based on graph neural networks and a joint bank knowledge graph, characterized in that, include: Multi-business customer data is collected from the bank's integrated business management system, corporate business system and transaction risk assessment system. Customer entity set, business event set and cross-business relationship set are extracted to construct a multi-source heterogeneous original dataset. A cross-business joint knowledge graph is constructed by using the customer entity set and the business event set as graph nodes and the cross-business relationship set as graph edges, and configuring time decay weights for the graph edges. For the cross-business joint knowledge graph, three types of business meta-paths are preset: guarantee chain, equity penetration and fund transfer. Node-level attention coefficient and semantic-level attention coefficient are calculated to obtain customer fusion embedding vector and candidate risk transmission path set. The candidate risk transmission path set is compared by path masking, the difference in risk representation before and after path masking is calculated, and the path risk confirmation coefficient is generated by combining the time decay weight. The path risk confirmation coefficient is used as the risk transmission calibration parameter. The customer is fused and embedded into a vector input graph neural network model, and cross-business risk is aggregated according to the path risk confirmation coefficient to obtain a cross-business risk feature vector. The customer credit risk score is calculated and a graded early warning result is generated. The process of generating the customer fusion embedding vector includes: For each customer-class node in the cross-business joint knowledge graph, the node-level attention coefficient between this customer-class node and its direct neighbor nodes is calculated under the guarantee chain meta-path, the equity penetration meta-path, and the fund transfer meta-path, respectively. Based on the node-level attention coefficient, the feature vectors of all direct neighbor nodes of each customer-type node in the cross-business joint knowledge graph are weighted and aggregated respectively under the guarantee chain meta-path, the equity penetration meta-path and the fund transfer meta-path to obtain the path semantic embedding vector of this customer-type node under the guarantee chain meta-path, the path semantic embedding vector under the equity penetration meta-path and the path semantic embedding vector under the fund transfer meta-path. For the path semantic embedding vector under the guarantee chain meta-path, the path semantic embedding vector under the equity penetration meta-path, and the path semantic embedding vector under the fund transfer meta-path, calculate the semantic-level attention coefficients corresponding to the three types of business meta-paths. Using the semantic-level attention coefficient as the weight, the path semantic embedding vectors under the guarantee chain meta-path, the path semantic embedding vectors under the equity penetration meta-path, and the path semantic embedding vectors under the fund transfer meta-path are weighted and summed to obtain the customer fusion embedding vectors of each customer class node. The process of generating the path risk confirmation coefficient includes: Read all candidate path instances from the candidate risk transmission path set, and divide the candidate risk transmission path set into a guarantee chain candidate path subset, an equity penetration candidate path subset, and a fund transfer candidate path subset according to the business element path type tag carried by each candidate path instance; From the customer fusion embedding vector, according to the unified global customer identifier of the starting customer class node of each candidate path instance, the customer fusion embedding vector of the corresponding customer class node is read, and the customer fusion embedding vector is used as the benchmark risk representation vector for path occlusion comparison. For each candidate path instance in the subset of candidate paths for guarantee chains, the subset of candidate paths for equity penetration, and the subset of candidate paths for fund transfers, a path masking operation is performed one by one; Using the masked cross-business joint knowledge graph as input, according to the calculation method of the node-level attention coefficient and the semantic-level attention coefficient, the neighbor node weighted aggregation and the three types of business meta-path semantic-level weighted fusion are re-executed on the starting customer class node to obtain the masking risk representation vector of the candidate path instance after masking; For each candidate path instance, calculate the risk representation difference between the baseline risk representation vector and the occlusion risk representation vector; By combining the time decay weight sequence, the timeliness correction is performed on the normalized risk representation difference value of each candidate path instance to generate the path risk confirmation coefficient.

2. The cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as described in claim 1, characterized in that, For each customer category node, the customer credit risk score is classified according to a preset risk threshold range, and a classified early warning result is generated, including: The system obtains a preset risk threshold range uniformly configured by the risk control parameter module of the bank's big data analysis and early warning platform. This preset risk threshold range is divided into three levels: normal level, attention level, and early warning level. When a customer's credit risk score is less than the first preset threshold, the customer is classified as normal and normal level operations are performed. When a customer's credit risk score is greater than or equal to the first preset threshold and less than the second preset threshold, the customer is classified as attention level and attention level operations are performed. When a customer's credit risk score is greater than or equal to the second preset threshold, the customer is classified as early warning level and early warning level operations are performed.

3. The cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as described in claim 2, characterized in that, Also includes, The normal level operation is as follows: the unified global customer identifier of this customer type node, the customer credit risk score, and the normal level label are recorded to the normal level customer list without triggering any warning actions or generating risk transmission path records. The attention level operation involves recording the unified global customer identifier, customer credit risk score, and attention level label of this customer type node into the attention level customer list. Simultaneously, it extracts the effective candidate path instance with the highest path risk confirmation coefficient under this customer type node from the risk transmission calibration parameters, uses the customer type node sequence of this effective candidate path instance as the attention transmission path, and encapsulates the attention transmission path, customer credit risk score, and attention level label together into an attention prompt record. The attention prompt record is for manual review by bank risk control personnel, but no proactive warning is pushed out at this time. The early warning level operation involves recording the unified global customer identifier, customer credit risk score, and early warning level label of this customer type node into the early warning level customer list. Simultaneously, it extracts the effective candidate path instance with the highest path risk confirmation coefficient under this customer type node from the risk transmission calibration parameters, uses the customer type node sequence of this effective candidate path instance as the risk transmission path, and encapsulates the risk transmission path, customer credit risk score, and early warning level label together into an early warning detail record. The early warning detail record is used by the bank's big data analysis and early warning platform to trigger real-time early warning pushes.

4. The cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as described in claim 3, characterized in that, The calculation process for the customer credit risk score includes: Read the customer fusion embedding vector corresponding to all customer class nodes from the customer fusion embedding vector, read the business element path type label, customer class node sequence and path risk confirmation coefficient of all valid candidate path instances corresponding to all customer class nodes from the risk transmission calibration parameter, and associate the customer fusion embedding vector with the risk transmission calibration parameter according to the unified global customer identifier to obtain the fusion input data pair of each customer class node; The message passing layer of the customer fusion embedded vector input graph neural network model for each customer class node in the fused input data pair is executed according to the three types of business meta-paths: guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path, under the adjacency structure constraint of the cross-business joint knowledge graph. For each customer-type node, the path aggregation message vector of all valid candidate path instances under the three types of business meta-paths—guarantee chain meta-path, equity penetration meta-path, and fund transfer meta-path—is summed within the domain according to the business meta-path type to obtain the guarantee chain domain aggregation vector, equity penetration domain aggregation vector, and fund transfer domain aggregation vector. The customer fusion embedding vector, the guarantee chain domain aggregation vector, the equity penetration domain aggregation vector, and the fund transfer domain aggregation vector of the target customer node are concatenated column by column, and then dimensionality-reduced through the cross-business fusion linear transformation layer of the graph neural network model to obtain the cross-business risk feature vector of each customer node. The cross-business risk feature vector is input into the credit risk scoring layer to obtain the customer credit risk score for each customer class node.

5. The cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as described in claim 1, characterized in that, The set of candidate risk transmission paths includes: Read the feature vectors of customer-type nodes, feature vectors of event-type nodes, time decay weights, association type labels, and source system identifiers from the cross-business joint knowledge graph. According to the association type labels carried by the directed weighted edges, divide the directed weighted edges in the cross-business joint knowledge graph into three subsets of directed weighted edges: guarantee chain edges, equity penetration edges, and fund transfer edges. Based on the guarantee chain edge subset, the equity penetration edge subset, and the fund transfer edge subset, three types of business meta-paths are preset in the cross-business joint knowledge graph, wherein the three preset business meta-paths include the guarantee chain meta-path, the equity penetration meta-path, and the fund transfer meta-path. For each customer-class node in the cross-business joint knowledge graph, enumerate all directed connection sequences that satisfy the corresponding maximum path length constraint along the guarantee chain meta-path, the equity penetration meta-path, and the fund transfer meta-path. Record the customer-class node sequence and the corresponding time decay weight sequence in each directed connection sequence as candidate path instances to obtain a set of candidate risk transmission paths.

6. A cross-business credit risk assessment system based on graph neural networks and a joint bank knowledge graph, based on the cross-business credit risk assessment method based on graph neural networks and a joint bank knowledge graph as described in any one of claims 1 to 5, characterized in that, include: The multi-source data construction module is used to collect multi-business customer data from the bank's integrated business management system, corporate business system and transaction risk assessment system, extract customer entity sets, business event sets and cross-business relationship sets, and construct multi-source heterogeneous raw datasets. The joint graph construction module is used to construct a cross-business joint knowledge graph by using the customer entity set and the business event set as graph nodes, the cross-business relationship set as graph edges, and configuring time decay weights for the graph edges. The risk path identification module is used to pre-set three types of business meta-paths—guarantee chain, equity penetration, and fund transfer—for the cross-business joint knowledge graph, calculate node-level attention coefficients and semantic-level attention coefficients, and obtain customer fusion embedding vectors and candidate risk transmission path sets. The path risk calibration module is used to compare the path masking of the candidate risk transmission path set, calculate the difference in risk representation before and after path masking, and generate a path risk confirmation coefficient in combination with the time decay weight, and use the path risk confirmation coefficient as a risk transmission calibration parameter. The risk scoring and early warning module is used to integrate the customer into a vector input graph neural network model, aggregate cross-business risks according to the path risk confirmation coefficient, obtain cross-business risk feature vectors, calculate customer credit risk scores, and generate graded early warning results.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the cross-business credit risk assessment method based on graph neural networks and bank joint knowledge graphs as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Industrial chain risk assessment model and method based on graph neural network, and medium

    CN117236698A

  • Multi-modal enterprise credit risk assessment method and device based on knowledge graph

    CN120509958A