A supply chain whole-process risk control method and system based on a graph neural network
Patent Information
- Application Number
- CN202611071125.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-20
AI Technical Summary
[0005]为了解决未对企业节点对在每个业务通道上的因果传导方向作出显式判定,进而导致各企业的风险评判结果不准确的技术问题,本申请提供了一种基于图神经网络的供应链全流程风控方法及系统
本申请构建通道拆分的供应链异构图,通过双向互相关比较提取节点间风险传导方向与有效时延,结合传导稳定性计算权重,并依据时延对前驱邻居执行逐通道异步聚合,克服了传统聚合未区分因果传导方向的缺陷,准确隔离上游风险源与下游受累方,还原风险在多维度的时空演变路径;防止评估结果被反向污染,提升预警准确度与溯源能力。
Smart Images

Figure CN122596681B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of supply chain risk control technology, and in particular to a supply chain full-process risk control method and system based on graph neural networks. Background Technology
[0002] Supply chain finance services facilitate the flow of funds and logistics between core manufacturing enterprises and their multi-tiered suppliers. The number of participating nodes is enormous and the business connections are complex. Small disturbances at a single upstream node often do not immediately affect the core enterprise. Instead, they are transmitted step by step over several days before financial risks such as raw material shortages or accounts payable backlogs emerge. Therefore, a full-process risk control method is needed that can identify the risk transmission path early and issue timely risk control warnings.
[0003] Currently, Chinese patent application CN120509958A discloses a multimodal enterprise credit risk assessment method based on knowledge graphs. This method collects relevant structured and unstructured data about enterprises to construct an enterprise financial knowledge graph. Static attributes and related information of entities are embedded and fused into multimodal initial feature vectors, which are then fed into a heterogeneous graph neural network. For each node, neighbor feature vectors are aggregated based on the attention scores of neighboring nodes through a heterogeneous message passing mechanism. Simultaneously, time-series attribute data is extracted from the knowledge graph, and each time step is weighted by a time attention mechanism, then convolved with a spatial graph to generate dynamic spatiotemporal feature vectors. Finally, the graph-level features and dynamic spatiotemporal features are concatenated and fed into a multilayer perceptron classifier to output the enterprise's credit risk probability.
[0004] However, the above method calculates the attention score of neighboring nodes based solely on the similarity of the node features themselves, without making an explicit determination of the causal transmission direction of enterprise nodes on each business channel. During the aggregation stage, it cannot distinguish the directional differences between causal transmissions, which causes the features of downstream victim enterprise nodes to be applied in reverse to upstream enterprise nodes, resulting in inaccurate risk assessment results for each enterprise. Summary of the Invention
[0005] To address the technical problem of inaccurate risk assessment results for enterprises due to the lack of explicit determination of the causal transmission direction of enterprise nodes in each business channel, this application provides a supply chain end-to-end risk control method and system based on graph neural networks.
[0006] In its first aspect, this application provides a supply chain end-to-end risk control method based on graph neural networks, comprising: S101, collecting enterprise node information, edge relationships, and multi-dimensional supply chain sequences, constructing a supply chain heterogeneous graph, and performing channel decomposition and standardization on the multi-dimensional supply chain sequence to obtain standardized feature sequences for each feature channel; S102, for each feature channel of any central node in the supply chain heterogeneous graph, performing forward cross-correlation calculations led by neighboring nodes and reverse cross-correlation calculations led by the central node, respectively, and when the forward peak value is greater than the reverse peak value, marking the neighboring node as a predecessor and determining the optimal forward delay as the effective delay, thereby obtaining a predecessor neighbor set; S103 103. Based on the standardized feature sequence and the effective delay, and considering both the transmission stability and the channel correlation degree, calculate the transmission weight of each neighbor node in the predecessor neighbor set to the central node on each feature channel. The channel correlation degree is related to the mutual information between feature channels. 104. After each feature channel backtracks the standardized feature sequence according to the effective delay, it uses the transmission weight to asynchronously aggregate each neighbor node in the predecessor neighbor set channel by channel to obtain the risk state vector of each feature channel. 105. The risk state vectors of each feature channel are concatenated and mapped by a classifier to obtain the risk score of each enterprise node. An early warning report is output according to the risk score.
[0007] By comparing the peak values of the multidimensional sequence through bidirectional cross-correlation, the adjacent nodes with larger positive peak values are marked as predecessors and their effective delays are extracted. Based on this delay, the sequence is backtracked and asynchronously aggregated channel by channel. This integrates the network topology with the business time lag, identifies the real trajectory of risk transmission across levels, avoids information misalignment caused by indiscriminate aggregation, and ensures the accuracy of risk scoring for each enterprise node.
[0008] Preferably, the steps of calculating the positive cross-correlation include: shifting the standardized feature sequence of the neighboring node in the feature channel along the negative time axis by several steps, multiplying it with the standardized feature sequence of the center node in the same feature channel step by step and accumulating the results to obtain the positive cross-correlation value for the number of shift steps; within a preset range of the number of shift steps, taking the number of shift steps corresponding to the maximum positive cross-correlation value as the optimal positive delay; the maximum value of the positive cross-correlation value corresponds to the positive peak value.
[0009] Preferably, the steps of calculating the reverse cross-correlation include: shifting the standardized feature sequence of the central node in the feature channel along the negative time axis by several steps, multiplying it with the standardized feature sequence of the neighboring node in the same feature channel step by step and accumulating the results to obtain the reverse cross-correlation value at the number of shift steps; within a preset range of the number of shift steps, taking the number of shift steps corresponding to the maximum reverse cross-correlation value as the optimal reverse delay; the maximum value of the reverse cross-correlation value corresponds to the reverse peak value.
[0010] Preferably, in the step of obtaining the predecessor neighbor set, the risk control method further includes: when both the forward peak and the reverse peak are lower than the effective threshold, deleting the corresponding neighbor node from the predecessor neighbor set; the effective threshold is determined by the 95th quantile of the cross-correlation peak distribution of unrelated node pairs in the statistical historical data.
[0011] By pre-statistically determining the effective threshold based on the peak distribution of cross-correlation between unrelated node pairs, and removing the corresponding neighbor node when both the positive and negative peaks are below the threshold, a reasonable correlation benchmark is established using historical statistical patterns. This filters out false correlation noise caused by random fluctuations in complex business environments, ensuring that the predecessor neighbor set has real business connections and improving the anti-interference capability of the entire process risk control.
[0012] Preferably, the calculation steps for the transmission stability are as follows: at each historical time point within the historical verification window, construct a first window vector using the local subsequences of any neighboring node on any feature channel after effective time delay alignment, and construct a second window vector using the local subsequences of the central node in the corresponding time period; calculate the cosine similarity between the first window vector and the second window vector, and use the mean of the cosine similarity at each historical time point within the historical verification window as the transmission stability of the neighboring node to the central node on the feature channel.
[0013] By extracting local subsequences aligned with effective time delays within the historical verification window and calculating the mean cosine similarity as an evaluation index of transmission stability, the long-term consistency of fluctuation trends between adjacent nodes is verified based on the evolution law of time series, and interference from high cross-correlation peaks caused by accidental factors is eliminated, ensuring the reliability and accuracy of transmission weights.
[0014] Preferably, the calculation steps for the channel correlation degree are as follows: determine the prior correlation strength between each feature channel according to the supply chain business logic to obtain the prior correlation matrix; statistically analyze the mutual information of any two feature channel sequences from historical data to obtain the mutual information matrix; and then weight and sum the prior correlation matrix and the mutual information matrix according to the fusion coefficient and map them through the Sigmoid activation function to obtain the channel correlation degree matrix including the channel correlation degree between each feature channel.
[0015] Preferably, calculating the propagation weights of each neighbor node in the predecessor neighbor set to the central node in each feature channel includes: multiplying the propagation stability by the standardized feature sequence of the neighbor node backtracked with effective time delay in the current feature channel to obtain an independent contribution term; for other feature channels of the neighbor node that are marked as predecessors outside the current feature channel, backtracking the standardized feature sequence according to the effective time delay corresponding to the other feature channels, and weighting and summing the backtracked standardized feature sequence with the corresponding elements in the channel correlation matrix to obtain a collaborative contribution term; calculating the sum of the independent contribution term and the collaborative contribution term, and then weighting it with a learnable attention parameter vector and mapping it through the ReLU activation function to obtain a marginal contribution score; normalizing the marginal contribution scores of each neighbor node in the predecessor neighbor set to obtain the propagation weights of each neighbor node to the central node in the current feature channel.
[0016] The calculation process of transmission weights is consistent with the objective situation that different characteristic channels in the supply chain have different time differences in risk transmission, and realizes the time alignment of asynchronous risk signals across channels, thereby more accurately and objectively quantifying the comprehensive risk push exerted by each upstream node.
[0017] Preferably, the asynchronous aggregation of each neighbor node in the predecessor neighbor set using the propagation weights includes: backtracking the standardized feature sequences of each neighbor node in the predecessor neighbor set on any feature channel with an effective time delay, then performing a linear transformation using a learnable transformation matrix, and multiplying them with the propagation weights of the neighbor nodes; summing the multiplication results of each neighbor node in the predecessor neighbor set, and mapping them using the ReLU activation function to obtain the risk state vector of the central node in the feature channel; when the predecessor neighbor set is empty, the risk state vector of the corresponding feature channel is a zero vector.
[0018] Preferably, the output of the early warning report based on the risk score includes: comparing the risk score with a high-risk threshold and a medium-risk threshold, and classifying it into three levels: high-risk, medium-risk, and normal; marking trigger nodes whose predecessor neighbor set is empty as risk sources in the early warning report; marking neighbor nodes whose reverse peak value is not lower than the forward peak value as risk diffusion receptors in the early warning report; and simultaneously outputting the predecessor neighbor set and transmission weight of the trigger node as traceable attribution information, wherein the trigger node is a high-risk or medium-risk enterprise node.
[0019] By identifying nodes whose predecessor neighbor set is empty as risk sources and marking nodes with higher reverse peak values as risk diffusion receptors, and simultaneously outputting transmission weights as attribution information, we can accurately locate the initial source of the crisis and the downstream enterprises that are vulnerable to its impact, providing direct guidance for risk prevention and rapid response in actual risk control operations.
[0020] In a second aspect, this application also provides a supply chain end-to-end risk control system based on graph neural networks, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a supply chain end-to-end risk control method based on graph neural networks as described in the first aspect of this application is implemented.
[0021] The technical solution of this application has the following beneficial technical effects: This application constructs a heterogeneous supply chain graph with channel splitting, extracts the risk transmission direction and effective delay between nodes through bidirectional cross-correlation comparison, calculates weights based on transmission stability, and performs asynchronous channel-by-channel aggregation on predecessor neighbors based on delay. This overcomes the defect of traditional aggregation that does not distinguish the causal transmission direction, accurately isolates upstream risk sources and downstream affected parties, and restores the spatiotemporal evolution path of risks in multiple dimensions; it prevents the assessment results from being contaminated in reverse, and improves the accuracy of early warning and traceability capabilities. Attached Figure Description
[0022] Figure 1 This is a flowchart of a supply chain end-to-end risk control method based on graph neural networks, according to an embodiment of this application.
[0023] Figure 2 This is a schematic diagram of the comparison of peak values of positive and negative cross-correlation and the determination of the predecessor neighbor set according to an embodiment of this application.
[0024] Figure 3 This is a structural block diagram of a supply chain end-to-end risk control system based on graph neural networks, according to an embodiment of this application. Detailed Implementation
[0025] According to the first aspect of this application, a supply chain end-to-end risk control method based on graph neural networks is provided, which is applied to the end-to-end risk control scenario of supply chain finance. For example, a manufacturing company's supply chain network includes approximately 1200 enterprise nodes across first-, second-, and third-tier suppliers. These nodes form complex directed connections through fund transfers and logistics fulfillment. Risk control of the financial risks of all enterprise nodes in the supply chain network is required, ultimately clearly distinguishing between proactive defaulters and passively affected parties in the early warning report. Figure 1 This is a flowchart illustrating a supply chain end-to-end risk control method based on graph neural networks, according to an embodiment of this application. Figure 1 As shown, the supply chain end-to-end risk control method based on graph neural networks includes steps S101 to S105, which are described in detail below.
[0026] S101, collect enterprise node information, connection relationships and multi-dimensional supply chain sequence, construct supply chain heterogeneous graph, and perform channel splitting and standardization on the multi-dimensional supply chain sequence to obtain standardized feature sequence.
[0027] In one embodiment, all enterprise entities participating in supply chain collaboration are extracted from the business data platform, and each enterprise is mapped as a node, with a total of 1200 nodes. Fund transfer and logistics fulfillment records between nodes are extracted from the order system, fund settlement system, and logistics fulfillment system. A directed edge is established between any two nodes that have any of the above relationships, with the direction from the fund payer to the payee and from the logistics shipper to the recipient, forming an edge set to obtain a supply chain heterogeneous graph.
[0028] Each node stores a multi-dimensional supply chain sequence. The financial, logistics, and operational business parameters differ significantly in scale and volatility. The business data is divided into five independent feature channels based on business attributes, with the time step at the daily level. In this embodiment, the multi-dimensional supply chain sequence includes five feature channels: logistics volatility, fund settlement days, order completion rate, account balance change rate, and warehouse turnover rate. These five independent feature channels are merely preferred examples; in other embodiments, other dimensions of feature data can be selected according to actual business needs, which does not constitute a limitation of this application.
[0029] The number of days for fund settlement ranges from 0 to 90, and the rate of change in account balance reaches the level of millions of yuan. The dimensional span between feature channels directly affects the subsequent calculation of cross-correlation and attention. To eliminate dimensional differences, standardization is performed independently for each channel. Specifically, each original feature value is subtracted from the mean of the corresponding channel over all nodes and the entire historical time period, and then divided by the standard deviation of the corresponding channel. The resulting dimensionless scalar is the standardized feature value. This transformation is the classic Z-score transformation in statistics.
[0030] It should be noted that this embodiment requires that the length of the effective historical data for each node be no less than 60 days, in order to meet the minimum requirements for backtracking depth for subsequent bidirectional cross-correlation calculations and propagation stability.
[0031] S102, for each feature channel of any central node in the supply chain heterogeneous graph, perform forward cross-correlation calculation with neighbor nodes leading and reverse cross-correlation calculation with central node leading respectively. When the forward peak value is greater than the reverse peak value, mark the neighbor node as the predecessor and determine the forward optimal delay as the effective delay to obtain the predecessor neighbor set.
[0032] In one embodiment, the central node is any enterprise node in the supply chain heterogeneous graph. For each pair of connected nodes and each characteristic channel in the supply chain heterogeneous graph, bidirectional cross-correlation calculations are performed independently to clarify the direction of risk transmission between enterprise nodes.
[0033] Cross-correlation is used to measure the degree of linear correlation between two sequences under different time shifts. Its traditional form can be expressed as:
[0034] In the formula, For time shift The cross-correlation value at the location; and For sequence and sequence Mid-moment and time The value; For summation indexing. Traditional forms only provide unidirectional computation, making it difficult to distinguish sequences. and sequence The causal and lag relationships between nodes are investigated. This application constructs cross-correlation in both positive and negative directions for each pair of connected nodes and introduces a finite backtracking window length. Only the most recent ones are accumulated. The product of steps.
[0035] The formula for positive cross-correlation is:
[0036] In the formula, For feature channels Upper neighbor node In time shift Under the condition relative to the central node The positive cross-correlation value; As the central node In the feature channel Mid-time step Standardized eigenvalues; For neighboring nodes In the feature channel Mid-time step Standardized eigenvalues; To backtrack and calculate the window length; This is the number of forward time shift steps, ranging from 1 to... ; For summation index.
[0037] Symmetrically, the formula for inverse cross-correlation is:
[0038] In the formula, For feature channels Upper central node In time shift Under the condition relative to neighboring nodes The inverse cross-correlation value; For neighboring nodes In the feature channel Mid-time step Standardized eigenvalues; As the central node In the feature channel Mid-time step The standardized eigenvalues.
[0039] It should be noted that in the forward formula, the node in the leading position is the neighbor node. Its eigenvalues are backtracked additionally. After the step, the central node at the response position Alignment corresponds to the propagation assumption that "neighboring nodes are ahead of the central node"; in the reverse formula, the central node is in the leading position. Its eigenvalues are backtracked additionally. After the step, the neighboring nodes at the response position Alignment corresponds to the opposite propagation hypothesis that "the central node leads the neighboring nodes." The two formulas quantify the linear correlation strength of the two opposite propagation directions, providing a data foundation for subsequent determination of causal direction.
[0040] The preferred value for the backtracking window length is 25 days. Since the complete fluctuation cycle of orders and settlements in historical business data is 18 to 22 days, a backtracking window length of 25 days can fully cover a business cycle. The preferred value for the maximum delay limit is 15 days, which covers the maximum value of the two standard settlement cycles in this scenario. This can accommodate most actual transmission paths and avoid introducing pseudo-peaks in an excessively long search range.
[0041] The forward optimal delay is the number of time shift steps that maximizes the positive cross-correlation value within the range of time shift values, and the corresponding maximum value is the forward peak value; the reverse optimal delay and the reverse peak value are obtained by taking the maximum value operation in the same way.
[0042] Figure 2 This is a schematic diagram illustrating the peak value comparison and predecessor neighbor set determination of forward cross-correlation calculation and reverse cross-correlation calculation according to an embodiment of this application. Figure 2 As shown, the curve on the left represents the result of the forward cross-correlation calculation, with the horizontal axis representing the number of time steps and the vertical axis representing the cross-correlation value. The peak of the curve corresponds to the forward peak and the forward optimal delay. The curve on the right represents the result of the reverse cross-correlation calculation, with the peak corresponding to the reverse peak and the reverse optimal delay. The dashed line marks the position of the effective threshold. When the forward peak is greater than the reverse peak and both are higher than the effective threshold, the corresponding neighbor node is included in the predecessor neighbor set.
[0043] After obtaining the cross-correlation peaks in both directions, comparing their relative magnitudes determines the causal direction of the two node pairs on the corresponding feature channel. When the positive peak is greater than the negative peak, the causal direction marker of the neighbor node is incremented by 1, indicating that the neighbor node is upstream of the central node in this feature channel, and its historical state has a causal predictive effect on the current state of the central node. The corresponding effective delay is the optimal positive delay. When the positive peak is not greater than the negative peak, the causal direction marker is decremented by 1, indicating that the neighbor node is downstream of the central node. The corresponding effective delay is set to null, and such neighbors will be excluded from subsequent attention calculations. All neighbor nodes with a causal direction marker of +1 are gathered to form the predecessor neighbor set of the central node on the channel, which is the range of neighbors participating in subsequent weight propagation calculations and asynchronous aggregation.
[0044] For example, consider a scenario where the neighboring node is a first-tier supplier and the central node is the core enterprise. On the logistics channel, the forward peak is 0.71, and the reverse peak is 0.24, indicating a causal direction of +1 and an effective delay of 3 days. This means the supplier's logistics anomaly precedes the core enterprise's by 3 days. Similarly, for the same two enterprises, on the funding channel, the forward peak is 0.18, and the reverse peak is 0.55, indicating a causal direction of -1. This signifies a reversal in the transmission direction of funding, with the core enterprise's payment delay preceding the supplier's financial difficulties. In this scenario, the neighboring node on this channel represents the victim, not the originator.
[0045] For some newly established node pairs, due to limited historical data, the causal direction cannot be determined, or there is indeed no causal direction on certain feature channels. Therefore, in order to avoid errors in risk scoring caused by forcibly introducing a causal direction, when both the positive peak and the negative peak are below the effective threshold, the corresponding neighbor node is marked as a non-predecessor, that is, the corresponding neighbor node is removed from the predecessor neighbor set.
[0046] The effective threshold is set to 0.15. The calibration process for this value is as follows: samples of unrelated node pairs that are known to have no business relationship are selected from historical business data. For example, isolated enterprises belonging to different industries and having no financial or logistical exchanges are selected. The bidirectional cross-correlation peaks of these unrelated node pairs are summarized to obtain the peak distribution of the unrelated node pairs. The 95th percentile of this distribution is then taken. This value represents the correlation between enterprise nodes on the feature channel when there is no business relationship. In other words, correlations below this value are likely due to noise and cannot represent the true causal direction between feature channels.
[0047] Through bidirectional cross-correlation calculation and causal direction determination, each pair of connected nodes outputs directional causal delay information carrying directional information on each channel, so that the subsequent attention calculation has directional discrimination capability in structure, and the attention calculation avoids the influence of reverse causality of downstream victim nodes.
[0048] S103, based on the standardized feature sequence and the effective time delay, and combining the transmission stability and the channel correlation degree, calculate the transmission weight of each neighbor node in the predecessor neighbor set to the central node on each feature channel. The channel correlation degree is related to the mutual information between feature channels.
[0049] In one embodiment, a propagation weight between 0 and 1 is calculated for each neighbor node in the predecessor neighbor set, which serves as the basis for attention allocation in subsequent asynchronous aggregation.
[0050] First, the calculation steps for the channel correlation degree are as follows: determine the prior correlation strength between each characteristic channel according to the supply chain business logic to obtain the prior correlation matrix; calculate the mutual information of any two characteristic channel sequences from historical data to obtain the mutual information matrix; and then map the prior correlation matrix and the mutual information matrix by weighting them according to the fusion coefficient and using the Sigmoid activation function to obtain the channel correlation degree matrix that includes the channel correlation degree between each characteristic channel.
[0051] Among them, the prior correlation matrix reflects the hard logical constraints in supply chain operations. Business experts assign values ranging from 0 to 1 to the correlation strength between each pair of the five feature channels. "No payment without goods" is the basic rule of supply chain settlement, so the correlation strength between "settlement days" and "logistics volatility" is assigned to 0.8. "Warehouse turnover rate" and "account balance change rate" also show obvious synergistic changes in most cases, and are assigned a value of 0.6. The prior correlation matrix can ensure that the subsequent fusion results are not dominated by data randomness.
[0052] Mutual information is a standard statistic in the field of measurement for the degree of nonlinear correlation between two random variables. This application directly uses the `mutual_info_regression` function from the sklearn library for implementation. The larger the elements of the mutual information matrix, the stronger the ability of observing the state of one channel at the data level to reduce the uncertainty of the other channel.
[0053] The elements of the channel correlation matrix satisfy:
[0054] In the formula, For feature channels With feature channels The channel correlation degree, with a value ranging from 0 to 1; The elements in the prior association matrix are assigned values by business experts according to the supply chain business logic. These are elements in the mutual information matrix, obtained from historical data statistics; The fusion coefficient; This is a commonly used nonlinear activation function in this field. In the early stages of operation, when the accumulated data is less than one year old, the mutual information estimation itself has a large variance, and the fusion coefficient should be increased to 0.7 to focus more on the hard logic of business experts; after the accumulated data exceeds one year, the mutual information estimation tends to stabilize, and the fusion coefficient can be reduced to 0.3 to focus more on the hidden correlations discovered by the data itself.
[0055] In this embodiment, after confirming that a neighboring node is the predecessor neighbor of the central node on the channel, it is still necessary to verify whether the transmission relationship has been repeatedly established in history, thereby obtaining the transmission stability. The calculation steps of the transmission stability are as follows: at each historical time point within the historical verification window, construct a first window vector with the local subsequence of any neighboring node on any feature channel after effective time delay alignment, and construct a second window vector with the local subsequence of the central node in the corresponding time period; calculate the cosine similarity between the first window vector and the second window vector, and take the mean of the cosine similarity at each historical time point within the historical verification window as the transmission stability of the neighboring node to the central node on the feature channel.
[0056] Conduction stability satisfies:
[0057] In the formula, For feature channels Upper neighbor node For the central node The conduction stability, with values ranging from -1 to +1; This is the length of the historical verification window; This is an index of historical time points within the historical verification window; For the first A historical point in time; For feature channels Above, based on neighboring nodes Effective delay alignment of historical time points The first window vector is formed by neighboring nodes. In the feature channel Above, from time step to The standardized eigenvalues are composed of, where For feature channels Upper neighbor node For the central node Effective delay, This is the length of the local sub-window; similarly, The second window vector is formed by the central node. In the feature channel Above, from time step to The standardized eigenvalues are composed of the numerator; the numerator is the dot product of the first window vector and the second window vector. For the second window vector The L2 norm; the denominator is the product of the L2 norms of the first window vector and the second window vector. For historical time points The cosine similarity between the first window vector and the second window vector.
[0058] The preferred length of the local sub-window is 7 days, covering the shortest period of 5 to 7 days from order confirmation to the first fund settlement in this embodiment, which can fully accommodate one working week. The preferred length of the historical verification window is 40 days, which is twice the longest complete period, based on the average interval of approximately 15 to 20 days between abnormal transmission events in the historical data of this embodiment, and can cover at least two complete transmission periods. It should be noted that when the L2 norm product of the first window vector and the second window vector is zero, since at least one window vector is a zero vector, the cosine similarity of that historical time point is directly assigned a value of 0. If the L2 norm product of all historical time points is zero within the complete historical verification window W, then the transmission stability of the neighboring node to the central node on the feature channel is directly assigned a value of 0.
[0059] When the local trends of neighboring nodes and the central node repeatedly move in the same direction within the historical window, the transmission stability approaches 1, and the transmission relationship is considered stable. If the trends of the two nodes are randomly distributed, the transmission stability approaches 0. Even if the cross-correlation peak is high at this moment, the transmission stability will still significantly reduce the independent contribution of the neighboring node in subsequent steps.
[0060] After obtaining the propagation stability and channel correlation, the propagation weights of each neighbor node in the predecessor neighbor set to the center node in each feature channel are calculated. Specifically, calculating the propagation weights of each neighbor node in the predecessor neighbor set to the center node in each feature channel includes: multiplying the propagation stability by the standardized feature sequence of the neighbor node backtracked with effective time delay in the current feature channel to obtain the independent contribution term; for other feature channels of the neighbor node that are marked as predecessors outside the current feature channel, backtracking the standardized feature sequence according to the effective time delay of the other feature channels, and weighting and summing the backtracked standardized feature sequence with the corresponding elements in the channel correlation matrix to obtain the collaborative contribution term; calculating the sum of the independent contribution term and the collaborative contribution term, and weighting it with a learnable attention parameter vector and mapping it through the ReLU activation function to obtain the marginal contribution score; normalizing the marginal contribution scores of each neighbor node in the predecessor neighbor set to obtain the propagation weights of each neighbor node to the center node in the current feature channel.
[0061] It should be noted that the predecessor neighbor set includes multiple neighbor nodes, and each neighbor node corresponds to a transmission weight on each feature channel. The transmission weight can be regarded as the attention of the central node to each neighbor node on each feature channel during the subsequent aggregation process.
[0062] Before giving the formula for calculating the transmission weight, let's first introduce the scoring formula for attention score in the standard graph attention mechanism. The specific relationship is as follows:
[0063] In the formula, For nodes in the graph result For nodes Attention score; A learnable attention parameter vector; This is the transpose of the learnable attention parameter vector; To share the linear transformation matrix; and For nodes With nodes eigenvectors; This is a vector concatenation operation; It is a non-linear activation function.
[0064] To match the causal transmission characteristics between various feature channels in the supply chain, this application modifies the attention score in the standard graph attention mechanism to additive attention, so that the transmission weight of the neighbor node to the central node on the current feature channel simultaneously considers the current feature channel and other feature channels with causal transmission relationships.
[0065] Using the current feature channel as the feature channel For example, feature channels Upper neighbor node For the central node Marginal contribution score Satisfying the relation:
[0066] In the formula, For feature channels The learnable attention parameter vector, The dimension is equal to the backtracking computation window length, which is updated and determined by backpropagation during the training phase; Transpose it. It is a non-linear activation function. For feature channels Other feature channels The degree of channel correlation.
[0067] In the above formula, As an independent contribution, For feature channels Upper neighbor node For the central node Conductivity stability, For neighboring nodes In the feature channel Above, by feature channel Upper neighbor node For the central node Effective delay Standardized feature values from backtracking.
[0068] In the above formula, For collaborative contributions, For feature channels Indexes of other feature channels, For neighboring nodes Other feature channels Above, by feature channel Effective delay Standardized feature values for backtracking; For indicator functions, parameters This indicates that neighboring nodes are in other feature channels. It was also determined to be the predecessor of the central node.
[0069] The independent contribution term modulates the backtracking feature value of neighboring nodes on the current feature channel using the transmission stability as a credibility weight, representing the risk thrust exerted by neighboring nodes on the central node through the current feature channel; the collaborative contribution term weights and accumulates the feature values of neighboring nodes backtracking through their respective causal alignment delays on all other channels according to the channel correlation, representing the additional risk thrust of collaborative transmission between feature channels; the sum of the two terms is projected as a scalar by the learnable attention parameter vector, and then activated to obtain the marginal contribution score.
[0070] It should be noted that the propagation delays of the same pair of nodes are not equal on different feature channels. A neighboring node in the logistics channel may be 3 days ahead of the central node, while a neighboring node in the funds channel may be 7 days ahead. If all other feature channels in the collaborative contribution item are backtracked by a uniform 3 days, the features of the funds channel will be sampled out of order. Therefore, each other feature channel uses its own effective delay for backtracking. Furthermore, only when a neighboring node is also determined to be a predecessor in other feature channels is the collaborative contribution generated by that other feature channel included in the collaborative contribution item, ensuring that the collaborative contribution item only collects the true upstream collaborative signals.
[0071] Subsequently, the marginal contribution scores of all neighbors in the predecessor neighbor set are normalized using Softmax to obtain the propagation weights of each neighbor node to the center node in the current feature channel. Upper neighbor node For the central node The transmission weight is denoted as This indicates that among all neighbor nodes in the predecessor neighbor set, the neighbor node is... For the central node In the feature channel In the risk contribution, neighboring nodes The share of responsibility.
[0072] When the set of predecessor neighbors of any enterprise node on a feature channel is empty, it means that the enterprise node is the upstream risk source of the feature channel. At this time, there is no need to calculate the transmission weight. In the subsequent asynchronous aggregation process of each channel, the risk state vector of the enterprise node in the feature channel is set to zero.
[0073] S104, after each feature channel backtracks the standardized feature sequence based on the effective time delay, the transmission weight is used to asynchronously aggregate each neighbor node in the predecessor neighbor set channel by channel to obtain the risk state vector of each feature channel.
[0074] In one embodiment, the risk thrust contributed by each neighbor node in the predecessor neighbor set is aggregated into the final state of the central node on each feature channel. In graph neural networks, traditional aggregation operations perform weighted summation of the channel features of all neighbor nodes, ignoring the causal transmission relationship between neighbor nodes. This application combines transmission weights to perform asynchronous aggregation only on a channel-by-channel basis on the neighbor nodes in the predecessor neighbor set.
[0075] Central Node In the feature channel Risk state vector on Satisfying the relation:
[0076] In the formula, For feature channels Upper central node The predecessor neighbor set; For feature channels Upper neighbor node For the central node Transmission weight; For feature channels A learnable transformation matrix is used to linearly map one-dimensional normalized eigenvalues to... 3D state space; For neighboring nodes In the feature channel Above, by feature channel Upper neighbor node For the central node Effective delay Standardized feature values for backtracking; This is a conventional activation function in this field.
[0077] Within the summation symbol, feature extraction is performed on the standardized eigenvalues of neighboring nodes after backtracking using a learnable transformation matrix to obtain a risk state fragment. This is further multiplied by the propagation weight, and the state fragment is discounted according to the allocated responsibility share; the summation is performed on the predecessor neighbor set, and the state fragments of all upstream neighbors are aggregated into a risk tensor, which is the central node. In the feature channel The risk state vector on the feature channel is used to characterize the risk state vector of all neighbor nodes in the predecessor neighbor set. The impact on the risk status of the central node; the central node can be obtained using the same method. The risk state vector across all feature channels. The preferred dimension of the risk state vector is 64, but it can also be set to other commonly used values such as 32 or 128; this application does not impose any restrictions.
[0078] Each feature channel corresponds to a learnable transformation matrix. The learnable transformation matrix is uniformly initialized by Xavier during the training initialization phase, and then updated by gradient descent according to the loss function through backpropagation.
[0079] It should be noted that when the set of predecessor neighbors of any enterprise node on a feature channel is empty, the risk state vector of the corresponding feature channel is a vector of all zeros.
[0080] Ultimately, the central node obtains five risk state vectors, each with 64 dimensions. The combination of independent aggregation per channel and directed constraints on the predecessor neighbor set preserves an independent and readable risk evolution trajectory for each business dimension. Small but real disturbances on any channel are neither masked by normal values in other channels nor contaminated by reverse signals from downstream victim nodes.
[0081] S105, the risk state vectors of each feature channel are concatenated, and the risk score of each enterprise node is obtained by mapping through the trained classifier. The early warning report is output according to the risk score.
[0082] In one embodiment, the risk state vectors of each feature channel are concatenated to obtain the concatenated feature vector of the central node. The concatenation adopts the vector head-to-tail connection operation, and the risk state vectors of each 64-dimensional channel are connected head-to-tail in channel order. In this embodiment, a 320-dimensional concatenated feature vector is obtained, with each of the 5 feature channels occupying 64 dimensions.
[0083] It should be noted that the process of obtaining the risk state vector in steps S101 to S104 constitutes the inference and prediction stage of the graph neural network. The learnable attention parameter vector involved in step S103 and the learnable transformation matrix involved in step S104 have already had their specific values determined during the training process. By inputting the concatenated feature vector into the trained classifier for mapping, the risk score of each enterprise node can be obtained.
[0084] The classifier is a fully connected network, and in this embodiment, the fully connected network has two layers. The concatenated feature vector is first fed into the first fully connected layer, which linearly maps the 320-dimensional input to a 128-dimensional hidden vector, and then... The output of the nonlinear activation function is fed into the second fully connected layer. This second fully connected layer linearly maps the 128-dimensional hidden layer vector to a numerical scalar. Then, a Sigmoid activation function maps this numerical scalar to a range of 0 to 1. The output value is the risk score of the central node; the closer it is to 1, the higher the default risk of the enterprise corresponding to the central node at the current time step. The number of layers in the fully connected network, the nonlinear activation function, and the Sigmoid activation are all well-known techniques to those skilled in the art. The number of layers in the fully connected network can be adjusted according to actual needs, and this application does not impose any restrictions.
[0085] This section describes the training process of the classifier and learnable parameters. Risk state vectors for each feature channel of enterprise nodes are collected over historical time periods. The classifier and learnable parameters are trained using the historical default labels of enterprise nodes as supervision signals. The loss function is binary cross-entropy loss, and the supervision signal is the node's actual default label, where a label of 1 indicates a genuine default and 0 indicates no default record. Both the classifier and learnable parameters are updated through backpropagation of the loss function and gradient descent by the Adam optimizer until convergence is achieved, yielding the specific values of the trained classifier and learnable parameters. The learnable parameters include the learnable attention parameter vector and the learnable transformation matrix for each feature channel.
[0086] The risk score is compared with the high-risk and medium-risk thresholds and divided into three levels: high-risk, medium-risk, and normal. Nodes with an empty predecessor neighbor set are marked as chain-head risk sources in the graded early warning report. Neighbor nodes with a reverse peak value not lower than the forward peak value are marked as risk diffusion receptors in the early warning report. At the same time, the predecessor neighbor set and transmission weight of the triggering node are output as traceable attribution information. The triggering node is a high-risk or medium-risk enterprise node.
[0087] In this embodiment, the high-risk threshold is set to 0.7 and the medium-risk threshold is set to 0.4. The calibration method for the two thresholds is as follows: on the validation set, candidate threshold combinations are scanned with a step size of 0.05, and the F1 score and recall rate for each combination are calculated. The combination that results in the highest F1 score and a recall rate of not less than 0.85 is selected. This threshold can be adjusted according to the different tolerances for false negatives and false positives in risk control operations. For scenarios sensitive to false negatives, the medium-risk threshold should be lowered, and for scenarios sensitive to false positives, the high-risk threshold should be raised.
[0088] The early warning report simultaneously outputs four types of structured attribution information: The first type is trigger node information, including node number, business channel, risk score, and corresponding risk level; the second type is the predecessor neighbor set and transmission path, listing all neighbor nodes in the predecessor neighbor set for each triggering channel, along with their transmission weights, so that risk control personnel can directly see which upstream neighbor the risk originates from and what share of responsibility they currently bear; the third type is the labeling of the risk source at the beginning of the chain, marking the risk source at the beginning of the chain in the report, indicating to risk control personnel that the enterprise node is the starting point of the entire risk chain on this characteristic channel, and its operating data should be checked first; the fourth type is the labeling of the risk diffusion receptor, listing neighbor nodes that have been determined to have a reverse peak value not lower than the forward peak value separately. These enterprise nodes are the recipients of the risk spreading outward from the central node, and after labeling, risk control personnel can take remedial measures in a timely manner.
[0089] According to a second aspect of this application, this application also provides a supply chain end-to-end risk control system based on graph neural networks. Figure 3 This is a structural block diagram of a supply chain end-to-end risk control system based on a graph neural network, according to an embodiment of this application. Figure 3 As shown, the system 50 includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a supply chain end-to-end risk control method based on a graph neural network as described in the first aspect of this application. The system also includes other components well-known to those skilled in the art, such as a communication bus and communication interfaces. Their configurations and functions are known in the art and will not be described further here.
[0090] It should be noted that any modifications and improvements made by those skilled in the art without departing from the concept of this application are within the scope of protection of this application.
Claims
1. A supply chain end-to-end risk control method based on graph neural networks, characterized in that, include: S101: Collect enterprise node information, connection relationships and multi-dimensional supply chain sequences, construct a supply chain heterogeneous graph, and perform channel decomposition and standardization on the multi-dimensional supply chain sequence to obtain the standardized feature sequence of each feature channel; S102, For each feature channel of any central node in the supply chain heterogeneous graph, perform forward cross-correlation calculation with neighbor nodes leading and reverse cross-correlation calculation with central node leading respectively. When the forward peak value is greater than the reverse peak value, mark the neighbor node as the predecessor and determine the optimal forward delay as the effective delay to obtain the predecessor neighbor set. The positive cross-correlation calculation includes: shifting the standardized feature sequences of neighboring nodes in the feature channel along the negative time axis by several steps, multiplying them step-by-step with the standardized feature sequences of the center node in the same feature channel, and then summing the results to obtain the positive cross-correlation value for that number of shift steps; within a preset range of shift steps, the shift step number corresponding to the maximum positive cross-correlation value is taken as the optimal positive delay; the maximum positive cross-correlation value corresponds to the positive peak value; The reverse cross-correlation calculation includes: shifting the standardized feature sequence of the central node in the feature channel along the negative time axis by several steps, multiplying it with the standardized feature sequence of the neighboring nodes in the same feature channel step by step and accumulating the results to obtain the reverse cross-correlation value at the number of shift steps; within a preset range of the number of shift steps, the number of shift steps corresponding to the maximum reverse cross-correlation value is taken as the optimal reverse delay; the maximum value of the reverse cross-correlation value corresponds to the reverse peak value; S103, based on the standardized feature sequence and effective time delay, and combining the transmission stability and channel correlation, calculates the transmission weight of each neighbor node in the predecessor neighbor set to the central node on each feature channel. The propagation stability is calculated as follows: at each historical time point within the historical verification window, a first window vector is constructed using the local subsequences of any neighboring node on any feature channel after effective time delay alignment, and a second window vector is constructed using the local subsequences of the central node in the corresponding time period; the cosine similarity between the first window vector and the second window vector is calculated, and the mean of the cosine similarity at each historical time point within the historical verification window is used as the propagation stability of the neighboring node on the feature channel to the central node; The channel correlation degree is calculated as follows: the prior correlation strength between each characteristic channel is determined according to the supply chain business logic to obtain the prior correlation matrix; the mutual information of any two characteristic channel sequences is statistically analyzed from historical data to obtain the mutual information matrix; the prior correlation matrix and the mutual information matrix are weighted and summed according to the fusion coefficient and then mapped by the Sigmoid activation function to obtain the channel correlation degree matrix including the channel correlation degree between each characteristic channel. S104. After backtracking the standardized feature sequence based on the effective time delay, each feature channel uses the transmission weight to asynchronously aggregate each neighbor node in the predecessor neighbor set to obtain the risk state vector of each feature channel. S105 concatenates the risk state vectors of each feature channel, maps them to the trained classifier to obtain the risk score of each enterprise node, and outputs an early warning report according to the risk score.
2. The supply chain end-to-end risk control method based on graph neural networks according to claim 1, characterized in that, In the step of obtaining the predecessor neighbor set, the risk control method further includes: When both the forward peak and the reverse peak are below the effective threshold, the corresponding neighbor node is removed from the predecessor neighbor set; The effective threshold is determined by the 95th quantile of the peak cross-correlation distribution of unrelated node pairs in statistical historical data.
3. The supply chain end-to-end risk control method based on graph neural networks according to claim 1, characterized in that, The calculation of the propagation weights of each neighbor node in the predecessor neighbor set to the center node in each feature channel includes: multiplying the propagation stability by the standardized feature sequence of the neighbor node backtracked with effective time delay in the current feature channel to obtain the independent contribution term; for other feature channels of the neighbor node that are marked as predecessors outside the current feature channel, backtracking the standardized feature sequence according to the effective time delay of the other feature channels, and weighting and summing the backtracked standardized feature sequence with the corresponding elements in the channel correlation matrix to obtain the collaborative contribution term; calculating the sum of the independent contribution term and the collaborative contribution term, and then weighting it with a learnable attention parameter vector and mapping it through the ReLU activation function to obtain the marginal contribution score; normalizing the marginal contribution scores of each neighbor node in the predecessor neighbor set to obtain the propagation weights of each neighbor node to the center node in the current feature channel.
4. The supply chain end-to-end risk control method based on graph neural networks according to claim 1, characterized in that, The asynchronous aggregation of neighbor nodes in the predecessor neighbor set using propagation weights includes: After backtracking the standardized feature sequences of each neighbor node in the predecessor neighbor set on any feature channel with an effective time delay, the sequence is linearly transformed by a learnable transformation matrix and then multiplied with the propagation weights of the neighbor nodes. The multiplication results of each neighbor node in the predecessor neighbor set are summed and mapped by the ReLU activation function to obtain the risk state vector of the central node in the feature channel. When the predecessor neighbor set is empty, the risk state vector of the corresponding feature channel is a vector of all zeros.
5. The supply chain end-to-end risk control method based on graph neural networks according to claim 1, characterized in that, The risk-based early warning report output includes: The risk score is compared with the high-risk threshold and the medium-risk threshold, and divided into three levels: high-risk, medium-risk and normal. For triggering nodes whose predecessor neighbor set is empty, they are marked as risk sources in the early warning report; for neighboring nodes whose reverse peak value is not lower than the forward peak value, they are marked as risk diffusion receptors in the early warning report; at the same time, the predecessor neighbor set and transmission weight of the triggering node are output as traceable attribution information, and the triggering node is a high-risk or medium-risk enterprise node.
6. A supply chain end-to-end risk control system based on graph neural networks, characterized in that, It includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a supply chain end-to-end risk control method based on a graph neural network according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Multi-modal enterprise credit risk assessment method and device based on knowledge graph
CN120509958A
Financial risk control and anomaly detection method and system based on graph neural network
CN121329708A
Supply chain demand prediction and risk early warning method and system based on multi-source data
CN121544028A