Artificial intelligence based fraud detection method and system for financing

By constructing a bidirectional causal graph and a risk propagation adjacency matrix for enterprises, and combining Wasserstein distance and ensemble learning framework, the problem of difficulty in modeling risk propagation relationships between enterprises in existing technologies is solved, thereby achieving accuracy and robustness in financing guarantee anti-fraud and reducing the risks and losses of financial institutions.

CN121481707BActive Publication Date: 2026-06-05QINGDAO FINANCING GUARANTEE GROUP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO FINANCING GUARANTEE GROUP CO LTD
Filing Date
2025-11-17
Publication Date
2026-06-05

Smart Images

  • Figure CN121481707B_ABST
    Figure CN121481707B_ABST
Patent Text Reader

Abstract

The application provides an artificial intelligence-based financing guarantee anti-fraud method and system, relates to the technical field of artificial intelligence, and comprises the following steps: constructing a feature vector by acquiring enterprise multi-source heterogeneous data, establishing a bidirectional causal graph structure to represent the risk transmission relationship between enterprises, acquiring a risk transmission adjacency matrix by using a NOTEARS algorithm, constructing a risk distribution metric space by using a Wasserstein distance to perform distribution robustness optimization, training to obtain an enterprise risk representation vector, and combining an ensemble learning framework to evaluate fraud risk. The application can effectively identify fraud in a complex network and improve the accuracy and reliability of financing guarantee risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to artificial intelligence technology, and more particularly to a method and system for preventing fraud in financing guarantees based on artificial intelligence. Background Technology

[0002] With the development of financial markets and the expansion of financing guarantee business, fraudulent activities have become increasingly complex and covert, posing significant risks and losses to financial institutions. Existing anti-fraud methods for financing guarantees mainly rely on manual review and simple rule-based judgments, which are insufficient to address the current complex and ever-changing fraud patterns. Existing anti-fraud technologies for financing guarantees still have the following shortcomings and deficiencies:

[0003] Existing technologies are insufficient to effectively capture the risk transmission relationships between enterprises, especially in terms of accurately modeling the two-way risk transmission mechanism between enterprises. This leads to the neglect of the risk transmission effect between related enterprises, affecting the accuracy of fraud risk assessment.

[0004] Existing risk assessment models are highly sensitive to changes in data distribution. When there is a discrepancy between the distribution of test data and training data, the model's generalization ability decreases significantly, making it difficult to cope with the constantly changing data distribution in actual business and lacking robustness.

[0005] Existing technologies typically treat enterprises as independent entities for risk assessment, failing to effectively integrate enterprise-related network structure information with individual characteristics. They lack a unified framework to integrate enterprise characteristics with network propagation effects, resulting in a lack of holistic consideration in the final fraud risk score. Summary of the Invention

[0006] This invention provides an artificial intelligence-based method and system for preventing fraud in financing guarantees, which can solve the problems in the prior art.

[0007] A first aspect of this invention provides an artificial intelligence-based anti-fraud method for financing guarantees, comprising:

[0008] A multi-source heterogeneous data set of enterprises is acquired to construct enterprise feature vectors; a bidirectional causal graph structure is constructed based on the enterprise feature vectors, and the bidirectional causal graph structure uses a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises; for each enterprise node, the risk status of the node is calculated based on the risk status of its directly associated enterprises; the NOTEARS optimization algorithm is used to calculate the risk propagation adjacency matrix between enterprises based on the risk status of the nodes.

[0009] Based on the enterprise feature vector and the risk propagation adjacency matrix, a metric space for enterprise risk distribution is constructed using Wasserstein distance. Distributed robustness optimization is performed in this metric space, and a triplet loss function is used to train and obtain an enterprise risk representation vector. This enterprise risk representation vector is then input into an ensemble learning framework to obtain an initial score for enterprise fraud risk.

[0010] Based on the risk propagation adjacency matrix and the initial score, the final fraud risk score considering the risk propagation effect is calculated, and the anti-fraud decision result for financing guarantee is generated.

[0011] A bidirectional causal graph structure is constructed based on the enterprise feature vectors. This structure uses a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises. For each enterprise node, its risk status is calculated based on the risk status of its directly associated enterprises. The steps for calculating the risk propagation adjacency matrix between enterprises using the NOTEARS optimization algorithm based on the node's risk status include:

[0012] Based on the enterprise feature vectors, the risk correlation between enterprises is calculated, and an enterprise risk correlation matrix is ​​constructed. A directed acyclic graph structure is constructed based on the enterprise risk correlation matrix, and the directed acyclic graph structure represents the bidirectional risk propagation relationship between enterprises.

[0013] For each enterprise node in the directed acyclic graph structure, the risk status of the enterprise node is calculated. The risk status of the enterprise node is obtained by a weighted combination of its own characteristic risk assessment value and the influence value of related enterprises. The influence value of related enterprises is the weighted sum of the risk status of the enterprise node's directly related enterprises and the corresponding risk propagation weight.

[0014] The risk states of the enterprise nodes are combined into an enterprise risk state matrix. An optimization objective function is constructed based on the enterprise risk state matrix. The optimization objective function includes a risk state fitting term, a risk propagation time-series dynamic term, and an acyclic constraint term. The NOTEARS optimization algorithm is used to solve the optimization objective function to obtain a risk propagation adjacency matrix between enterprises that satisfies the directed acyclic constraint. The risk propagation adjacency matrix represents the topological structure of risk propagation between enterprises.

[0015] The steps for constructing and solving the objective function include:

[0016] The risk status fitting term is calculated by combining the differences in risk status among enterprises with their importance weights; the risk propagation time series dynamic term is calculated by combining the time window decay weight with the historical risk propagation intensity; and the acyclicity constraint term is calculated by the cyclic connection of risk propagation relationships among enterprises.

[0017] Construct an augmented Lagrange function that includes the optimization objective function, wherein the augmented Lagrange function introduces Lagrange multipliers and penalty terms;

[0018] The augmented Lagrangian function is optimized by using the L-BFGS algorithm, which incorporates a second-order Hessian correction term. The second-order Hessian correction term is used to construct a diagonal block matrix based on the risk propagation time series characteristics, and the gradient update direction is adjusted through the diagonal block matrix.

[0019] When the change value of the optimization objective is less than the set threshold and the acyclic constraint is satisfied, the optimized temporal risk propagation adjacency matrix is ​​obtained; the propagation intensity in the optimized temporal risk propagation adjacency matrix is ​​normalized, and the risk propagation relationship between enterprises is determined by temporal confidence score.

[0020] Based on the enterprise feature vector and the risk propagation adjacency matrix, the steps of constructing a metric space for enterprise risk distribution using Wasserstein distance, performing disjoint robustness optimization in the metric space, and training the enterprise risk representation vector using a triplet loss function include:

[0021] The enterprise feature vector is weighted by the risk propagation adjacency matrix, and the weighted enterprise feature vector is mapped to the probability density space using a Gaussian kernel function to obtain the enterprise risk probability distribution.

[0022] Based on the enterprise risk probability distribution, a Wasserstein distance metric space is constructed, where the Wasserstein distance is obtained by solving the p-order norm of the optimal transmission plan between enterprise risk probability distributions;

[0023] A triplet loss function is constructed based on the Wasserstein distance. The risk distribution pairs of enterprises are classified into categories according to a preset Wasserstein distance threshold, and a training sample set is constructed. By optimizing the triplet loss function, a representation vector reflecting the characteristics of enterprise risk distribution is learned.

[0024] An adversarial perturbation is introduced into the representation vector, and a two-layer optimization objective is set. The adversarial perturbation is maximized in the inner layer, and the triple loss is minimized in the outer layer to enhance the robustness of the representation vector and obtain the final enterprise risk representation vector. Based on the final enterprise risk representation vector, the distance between enterprise pairs is calculated to obtain the consistency index and the discriminative index, and the enterprise risk representation results are evaluated.

[0025] The steps to introduce adversarial perturbations into the representation vector, set a two-layer optimization objective (maximizing the adversarial perturbation in the inner layer and minimizing the triplet loss in the outer layer) to enhance the robustness of the representation vector and obtain the final enterprise risk representation vector include:

[0026] The sensitive direction in the representation vector space is calculated based on the gradient information of the representation vector; an initial adversarial perturbation is generated in the sensitive direction, and the magnitude of the initial adversarial perturbation is determined based on the enterprise's historical risk fluctuation range;

[0027] Design a triplet loss function with dynamic weights, wherein the dynamic weights are determined based on the Euclidean distance between pairs of enterprise risk representation vectors; adjust the direction and magnitude of the initial adversarial perturbation according to the Euclidean distance to generate the final adversarial perturbation;

[0028] A two-layer optimization framework for adversarial training is constructed. The inner layer optimization finds the optimal perturbation direction by maximizing the triplet loss function under adversarial perturbation. The outer layer optimization improves the robustness of the representation vector by minimizing the weighted combination of the original triplet loss function and the adversarial triplet loss function. The weight coefficients of the weighted combination are dynamically adjusted according to the degree of perturbation influence during the training process.

[0029] The robustness of the representation vector is evaluated based on the validation sample set. The stability of the representation results under different perturbation amplitudes is calculated. When the stability meets the preset threshold, the enterprise risk representation vector with anti-robustness is obtained.

[0030] The steps for inputting the enterprise risk representation vector into the ensemble learning framework to obtain an initial score for enterprise fraud risk include:

[0031] A feature importance matrix is ​​constructed based on the enterprise risk representation vector of historical fraud samples. Principal component analysis is used to reduce the dimensionality of the feature importance matrix to obtain a feature combination that can identify fraud.

[0032] The enterprise risk representation vector is projected onto the subspace corresponding to the feature combination method to construct a gradient boosting-based decision tree ensemble structure, with each decision tree corresponding to a feature combination method; the prediction results of each decision tree are weighted and fused to obtain the initial score of enterprise fraud risk.

[0033] The steps for calculating the final fraud risk score after considering the risk propagation effect, and generating the anti-fraud decision result for financing guarantees, based on the risk propagation adjacency matrix and the initial score, include:

[0034] Construct an enterprise risk propagation network, determine the risk propagation path between enterprises based on the risk propagation adjacency matrix, and calculate the risk propagation time window based on the degree of business correlation between enterprises and historical transaction data to obtain the risk propagation intensity under different time windows;

[0035] For each enterprise node, a first-order risk propagation impact value is calculated based on the initial scores of its neighboring enterprise nodes and the corresponding risk propagation intensity. A distance attenuation coefficient is introduced, which decreases as the propagation path increases, to calculate a second-order risk propagation impact value. The first-order and second-order risk propagation impact values ​​are weighted and combined to obtain the enterprise's cumulative risk impact coefficient.

[0036] The initial score is adjusted based on the cumulative risk impact coefficient to obtain a final fraud risk score that takes into account the risk propagation effect; based on the final fraud risk score and the company's guarantee limit, a decision result for financing guarantee anti-fraud is generated.

[0037] A second aspect of the present invention provides an artificial intelligence-based anti-fraud system for financing guarantees, comprising:

[0038] The first unit is used to acquire multi-source heterogeneous data of enterprises to construct enterprise feature vectors; construct a bidirectional causal graph structure based on the enterprise feature vectors, and use a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises; for each enterprise node, calculate the risk status of the node according to the risk status of its directly associated enterprises; use the NOTEARS optimization algorithm to calculate the risk propagation adjacency matrix between enterprises based on the risk status of the nodes.

[0039] The second unit is used to construct a metric space for enterprise risk distribution based on the enterprise feature vector and the risk propagation adjacency matrix using Wasserstein distance, perform dissimilar robustness optimization in the metric space, and train the enterprise risk representation vector using a triplet loss function; input the enterprise risk representation vector into the ensemble learning framework to obtain an initial score for enterprise fraud risk;

[0040] The third unit is used to calculate the final fraud risk score after considering the risk propagation effect based on the risk propagation adjacency matrix and the initial score, and generate the financing guarantee anti-fraud decision result.

[0041] A third aspect of the present invention,

[0042] An electronic device is provided, comprising:

[0043] processor;

[0044] Memory used to store processor-executable instructions;

[0045] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0046] Fourth aspect of the embodiments of the present invention,

[0047] A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0048] The beneficial effects of this application are as follows:

[0049] This invention accurately captures the complex risk propagation relationships between enterprises by constructing a two-way causal graph structure and a risk propagation adjacency matrix, effectively identifies related fraudulent activities, and improves the accuracy of anti-fraud measures for financing guarantees.

[0050] This invention introduces Wasserstein distance to construct a corporate risk distribution metric space and performs robust optimization, enabling the model to adapt to changes in data distribution and the evolution of corporate behavior, thereby enhancing the robustness and adaptability of the anti-fraud system.

[0051] This invention employs an integrated learning framework combined with risk propagation effect analysis, comprehensively considering the characteristics of the enterprise itself and related risk factors, to achieve accurate assessment and early warning of financing guarantee fraud, thereby reducing the credit risk and loss rate of financial institutions. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the anti-fraud method for financing guarantees based on artificial intelligence, as described in an embodiment of the present invention.

[0053] Figure 2 A flowchart for constructing enterprise risk representation vectors based on Wasserstein distance and adversarial robustness optimization. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0056] Figure 1 This is a flowchart illustrating an artificial intelligence-based anti-fraud method for financing guarantees according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0057] A multi-source heterogeneous data set of enterprises is acquired to construct enterprise feature vectors; a bidirectional causal graph structure is constructed based on the enterprise feature vectors, and the bidirectional causal graph structure uses a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises; for each enterprise node, the risk status of the node is calculated based on the risk status of its directly associated enterprises; the NOTEARS optimization algorithm is used to calculate the risk propagation adjacency matrix between enterprises based on the risk status of the nodes.

[0058] Based on the enterprise feature vector and the risk propagation adjacency matrix, a metric space for enterprise risk distribution is constructed using Wasserstein distance. Distributed robustness optimization is performed in this metric space, and a triplet loss function is used to train and obtain an enterprise risk representation vector. This enterprise risk representation vector is then input into an ensemble learning framework to obtain an initial score for enterprise fraud risk.

[0059] Based on the risk propagation adjacency matrix and the initial score, the final fraud risk score considering the risk propagation effect is calculated, and the anti-fraud decision result for financing guarantee is generated.

[0060] In one optional implementation, a bidirectional causal graph structure is constructed based on the enterprise feature vectors. This bidirectional causal graph structure uses a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises. For each enterprise node, the risk state of that node is calculated based on the risk state of its directly associated enterprises. The step of calculating the risk propagation adjacency matrix between enterprises based on the node's risk state using the NOTEARS optimization algorithm includes:

[0061] Based on the enterprise feature vectors, the risk correlation between enterprises is calculated, and an enterprise risk correlation matrix is ​​constructed. A directed acyclic graph structure is constructed based on the enterprise risk correlation matrix, and the directed acyclic graph structure represents the bidirectional risk propagation relationship between enterprises.

[0062] For each enterprise node in the directed acyclic graph structure, the risk status of the enterprise node is calculated. The risk status of the enterprise node is obtained by a weighted combination of its own characteristic risk assessment value and the influence value of related enterprises. The influence value of related enterprises is the weighted sum of the risk status of the enterprise node's directly related enterprises and the corresponding risk propagation weight.

[0063] The risk states of the enterprise nodes are combined into an enterprise risk state matrix. An optimization objective function is constructed based on the enterprise risk state matrix. The optimization objective function includes a risk state fitting term, a risk propagation time-series dynamic term, and an acyclic constraint term. The NOTEARS optimization algorithm is used to solve the optimization objective function to obtain a risk propagation adjacency matrix between enterprises that satisfies the directed acyclic constraint. The risk propagation adjacency matrix represents the topological structure of risk propagation between enterprises.

[0064] For example, a company's feature vector is obtained, which includes multi-dimensional feature information such as the company's financial indicators, operating status, and industry classification. These features may include financial indicators such as debt-to-equity ratio, current ratio, net profit margin, and revenue growth rate, as well as basic information such as company size, years of establishment, and number of employees. For example, for a manufacturing company A, its feature vector includes multi-dimensional features such as a debt-to-equity ratio of 0.45, a current ratio of 1.2, a net profit margin of 0.08, a revenue growth rate of 0.12%, a company size (large), a history of 15 years, and 1000 employees.

[0065] Risk correlation can be determined by measuring the similarity between enterprise feature vectors, such as cosine similarity or Euclidean distance. For example, for two enterprises A and B, if they have 5-dimensional feature vectors [0.45, 1.2, 0.08, 0.12, 5] and [0.48, 1.1, 0.07, 0.15, 4] respectively, their risk correlation is calculated to be 0.92. The risk correlation results between all pairs of enterprises are combined into an enterprise risk correlation matrix, which is an n x n square matrix, where n is the total number of enterprises.

[0066] A directed acyclic graph (DAG) structure is constructed, and a risk correlation threshold is set, for example, 0.8. When the risk correlation between two companies exceeds this threshold, a connection is established in the graph. For company pairs with correlations higher than the threshold, the direction of risk propagation is determined based on factors such as company size and industry influence. For example, if company A is larger than company B, a risk propagation direction from A to B is defined. This method constructs a directed graph, and topological sorting ensures that the graph structure is acyclic, forming a DAG.

[0067] For each enterprise node in the constructed directed acyclic graph, its risk status is calculated. The risk status of an enterprise node consists of two parts: its own risk assessment value and the impact value of related enterprises. The own risk assessment value is calculated by assigning weights to key enterprise indicators. For example, weights of 0.4, 0.3, and 0.3 are assigned to indicators such as debt-to-equity ratio, current ratio, and net profit margin, respectively, and the weighted average is used to obtain the risk assessment value. If an enterprise's debt-to-equity ratio risk score is 0.5, its current ratio risk score is 0.3, and its net profit margin risk score is 0.2, then its own risk assessment value is 0.5 × 0.4 + 0.3 × 0.3 + 0.2 × 0.3 = 0.35. The impact value of related enterprises is calculated by multiplying the risk status of directly related enterprises by their corresponding risk propagation weights and then summing the results. For example, if enterprise A is directly related to enterprises B and C, and their risk statuses are 0.4 and 0.5 respectively, and their propagation weights are 0.3 and 0.2 respectively, then the impact value of related enterprises is 0.4 × 0.3 + 0.5 × 0.2 = 0.22. Company A's risk status is a weighted combination of its own risk assessment value and the impact value of related companies, such as 0.35×0.7+0.22×0.3=0.311.

[0068] The risk states of all enterprise nodes are organized into an enterprise risk state matrix, which is an n-dimensional vector where n is the total number of enterprises. When constructing the optimization objective function based on the enterprise risk state matrix, the risk state matrix is ​​directly used as the data input for the optimization objective. The risk state fitting term calculates the error by comparing the risk states predicted by the model with the actual values ​​in the risk state matrix; the risk propagation time-series dynamic term uses the time-series change data of the risk state matrix; and the acyclicity constraint term imposes constraints on the adjacency matrix generated by the model to ensure that the graph structure it represents is free of cycles. These three terms together constitute the optimization objective function, guiding the algorithm to find the optimal risk propagation structure.

[0069] The NOTEARS optimization algorithm is used to solve the above-mentioned objective function. The core innovation of the NOTEARS algorithm lies in transforming the discrete constraints of a directed acyclic graph (DAG) into a continuously differentiable function, allowing the application of standard continuous optimization techniques. The algorithm first defines a special continuous function to represent whether the graph structure contains cycles; when this function is zero, it indicates that the graph structure does not contain any cycles. Then, the algorithm transforms the original problem into a constrained optimization problem, further transforming it into an unconstrained optimization problem by introducing Lagrange multipliers. During the solution process, the algorithm uses gradient descent to iteratively update the elements of the risk propagation adjacency matrix, checking the acyclic constraint after each update. The algorithm stops iterating when the change in the objective function value is less than a preset threshold and the acyclic constraint is satisfied. In practical applications, an appropriate learning rate can be set to control the step size of each update, a maximum number of iterations can be set to avoid excessive computation time, and a convergence threshold can be used to determine whether the algorithm has reached convergence. This optimization method can efficiently solve for the inter-firm risk propagation adjacency matrix that satisfies the directed acyclic constraint. For example, the learning rate can be set to 0.01, the maximum number of iterations to 1000, and the convergence threshold to 1e-6.

[0070] After optimization, a risk propagation adjacency matrix is ​​obtained between enterprises. Each element in this matrix represents the risk propagation intensity from enterprise i to enterprise j. For example, in a case involving 5 enterprises, the final risk propagation adjacency matrix is: [[0,0.3,0.5,0,0],[0,0,0.2,0.4,0],[0,0,0,0.1,0.3],[0,0,0,0,0.2],[0,0,0,0,0]]. This matrix indicates that enterprise 1 propagates risk to enterprises 2 and 3 with intensities of 0.3 and 0.5 respectively, enterprise 2 propagates risk to enterprises 3 and 4 with intensities of 0.2 and 0.4 respectively, and so on. By analyzing this matrix, the critical paths of risk propagation and core enterprises can be identified, providing decision support for risk prevention and control.

[0071] This method accurately characterizes the risk propagation relationships between enterprises by constructing a bidirectional causal graph structure and a risk correlation matrix, and solves the directed acyclic constraint problem using the NOTEARS optimization algorithm. This approach overcomes the limitations of traditional models in representing complex relationships between enterprises, enabling more accurate identification of risk propagation paths and key node enterprises, providing financial institutions with a comprehensive view of the risk network. By calculating the risk status of enterprise nodes and constructing a risk propagation adjacency matrix, it can capture hidden risk transmission patterns, improve the accuracy of fraud risk identification, and reduce missed detections and false positives.

[0072] In one alternative implementation, the steps of constructing and solving the optimization objective function include:

[0073] The risk status fitting term is calculated by combining the differences in risk status among enterprises with their importance weights; the risk propagation time series dynamic term is calculated by combining the time window decay weight with the historical risk propagation intensity; and the acyclicity constraint term is calculated by the cyclic connection of risk propagation relationships among enterprises.

[0074] Construct an augmented Lagrange function that includes the optimization objective function, wherein the augmented Lagrange function introduces Lagrange multipliers and penalty terms;

[0075] The augmented Lagrangian function is optimized by using the L-BFGS algorithm, which incorporates a second-order Hessian correction term. The second-order Hessian correction term is used to construct a diagonal block matrix based on the risk propagation time series characteristics, and the gradient update direction is adjusted through the diagonal block matrix.

[0076] When the change value of the optimization objective is less than the set threshold and the acyclic constraint is satisfied, the optimized temporal risk propagation adjacency matrix is ​​obtained; the propagation intensity in the optimized temporal risk propagation adjacency matrix is ​​normalized, and the risk propagation relationship between enterprises is determined by temporal confidence score.

[0077] For example, risk status data for multiple enterprises at different time windows can be acquired to construct a time-series risk status matrix. Each row in this matrix represents an enterprise, each column represents a point in time, and the matrix elements represent the risk status value of the corresponding enterprise at that point in time. For instance, monthly risk status data for the past 12 months can be collected to form a time-series risk status matrix containing all target enterprises.

[0078] The objective function is constructed based on the time-series risk state matrix, which includes three key components: risk state fitting term, risk propagation time-series dynamic term, and acyclicity constraint term.

[0079] The risk state fitting term is calculated as follows: For any two firms i and j, the Euclidean distance between their corresponding row vectors in the time-series risk state matrix is ​​calculated as the risk state difference. Then, weights are assigned to each pair of firm relationships based on firm importance. The importance weights are determined according to factors such as firm asset size and industry influence. For example, firms are divided into three categories: large, medium, and small, and assigned weights of 0.8, 0.5, and 0.2, respectively. The importance weight between two firms can be the average of their respective weights. The risk state fitting term is obtained by summing the products of the risk state differences of all firm pairs and their corresponding importance weights.

[0080] The calculation of the risk propagation time-series dynamics involves the time window decay weight and the historical risk propagation intensity. The time window decay weight uses an exponential decay function, in the form: weight = α (T-t)Where α is the attenuation coefficient (e.g., 0.9), T is the total number of time points, and t is the index of a specific time point. For example, if there are 12 time points, the weight of the first time point is 0.9. 11 The weight of the 12th time point is 0.9. 0 =1. The intensity of historical risk propagation is determined by calculating the cross-correlation coefficient r(τ) between the risk indicators of the two companies, where τ represents the time delay. The maximum correlation coefficient r is selected. max and the corresponding delay τ max If r max >0.6 and τ max If the value is greater than 0, it is considered that risk propagation exists from leading firms to lagging firms, and the initial propagation strength is set as r. max ×(1-e^(-τ max The risk propagation time-series dynamics are obtained by calculating the weighted sum of the time window decay weights at each time point and the corresponding historical risk propagation intensity.

[0081] The purpose of the acyclicity constraint is to ensure that the final risk propagation network does not contain circular dependencies. A circular connection refers to a closed loop formed in the enterprise risk propagation path, such as enterprise A influencing enterprise B, enterprise B influencing enterprise C, and enterprise C influencing enterprise A. Circular connections can be detected by analyzing the characteristics of the risk propagation adjacency matrix. Specifically, this involves checking the results of successive powers of the adjacency matrix; when a non-zero value appears on the diagonal element of a power, it indicates the existence of a circular path. The acyclicity constraint is calculated by summing the penalty values ​​for all cycles.

[0082] When constructing the augmented Lagrangian function, the original objective function f(W) is combined with the acyclic constraint h(W)=0, in the form L(W,λ,μ)=f(W)+λh(W)+μ / 2×h(W). 2 Here, W is the risk propagation adjacency matrix, λ is the Lagrange multiplier, and μ is the penalty factor. The Lagrange multiplier λ is initially set to 0 and updated to λ + μh(W) after each iteration. The penalty factor μ is initially set to 1.0 and multiplied by 1.5 after each iteration until it reaches an upper limit of 100. This method ensures that the algorithm gradually increases the penalty for acyclic constraints, guiding the solution to converge to the feasible region that satisfies the constraints.

[0083] When using the L-BFGS algorithm to solve for the augmented Lagrangian function, a second-order Hessian correction term is introduced to improve convergence speed and stability. For n companies, they can be divided into k groups, each group being approximately n / k in size. For example, 20 companies can be divided into 4 groups of 5 companies each, with grouping based on industry category or historical transaction correlation between companies. A Hessian submatrix is ​​constructed for each company within each group, and this is achieved by storing the gradient difference (yi) from the most recent m=10 iterations. k ) and parameter difference (s kThe Hessian matrix is ​​approximated using the BFGS formula. The Hessian matrix for each iteration consists of three parts: retaining the Hessian matrix from the previous iteration, adding a positive correction term (the ratio of the outer product of the gradient difference vectors to the inner product of the gradient difference vectors and the parameter difference vectors), and subtracting a negative correction term (the ratio of the inner product of the previous Hessian matrix, the parameter difference vectors, and their transposes to the inner product of the parameter difference vectors and the Hessian matrix and the parameter difference vectors). This iterative calculation method avoids directly calculating the second derivative, improving algorithm efficiency. These block-wise Hessian matrices form a diagonal block matrix, used to adjust the gradient update direction and accelerate the optimization process.

[0084] The optimization process is terminated when the change in the objective function value of two consecutive iterations is less than the preset threshold of 0.0001 and the acyclic constraint value is less than 0.01. The algorithm is considered to have converged and the optimized temporal risk propagation adjacency matrix is ​​obtained.

[0085] The propagation intensity in the optimized adjacency matrix is ​​normalized by dividing the risk propagation intensity emitted by each enterprise by its maximum value, mapping all intensity values ​​to the interval between zero and one, facilitating subsequent analysis and comparison. Finally, the risk propagation relationship between enterprises is determined through time-series confidence scoring, which considers both the stability and persistence of the propagation intensity.

[0086] Stability is assessed by calculating the coefficient of variation (CV) of the transmission intensity over a historical time window. The CV is defined as the standard deviation divided by the mean, used to measure the relative dispersion of the data. For example, if a company's average transmission intensity over a 12-month window is 0.45 and the standard deviation is 0.09, then its CV = 0.09 / 0.45 = 0.2, indicating that the transmission intensity is relatively stable. Persistence is assessed by counting the number of time windows in which the transmission intensity exceeds a specific threshold. The threshold is set at 0.3, meaning that when the monthly transmission intensity is greater than 0.3, the month is considered to have a significant risk of transmission. The persistence ratio is defined as the ratio of the number of months exceeding the threshold to the total number of months. For example, if the transmission intensity exceeds 0.3 in 9 out of 12 months, then the persistence ratio is 9 / 12 = 0.75. The time series confidence score comprehensively considers stability and persistence, and the calculation formula is: Time Series Confidence Score = 0.5 × (1 - min(1, CV / 0.5)) + 0.5 × min(1, N) s / N total ); where CV is the coefficient of variation, N s N is the number of time windows in which the propagation intensity exceeds the threshold. total This represents the total number of time windows.

[0087] Taking the above example, the stability score = 0.5 × (1 - 0.2 / 0.5) = 0.3, the persistence score = 0.5 × 0.75 = 0.375, and the total score = 0.675. When the confidence score exceeds 0.6, a significant risk transmission relationship is confirmed.

[0088] A complete real-world case study: Suppose we analyze the risk propagation relationships of 5 companies (A, B, C, D, and E) over 12 months. In the initial adjacency matrix before optimization, the propagation strength from company A to B is estimated to be 0.4. After 35 iterations of the optimization algorithm, the final propagation strength from company A to B is adjusted to 0.65, with a coefficient of variation of 0.15. The propagation strength exceeds 0.3 in 10 out of the 12 months. The time-series confidence score is 0.5×(1-0.15 / 0.5)+0.5×(10 / 12)≈0.35+0.42=0.77, which exceeds the threshold of 0.6. Therefore, a significant risk propagation relationship is confirmed between A and B.

[0089] This method innovatively constructs an optimization objective function that includes a risk state fitting term, a risk propagation time-series dynamic term, and an acyclicity constraint term, and solves it using an augmented Lagrangian function. The L-BFGS algorithm, which introduces a second-order Hessian correction term, significantly improves computational efficiency and optimization accuracy, effectively handling the complex calculations of large-scale enterprise risk networks. By normalizing the optimized adjacency matrix and applying time-series confidence scores, this method provides analysis of the temporal changes in risk propagation intensity, helping financial institutions predict risk transmission trends and achieve proactive risk management and accurate decision-making.

[0090] In one optional implementation, the steps of constructing a metric space for enterprise risk distribution using Wasserstein distance based on the enterprise feature vector and the risk propagation adjacency matrix, performing dispersive robustness optimization in the metric space, and training the enterprise risk representation vector using a triplet loss function include:

[0091] The enterprise feature vector is weighted by the risk propagation adjacency matrix, and the weighted enterprise feature vector is mapped to the probability density space using a Gaussian kernel function to obtain the enterprise risk probability distribution.

[0092] Based on the enterprise risk probability distribution, a Wasserstein distance metric space is constructed, where the Wasserstein distance is obtained by solving the p-order norm of the optimal transmission plan between enterprise risk probability distributions;

[0093] A triplet loss function is constructed based on the Wasserstein distance. The risk distribution pairs of enterprises are classified into categories according to a preset Wasserstein distance threshold, and a training sample set is constructed. By optimizing the triplet loss function, a representation vector reflecting the characteristics of enterprise risk distribution is learned.

[0094] An adversarial perturbation is introduced into the representation vector, and a two-layer optimization objective is set. The adversarial perturbation is maximized in the inner layer, and the triple loss is minimized in the outer layer to enhance the robustness of the representation vector and obtain the final enterprise risk representation vector. Based on the final enterprise risk representation vector, the distance between enterprise pairs is calculated to obtain the consistency index and the discriminative index, and the enterprise risk representation results are evaluated.

[0095] Combination Figure 2 The flowchart illustrating the construction of enterprise risk representation vectors based on Wasserstein distance and adversarial robustness optimization is provided below. For example, enterprise feature vectors are weighted using a risk propagation adjacency matrix. Specifically, the enterprise feature vector is multiplied by the corresponding row vector of the risk propagation adjacency matrix to obtain a weighted feature vector that comprehensively considers the impact of risk propagation. The element values ​​in the risk propagation adjacency matrix reflect the intensity of risk transmission between enterprises. The weighting process essentially integrates the risk characteristics of related enterprises into the feature representation of the target enterprise according to the intensity of risk propagation. For example, the original feature vector of enterprise A contains 20 feature dimensions such as debt-to-equity ratio, current ratio, and profitability, each standardized to the 0-1 range. Assuming that the weights of related enterprises B, C, and D in the risk propagation adjacency matrix are 0.6, 0.3, and 0.1 respectively, then the weighted feature vector of enterprise A is calculated as 0.6 times the original feature vector of A, plus 0.3 times the original feature vector of B, plus 0.1 times the original feature vector of C.

[0096] A Gaussian kernel function is applied to the weighted enterprise feature vectors for nonlinear transformation, mapping the feature vectors to a probability density space. The Gaussian kernel function has the form K(x,y)=exp(-||xy|| ² / 2σ ²The vector is defined as follows: x and y are eigenvectors, ||xy|| is the Euclidean distance, and σ is the bandwidth parameter. The bandwidth parameter is determined based on the distribution characteristics of the eigenvectors and can be calculated using the Silverman rule: σ = 0.9 × min(standard deviation, interquartile range / 1.34) × n^(-1 / 5), where n is the number of samples. For example, for a company's 20-dimensional weighted eigenvector, each dimension can be considered as a data point. If the standard deviation of these eigenvalues ​​is 0.15 and the sample size is 20, then the bandwidth parameter σ = 0.9 × 0.15 × 20^(-1 / 5) ≈ 0.102. After applying the Gaussian kernel function, the original eigenvector is transformed into a probability density function, which can more comprehensively characterize the company's risk distribution characteristics.

[0097] A Wasserstein distance metric space is constructed based on the enterprise risk probability distributions. The Wasserstein distance measures the difference between two probability distributions by calculating the optimal transmission scheme between them. For the risk probability distributions P and Q of two enterprises, their support sets can be obtained by uniformly sampling in the feature space, typically sampling 50-100 points as discretized support points. The element c of the transmission cost matrix C... ij The cost of moving support point i from distribution P to support point j from distribution Q is expressed as the square of the Euclidean distance between the two points. The optimal transport plan can be solved using the Sinkhorn iterative algorithm, with a regularization parameter ε = 0.01 and 100 iterations, yielding the transport matrix T. The Wasserstein distance is calculated as the sum of the corresponding product of the elements of the transport matrix T and the cost matrix C. For example, for the risk probability distributions of two companies, if each support set contains 50 points, and the elements of the constructed 50×50 transport cost matrix range from 0 to 1, the sum of the elements of the transport matrix obtained by the Sinkhorn algorithm is 1. Multiplying this by the cost matrix yields a Wasserstein distance of 0.35, indicating a high degree of similarity between the risk distributions of the two companies.

[0098] A triplet loss function is constructed using the calculated Wasserstein distance. A triple consists of an anchor firm *a*, a positive sample firm *p*, and a negative sample firm *n*. Positive sample firms are those with risk distributions similar to the anchor firm, while negative sample firms have significantly different risk distributions. Specifically, a Wasserstein distance threshold of 0.5 is set; firms with a Wasserstein distance less than 0.5 are considered similar, while those with a distance greater than 0.5 are considered dissimilar. Firms with known risk profiles are selected from historical data as training samples to construct a large number of triplet samples. For example, 100 firms are selected from historical financial data, an anchor firm is randomly chosen, and firms with Wasserstein distances less than 0.5 and greater than 0.5 are selected as positive and negative samples, respectively, constructing 10,000 triplet training samples.

[0099] The design goal of the triplet loss function is to ensure that the distance between similar firm pairs in the representation space is less than the distance between dissimilar firm pairs, and the difference exceeds a specified margin value. Specifically, it takes the form max(0, d(f...). a ,f p )-d(f a ,f n f(d) + margin), where d represents the Euclidean distance between the vectors, f(d) + margin). a f p f n Let represent the representation vectors of the anchor point, positive sample, and negative sample enterprises, respectively. The margin represents the expected marginal value, set to 0.2. A loss value greater than 0 indicates that the current representation vector does not meet the expected distance relationship and needs adjustment. The loss function is optimized using gradient descent, with an initial learning rate of 0.01. Every 5000 batches, the learning rate is reduced to 0.8 times the original rate, for a total of 30000 batches, resulting in representation vectors that reflect the risk distribution characteristics of the enterprises.

[0100] Adversarial perturbations are introduced into the learned representation vectors to enhance robustness. The adversarial perturbation δ is generated by calculating the gradient ∇_fL of the representation vector f with respect to the loss function L. Its direction is consistent with the gradient direction, and its magnitude is limited to 5% of the representation vector norm, i.e., ||δ||≤0.05×||f||. Adversarial examples are generated as f'=f+δ. A two-layer optimization objective is set: the inner optimization adjusts the perturbation δ through 5 steps of gradient ascent, with a step size of 0.01, aiming to maximize the loss function L(f+δ); the outer optimization adjusts the model parameters θ through 20 steps of gradient descent, with a step size of 0.005, aiming to minimize the weighted sum of the original loss function L(f) and the adversarial loss function L(f+δ), 0.5×L(f)+0.5×L(f+δ). In practical applications, adversarial perturbation is applied to the representation vector of a certain enterprise. A perturbation with an amplitude of 0.03 is added to the original representation vector. The representation vector obtained after two-layer optimization has stronger robustness to data perturbation while maintaining the original semantic information.

[0101] When evaluating the results of enterprise risk characterization, consistency and discriminability indices are calculated. The consistency index measures the clustering of enterprises with similar risk distributions in the characterization space, calculated as the average Euclidean distance of all enterprise pairs with a Wasserstein distance less than 0.5 in the characterization space. The discriminability index measures the separation of enterprises with different risk distributions in the characterization space, calculated as the average Euclidean distance of all enterprise pairs with a Wasserstein distance greater than 0.5 in the characterization space. The Euclidean distance between characterization vectors is calculated as the square root of the sum of the squares of the differences between the two vectors. Ideally, the consistency index should be small, and the discriminability index should be large. For example, in an experimental evaluation analyzing the characterization results of 100 enterprises, the average characterization distance for similar enterprise pairs was 0.15, and the average characterization distance for dissimilar enterprise pairs was 0.85, with a distance ratio of 0.15 / 0.85 = 0.18, indicating that the characterization vectors have good discriminative ability.

[0102] For example, a risk characterization analysis was performed on enterprises in a financial dataset containing historical risk data for 200 enterprises. Each enterprise's original feature vector contained 30 financial indicators and related-party transaction features. After weighting using a risk propagation adjacency matrix, the feature vectors were mapped to a probability density space using a Gaussian kernel function with a bandwidth parameter of 0.1. 75 support points were sampled to represent the risk probability distribution of each enterprise, and the Wasserstein distance between enterprise pairs was calculated, forming a 200×200 distance matrix with an average distance of 0.62. A distance threshold of 0.4 was set, and approximately 15,000 triplet training samples were constructed. By optimizing the triplet loss function, a 128-dimensional enterprise risk characterization vector was learned. An adversarial perturbation with an amplitude of 3% of the vector norm was introduced. After two-layer optimization, the robustness of the characterization vector was significantly improved. The consistency index on the validation set was 0.18, the discriminancy index was 0.79, and the distance ratio was 0.23, indicating that the model can effectively distinguish enterprises with different risk patterns.

[0103] This method, through a metric space constructed using Wasserstein distance, accurately captures subtle differences in enterprise risk distribution. The triplet loss function effectively learns the representation vector of the risk distribution, and the adversarial training mechanism significantly enhances the robustness of the representation. Compared to traditional methods, this approach can more accurately identify risk similarities among enterprises, exhibits stronger adaptability to data noise and distribution variations, provides a more reliable risk representation basis for fraud prevention decisions in financing guarantees, reduces the false positive rate of fraud risk, and improves the accuracy and stability of financial risk management.

[0104] In one optional implementation, the steps of introducing adversarial perturbations into the representation vector, setting a two-layer optimization objective—maximizing the adversarial perturbation in the inner layer and minimizing the triplet loss in the outer layer—to enhance the robustness of the representation vector and obtain the final enterprise risk representation vector include:

[0105] The sensitive direction in the representation vector space is calculated based on the gradient information of the representation vector; an initial adversarial perturbation is generated in the sensitive direction, and the magnitude of the initial adversarial perturbation is determined based on the enterprise's historical risk fluctuation range;

[0106] Design a triplet loss function with dynamic weights, wherein the dynamic weights are determined based on the Euclidean distance between pairs of enterprise risk representation vectors; adjust the direction and magnitude of the initial adversarial perturbation according to the Euclidean distance to generate the final adversarial perturbation;

[0107] A two-layer optimization framework for adversarial training is constructed. The inner layer optimization finds the optimal perturbation direction by maximizing the triplet loss function under adversarial perturbation. The outer layer optimization improves the robustness of the representation vector by minimizing the weighted combination of the original triplet loss function and the adversarial triplet loss function. The weight coefficients of the weighted combination are dynamically adjusted according to the degree of perturbation influence during the training process.

[0108] The robustness of the representation vector is evaluated based on the validation sample set. The stability of the representation results under different perturbation amplitudes is calculated. When the stability meets the preset threshold, the enterprise risk representation vector with anti-robustness is obtained.

[0109] For example, the gradient information of the representation vector is calculated by differentiating the triplet loss function. The sensitive direction of the representation vector refers to the gradient direction of the loss function with respect to the representation vector; a small change in this direction will lead to a significant change in the value of the loss function. During calculation, a batch of triplet samples is selected, and the gradient of the loss function with respect to the representation vector is calculated for each sample. The average value of the gradient vector is taken as the sensitive direction. The quantized representation of the sensitive direction is a vector with the same dimension as the representation vector, and the absolute value of the vector elements indicates the sensitivity of the corresponding dimension. For example, for a 128-dimensional enterprise risk representation vector, the values ​​of the 57th and 93rd dimensions in the calculated sensitive direction vector are 0.23 and -0.31, respectively, indicating that these two dimensions have a significant impact on changes in the loss function.

[0110] Initial adversarial disturbances are generated in sensitive directions, and the magnitude of the disturbance is determined based on the enterprise's historical risk fluctuation range. The historical risk fluctuation range is obtained by calculating the standard deviation of the enterprise's risk indicators over the past 12 months. Risk indicators refer to the risk status values ​​of enterprise nodes recorded in the enterprise risk status matrix. These values ​​are obtained by weighting the enterprise's own characteristic risk assessment value with the influence values ​​of related enterprises, ranging from 0 to 1, with higher values ​​indicating higher risk. Specifically, the calculation method involves extracting the monthly risk status values ​​for each enterprise over the past 12 months, calculating the standard deviation of these values, and obtaining the enterprise's historical risk fluctuation range. For example, if an enterprise's risk status values ​​for the past 12 months are [0.35, 0.38, 0.42, 0.45, 0.47, 0.50, 0.48, 0.46, 0.43, 0.41, 0.39, 0.37], the calculated standard deviation is 0.047, which represents the enterprise's historical risk fluctuation range. For each enterprise, a personalized perturbation magnitude is set, ranging from 30% to 50% of the historical risk standard deviation. This ensures the perturbation is small enough not to alter the semantic information of the representation, while effectively testing the model's robustness. Specifically, the direction of the perturbation vector is set as the sensitive direction, and the perturbation magnitude is set to 5% of the representation vector norm. For example, if an enterprise's historical risk standard deviation is 0.047, taking 40% of that is 0.0188, which is approximately 5% of the enterprise's representation vector norm, then its corresponding adversarial perturbation magnitude can be set to 0.05.

[0111] The dynamic weighted triplet loss function is designed based on the Euclidean distance between the enterprise risk representation vectors. The smaller the Euclidean distance between the anchor enterprise and the positive sample enterprises in the triplet, the more similar their risk profiles; conversely, the larger the Euclidean distance between the anchor enterprise and the negative sample enterprises, the more different their risk profiles. The Euclidean distance is calculated as the square root of the sum of the squares of the differences between corresponding elements of the two representation vectors. For example, the Euclidean distance between two 128-dimensional representation vectors is calculated as the square root of the sum of the squares of the differences between corresponding elements across all 128 dimensions, typically ranging from 0 to 2 for normalized representation vectors. Dynamic weights are calculated based on distance ratios: when the distance between the anchor and the positive sample is close to the distance between the anchor and the negative sample, the triplet is given a higher weight, as such samples are more crucial for improving the model's discriminative ability. The weight calculation formula is the inverse of the distance ratio, where the distance ratio is the distance between the anchor and the positive sample divided by the distance between the anchor and the negative sample. For example, if the distance between the anchor point and the positive sample is 0.2 and the distance between the anchor point and the negative sample is 0.8, then the weight of the triple is 1-(0.2 / 0.8)=0.75.

[0112] The initial adversarial perturbation is adjusted based on Euclidean distance to generate the final adversarial perturbation. The adjustment methods include direction adjustment and amplitude adjustment. Direction adjustment updates the sensitive direction by calculating the gradient of the representation vector under triplet loss, updating the direction once after each iteration. The amplitude coefficient for direction update is set to 0.8, indicating that the new direction consists of 80% of the current direction and 20% of the newly calculated gradient direction, ensuring the smoothness of the direction adjustment. Amplitude adjustment is performed based on the difficulty of the triplet samples, which is measured by the ratio of the distance between the anchor point and the positive sample to the distance between the anchor point and the negative sample. Specifically, the difficulty coefficient is defined as r = p / n, where p is the distance between the anchor point and the positive sample, and n is the distance between the anchor point and the negative sample. The perturbation amplitude adjustment factor is set according to the difficulty coefficient: when r>0.8, it is judged as a difficult triplet, the perturbation amplitude is increased, and the adjustment factor is 1+(r-0.8)×2.5; when r<0.4, it is judged as an easy triplet, the perturbation amplitude is decreased, and the adjustment factor is 0.7+(r-0.2)×0.75; when 0.4≤r≤0.8, it belongs to a medium difficulty triplet, the original perturbation amplitude is maintained, and the adjustment factor is 1.0. For example, in a triplet, the distance between the anchor point and the positive sample is 0.35, and the distance between the anchor point and the negative sample is 0.40. The difficulty coefficient is 0.35 / 0.40=0.875>0.8, which belongs to a difficult triplet. The adjustment factor is 1+(0.875-0.8)×2.5=1.1875. Therefore, the initial perturbation amplitude is increased by 18.75%, from 0.05 to 0.059. In another example, in a triplet, the distance between the anchor point and the positive sample is 0.15, and the distance between the anchor point and the negative sample is 0.70. The difficulty coefficient is 0.15 / 0.70=0.214<0.4, which belongs to an easily distinguishable triplet. The adjustment factor is 0.7+(0.214-0.2)×0.75=0.71. Therefore, the initial perturbation amplitude is reduced by 29%, from 0.05 to 0.036.

[0113] The adversarial training two-layer optimization framework comprises inner and outer optimization layers. The inner optimization maximizes the triplet loss under adversarial perturbation using gradient ascent, iteratively updating the direction and magnitude of the perturbation. The triplet loss under adversarial perturbation is the triplet loss calculated by adding the adversarial perturbation to the original representation vector. The inner optimization is set to 5 iterations, with each iteration moving 0.01 steps along the gradient direction, while ensuring the perturbation magnitude does not exceed a set upper limit, which is 10% of the representation vector norm. The outer optimization minimizes the weighted combination of the original triplet loss and the adversarial triplet loss using gradient descent, updating the model parameters. The model parameters include the weights and biases of the representation vector learning network, typically ranging from thousands to tens of thousands, depending on the network complexity. The outer optimization is set to 20 iterations with a learning rate of 0.005, using the Adam optimizer, where the hyperparameters β1 are set to 0.9, β2 to 0.999, and ε to 10. -8The weighted combination is initially set to 0.5:0.5, meaning the original loss and adversarial loss each account for 50%. This is dynamically adjusted during training: if the adversarial perturbation significantly impacts model performance (e.g., a performance drop exceeding 8% on the validation set), the weight of the adversarial loss is increased to 60%-70%; if the impact is minor (e.g., a performance drop less than 3% on the validation set), the weight of the adversarial loss is decreased to 30%-40%. For example, if both the adversarial and original losses have a weight of 0.5 at the beginning of training, after 10,000 iterations, if the adversarial perturbation causes a 10% performance drop on the validation set, the weights are adjusted to 0.3:0.7, meaning the original loss has a weight of 0.3 and the adversarial loss has a weight of 0.7.

[0114] Robustness evaluation of the representation vectors is performed on a validation set. The validation set consists of enterprise data not used in training, typically representing 20% ​​of the total dataset. Random perturbations of varying magnitudes are added to the enterprise representation vectors in the validation set, increasing from 1% to 10% of the representation vector norm, for a total of 10 perturbation levels. The random perturbation is achieved by adding a normally distributed N(0,σ) perturbation vector to each dimension of the representation vector. 2 Random noise is generated, where σ is the perturbation amplitude. For each perturbation level, the cosine similarity between the representation vectors before and after the perturbation is calculated. The cosine similarity is calculated as the inner product of the two vectors divided by the product of their respective norms, with a value between -1 and 1. A larger value indicates that the two vectors are more similar in direction. Simultaneously, the rate of change of the triplet loss function value calculated based on the perturbation representation vector is calculated. The rate of change is calculated as the difference between the perturbation loss value and the original loss value divided by the original loss value. The stability of the representation result is defined as the weighted average of the cosine similarity and the rate of change of loss under different perturbation levels, with weights of 0.6 and 0.4, respectively. The stability score is calculated as 0.6 × average cosine similarity + 0.4 × (1 - average rate of change of loss), with a value between 0 and 1, where a larger value indicates better stability. When the average cosine similarity is higher than 0.9 and the rate of change of loss is lower than 15%, i.e., the stability score is higher than 0.84, the representation vector is considered to have sufficient robustness. For example, for the representation vectors of 100 companies in the validation set, with a 5% perturbation, the average cosine similarity is 0.94, the loss change rate is 8%, and the stability score is 0.6×0.94+0.4×(1-0.08)=0.564+0.368=0.932, indicating that the model has good robustness.

[0115] In this application example, adversarial robustness enhancement was performed on the risk representation vectors of 200 companies. These companies come from various industries, including manufacturing, services, and finance, and each company has 30 original features. A 128-dimensional risk representation vector was obtained using the aforementioned method. The average risk state value of each company over the past 12 months ranges from 0.2 to 0.8, and the average historical risk standard deviation is 0.06. The initial representation vector is 128-dimensional. Sensitive directions are calculated using the above method, and initial adversarial perturbations are generated. The average Euclidean distance between the anchor company and positive sample companies in the triplet sample is calculated to be 0.25, and the average Euclidean distance with negative sample companies is calculated to be 0.73. Based on this, a dynamic weighted triplet loss function was designed. In the two-layer optimization framework, the inner layer optimization generates the optimal adversarial perturbation through 5 iterations, with an average perturbation amplitude of 6.5% of the representation vector norm. The outer layer optimization updates the model parameters through 20 iterations. The model consists of a three-layer fully connected neural network with approximately 50,000 parameters, and undergoes 15,000 training batches, each containing 64 triplet samples. During training, the weights of the weighted combination are adjusted from the initial 0.5:0.5 to 0.4:0.6. After training, the robustness of the representation vector is evaluated on the validation set. At a 5% perturbation amplitude, the average cosine similarity is 0.93, the loss change rate is 10.5%, and the stability score is 0.894, meeting the preset threshold requirements, thus yielding an adversarially robust enterprise risk representation vector.

[0116] This invention effectively enhances the robustness of enterprise risk representation vectors by introducing adversarial perturbations based on the historical risk fluctuation range of enterprises, combined with a dynamically weighted triplet loss function and a two-level optimization framework. Compared with traditional methods, this technique can better cope with data noise and distribution bias, enabling the risk representation to remain stable in the face of minor changes in enterprise financial data. It improves the model's generalization ability in different scenarios, provides a more reliable risk measurement basis for fraud prevention in financing guarantees, reduces the false judgment rate, and enhances the accuracy and reliability of financial risk control.

[0117] In one optional implementation, the step of inputting the enterprise risk representation vector into an ensemble learning framework to obtain an initial score for enterprise fraud risk includes:

[0118] A feature importance matrix is ​​constructed based on the enterprise risk representation vector of historical fraud samples. Principal component analysis is used to reduce the dimensionality of the feature importance matrix to obtain a feature combination that can identify fraud.

[0119] The enterprise risk representation vector is projected onto the subspace corresponding to the feature combination method to construct a gradient boosting-based decision tree ensemble structure, with each decision tree corresponding to a feature combination method; the prediction results of each decision tree are weighted and fused to obtain the initial score of enterprise fraud risk.

[0120] For example, labeled fraudulent enterprise samples and normal enterprise samples are extracted from a historical database. Risk representation vectors have been constructed for all these samples. Each enterprise's risk representation vector includes features across dimensions such as financial indicators, operating conditions, and legal risks, totaling up to 300 feature dimensions. For these features, feature importance is assessed by calculating the correlation between each feature and the fraud label. Specifically, for numerical features, a point-to-binary correlation coefficient with the fraud label is calculated; for categorical features, their information gain value is calculated. In this way, a feature importance matrix of size m×n is generated, where m represents the number of samples, n represents the number of features, and the elements in the matrix represent the contribution of that feature to fraud detection.

[0121] Suppose there are 1000 samples of companies known to be in a fraudulent state, each with 200 valid features. The feature importance matrix calculated using the method described above is a 1000×200 matrix. Values ​​in the matrix range from 0 to 1; a higher value indicates greater importance of the feature in identifying fraud. For example, the importance of "financial anomaly indicators" is 0.85, while the importance of "company age" is 0.32.

[0122] The feature importance matrix is ​​dimensionality reduced by applying principal component analysis (PCA) to standardize it. The covariance matrix between each feature is then calculated, followed by the eigenvalues ​​and eigenvectors. The eigenvalues ​​are sorted from largest to smallest, and the top k principal components with a cumulative contribution rate of 95% are selected as new feature combinations. These principal components reflect the most effective directions for fraud detection in the original feature space.

[0123] Suppose that principal component analysis yields 15 main feature combinations (principal components). These principal components capture the most critical directions of change for fraud detection in the original 200-dimensional feature space. Each principal component is a linear combination of the original features, and the weights reflect the importance of each original feature in that principal component. For example, the first principal component mainly consists of three original features: "abnormal cash flow," "abnormal debt-to-equity ratio," and "frequent related-party transactions," with weights of 0.6, 0.5, and 0.4, respectively.

[0124] In the decision tree ensemble construction phase, the risk representation vectors of the companies to be evaluated are projected onto the previously obtained k principal component spaces. For each principal component direction, a decision tree based on the gradient boosting algorithm is constructed. Each decision tree focuses on fraud detection based on features extracted from the corresponding principal component direction. During training, the fraud labels of historical samples are used as target variables, and the structure and parameters of the decision tree are continuously optimized through gradient boosting. The decision tree generation process considers parameters such as feature splitting gain, tree depth control, and the minimum number of samples per leaf node to prevent overfitting.

[0125] The configuration of each decision tree is as follows: maximum depth is set to 6, minimum number of leaf node samples is 20, learning rate is set to 0.1, and log loss function is used to evaluate split quality. For each principal component direction, the optimal tree parameter configuration is determined through cross-validation.

[0126] Finally, a weighted fusion is performed to obtain the initial score. Each decision tree corresponds to a principal component direction, assigning a fraud risk score (between 0 and 1) to the company being evaluated. The weight is determined based on the proportion of explained variance of each principal component (i.e., feature combination method); the principal component with greater explained variance has a higher weight. The sum of the weighted scores of all decision trees is the initial score for the company's fraud risk.

[0127] Suppose that the risk representation vector of the company A to be evaluated is projected into 15 principal component spaces after principal component analysis. The fraud scores given by the decision trees corresponding to these 15 directions are [0.75, 0.82, 0.68, 0.71, 0.65, 0.60, 0.58, 0.55, 0.52, 0.48, 0.45, 0.42, 0.40, 0.38, 0.35], and the proportion of variance explained by these 15 principal components is [0.20, 0.15, 0.12, 0.10, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.03, 0.02, 0.02, 0.02, 0.01]. Through weighted fusion calculation: 0.75×0.20+0.82×0.15+0.68×0.12+...+0.35×0.01, the initial fraud risk score for Company A is 0.65. This score indicates that Company A has a high fraud risk and requires further review.

[0128] To improve accuracy, sample resampling techniques can be used to address data imbalance during implementation, and cross-validation can be employed to fine-tune model parameters, ensuring the stability and reliability of the scoring results.

[0129] This method utilizes principal component analysis (PCA) to reduce the dimensionality of the feature importance matrix, extracting the most valuable feature combinations for fraud detection, effectively reducing feature redundancy and computational complexity. The gradient boosting-based decision tree ensemble structure enables specialized learning for different feature combinations, capturing the multidimensional features and nonlinear relationships of enterprise risk. By weighted fusion of the prediction results from each decision tree, this method fully leverages the advantages of ensemble learning, improving the accuracy and generalization ability of fraud risk assessment, providing more precise initial risk scores for anti-fraud decisions, and reducing business losses caused by misjudgments.

[0130] In one optional implementation, the step of calculating the final fraud risk score considering the risk propagation effect based on the risk propagation adjacency matrix and the initial score, and generating the financing guarantee anti-fraud decision result, includes:

[0131] Construct an enterprise risk propagation network, determine the risk propagation path between enterprises based on the risk propagation adjacency matrix, and calculate the risk propagation time window based on the degree of business correlation between enterprises and historical transaction data to obtain the risk propagation intensity under different time windows;

[0132] For each enterprise node, a first-order risk propagation impact value is calculated based on the initial scores of its neighboring enterprise nodes and the corresponding risk propagation intensity. A distance attenuation coefficient is introduced, which decreases as the propagation path increases, to calculate a second-order risk propagation impact value. The first-order and second-order risk propagation impact values ​​are weighted and combined to obtain the enterprise's cumulative risk impact coefficient.

[0133] The initial score is adjusted based on the cumulative risk impact coefficient to obtain a final fraud risk score that takes into account the risk propagation effect; based on the final fraud risk score and the company's guarantee limit, a decision result for financing guarantee anti-fraud is generated.

[0134] For example, the construction of the enterprise risk propagation network is based on a risk propagation adjacency matrix. This adjacency matrix is ​​an n×n matrix, where n is the number of enterprises. The element aij in the matrix represents the risk propagation intensity from enterprise i to enterprise j, with values ​​ranging from 0 to 1; a larger value indicates a stronger risk propagation intensity. The risk propagation path between enterprises is determined by the non-zero elements in the matrix. For example, when a12=0.4 and a23=0.5, it indicates that a risk propagation path exists from enterprise 1 to enterprise 2, and then to enterprise 3. The degree of business association between enterprises is calculated by analyzing historical transaction data, mainly considering three aspects: transaction amount ratio, transaction frequency, and transaction continuity. Transaction amount ratio refers to the proportion of transaction amount with a particular enterprise to one's total transaction amount; transaction frequency refers to the number of transactions per unit time; and transaction continuity refers to the time span of continuous transactions. The weighted combination of these three indicators is used to measure the degree of business association, with weights set to 0.5, 0.3, and 0.2, respectively. The risk propagation time window is determined by analyzing the propagation time of historical risk events, and is divided into short-term (1-3 months), medium-term (4-6 months), and long-term (7-12 months). For each time window, a decay coefficient for the risk propagation intensity is calculated: 1.0 for the short-term window, 0.8 for the medium-term window, and 0.6 for the long-term window. For example, if the business association degree between company A and company B is 0.7, and the corresponding element value in the risk propagation adjacency matrix is ​​0.6, then the risk propagation intensity is 0.6 × 1.0 = 0.6 in the short-term window, 0.6 × 0.8 = 0.48 in the medium-term window, and 0.6 × 0.6 = 0.36 in the long-term window.

[0135] The calculation of the first-order risk propagation impact value is based on the directly connected neighbor nodes of the enterprise. For enterprise i, its first-order risk propagation impact value is the sum of the products of the initial scores of all directly connected enterprises j and their corresponding risk propagation intensities. The initial score of enterprise j refers to the initial fraud risk score obtained through the aforementioned ensemble learning framework, with a value between 0 and 1. The corresponding risk propagation intensity is the element value aji in the adjacency matrix, representing the risk propagation intensity from enterprise j to enterprise i. For example, if enterprise A has three directly connected enterprises B, C, and D, with initial scores of 0.8, 0.6, and 0.4 respectively, and corresponding risk propagation intensities of 0.7, 0.5, and 0.3 respectively, then the first-order risk propagation impact value of enterprise A is 0.8×0.7+0.6×0.5+0.4×0.3=0.87. The second-order risk propagation impact value considers indirectly connected enterprise nodes, i.e., the case where enterprise i is connected to enterprise k through an intermediate enterprise j. A distance attenuation coefficient is introduced in the calculation, which decreases as the propagation path increases, using an exponential attenuation form, and the attenuation coefficient is set to 0.5. The second-order risk propagation impact value is the sum of the products of the initial scores of all second-order connected firms k multiplied by their corresponding second-order risk propagation strengths, where the second-order risk propagation strength is ajk × aji × 0.5. For example, firm A is indirectly connected to firms E and F through firm B, with initial scores of 0.9 and 0.7 for E and F respectively. The risk propagation strength from B to E is 0.8, and from B to A is 0.7. B is indirectly connected to firm G through C, with an initial score of 0.5 for G. The risk propagation strength from C to G is 0.6, and from C to A is 0.5. Therefore, the second-order risk propagation impact value for firm A is 0.9 × 0.8 × 0.7 × 0.5 + 0.7 × 0.8 × 0.7 × 0.5 + 0.5 × 0.6 × 0.5 × 0.5 = 0.347.

[0136] The cumulative risk impact coefficient of a company is obtained by weighting the first-order and second-order risk propagation impact values. The weight of the first-order impact is set to 0.7, and the weight of the second-order impact is set to 0.3, to reflect that the risk of directly related companies has a greater impact on the current company. Continuing the previous example, the cumulative risk impact coefficient of company A is 0.87 × 0.7 + 0.347 × 0.3 = 0.713. When the risk propagation network is more complex or when a specific industry needs to consider risk propagation over longer distances, the third-order or higher-order risk propagation impact values ​​can also be calculated, and the weights of each order of impact value can be adjusted appropriately. For example, in industries with highly concentrated supply chains, the weights can be set to 0.6 for the first order, 0.3 for the second order, and 0.1 for the third order.

[0137] The final fraud risk score is calculated based on an adjustment to the initial score and the cumulative risk impact coefficient. The adjustment method uses a weighted average, meaning the final score equals the product of the initial score and (1 + cumulative risk impact coefficient), then normalization is applied to ensure the final score is between 0 and 1. The normalization process uses a min-max scaling method: if the adjusted score exceeds 1, it is set to 1; if it is below 0, it is set to 0. For example, if Company A's initial score is 0.6 and its cumulative risk impact coefficient is 0.713, the adjusted score is 0.6 × (1 + 0.713) = 1.028, and after normalization, it is 1.0. To avoid excessive risk amplification, a risk impact cap can be set, such as limiting the maximum value of the cumulative risk impact coefficient to 1.5; even if the calculated value exceeds 1.5, it is still calculated as 1.5. Furthermore, the risk impact can be adjusted based on the company's own risk resistance capabilities, which can be comprehensively assessed through factors such as company size, asset status, and industry position.

[0138] The anti-fraud decision-making process for financing guarantees is based on the final fraud risk score and the company's guarantee limit. The fraud risk score is divided into four levels: low risk (0-0.3), low-to-medium risk (0.3-0.5), medium-to-high risk (0.5-0.7), and high risk (above 0.7). For different risk levels, corresponding guarantee limit adjustment strategies are set. For low-risk companies, 100% of the applied guarantee limit can be provided; for low-to-medium risk companies, 80%; for medium-to-high risk companies, 50% with additional collateral required; and for high-risk companies, the guarantee application is rejected. For example, if company A's final fraud risk score is 0.65, classifying it as medium-to-high risk, and its applied guarantee limit is 1 million yuan, then according to the decision-making strategy, a guarantee limit of 500,000 yuan can be provided, with additional collateral valued at no less than 300,000 yuan required. In practice, other factors can be considered to fine-tune the decision-making results, such as the company's historical credit record, industry prosperity, and macroeconomic environment. In addition, for companies with risk scores near the critical value, such as 0.49 or 0.51, fuzzy decision-making methods can be used to determine the final decision after comprehensively considering multiple factors.

[0139] This method constructs a corporate risk propagation network and calculates the risk propagation time window, enabling precise measurement of the intensity of risk propagation at different time scales. By adjusting the initial score through a cumulative risk impact coefficient, this method considers the amplifying effect of risk propagation on corporate fraud risk, resulting in a more comprehensive and accurate final fraud risk score. Combined with the corporate guarantee limit to generate decision-making results, it provides financial institutions with risk-sensitive guarantee strategies, effectively controlling business risks.

[0140] A second aspect of the present invention provides an artificial intelligence-based anti-fraud system for financing guarantees, comprising:

[0141] The first unit is used to acquire multi-source heterogeneous data of enterprises to construct enterprise feature vectors; construct a bidirectional causal graph structure based on the enterprise feature vectors, and use a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises; for each enterprise node, calculate the risk status of the node according to the risk status of its directly associated enterprises; use the NOTEARS optimization algorithm to calculate the risk propagation adjacency matrix between enterprises based on the risk status of the nodes.

[0142] The second unit is used to construct a metric space for enterprise risk distribution based on the enterprise feature vector and the risk propagation adjacency matrix using Wasserstein distance, perform dissimilar robustness optimization in the metric space, and train the enterprise risk representation vector using a triplet loss function; input the enterprise risk representation vector into the ensemble learning framework to obtain an initial score for enterprise fraud risk;

[0143] The third unit is used to calculate the final fraud risk score after considering the risk propagation effect based on the risk propagation adjacency matrix and the initial score, and generate the financing guarantee anti-fraud decision result.

[0144] A third aspect of the present invention provides an electronic device, comprising:

[0145] processor;

[0146] Memory used to store processor-executable instructions;

[0147] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0148] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0149] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A financing guarantee anti-fraud method based on artificial intelligence, characterized in that, include: Construct enterprise feature vectors by acquiring multi-source heterogeneous data from enterprises; A bidirectional causal graph structure is constructed based on the enterprise feature vectors. The bidirectional causal graph structure uses a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises. For each enterprise node, calculate the risk status of that node based on the risk status of its directly associated enterprises; The NOTEARS optimization algorithm is used to calculate the risk propagation adjacency matrix between enterprises based on the risk status of nodes. Specifically, this includes: calculating the risk correlation degree between enterprises based on their feature vectors, and constructing an enterprise risk correlation degree matrix; constructing a directed acyclic graph (DAG) structure based on the DAG matrix, where the DAG structure represents the bidirectional risk propagation relationship between enterprises; calculating the risk status of each enterprise node in the DAG structure, where the risk status of an enterprise node is obtained by a weighted combination of its own feature risk assessment value and the influence value of its associated enterprises, where the influence value of the associated enterprises is the weighted sum of the risk status of the enterprise node's directly associated enterprises and their corresponding risk propagation weights; forming an enterprise risk status matrix from the risk status of the enterprise nodes; constructing an optimization objective function based on the enterprise risk status matrix, where the optimization objective function includes a risk status fitting term, a risk propagation time-series dynamic term, and an acyclicity constraint term; and solving the optimization objective function using the NOTEARS optimization algorithm to obtain an inter-enterprise risk propagation adjacency matrix that satisfies the directed acyclicity constraint, where the risk propagation adjacency matrix represents the topological structure of risk propagation between enterprises. Based on the enterprise feature vector and the risk propagation adjacency matrix, a metric space for enterprise risk distribution is constructed using Wasserstein distance. Distributed robustness optimization is performed in this metric space, and a triplet loss function is used to train and obtain an enterprise risk representation vector. This enterprise risk representation vector is then input into an ensemble learning framework to obtain an initial score for enterprise fraud risk. Based on the risk propagation adjacency matrix and the initial score, the final fraud risk score considering the risk propagation effect is calculated, and the anti-fraud decision result for financing guarantee is generated.

2. The method according to claim 1, characterized in that, The steps for constructing and solving the objective function include: The risk status fitting term is calculated by combining the differences in risk status among enterprises with their importance weights; the risk propagation time series dynamic term is calculated by combining the time window decay weight with the historical risk propagation intensity; and the acyclicity constraint term is calculated by the cyclic connection of risk propagation relationships among enterprises. Construct an augmented Lagrange function that includes the optimization objective function, wherein the augmented Lagrange function introduces Lagrange multipliers and penalty terms; The augmented Lagrangian function is optimized by using the L-BFGS algorithm, which incorporates a second-order Hessian correction term. The second-order Hessian correction term is used to construct a diagonal block matrix based on the risk propagation time series characteristics, and the gradient update direction is adjusted through the diagonal block matrix. When the change value of the optimization objective is less than the set threshold and the acyclic constraint is satisfied, the optimized temporal risk propagation adjacency matrix is ​​obtained; the propagation intensity in the optimized temporal risk propagation adjacency matrix is ​​normalized, and the risk propagation relationship between enterprises is determined by temporal confidence score.

3. The method according to claim 1, characterized in that, Based on the enterprise feature vector and the risk propagation adjacency matrix, the steps of constructing a metric space for enterprise risk distribution using Wasserstein distance, performing disjoint robustness optimization in the metric space, and training the enterprise risk representation vector using a triplet loss function include: The enterprise feature vector is weighted by the risk propagation adjacency matrix, and the weighted enterprise feature vector is mapped to the probability density space using a Gaussian kernel function to obtain the enterprise risk probability distribution. Based on the enterprise risk probability distribution, a Wasserstein distance metric space is constructed, where the Wasserstein distance is obtained by solving the p-order norm of the optimal transmission plan between enterprise risk probability distributions; A triplet loss function is constructed based on the Wasserstein distance. The risk distribution pairs of enterprises are classified into categories according to a preset Wasserstein distance threshold, and a training sample set is constructed. By optimizing the triplet loss function, a representation vector reflecting the characteristics of enterprise risk distribution is learned. An adversarial perturbation is introduced into the representation vector, and a two-layer optimization objective is set. The adversarial perturbation is maximized in the inner layer, and the triple loss is minimized in the outer layer to enhance the robustness of the representation vector and obtain the final enterprise risk representation vector. Based on the final enterprise risk representation vector, the distance between enterprise pairs is calculated to obtain the consistency index and the discriminative index, and the enterprise risk representation results are evaluated.

4. The method according to claim 3, characterized in that, The steps to introduce adversarial perturbations into the representation vector, set a two-layer optimization objective (maximizing the adversarial perturbation in the inner layer and minimizing the triplet loss in the outer layer) to enhance the robustness of the representation vector and obtain the final enterprise risk representation vector include: The sensitive direction in the representation vector space is calculated based on the gradient information of the representation vector; an initial adversarial perturbation is generated in the sensitive direction, and the magnitude of the initial adversarial perturbation is determined based on the enterprise's historical risk fluctuation range; Design a triplet loss function with dynamic weights, wherein the dynamic weights are determined based on the Euclidean distance between pairs of enterprise risk representation vectors; adjust the direction and magnitude of the initial adversarial perturbation according to the Euclidean distance to generate the final adversarial perturbation; A two-layer optimization framework for adversarial training is constructed. The inner layer optimization finds the optimal perturbation direction by maximizing the triplet loss function under adversarial perturbation. The outer layer optimization improves the robustness of the representation vector by minimizing the weighted combination of the original triplet loss function and the adversarial triplet loss function. The weight coefficients of the weighted combination are dynamically adjusted according to the degree of perturbation influence during the training process. The robustness of the representation vector is evaluated based on the validation sample set. The stability of the representation results under different perturbation amplitudes is calculated. When the stability meets the preset threshold, the enterprise risk representation vector with anti-robustness is obtained.

5. The method according to claim 1, characterized in that, The steps for inputting the enterprise risk representation vector into the ensemble learning framework to obtain an initial score for enterprise fraud risk include: A feature importance matrix is ​​constructed based on the enterprise risk representation vector of historical fraud samples. Principal component analysis is used to reduce the dimensionality of the feature importance matrix to obtain a feature combination that can identify fraud. The enterprise risk representation vector is projected onto the subspace corresponding to the feature combination method to construct a gradient boosting-based decision tree ensemble structure, with each decision tree corresponding to a feature combination method; the prediction results of each decision tree are weighted and fused to obtain the initial score of enterprise fraud risk.

6. The method according to claim 1, characterized in that, The steps for calculating the final fraud risk score after considering the risk propagation effect, and generating the anti-fraud decision result for financing guarantees, based on the risk propagation adjacency matrix and the initial score, include: Construct an enterprise risk propagation network, determine the risk propagation path between enterprises based on the risk propagation adjacency matrix, and calculate the risk propagation time window based on the degree of business correlation between enterprises and historical transaction data to obtain the risk propagation intensity under different time windows; For each enterprise node, a first-order risk propagation impact value is calculated based on the initial scores of its neighboring enterprise nodes and the corresponding risk propagation intensity. A distance attenuation coefficient is introduced, which decreases as the propagation path increases, to calculate a second-order risk propagation impact value. The first-order and second-order risk propagation impact values ​​are weighted and combined to obtain the enterprise's cumulative risk impact coefficient. The initial score is adjusted based on the cumulative risk impact coefficient to obtain a final fraud risk score that takes into account the risk propagation effect; based on the final fraud risk score and the company's guarantee limit, a decision result for financing guarantee anti-fraud is generated.

7. An AI-based financing guarantee anti-fraud system, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to acquire multi-source heterogeneous data of an enterprise and construct an enterprise feature vector. A bidirectional causal graph structure is constructed based on the enterprise feature vectors. The bidirectional causal graph structure uses a directed acyclic graph to represent the bidirectional risk propagation relationship between enterprises. For each enterprise node, calculate the risk status of that node based on the risk status of its directly associated enterprises; The NOTEARS optimization algorithm is used to calculate the risk propagation adjacency matrix between enterprises based on the risk status of nodes; The second unit is used to construct a metric space for enterprise risk distribution based on the enterprise feature vector and the risk propagation adjacency matrix using Wasserstein distance, perform pluralistic optimization in the metric space, and train the enterprise risk representation vector using a triplet loss function. The enterprise risk representation vector is input into the ensemble learning framework to obtain an initial score for enterprise fraud risk; The third unit is used to calculate the final fraud risk score after considering the risk propagation effect based on the risk propagation adjacency matrix and the initial score, and generate the financing guarantee anti-fraud decision result.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Anti-fraud method and system based on complex relation network

    CN120579974A