Multi-mechanism federal anti-fraud modeling method and system based on risk cost sensitivity and credibility evaluation, and computer readable storage medium
By deploying federated learning clients among banks and using risk cost matrix loss function and risk interval division to assess node credibility, multi-institutional federated anti-fraud modeling was able to efficiently, accurately, and compliantly identify user fraud risks, thereby improving the model's identification accuracy and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-14
AI Technical Summary
Existing anti-fraud models in banks rely on transaction data and customer behavior data for modeling, resulting in low accuracy and violating privacy regulations, thus failing to effectively identify user fraud risks.
A multi-institutional federated anti-fraud modeling approach based on risk cost sensitivity and credibility assessment is adopted. By deploying a federated learning client, introducing a risk cost matrix loss function, dividing risk intervals and calculating credibility weights, adaptive weighted federated aggregation is performed to generate predicted risk probabilities that meet the risk management needs of banks.
Without sharing the original data, it improves the identification accuracy and stability of the multi-agency joint anti-fraud model, meets data privacy protection and regulatory compliance requirements, and performs particularly well in the identification of high-risk transactions.
Smart Images

Figure CN121860747A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a multi-agency federal anti-fraud modeling method, system, and computer-readable storage medium based on risk cost sensitivity and credibility assessment. Background Technology
[0002] Currently, anti-fraud work is an important part of the daily work of banks. However, banks currently rely on anti-fraud models to determine whether users are at risk of being defrauded. The construction of anti-fraud models mainly depends on transaction data, customer behavior data, and historical risk tags for modeling. Customer behavior data and user transaction data involve user privacy, and aggregating user information violates privacy regulatory requirements. Therefore, anti-fraud models can all rely on data within the bank's own nodes for training, resulting in low accuracy. Summary of the Invention
[0003] The present invention aims to at least partially solve one of the technical problems in the related art.
[0004] Therefore, the first objective of this invention is to propose a multi-agency federal anti-fraud modeling method based on risk cost sensitivity and credibility assessment, comprising: S1, deploy the federated learning client to each participating institution, access the local anti-fraud dataset, and process the dataset; S2, the model is trained locally based on the loss function structure of the risk cost matrix. By introducing the risk cost matrix, the fraud identification error is differentially penalized, and the predicted risk probability that meets the bank's risk management needs is generated. S3, divide the risk probability output by the model into multiple risk intervals, calculate the identification performance index of each institutional node in different risk intervals and construct a performance vector, and determine the node credibility weight based on the performance vector and the preset risk interval weight coefficient. S4: Under the protection of the encrypted parameter transmission mechanism, the central coordination server obtains the sample quantity and credibility weight of each node, performs adaptive weighted federated aggregation based on the sample quantity and credibility weight of each node, updates the global model parameters and distributes them to each institutional node to improve the effectiveness of the joint model in identifying high-risk transactions.
[0005] In one embodiment of the present invention, the loss function based on the risk cost matrix in step S2 is: = ; Where N is the number of local training samples. For binary labels representing the authenticity of a sample, when A value of 1 indicates that the sample is a fraudulent sample. A value of 1 indicates that the sample is a normal sample. This represents the probability of fraud occurring as predicted by the model. These represent the risk cost coefficients for missed fraud detection and, respectively. This represents the risk cost coefficient for misjudging a normal transaction, and > .
[0006] In one embodiment of the present invention, S3 further includes: S31, divide the high-risk interval, medium-risk interval and low-risk interval based on the first probability threshold and the second probability threshold, wherein the first probability threshold is greater than the second probability threshold; S32, calculate the identification performance index in different risk areas respectively, and calculate the performance vector of the corresponding risk area based on the identification performance index.
[0007] In one embodiment of the present invention, the formula for updating the global model parameters in step S4 is: , ; in, Let be the number of local samples at the k-th node. Let k be the credibility weight of the k-th node. Let be the weight coefficient of the j-th node.
[0008] In one embodiment of the present invention, it further includes: Repeat steps S2-S4 until the global model performance meets the preset convergence condition.
[0009] In one embodiment of the present invention, the dataset in step S1 includes transaction records, login behavior, device fingerprints, and risk tags.
[0010] In one embodiment of the present invention, the formula for calculating the performance vector is: ; in, This represents the recognition performance of node k within the risk interval r. The risk interval weighting coefficient, and satisfies This reflects the banking industry's focus on high-risk transactions.
[0011] To achieve the above objectives, a second aspect of the present invention provides a multi-agency federal anti-fraud modeling apparatus based on risk cost sensitivity and credibility assessment, comprising: The Federated Learning Client Deployment Module is used to deploy Federated Learning Clients in participating institutions, access local anti-fraud datasets, and store and process data only locally, avoiding the transfer of raw customer data across institutions. The local model training module is used to train the model locally based on the loss function structure of the risk cost matrix. By introducing the risk cost matrix, it differentiates the penalty for fraud identification error and generates predicted risk probabilities that meet the risk management needs of banks. The risk interval division and credibility assessment module is used to divide the risk probability output by the model into multiple risk intervals, calculate the identification performance index of each institutional node in different risk intervals and construct a performance vector, and determine the node credibility weight based on the performance vector and the preset risk interval weight coefficient. The adaptive federated aggregation module is used by the central coordination server to obtain the sample quantity and credibility weight of each node under the protection of the encrypted parameter transmission mechanism. Based on the sample quantity and credibility weight of each node, it performs adaptive weighted federated aggregation, updates the global model parameters, and distributes them to each institutional node to improve the effectiveness of the joint model in identifying high-risk transactions.
[0012] In one embodiment of the present invention, the risk interval division and credibility assessment module is further used for: The high-risk, medium-risk and low-risk intervals are divided based on a first probability threshold and a second probability threshold, wherein the first probability threshold is greater than the second probability threshold. Calculate the identification performance index for different risk areas, and calculate the performance vector for the corresponding risk area based on the identification performance index.
[0013] To achieve the above objectives, a third aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0014] The methods, systems, and storage media of this invention effectively improve the identification accuracy and stability of multi-agency joint anti-fraud models in heterogeneous data environments, while meeting data privacy protection and regulatory compliance requirements.
[0015] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0016] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a multi-agency federal anti-fraud modeling method based on risk cost sensitivity and credibility assessment according to an embodiment of the present invention; Figure 2 This is a structural diagram of a multi-agency federal anti-fraud modeling system based on risk cost sensitivity and credibility assessment according to an embodiment of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0019] The following describes, with reference to the accompanying drawings, a multi-agency federal anti-fraud modeling method based on risk cost sensitivity and credibility assessment proposed according to an embodiment of the present invention.
[0020] Example 1 Figure 1 This is a flowchart of a multi-agency federal anti-fraud modeling method based on risk cost sensitivity and credibility assessment, according to an embodiment of the present invention.
[0021] like Figure 1 As shown, the multi-agency federal anti-fraud modeling approach based on risk cost sensitivity and credibility assessment includes the following steps: S1 deploys the federated learning client to each participating institution, accesses the local anti-fraud dataset, and processes the dataset.
[0022] Specifically, in some implementations, step 1, "multi-institution node access," is a fundamental component of the entire federated learning anti-fraud modeling system. Its implementation involves client deployment, data access, and localized processing mechanisms. Each participating institution (such as banks, payment platforms, and credit reporting companies) needs to deploy a federated learning client locally. This client is typically based on a distributed computing framework (such as TensorFlow Federated, PySyft, or a self-developed federated engine), possesses the ability to communicate with a central coordination server, and supports local model training and encrypted parameter uploading.
[0023] Specifically, after deployment, the federated learning client first accesses the local anti-fraud dataset through standardized interfaces (such as REST API or gRPC). The dataset typically includes structured transaction data (such as transaction time, amount, and merchant type), user behavior data (such as login frequency and geographic location changes), device fingerprint information (such as IMEI and MAC address), and historical fraud tags. To ensure data privacy, the client only stores and processes data locally during training, avoiding cross-institutional transmission of raw customer data, thus meeting the requirements for localized data processing under the Personal Information Protection Law and the Data Security Law.
[0024] In practical applications, this step is typically deployed in each organization's private computing environment, such as a local server cluster or a containerized deployment platform. Through this step, organizations can access their own data resources without exposing customer privacy, providing foundational support for subsequent collaborative modeling. Its technical value lies in building a secure, compliant, and scalable federated learning collaboration framework, providing a standardized path for data access and processing for cross-organizational anti-fraud modeling.
[0025] S2, based on the loss function structure of the risk cost matrix, trains the model locally. By introducing the risk cost matrix, it differentiates the penalty for fraud identification error and generates a predicted risk probability that meets the bank's risk management needs.
[0026] The loss function based on the risk cost matrix is: = ; Where N is the number of local training samples. For binary labels representing the authenticity of a sample, when A value of 1 indicates that the sample is a fraudulent sample. A value of 1 indicates that the sample is a normal sample. This represents the probability of fraud occurring as predicted by the model. These represent the risk cost coefficients for missed fraud detection and, respectively. This represents the risk cost coefficient for misjudging a normal transaction, and > .
[0027] Specifically, in the local model training step, each participating institution independently trains its model on its local dataset based on a unified model structure (such as deep neural networks, gradient boosting decision trees, etc.). The core objective is to output the predicted risk probability that a transaction sample belongs to the fraud category. The key to this step lies in introducing a risk-cost-sensitive loss function to address the issues of scarce fraud samples and asymmetric risk costs in bank anti-fraud scenarios.
[0028] At the implementation level, this loss function is embedded into the local model's training process, typically used in conjunction with gradient descent optimizers such as Adam or SGD. The model adjusts its parameters based on this loss function in each iteration, thereby enhancing its ability to identify high-risk samples.
[0029] This step plays a fundamental role in the federated learning framework, and its output is the risk probability. This provides crucial input for subsequent risk range segmentation and node credibility assessment. By introducing a risk cost matrix into local training, the model's sensitivity to fraudulent activities is improved, as is the business interpretability and risk controllability of the model's output. This lays a solid foundation for building a joint anti-fraud model that meets the risk management needs of banks.
[0030] S3 divides the risk probability output by the model into multiple risk intervals, calculates the identification performance index of each institutional node in different risk intervals and constructs a performance vector, and determines the node credibility weight based on the performance vector and the preset risk interval weight coefficient.
[0031] Further, step S3 includes: S31, divide the high-risk interval, medium-risk interval and low-risk interval based on the first probability threshold and the second probability threshold, wherein the first probability threshold is greater than the second probability threshold; S32, calculate the identification performance index in different risk areas respectively, and calculate the performance vector of the corresponding risk area based on the identification performance index.
[0032] Specifically, in step 3, this invention proposes a credibility weighting mechanism based on risk interval division and node identification performance evaluation, used to differentiate the weighting of the contributions of each agency node during the federated aggregation stage. The core of this step lies in calculating the risk probability output by the model. According to the preset risk threshold and Divided into three risk zones: low risk zone Medium-risk area and high-risk areas Among them, the second probability threshold With the first probability threshold It can be configured according to business needs. For example, in actual bank anti-fraud scenarios, it can be... , This is to highlight the ability to identify high-risk transactions.
[0033] Within each risk interval, each institutional node independently calculates its model's recognition performance metrics, such as AUC (Area Under Curve), KS (Kolmogorov-Smirnov) statistic, or fraud recall. These metrics reflect the model's ability to discriminate at different risk levels. Particularly in high-risk intervals, the accuracy of fraud sample identification has a decisive impact on the overall risk control effect. This is based on the performance vector and preset risk interval weight coefficients. Calculate the credibility weight of the node Its formula is: ; in, This represents the recognition performance of node k within the risk interval r. The risk interval weighting coefficient, and satisfies This reflects the banking industry's focus on high-risk transactions.
[0034] In practical applications, this step is suitable for anti-fraud systems that utilize multi-agency collaborative modeling, especially in scenarios with heterogeneous data distribution and inconsistent risk preferences. It effectively quantifies the differences in identification capabilities of each node across different risk ranges, thereby improving the robustness and business adaptability of the overall model. Its technical value lies in introducing a dynamic evaluation mechanism for risk range performance, enabling adaptive adjustment of node weights in federated learning, and enhancing the model's focus and overall performance in high-risk fraud identification.
[0035] S4: Under the protection of the encrypted parameter transmission mechanism, the central coordination server obtains the sample quantity and credibility weight of each node, performs adaptive weighted federated aggregation based on the sample quantity and credibility weight of each node, updates the global model parameters and distributes them to each institutional node to improve the effectiveness of the joint model in identifying high-risk transactions.
[0036] Specifically, in some implementations, the central coordination server performs adaptive weighted federated aggregation under the protection of an encrypted parameter transmission mechanism. This technical implementation is based on parameter aggregation strategies and a risk-sensitive modeling framework within federated learning. The core of this step lies in constructing dynamic trust weights by introducing the identification performance of nodes within different risk ranges. and combined with local sample size The encrypted model parameters uploaded by each institutional node are weighted and fused to generate more representative and robust global model parameters.
[0037] Specifically, the central coordination server first receives the local model parameters transmitted by each node through homomorphic encryption or secure multi-party computation (SMPC) mechanisms. ,in This indicates the current communication round. In decryption or secure computing environments, the server determines the round based on the number of samples from each node. With credibility weight Weighted aggregation is performed according to the following formula: , ; in, Let be the number of local samples at the k-th node. Let k be the credibility weight of the k-th node. Let be the weight coefficient of the j-th node.
[0038] This formula ensures that nodes with a large sample size and stronger identification capabilities in high-risk regions have a greater influence on global model updates.
[0039] This step has significant value in real-world anti-fraud scenarios. For example, in a multi-bank joint risk control system, the performance of the models of different banks varies across different risk ranges due to differences in customer groups and transaction characteristics. Through adaptive weighted aggregation, the system can dynamically adjust the contribution ratio of each node, making the global model more closely match the actual business needs of high-risk fraud identification, thereby improving the model's generalization ability and risk identification accuracy. Furthermore, this mechanism enhances the robustness of the federated learning system, preventing low-quality nodes from interfering with the global model.
[0040] The multi-agency federated anti-fraud modeling method based on risk cost sensitivity and credibility assessment in this invention improves the stability and practicality of the multi-agency joint anti-fraud model in high-risk transaction identification by introducing a risk cost sensitive loss function and node credibility weights based on risk interval identification performance, without sharing the original data.
[0041] Example 2 Figure 2 This is a structural diagram of a multi-agency federal anti-fraud modeling system based on risk cost sensitivity and credibility assessment, according to an embodiment of the present invention.
[0042] like Figure 2 As shown, a multi-agency federal anti-fraud modeling system based on risk cost sensitivity and credibility assessment includes: The Federated Learning Client Deployment Module is used to deploy Federated Learning Clients in participating institutions, access local anti-fraud datasets, and store and process data only locally, avoiding the transfer of raw customer data across institutions. The local model training module is used to train the model locally based on the loss function structure of the risk cost matrix. By introducing the risk cost matrix, it differentiates the penalty for fraud identification error and generates predicted risk probabilities that meet the risk management needs of banks. The risk interval division and credibility assessment module is used to divide the risk probability output by the model into multiple risk intervals, calculate the identification performance index of each institutional node in different risk intervals and construct a performance vector, and determine the node credibility weight based on the performance vector and the preset risk interval weight coefficient. The adaptive federated aggregation module is used by the central coordination server to obtain the sample quantity and credibility weight of each node under the protection of the encrypted parameter transmission mechanism. Based on the sample quantity and credibility weight of each node, it performs adaptive weighted federated aggregation, updates the global model parameters, and distributes them to each institutional node to improve the effectiveness of the joint model in identifying high-risk transactions.
[0043] Furthermore, the risk range segmentation and credibility assessment module is also used for: The high-risk, medium-risk and low-risk intervals are divided based on a first probability threshold and a second probability threshold, wherein the first probability threshold is greater than the second probability threshold. Calculate the identification performance index for different risk areas, and calculate the performance vector for the corresponding risk area based on the identification performance index.
[0044] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described multi-agency federal anti-fraud modeling method based on risk cost sensitivity and credibility assessment.
[0045] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0046] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A multi-agency federal anti-fraud modeling method based on risk cost sensitivity and credibility assessment, characterized in that, include: S1, deploy the federated learning client to each participating institution, access the local anti-fraud dataset, and process the dataset; S2, the model is trained locally based on the loss function structure of the risk cost matrix. By introducing the loss function of the risk cost matrix, a predicted risk probability that meets the risk management needs of banks is generated. S3, divide the risk probability output by the model into multiple risk intervals, calculate the identification performance index of each institutional node in different risk intervals and construct a performance vector, and determine the node credibility weight based on the performance vector and the preset risk interval weight coefficient. S4: Under the protection of the encrypted parameter transmission mechanism, the central coordination server obtains the sample quantity and credibility weight of each node, performs adaptive weighted federated aggregation based on the sample quantity and credibility weight of each node, updates the global model parameters and distributes them to each institutional node to improve the effectiveness of the joint model in identifying high-risk transactions.
2. The method as described in claim 1, characterized in that, The loss function based on the risk cost matrix mentioned in step S2 is: = ; Where N is the number of local training samples. For binary labels representing the authenticity of a sample, when A value of 1 indicates that the sample is a fraudulent sample. A value of 1 indicates that the sample is a normal sample. The probability of fraud occurring is predicted by the model. These represent the risk cost coefficients for missed fraud detection and, respectively. This represents the risk cost coefficient for misjudging a normal transaction, and > .
3. The method as described in claim 1, characterized in that, S3 further includes: S31, divide the high-risk interval, medium-risk interval and low-risk interval based on the first probability threshold and the second probability threshold, wherein the first probability threshold is greater than the second probability threshold; S32, calculate the identification performance index in different risk areas respectively, and calculate the performance vector of the corresponding risk area based on the identification performance index.
4. The method as described in claim 1, characterized in that, The formula for updating the global model parameters in step S4 is: , ; in, Let be the number of local samples at the k-th node. Let k be the credibility weight of the k-th node. Let be the weight coefficient of the j-th node.
5. The method as described in claim 1, characterized in that, Also includes: Repeat steps S2-S4 until the global model performance meets the preset convergence condition.
6. The method as described in claim 1, characterized in that, The dataset mentioned in step S1 includes transaction records, login behavior, device fingerprints, and risk tags.
7. The method as described in claim 3, characterized in that, The formula for calculating the performance vector is: ; in, This represents the recognition performance of node k within the risk interval r. The risk interval weighting coefficient, and satisfies This reflects the banking industry's focus on high-risk transactions.
8. A multi-agency federal anti-fraud modeling system based on risk cost sensitivity and credibility assessment, characterized in that, include: The Federated Learning Client Deployment Module is used to deploy Federated Learning Clients in participating institutions, access local anti-fraud datasets, and store and process data only locally, avoiding the transfer of raw customer data across institutions. The local model training module is used to train the model locally based on the loss function structure of the risk cost matrix. By introducing the risk cost matrix, it differentiates the penalty for fraud identification error and generates predicted risk probabilities that meet the risk management needs of banks. The risk interval division and credibility assessment module is used to divide the risk probability output by the model into multiple risk intervals, calculate the identification performance index of each institutional node in different risk intervals and construct a performance vector, and determine the node credibility weight based on the performance vector and the preset risk interval weight coefficient. The adaptive federated aggregation module is used by the central coordination server to obtain the sample quantity and credibility weight of each node under the protection of the encrypted parameter transmission mechanism. Based on the sample quantity and credibility weight of each node, it performs adaptive weighted federated aggregation, updates the global model parameters, and distributes them to each institutional node to improve the effectiveness of the joint model in identifying high-risk transactions.
9. The system as described in claim 8, characterized in that, The risk interval division and credibility assessment module is also used for: The high-risk, medium-risk and low-risk intervals are divided based on a first probability threshold and a second probability threshold, wherein the first probability threshold is greater than the second probability threshold. Calculate the identification performance index for different risk areas, and calculate the performance vector for the corresponding risk area based on the identification performance index.
10. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method as claimed in any one of claims 1-7.