A multi-party risk label real-time data processing method and system

By splitting the training process into basic model and incremental model in multi-party secure computing model training, and using secret sharing and homomorphic encryption technology for data encryption processing, the problems of low computing efficiency and high latency requirements in the existing technology are solved, and efficient real-time data processing and model fusion of risk tags are achieved.

CN119766418BActive Publication Date: 2025-05-27FUZHOU PUBLIC SECURITY BUREAU +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510253527.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-27
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The prior art has problems such as low computing efficiency, large communication overhead, poor scalability, and inability to meet the low latency requirements in the model training of multi-party security computing.

Method used

A multi-party risk tag real-time data processing method is proposed. By splitting the model training process into the basic model training stage and the incremental model training stage, the data is encrypted using secret sharing and homomorphic encryption technology, and the model fusion method is used to fuse the basic model with the incremental model.

Benefits of technology

It realizes joint modeling without leaking the original data, reduces the calculation time, meets the needs of low-latency inference in anti-fraud scenarios, and improves the risk prediction ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119766418B_ABST
    Figure CN119766418B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for real-time data processing of multi-party risk tags. The method includes the following steps: The data party obtains local user risk tag data, generates an aligned data set through tag alignment processing for secret sharing or homomorphic encryption processing; The task participant obtains the encrypted data, performs joint training based on the encrypted data to generate a basic model, and deploys the trained basic model to the business system; Continuously collect new user risk tag data through the business system, perform tag alignment processing on the new data, and then perform encrypted state processing through secret sharing or homomorphic encryption; Perform joint training based on the encrypted state incremental data to generate an incremental model; Combine the basic model and the incremental model for model fusion, generate a fusion model to replace the basic model and deploy it to the business system. The present invention can improve the computational efficiency of real-time data processing of multi-party risk tags and meet the requirements of the business system for low latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to a method and system for real-time processing of multi-party risk label data. Background Art

[0002] Telecom fraud is a new type of crime that seriously endangers the property safety of citizens. With the continuous progress of communication technology, telecom fraud cases are currently emerging in an endless stream worldwide, with various means and are extremely difficult to guard against. To accurately identify the risks of telecom fraud, public security organs have jointly built an anti-fraud model with financial institutions and telecom operators (also known as Internet service providers, Internet Service Provider, ISP), effectively blocking the occurrence of fraud.

[0003] The Chinese patent application with the publication number CN115271930A discloses a method and device for identifying abnormal users. The method includes the following steps: obtaining the first data of the user on the bank side and the second data of the user on the operator side, where the first data includes the first identification information and the first feature, and the second data includes the second identification information and the second feature; based on the first identification information and the second identification information, matching and aligning the first data and the second data; obtaining the first parameters of the first neural network model on the bank side and the second parameters of the second neural network model on the operator side; using the first parameters and the second parameters to generate a neural network joint model; using the neural network joint model to identify whether the user logging in on the bank side is an abnormal user.

[0004] Due to the extremely high privacy and sensitivity of the data of public security organs, financial institutions, and telecom operators, the data that can be taken out for joint model construction is very limited. The above solution has insufficient data security and privacy protection. This method may involve the direct exposure of user identity information and has a relatively high risk of privacy leakage.

[0005] Multi-Party Computation (MPC) is a cryptographic technology that aims to allow multiple participating parties to jointly calculate a function and obtain the result without revealing their respective private data. Its core goal is to ensure data privacy and computational correctness, even if some participating parties behave improperly or do not trust other parties. Therefore, combining privacy computing technology to jointly train an anti-fraud model with multi-party data without leaving the domain is a very worthy exploration direction.

[0006] In view of this, a Chinese invention patent application with the publication number CN117195060A discloses a method for identifying telecommunications fraud and a method for training a model based on multi-party secure computing. The method includes the following steps: obtaining telecommunications data to be identified; using a secret sharing method with a second participating party and a third participating party, and using a target recognition model to process the telecommunications data to be identified, the bank statement data to be identified in the second participating party, and the network browsing behavior data to be identified in the third participating party to obtain the first fragment of the secret sharing of the target prediction value; receiving the second fragment and the third fragment of the secret sharing of the target prediction value; determining the telecommunications fraud recognition result according to the first, second, and third fragments; the target recognition model is obtained by combining multiple single-tree models; the single-tree model is constructed based on the fragments of the first-order derivative and the second-order derivative of the target loss function obtained by secret sharing; the target loss function includes a penalty parameter to suppress the imbalance in the number of positive and negative samples.

[0007] In the anti-fraud scenario, it is of crucial significance to promptly detect possible fraud behaviors and quickly implement blocking. The blocking success rate of fraud behaviors is highly correlated with the response time. The faster the response speed, the higher the probability of successful blocking. Therefore, to meet this requirement, the technical solution needs to have the ability of real-time dynamic learning and reasoning, and at the same time, it can continuously optimize the model to cope with complex and changeable fraud means. The above multi-party secure computing technology has significant problems in model training, including low computing efficiency, large communication overhead, poor scalability, and inability to meet the low-latency requirements. Especially in the scenario of large-scale data and high-dimensional features, frequent cryptographic operations and intermediate result synchronization greatly increase the computing and communication costs, and at the same time, its optimization for specific tasks is still relatively lacking. In addition, solving the trade-off problem between efficiency and security remains a key difficulty, resulting in it being difficult to meet industrial-level performance requirements in practical applications. These problems have posed huge obstacles to the large-scale implementation of multi-party secure computing in model training, and it is urgent to make breakthroughs in algorithm optimization, framework design, and system implementation. Summary of the Invention

[0008] The present invention provides a method and system for real-time data processing of multi-party risk labels, aiming to solve the problems of low computing efficiency, large communication overhead, poor scalability, and inability to meet the low-latency requirements in the model training of multi-party secure computing in the prior art.

[0009] To solve the above technical problems, the method for real-time data processing of multi-party risk labels proposed by the present invention includes the following steps:

[0010] Multiple data providers obtain local user risk label data, generate an aligned data set through label alignment processing, and perform secret sharing or homomorphic encryption processing on the aligned data set;

[0011] Task participants obtain encrypted data, conduct joint training based on the encrypted data to generate a basic model, and deploy the trained basic model to the business system;

[0012] The business system continuously collects new user risk label data, performs label alignment processing on the new data, and then performs encrypted state processing through secret sharing or homomorphic encryption;

[0013] Conduct joint training based on the encrypted incremental data to generate an incremental model;

[0014] Combine the basic model and the incremental model for model fusion, generate a fusion model to replace the basic model and deploy it to the business system.

[0015] Preferably, the label alignment processing is specifically as follows: Each participant broadcasts its local label categories to other participants for label alignment, and compares the received label categories with the locally stored label categories. For the missing label categories in the local storage, preset data is used for filling.

[0016] Preferably, the secret sharing is specifically as follows:

[0017] There are N data providers, and each provider holds a local user label data set;

[0018] For each user label data , construct a random polynomial of degree less than the set threshold T:

[0019]

[0020] In the formula, , represents the number of the th label of user in participant , are randomly selected coefficients, T is the minimum number of shares required to reconstruct the data, and it satisfies ;

[0021] Calculate the values of at N different points to obtain secret shares:

[0022]

[0023] Each participant is assigned a secret share .

[0024] Preferably, the homomorphic encryption is specifically as follows:

[0025] Select a security parameter, and a trusted party generates a public-private key pair , the public key For encrypting data, the private key For decrypting data;

[0026] Each participant Performs homomorphic encryption on the local user label data to generate encrypted label data.

[0027] Preferably, the homomorphic encryption adopts one of the Paillier algorithm, the ElGamal algorithm, and the Boneh - Goh - Nissim algorithm.

[0028] Preferably, the method of model fusion is specifically as follows:

[0029] Process the incremental data to make its feature distribution consistent with the input of the base model, and generate the prediction results of the base model and the incremental model;

[0030] Use the base model as the teacher model and the incremental model as the student model to train a fusion model;

[0031] Calculate the performance of the validation sets of the base model and the incremental model, assign weights to the base model and the incremental model according to the validation set performance, and use the weighted average method to fuse the outputs of the base model and the incremental model.

[0032] Preferably, when training a fusion model, the loss function is:

[0033]

[0034] In the formula, is the temperature constant, used to control the softening degree, represents the Softmax function, is the prediction result of the base model, is the prediction result of the incremental model, represents and the relative entropy of.

[0035] Preferably, in the weighted average method of model fusion, the weights of the base model and the incremental model are dynamically adjusted by Bayesian optimization.

[0036] Preferably, the joint training adopts a federated learning architecture, and each participant uses local data to train a local model and updates the global model through a secure aggregation mechanism.

[0037] Correspondingly, the present invention also proposes a multi - party risk label real - time data processing system, which is used to implement the above - mentioned real - time data processing method, and includes:

[0038] A data acquisition module, which is used to acquire local user risk label data, continuously collect new user risk label data, and perform label alignment processing to generate an aligned data set;

[0039] A data encryption module, which is used to perform secret sharing or homomorphic encryption on the aligned data set;

[0040] A joint training module, which is used to obtain encrypted data and perform joint training based on the encrypted data to generate a basic model;

[0041] A model deployment module, which is used to deploy the trained basic model to a business system and apply it in the business system;

[0042] An incremental training module, which is used to perform joint training based on encrypted incremental data to generate an incremental model;

[0043] A model fusion module, which is used to combine the basic model and the incremental model for model fusion to generate a fusion model, and dynamically adjust the weights according to the performance of the validation sets of the basic model and the incremental model;

[0044] A model update module, which is used to replace the basic model with the fusion model and deploy it to the business system.

[0045] Compared with the prior art, the present invention has the following technical effects:

[0046] 1. The multi-party risk label real-time data processing method proposed by the present invention splits the existing model training process into two stages, namely the basic model training stage and the incremental model training stage. The stage with the largest data volume and the longest training time is placed in the basic model stage, and this stage is not sensitive to computing time consumption. In the incremental learning stage, less incremental data is used to train the incremental model. Since the data volume is very small in this process, the communication volume in the multi-party secure computing protocol interaction is greatly reduced, the model training time is accelerated, and the computing time consumption can be greatly reduced, so as to meet the low-latency inference requirements of the business system in the anti-fraud scenario.

[0047] 2. The multi-party risk label real-time data processing method proposed by the present invention fuses the basic model and the incremental model through the model fusion method, effectively utilizes historical knowledge, and absorbs the latest information at the same time, avoiding the problem that the model degrades rapidly due to changes in data distribution.

[0048] 3. The multi-party risk label real-time data processing method proposed by the present invention uses secret sharing and homomorphic encryption technologies to encrypt multi-party user risk label data, and realizes joint modeling without leaking the original data, avoiding the privacy leakage risk caused by direct data sharing.

[0049] 4. The multi-party risk label real-time data processing method proposed by the present invention supports data collaboration among multiple institutions such as banks and operators through multi-party secure computing. On the premise of ensuring that data does not leave the local area and is not directly shared, it realizes cross-institutional joint modeling, improves the accuracy of risk identification, is applicable to multiple fields such as financial risk control, telecom fraud detection, and network security monitoring, meets the requirements of data security and compliance, and at the same time enhances the risk prediction ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a schematic flowchart of the real-time data processing method of the present invention;

[0051] Figure 2 is a schematic diagram of label alignment in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present application and with reference to the accompanying drawings.

[0053] Embodiment 1

[0054] This embodiment is a multi-party risk label real-time data processing method, as Figure 1 shown, including the following steps 1 to 5:

[0055] Step 1, multiple data providers obtain local user risk label data, generate an aligned data set through label alignment processing, and perform secret sharing or homomorphic encryption processing on the aligned data set.

[0056] The label alignment processing is specifically as follows: each participant broadcasts its local label categories to other participants for label alignment, and compares the received label categories with the locally stored label categories. For the label categories missing in the local storage, preset data is used for filling.

[0057] In this embodiment, the preset data for filling is the constant 0, that is, the constant 0 is used to fill the labels that the local side does not have to achieve label alignment.

[0058] An example of label alignment in this embodiment is as Figure 2As shown in the figure, assume there are three participating parties: P1, P2, and P3, each of which holds local user risk label data: P1 has label A and label C, P2 has label A and label B, and P3 has label A and label B, where labels A, B, and C are labels generated by each data party for the corresponding users according to their respective businesses. Each participating party broadcasts the local label categories to other participating parties. As shown in the figure, P1 sends A and C to P2 and P3, P2 sends A and B to P1 and P3, and P3 sends A and B to P1 and P2. After each party receives the labels of other participating parties, it fills in the missing labels with 0, that is, the user data without this label, and sets the default value to 0. After alignment, the labels of each data party all include A, B, and C, where the label B of P1 = 0, the label C of P2 = 0, and the label C of P3 = 0.

[0059] The specific secret sharing is as follows:

[0060] There are N data providers, and each provider holds a local user label data set;

[0061] For each user label data , construct a random polynomial of degree less than the set threshold T:

[0062]

[0063] In the formula, , represents the number of the th label of user in the participating party , are randomly selected coefficients, T is the minimum number of shares required to reconstruct the data, satisfying ;

[0064] Calculate the values of at N different points to obtain the secret shares:

[0065]

[0066] Each participating party is assigned a secret share .

[0067] The specific homomorphic encryption is as follows:

[0068] Select a security parameter, and a trusted party generates a public-private key pair , the public key is used to encrypt data, and the private key is used to decrypt data;

[0069] Each participating party performs homomorphic encryption on the local user label data to generate encrypted label data.

[0070] The homomorphic encryption adopts one of the Paillier algorithm, the ElGamal algorithm, and the Boneh-Goh-Nissim algorithm.

[0071] After label alignment and secret sharing / homomorphic encryption, the data of all participating parties have been standardized into the same label structure and encrypted to ensure data privacy. The next step is to perform joint training based on this encrypted data to generate a basic model. By using secret sharing and homomorphic encryption technologies to encrypt the risk label data of multiple-party users, joint modeling is achieved without revealing the original data, avoiding the risk of privacy leakage caused by direct data sharing.

[0072] Step 2: The task participating parties obtain the encrypted data, perform joint training based on the encrypted data to generate a basic model, and deploy the trained basic model to the business system.

[0073] In the previous step, all data providers (participating parties) completed label alignment and protected the data using secret sharing or homomorphic encryption. The goal of this step is to construct a basic model through joint training and deploy it to the business system without revealing the original data.

[0074] As Figure 2 shown in the label alignment task, an example participating party in this embodiment can be the original data provider (P1, P2, P3), who can directly use their own encrypted data for training; or a third-party computing service provider such as a trusted computing party, a cloud server, etc., who does not own the data but is responsible for secure computing.

[0075] If secret sharing is adopted, the share data provided by each party is jointly calculated by one or more computing parties; if homomorphic encryption is adopted, the computing party can directly calculate on the encrypted data without decryption.

[0076] For secret sharing, an embodiment of the present invention adopts a secure computing scheme based on secret sharing. Each participating party holds the secret share of the data but cannot recover the original data alone. During training, each calculation step is jointly calculated by multiple participating parties, and finally the calculation results are merged.

[0077] The merged calculation results can adopt secret sharing addition:

[0078]

[0079] In the formula, represents the number of the th label of user in participating party , , , represent the encrypted shares provided by the 1st, 2nd, …, Nth participants, that is, the encrypted parts of the 1st, 2nd, …, Nth participants for . After the calculation is completed, the gradients are merged through secret reconstruction, and the model parameters are updated.

[0080] For homomorphic encryption, an embodiment of the present invention uses additive homomorphic to accumulate gradients. After the calculation is completed, only the authorized party can use the private key to decrypt the gradients, and the model parameters are updated through a gradient descent algorithm (such as SGD).

[0081] In some other embodiments of the present invention, the joint training may also adopt a federated learning architecture. Each participant uses local data to train a local model and updates the global model through a secure aggregation mechanism. Specifically, each participant trains a sub-model locally, then encrypts the gradient parameters and uploads them to the aggregation server; the server performs encrypted aggregation, and the calculated model parameters are then distributed back to each participant.

[0082] So far, the model has been trained and can be deployed locally for inference. However, there is a problem that due to the bottleneck in communication efficiency of the MPC protocol itself, the training efficiency for large-scale datasets is very poor, which is obviously not suitable for the low-latency requirements of the anti-fraud scenario. Therefore, the following steps three to five are continued.

[0083] Step three, continuously collect new user risk label data through the business system, perform label alignment processing on the new data, and then perform encrypted processing through secret sharing or homomorphic encryption.

[0084] As the business runs, new user data is continuously generated. For example, various parties such as financial institutions, telecommunications operators, and public security organs may regularly or real-time generate new user risk label data. At this time, the business system needs to continuously collect this new data from each participant. Each participant regularly processes the newly added user data and sends it to the central system or computing platform for processing; the new data includes user risk label data, and each label corresponds to specific risk information.

[0085] Since different data providers may adopt different label sets, label alignment processing is crucial, which ensures that all participants use the same label set. The label alignment method is the same as in step one.

[0086] For the aligned data, perform encryption processing to ensure that the data is not leaked during the calculation process. Specifically, the same as in step one, the method of secret sharing or homomorphic encryption can be used, and the specific secret sharing or homomorphic encryption also selects the same method as in step one of this embodiment.

[0087] Step four, perform joint training based on the encrypted incremental data to generate an incremental model.

[0088] In this step, the main objective of the joint training is to generate an incremental model by processing encrypted incremental data. This incremental model will supplement and optimize the base model based on the latest data (such as recently added risk label data). Through the joint training method, all participating parties can jointly participate in the training using the incremental data to generate a more accurate incremental model.

[0089] Through joint training and secure aggregation, the incremental data will be used to update the global model to generate an incremental model. The incremental model is a further optimization of the base model on the incremental data. After the joint training is completed, the performance of the incremental model will be evaluated, usually using a validation set to evaluate the accuracy and performance of the incremental model. The performance of the incremental model is compared with that of the base model to ensure that the incremental training brings meaningful optimization.

[0090] Step Five, combine the base model and the incremental model for model fusion to generate a fused model to replace the base model and deploy it to the business system. The goal of this step is to fuse the base model and the incremental model to generate a new fused model and deploy it to the business system to replace the original base model. Through model fusion, the advantages of the two models are utilized to improve the overall prediction effect, enabling the system to better adapt to new data characteristics.

[0091] The specific method of the model fusion is as follows:

[0092] Process the incremental data to make its feature distribution consistent with the input of the base model, and generate the prediction results of the base model and the incremental model; Before performing model fusion, it is first necessary to ensure that the feature distribution of the incremental data is consistent with the input of the base model. This can be achieved by performing feature alignment, standardization, or normalization on the incremental data to ensure that the base model and the incremental model can output consistent results when processing the same type of input.

[0093] Use the base model as the teacher model and the incremental model as the student model to train a fused model.

[0094] Calculate the performance of the validation sets of the base model and the incremental model, assign weights to the base model and the incremental model according to the validation set performance, and use the weighted average method to fuse the outputs of the base model and the incremental model.

[0095] Specifically, during the model fusion process, when training a fused model, the loss function is:

[0096]

[0097] In the formula, is the temperature constant, which is used to control the softening degree, denotes the Softmax function, is the prediction result of the base model, is the prediction result of the incremental model, denotes and the relative entropy of

[0098] Using the prediction results of the base model and the incremental model, optimize through the above loss function to generate a fusion model. The fusion model will gradually improve its prediction ability by learning the outputs of the two models and ultimately combine the knowledge of the base model and the incremental model.

[0099] In the weighted average method of the model fusion described above, the weights of the base model and the incremental model are dynamically adjusted through Bayesian optimization.

[0100] Through this fusion process, the final fusion model can integrate the advantages of the base model and the incremental model, continuously optimize in a dynamically changing environment, and exhibit stronger prediction ability and adaptability in actual business.

[0101] In some other embodiments of the present invention, the fusion method of the base model and the incremental model is expressed as:

[0102]

[0103] In the formula, denotes the newly fused fusion model, denotes the base model, denotes the incremental model. It means that based on the prediction result of the existing base model, the incremental model is optimized through the learning brought by new data, and the prediction results of the two are accumulated and fused.

[0104] Or it means that the base model makes predictions on the encrypted data of all participating parties and obtains a sum of prediction values, and the incremental model makes predictions on the encrypted data of all participating parties and obtains a sum of prediction values. The final fusion result is the sum of the two, and the calculation method is expressed as follows:

[0105]

[0106] Similarly, in the above formula, denotes the newly fused fusion model, N denotes the number of data providers, denotes the encrypted label data of the j-th data provider regarding user u, that is, the risk label of user u at data provider j, denotes the prediction output of the base model on the encrypted label data, denotes the prediction output of the incremental model on the encrypted label data.

[0107] For the and , which is generally represented as follows:

[0108]

[0109] The j-th shard data representing the sum of the label vectors of user u, represents the j-th shard of the label vector of user u in participant i, and N is the number of data providers. It represents that each participant performs an encryption operation on the aligned data set through a secret sharing protocol. The encryption operation is represented as follows:

[0110]

[0111] L is the plaintext vector data. After sharding, each participant holds the shard data , and the sum of the shard data of all parties is L.

[0112] Example 2

[0113] This example is a multi-party risk label real-time data processing system. The system is used to implement the real-time data processing method as described in Example 1, including:

[0114] A data acquisition module, used to acquire local user risk label data, continuously collect new user risk label data, and perform label alignment processing to generate an aligned data set;

[0115] A data encryption module, used to perform secret sharing or homomorphic encryption on the aligned data set;

[0116] A joint training module, used to obtain encrypted data and perform joint training based on the encrypted data to generate a basic model;

[0117] A model deployment module, used to deploy the trained basic model to the business system and apply it in the business system;

[0118] An incremental training module, used to perform joint training based on the encrypted incremental data to generate an incremental model;

[0119] A model fusion module, used to combine the basic model and the incremental model for model fusion to generate a fusion model, and dynamically adjust the weights according to the performance of the validation sets of the basic model and the incremental model;

[0120] A model update module, used to replace the basic model with the fusion model and deploy it to the business system.

[0121] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A multi-party risk label real-time data processing method, characterized in that: The following steps are involved: Multiple data providers obtain local user risk label data, and generate aligned data sets through label alignment, and perform secret sharing or homomorphic encryption on the aligned data sets; The task participants obtain the encrypted data, perform joint training based on the encrypted data, generate a basic model, and deploy the trained basic model to the business system; Continuously collect new user risk label data through the business system, perform label alignment on the new data, and then perform confidential processing through secret sharing or homomorphic encryption; Joint training is performed based on the secret incremental data to generate an incremental model; the joint training adopts a federated learning architecture, where each participant uses local data to train a local model and updates the global model through a secure aggregation mechanism; The basic model and the incremental model are combined to perform model fusion, and a fusion model is generated to replace the basic model and deployed to the business system. The model fusion method is specifically as follows: Process the incremental data to make its feature distribution consistent with the basic model input, and generate prediction results for the basic model and incremental model; The basic model is used as the teacher model, the incremental model is used as the student model, and a fusion model is trained. The loss function for: In the formula, is the temperature constant used to control the degree of softening, represents the Softmax function, The prediction results of the basic model are is the prediction result of the incremental model, express and The relative entropy of Calculate the validation set performance of the basic model and the incremental model, assign weights to the basic model and the incremental model according to the validation set performance, and use the weighted average method to fuse the outputs of the basic model and the incremental model; in the weighted average method, the weights of the basic model and the incremental model are dynamically adjusted through Bayesian optimization.

2. A multi-party risk label real-time data processing method according to claim 1, characterized in that: The label alignment process is specifically as follows: each participant broadcasts its own local label categories to other participants for label alignment, and compares the label categories received by the party with the locally stored label categories, and fills in the missing label categories in the local storage with preset data.

3. The method for real-time data processing of multi-party risk labels according to claim 1 is characterized in that: The secret sharing is specifically: There are N data providers, each of which holds a local user label dataset; For each user label data , construct a random polynomial whose order is less than the set threshold T: In the formula, , indicating that the user On the Participants The The number of tags, is a randomly selected coefficient, T is the minimum number of shares required to reconstruct the data, satisfying ; calculate The value at N different points gives the secret share: Each participant Assign a secret share .

4. The method for real-time data processing of multi-party risk labels according to claim 1 is characterized in that: The homomorphic encryption is specifically: Select a security parameter and have a trusted party generate a public-private key pair , public key Used to encrypt data, private key Used to decrypt data; Each participant Perform homomorphic encryption on local user label data to generate encrypted label data.

5. A multi-party risk label real-time data processing method according to claim 4, characterized in that: The homomorphic encryption adopts one of the Paillier algorithm, the ElGamal algorithm and the Boneh-Goh-Nissim algorithm.

6. A multi-party risk label real-time data processing system, characterized in that: The system is used to implement the real-time data processing method according to any one of claims 1 to 5, comprising: The data acquisition module is used to obtain local user risk label data, continuously collect new user risk label data, and perform label alignment processing to generate an aligned data set; Data encryption module, used to perform secret sharing or homomorphic encryption on the aligned data set; A joint training module, used to obtain encrypted data, and perform joint training based on the encrypted data to generate a basic model; The model deployment module is used to deploy the trained basic model to the business system and apply it in the business system; The incremental training module is used to perform joint training based on dense incremental data to generate incremental models; The model fusion module is used to combine the basic model and the incremental model to generate a fusion model and dynamically adjust the weights according to the validation set performance of the basic model and the incremental model. The model update module is used to replace the basic model with the fusion model and deploy it to the business system.

Citation Information

Patent Citations

  • Abnormal user identification method and device

    CN115271930A

  • Telecommunication fraud identification method and model training method based on multi-party security computing

    CN117195060A

  • Multi-party joint risk recognition method and device

    CN111046425A

  • Enterprise electricity fee payment risk prediction method and system based on longitudinal federated logistic regression

    CN115392531A