Data subset-based vertical federated learning anomaly data debugging method and system

By designing a subset evaluation technology for privacy protection in vertical federated federal learning and multi-party security computing technology, automating the location and correction of abnormal data, the problem of data errors in vertical federated learning is solved, ensuring data privacy and model performance.

WO2025102899A1PCT designated stage expired Publication Date: 2025-05-22HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
PCT/CN2024/115413
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2024-08-29
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

In vertical federated learning scenarios, data errors or missing data lead to poor performance, but existing debugging technologies are difficult to automatically locate abnormal data, and privacy protection is difficult to ensure.

Method used

A vertical federated learning abnormal data debugging method based on data subsets was designed. Through the federal data subset evaluation technology protected by privacy, abnormal data is automatically located to ensure that data privacy is not leaked. This method includes using mask vectors for privacy protection, multi-party security computing technology for data subset evaluation and traceability, and machine learning models for problem data subset screening.

Benefits of technology

Without leaking participant data and sensitive information, automate the location and correction of abnormal data, improve the performance of federated learning models, and reduce the risk of manual participation and data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115413_22052025_PF_FP_ABST
    Figure CN2024115413_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a data subset-based vertical federated learning anomaly data debugging method and system. The method comprises: on the basis of vertical federated learning, an initiator performing modeling and training a federated model (S1); using the trained federated model to acquire problematic data subsets in a data set, the prediction accuracy of the problematic data subsets in the federated model being lower than the prediction accuracy of other data subsets in the federated model (S2); screening the problematic data subsets on the basis of a feature description combination, so as to acquire problematic data subsets having anomaly descriptions (S3); and the initiator or a participant performing data provenance and correction on the basis of the problematic data subsets having the anomaly descriptions, and re-training the federated model after correction (S4). As a federated data subset evaluation technology for privacy protection, the present invention correctly calculates federated data subset evaluation indexes while ensuring data privacy, so as to form the data subset-based federated learning debugging method, thus automatically positioning anomaly data, and solving the problem of abnormal performance of federated learning models.
Need to check novelty before this filing date? Find Prior Art

Description

Abnormal data debugging method and system for vertical federated learning based on data subsets Technical Field

[0001] The present invention belongs to the field of computer technology, and in particular relates to a method and system for debugging abnormal data in vertical federated learning based on data subsets. Background Art

[0002] Federated learning, which involves multiple data holders, is currently gaining increasing application in areas such as financial risk control and smart healthcare. However, debugging technology for federated learning models remains a niche area. The main reasons for this are as follows: 1) Privacy protection: Since the data required to train a federated learning model comes from two or more participants, the initiator of model training has difficulty accessing the data and feature information of other data providers for privacy reasons, making it difficult to identify data problems and debug. 2) Debugging technology: Existing federated debugging technology mainly targets centralized training scenarios, where the initiator of model training can access all data and perform debugging based on the data. The execution process of such debugging technology is difficult to adapt to the data distribution in federated learning. 3) Participant cooperation: When a federated learning model encounters a problem, the initiator and the data provider are usually required to manually identify the problem. This process is not only time-consuming and labor-intensive, but excessive manual participation may also further lead to privacy leaks.

[0003] Vertical federated learning targets vertical data distribution, also known as sample-aligned data distribution. In this scenario, data is distributed across multiple parties' databases. Each party has highly overlapping data IDs, but little or no overlap in the data's features. When performing data analysis, multiple data parties must first find the intersection of their respective data IDs and extract the data with the same ID for subsequent data analysis and query tasks. Vertical data distribution is primarily used when there is significant user overlap but little overlap in feature dimensions across the datasets. Multi-party vertical data distribution is common in cross-industry scenarios. For example, in financial risk management, banks, as data holders, hold pre-loan information on some customers. Banks can use this information to train pre-loan risk control models. In a typical application scenario of federated modeling, banks incorporate carrier data as supplementary data for pre-loan risk control. In this case, only data from customers that are both bank and carrier customers can participate in the federated modeling. In this scenario, the data held by the bank and carrier are multi-party vertically distributed. Participating in federated modeling, their data IDs are located in both the bank and carrier data. This type of data distribution qualifies as multi-party vertical data distribution.

[0004] In a vertical federated learning process, multiple participants use their own data to train a federated learning model. When the data held by all parties is normal and error-free, the performance of the federated learning model typically meets business requirements. However, in many cases, errors or omissions in the data held by one participant can lead to poor model performance and fail to meet business requirements. For example, in a vertical federated learning scenario between an insurance company and a medical institution, due to negligence, the medical institution mistakenly set the test results for a certain medical indicator for patients aged 30-40 as positive when operating the database. In the federated model, such a data error could affect the performance of the model. However, it is difficult for insurance company staff to identify the issue without accessing the medical institution's data. Existing solutions cannot guarantee that the data held by each participant will not be transferred locally during successful debugging, nor can they guarantee that debugging can be completed without manual data review. Both of these approaches pose a serious risk of data leakage.

[0005] Federated learning models are trained by two or more participants, and the data involved in model training must be prepared before model training. For example, the participant with labels is called the initiator, and is typically the party with actual business needs. The participant without labels is called the collaborator, and typically possesses a large amount of data features and hopes to profit by providing data services. However, in business scenarios, due to various privacy and legal reasons, the initiator only has the right to query, inspect, modify, and add to its own data; it cannot perform such operations on the data of collaborators across other participants. Therefore, in this scenario, neither personnel involved in federated learning can see data from non-participants, nor can any data participating in federated learning be sent from the local server to other participants.

[0006] In federated learning scenarios, data anomalies can cause a sharp increase in the federated model's prediction error rate for some test data, leading to model anomalies. However, existing research has difficulty locating data anomalies in federated learning or directly applying them to federated learning scenarios.

[0007] Summary of the Invention

[0008] In response to the above problems, the present invention provides a method and system for debugging abnormal data in vertical federated learning based on data subsets. By designing a privacy-preserving federated data subset evaluation technology, the relevant evaluation indicators of the federated data subset are correctly calculated while ensuring data privacy, forming a federated learning debugging framework based on federated data subsets, automatically locating abnormal data, and solving the problem of abnormal performance of the federated learning model.

[0009] According to a first aspect of an embodiment of the present disclosure, a method for debugging abnormal data in longitudinal federated learning based on a data subset is provided, the method comprising:

[0010] The initiator models and trains the federated model based on vertical federated learning;

[0011] Using the trained federated model to obtain a problem data subset in a data set, wherein the prediction accuracy of the problem data subset in the federated model is lower than the prediction accuracy of other data subsets in the federated model;

[0012] The problem data subset is screened based on a combination of feature descriptions to obtain a problem data subset with abnormal descriptions; the initiator performs data tracing and correction based on the problem data subset with abnormal descriptions, and retrains the federated model after correction.

[0013] In one embodiment, before screening the problem data subset based on the feature description combination, the discrete features are first classified according to the categories of discrete data, and the continuous features are divided into data intervals.

[0014] In one embodiment, the feature description combination-based screening uses a multi-party secure computing method to protect the data ID, and forms a data anomaly description by merging the data ID into intervals.

[0015] In one embodiment, the vertical federated learning abnormal data debugging method further includes a federated data subset privacy protection method based on a mask vector, specifically including:

[0016] Perform ID alignment on data IDs before modeling based on longitudinal federated learning;

[0017] The mask vector is an array of full intersection data. The ID set corresponding to a feature in each data subset corresponds to a separate mask vector. The separate mask vector only exists in the data set to which the corresponding feature belongs. When a data subset contains multiple features, the mask vectors held by each feature owner are required and multiplied using a multi-party secure computing method to determine the true mask vector of the current data subset.

[0018] An evaluation metric for the current data subset is calculated based on the true mask vector.

[0019] In one embodiment, the screening based on the combination of feature descriptions also includes training a machine learning model for screening problem data subsets. During the training process of the machine learning model, multiple public data sets are used to randomly generate rules to select data subsets for label destruction. The data subsets with real problems are marked as positive sample labels for model training, and the rest are negative sample labels. Feature data of different data subsets are collected through data subset discovery technology to form a data set for the machine learning model. After training with the data set, the machine learning model has the ability to finely distinguish problem data subsets. The problem data subset is input into the machine learning model after training to obtain a problem data subset with abnormal descriptions.

[0020] In one embodiment, the multi-party secure computing method is implemented using secret sharing technology as the underlying technology. The secret sharing technology includes: randomly splitting a number into two or more numbers, and the split numbers belong to different computing parties. Each computing party then performs arithmetic calculations under privacy protection based on the assigned data.

[0021] According to a second aspect of an embodiment of the present disclosure, a vertical federated learning abnormal data debugging system based on a data subset is provided, the system comprising:

[0022] The federated model training unit is used by the initiator to model and train the federated model based on vertical federated learning;

[0023] a problem data subset acquisition unit, configured to acquire a problem data subset from a data set using the trained federated model, wherein the prediction accuracy of the problem data subset in the federated model is lower than the prediction accuracy of other data subsets in the federated model;

[0024] a problem data subset screening unit, configured to screen the problem data subset based on a combination of feature descriptions to obtain a problem data subset with an abnormal description;

[0025] The problem data subset correction unit is used for the initiator to trace and correct the data based on the problem data subset with the abnormal description, and retrain the federated model after the correction.

[0026] In one embodiment, the screening based on the feature description combination in the problem data subset screening unit also includes training a machine learning model for problem data subset screening. During the training process of the machine learning model, multiple public data sets are used to randomly generate rules to select data subsets for label destruction. The data subsets with real problems are marked as positive sample labels for model training, and the rest are negative sample labels. Feature data of different data subsets are collected through data subset discovery technology to form a data set for the machine learning model. After training using the data set, the machine learning model has the ability to finely distinguish problem data subsets. The problem data subset is input into the machine learning model after training to obtain a problem data subset with abnormal descriptions.

[0027] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned method for debugging abnormal data in longitudinal federated learning based on a data subset is implemented.

[0028] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the above-mentioned vertical federated learning abnormal data debugging method based on data subsets is implemented.

[0029] The disclosed embodiments provide a method and system for debugging abnormal data in vertical federated learning based on data subsets, which is oriented to the application scenario of automated federated debugging with privacy protection. By designing a privacy-protected federated data subset evaluation technology, the relevant evaluation indicators of the federated data subset are correctly calculated while ensuring data privacy, forming a federated learning debugging framework based on the federated data subset, automatically locating abnormal data, and thus helping personnel solve the problem of abnormal performance of the federated learning model. It does not require the initiator to access the key privacy information such as data and features of other participants during the debugging process, nor does it require the participants of the federated learning (or federated learning applications) to leak or send the data, features and other information held by the participants to each other. Its beneficial effects include:

[0030] When a federated model performs poorly due to data errors, the federated model debugging process can automatically find the erroneous data without the data being exported locally. After deleting the erroneous data, the trained federated model can achieve normal performance.

[0031] Using privacy-preserving problem data subset search technology, without human intervention and without the data held by each participant being released locally, secure multi-party computing technology is used to identify the problem subset in the global dataset. For this problem subset, the federated model's prediction accuracy on this data subset is significantly lower than the model's prediction accuracy on other datasets other than the problem subset. The problem data subset found by this technology can be described by human-understandable interpretability conditions. A screening method for problem data subsets is introduced. After the machine learning model has screened the problem data subset, it uses error data tracing technology to locate the data in the dataset that caused the federated model to have label errors. Finally, the federated model is retrained by removing the traced error data to complete debugging.

[0032] This machine learning-based problem subset filtering technology screens the many problem data subsets identified, eliminating those that don't meet the actual requirements and retaining those that actually negatively impact the performance of the federated model. The machine learning model used in this technology extracts features and labels from the size and effect size of multiple existing problem data subsets. This data is then used as a training dataset to train the machine learning model, enabling it to finely distinguish problem data subsets.

[0033] This technology uses erroneous data tracing technology. This technology uses a filtered subset of problematic data and performs interval merging operations to form a subset of data describing the problem. Leveraging the underlying technology of multi-party secure computation, this technology ensures that all data from participants remains local throughout the entire execution process, eliminating the need for manual data review and ensuring the security of all participants' data and sensitive information.

[0034] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the present invention.

[0036] FIG1 is a flow chart of a method for debugging abnormal data in vertical federated learning based on a data subset according to an embodiment of the present invention;

[0037] FIG2 is a logic block diagram of a method for debugging abnormal data in vertical federated learning based on a data subset in an embodiment of the present invention;

[0038] FIG3 is a schematic diagram of a data subset definition according to an embodiment of the present invention;

[0039] FIG4 is a schematic diagram of a data subset screening process according to an embodiment of the present invention;

[0040] FIG5 is a flow chart of a method for protecting privacy of a federated data subset based on a mask vector according to an embodiment of the present invention;

[0041] FIG6 is a schematic diagram of error rate calculation of a federated data subset according to an embodiment of the present invention;

[0042] FIG7 is a flow chart of a method for screening a subset of questions based on machine learning according to an embodiment of the present invention;

[0043] FIG8 is a schematic diagram of the structure of a vertical federated learning abnormal data debugging system based on data subsets in an embodiment of the present invention;

[0044] FIG9 is a schematic diagram of an electronic device according to an embodiment of the present invention.

[0045] Implementation Method

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0047] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0048] The present invention provides a method and system for debugging abnormal data in vertical federated learning based on data subsets, and provides the following embodiments:

[0049] Example 1 is used to illustrate a method for debugging abnormal data in vertical federated learning based on a data subset. Referring to FIG1 , which is a flow chart of a method for debugging abnormal data in vertical federated learning based on a data subset, the specific steps include:

[0050] S1. The initiator models and trains the federated model based on vertical federated learning;

[0051] S2. Using the trained federated model to obtain a problem data subset in the data set, wherein the prediction accuracy of the problem data subset in the federated model is lower than the prediction accuracy of other data subsets in the federated model;

[0052] S3. Screening the problem data subset based on a combination of feature descriptions to obtain a problem data subset with abnormal descriptions;

[0053] S4. The initiator or participant performs data tracing and correction based on the problem data subset with the abnormal description, and retrains the federated model after correction.

[0054] During the specific implementation process, see Figure 2, which is a schematic diagram of the overall framework of the debugging method of the embodiment. On the basis of not affecting the training of the existing federated learning model, the existing training model is used to perform data debugging through test data. The specific steps of the debugging process are as follows:

[0055] Step 1: Participant A (the initiator of the federated learning process) and Participant B (the data provider of the federated learning process) use their respective datasets to train the federated model. This training process includes steps such as data preprocessing, federated feature engineering, and federated learning training. In the federated learning debugging framework based on the federated data subset, Step 1 is the same as the original federated learning training process, and there is no need to modify the original framework or code.

[0056] Step 2: Participant A and Participant B put the model into actual business use, using test data as input to obtain the prediction results of the federated learning model. The prediction results of the federated learning model are then verified in actual business use. If the prediction results of the federated learning model are normal, no subsequent federated debugging framework intervention is required. If the prediction results of the federated learning model are abnormal, the federated debugging phase begins.

[0057] Step 3: Party A initiates the federation debugging phase. Party B follows the steps of the federation debugging process. The two parties do not exchange raw data during the process. After the federation debugging phase is complete, Party A can obtain the output of the federation debugging phase.

[0058] Step 4: Participant A communicates with Participant B to resolve data issues based on the output of federated debugging. Due to the complexity of data anomalies, correcting data anomalies still requires the data holder to intervene and adjust with the help of abnormal data information, without the need for other data holders (participants) to intervene or observe data adjustments, thereby maintaining the ability to protect privacy. After resolving the abnormal data problem, the federated learning model is retrained.

[0059] The privacy-preserving federated learning automated debugging method supports multiple participants in the operation, with one participant acting as the initiator of the federated learning process and the others acting as collaborators. The specific technical process of privacy-preserving federated learning automated debugging is as follows: 1) Before debugging, the federated model is first trained. This model training needs to be performed by the initiator and is modeled using longitudinal federated learning technology. The embodiment uses logistic regression as an example. 2) After the federated model is trained, the problem data subsets in the entire dataset are calculated. For these problem data subsets, the prediction accuracy of the federated model on this data subset is much lower than the prediction accuracy of the model on other datasets except the problem subset. The trained federated model is not as good as the global effect on this problem data subset. 3) Due to technical characteristics, the problem subsets found in 2) are not necessarily caused by real problem data. Therefore, a problem subset filtering technology based on machine learning is further used. This technology screens the many problematic data subsets that have been found, eliminating those that don't meet the actual requirements and retaining those that actually have a negative impact on the performance of the federated model. It also generates anomaly descriptions for these data subsets. 4) After completing the tracing of the erroneous data, Participant A needs to communicate with Participant B using the output of the federated debugging to resolve the data issue. Due to the complexity of data anomalies, correcting the data anomaly still requires the data holder to intervene and make adjustments using the problematic data information. Other data holders (participants) do not need to intervene or observe data adjustments, thereby maintaining privacy protection capabilities. After resolving the anomaly data issue, the federated learning model is retrained. This concludes the entire privacy-preserving federated learning automated debugging process.

[0060] It should be noted that the definition of the federated data subset. As shown in Figure 3, the privacy-preserving federated data subset is a federated data subset across data participants. The federated data subset needs to be composed of descriptions of the characteristics distributed among different participants, for example: gender is male, age is between 35 and 40 years old, where the gender feature is located in participant A and the age feature is located in participant B. The federated data subset determines the sample entries that specifically belong to the federated data subset by the ID of the data sample, that is, each federated data subset needs to maintain a data ID list, which stores the data IDs that meet the description of the federated features, and the data corresponding to these data IDs constitute the federated data subset. For a given data subset, the model performance of the trained federated model on this data subset, in the preferred embodiment, the performance is defined using accuracy. When it is lower than the performance of the federated model on another part of the data in the entire data set except for the data subset, such a data subset is called a problem data subset. It is worth noting that since there are no strict restrictions on the size and performance index gaps of problem data subsets, not all problem data subsets can reveal the characteristics of problem data. In the next step, further screening and filtering using other indicators are required. The embodiment further uses problem subset filtering technology based on machine learning for screening and filtering.

[0061] Before screening the problem data subset based on the feature description combination, the discrete features are first classified according to the categories of discrete data, and the continuous features are divided into data intervals.

[0062] Specifically, as shown in Figure 4, the steps for screening the problem data subset are divided into four steps: 1) Classification for discrete features: For discrete features, the discrete features held by each participant are classified according to the category of discrete data; 2) Segmentation for continuous features: The continuous features held by each participant are binned, and the binned data intervals are used as the segmentation results; 3) Screening of problem data subsets based on feature description combinations; 4) Output of problem subsets.

[0063] Regarding the classification of discrete and continuous features, since discrete features, such as gender, job type, and whether a person owns a house, typically have only a limited number of data types, in this embodiment, the descriptions of discrete features can be separately classified and used as descriptive conditions for data subsets. However, for continuous features, such as age, savings, annual income, and social security contributions, since the number and span of values ​​involved are typically much greater than for discrete features, it is necessary to first use a binning method to divide a single continuous feature into several data segments through binning, and use these segments as descriptive conditions for data subsets.

[0064] After classifying all features, the present invention proposes a privacy-preserving data subset indicator calculation scheme. Using vertical distribution as an example, the potential data leakage risk in this scenario is illustrated: Assume two participants, A and B, each holding 10 features. A's first feature is called FA1, B's tenth feature can be called FB10, and so on. The features held by initiator A are called FA1 through FA10, and those held by collaborator B are called FB1 through FB10. One of these features is a gender feature (sex). When a data subset is described as sex = male, if initiator A knows which features belong to the IDs of this data subset, then the sex attribute data of these IDs belonging to this data subset has been accurately leaked! In other words, initiator A can obtain participant B's information by viewing the data IDs. To mitigate this data leakage risk, the present invention uses multi-party computation technology to protect data IDs during the problem data subset screening phase. Furthermore, during the problem data screening phase, the found data IDs are interval-merged to form a data error description.

[0065] Privacy protection is an important requirement for federated data subsets. The reason is that for one participant, the feature description in the federated data subset contains the numerical description of the features held by other participants. If this type of federated feature description corresponds to the data ID, it will cause a certain degree of data leakage risk. The data features in the embodiment are male and between 35 and 40 years old. The gender feature is located in participant A and the age feature is located in participant B. If the data ID list is open to participants A and B, then participant A can know that the age of the data corresponding to the data ID list is between 30 and 40 years old, which increases the risk of privacy information leakage. Participant B can know the gender information of the data corresponding to the data ID list, which is a direct data leakage. In summary, the data list of the federated data subset needs to be invisible to the participants, so as to solve the problem of avoiding direct or indirect privacy leakage.

[0066] In order to solve the problem of privacy leakage, the present invention proposes a federal data subset privacy protection method based on a mask vector. For the application scenario of vertical federated learning, before the vertical federated learning modeling, it is necessary to perform an ID alignment operation on the data ID. In the embodiment, the data ID of the default data set is after the ID alignment operation. The mask vector is defined as an array with a length of the full intersection data, and the elements are composed of 1 or 0. For a federated data subset, the ID belonging to this federated data subset has an element of 1 in the corresponding position of the mask vector, otherwise it is 0. At the same time, the content of this mask vector needs to be protected. Except for the participant holding the feature in the federated data subset, other participants should not know the specific value of each element in the mask vector. The corresponding relationship with the feature is shown in Figure 5.

[0067] Each feature in each data subset has a corresponding ID set associated with a separate mask vector. This mask vector exists only in the dataset to which the feature belongs. When a data subset contains multiple features, the mask vectors held by each feature owner are first multiplied using multi-party secure computing technology to determine the true mask vector of the current data subset. At this time, the multiplication result of the mask vectors remains encrypted. After obtaining the encrypted mask vector, subsequent calculations of indicators related to the data subset require the use of the mask vector, and the evaluation indicators of the current data subset can be calculated based on the true mask vector.

[0068] Taking the accuracy calculation for a data subset of 10 samples as an example, assuming the prediction accuracy for this data subset is represented by a vector [0, 1, 1, 1, 1, 1, 1, 1, 0], and the mask vector is an encrypted vector [1, 0, 1, 0, 1, 1, 1, 1, 0, 1], then the accuracy of this data subset is 5 / 7 * 100% = 71.4286%, where 7 is the number of samples in the data subset and 5 is the number of correctly predicted samples. Note that all operations in this calculation are performed using secure multi-party computation technology, eliminating the risk of data leakage. Similarly, by utilizing mask vectors and secure multi-party computation technology, the calculation method can be extended to calculate other evaluation metrics.

[0069] After obtaining the encrypted mask vector, subsequent calculations of error rates and evaluation metrics related to the federated data subset require the mask vector, employing privacy-preserving techniques. The following example illustrates the calculation of the error rate for a federated data subset. Only participant A possesses the label Y. Regardless of how the federated model is trained, the predicted value for a dataset is also obtained by participant A. This allows the initiator to determine whether the model's prediction for each sample is correct, generating a vector consisting of either 0 or 1. As shown in Figure 6, the accuracy calculation for a data subset with 10 samples is used as an example. Assuming the prediction accuracy for a data subset of 10 samples is represented by a vector [0, 1, 1, 1, 1, 1, 1, 0], and the mask vector is an encrypted vector [1, 0, 1, 0, 1, 1, 1, 1, 0, 1]. The accuracy for this data subset is 5 / 7 * 100% = 71.4286%, where 7 is the number of samples in the data subset and 5 is the number of correctly predicted samples.

[0070] Screening based on a combination of feature descriptions also includes training a machine learning model for screening problem data subsets. During the training of the machine learning model, multiple public data sets are used to randomly generate rules to select data subsets for label destruction. The data subsets with real problems are marked as positive sample labels for model training, and the rest are negative sample labels. Feature data of different data subsets are collected through data subset discovery technology to form a data set for the machine learning model; the problem data subset is input into the machine learning model after training to obtain a problem data subset with abnormal descriptions.

[0071] Specifically, as shown in FIG7 , the problem subset filtering technology based on machine learning is used. Since the number of problem data subsets found in the data set is usually large, and not all problem data subsets are caused by real problem data, many data subsets cannot fully reflect the situation of the problem data, nor can they help locate the problem data due to reasons such as too large a data volume, too small a data volume, and performance differences that are basically similar to the global data. Therefore, such data subsets need to be eliminated and filtered in subsequent processes. In summary, the present invention proposes a problem subset filtering technology based on machine learning.

[0072] The features of machine learning input include evaluation indicators not limited to data subsets, evaluation indicators of the entire set of data excluding the data subset, the number of data items in the data subset, the percentage of the number of data items in the data subset, the ratio of the evaluation indicators of the data subset and the entire set of data excluding the data subset, etc. The evaluation indicators include but are not limited to different indicators such as accuracy, precision, F-Score, and effect size.

[0073] During the model training phase, multiple public datasets are used. Data subsets are selected based on randomly generated rules for label destruction. The subsets containing problematic data are labeled as positive samples for model training, while the remaining samples are labeled as negative samples. Data subset discovery techniques are then used to collect feature data from different data subsets, forming a dataset for training the machine learning model. Once the data subsets are screened, the data subsets identified as positive samples by the machine learning model serve as input for the next technical step.

[0074] The multi-party secure computing method uses secret sharing technology as the underlying technology. The secret sharing technology includes: randomly splitting a number into two or more numbers, and the split numbers belong to different computing parties. Each computing party then performs arithmetic calculations under privacy protection based on the assigned data.

[0075] Specifically, secure and controllable multi-party secure computing technology is used to implement the problem data subset search and problem data tracing process involved in the entire federated debugging process, thereby ensuring that the result acquirer can correctly obtain the final analysis results, but cannot obtain sensitive information other than the analysis results, including but not limited to: the privacy data of other participants, the data IDs of other participants, etc.; in addition to providing data for calculation, the data provider cannot view or infer sensitive information owned by other participants from the intermediate results of the execution. The specific technical description used in the preferred embodiment is as follows.

[0076] Multi-party secure computation operations utilize secret sharing as the underlying technology. Secret sharing involves appropriately splitting a secret into shares managed by different participants. Each participant holds a share and collaborates to complete computational tasks (such as addition and multiplication). A single participant cannot recover the secret information; it can only be recovered by the collaborative efforts of multiple participants. Each participant independently performs addition and multiplication on the data in the shared shares. Each participant sends the results of the shared shares to the result-handler, who aggregates and restores the result. Throughout this process, no participant can access any secret information; the result-handler only has access to the result information. This effectively protects the original data from being leaked and allows the expected result to be calculated. In a secret sharing system, an attacker must simultaneously obtain a certain number of secret shares to obtain the key, thus ensuring system security. Furthermore, if some secret shares are lost or destroyed, the secret information can still be obtained using the remaining secret shares, thus ensuring system reliability.

[0077] The secret sharing scheme consists of a secret splitting algorithm and a secret reassembly algorithm. Since computational problems can always be represented as an arithmetic circuit consisting of addition gates and multiplication gates, if the secret sharing scheme can calculate addition and multiplication, then theoretically any complex problem can be calculated. The multi-party computing service node supports each participant in the case where they cannot obtain (or decrypt or infer) the original data of any other party, to achieve multi-party data collaborative computing, obtain model prediction results, and protect algorithm parameters, model parameters, and final results. It supports a variety of secure multi-party computing operators, including but not limited to arithmetic operations, comparison operations, logical operations, statistical operations, and function operations. Examples are as follows:

[0078] Four arithmetic operations (+, -, ×, ÷)

[0079] Comparison operations (>, ≥, =, ≠, <, ≤)

[0080] Logical operations (AND, OR, NOT, etc.)

[0081] Statistical operations (sum, count, mean, variance, etc.)

[0082] The main idea of ​​secret sharing is to randomly split a number into two or more numbers. The split numbers belong to different computing parties, and each computing party can perform arithmetic calculations under privacy protection based on the shared data.

[0083] Addition (in computer processing, subtraction is converted to addition, i.e., adding the subtrahend*(-1)): Assume that Party A and Party B each have numbers x and y, where x = x1 + x2 and y = y1 + y2. Each party shares x2 and y1. Party A shares x2 with Party B, and Party B shares y1 with Party A. Then Party A calculates z1 = x1 + y1, and Party B calculates z2 = x2 + y2.

[0084] Sharing z2 with participant A, it is obvious that: z = x + y = z1 + z2 = x1 + x2 + y1 + y2, and participant A can calculate z, which is the sum of x and y.

[0085] Multiplication (in computer processing, division is converted to multiplication, i.e., multiplying the reciprocal of the denominator): Participants A and B hold x and y respectively, and obtain a pair of random multiplication triples. Where a = [a]1 + [a]2, b = [b]1 + [b]2, ab = a*b = [ab]1 + [ab]2.

[0086] The product of x and y is calculated as follows:

[0087] Participant A and Participant B each share a shard of their own data with the other party, exchanging [x]2 and [y]1.

[0088] Participant A and Participant B each share and restore through addition, obtaining the blinded x d = xa and the blinded y e = yb, respectively, and then disclose d and e to each other. In this process, x, y, a, and b are not leaked.

[0089] Now, x*y = (d+a)*(e+b) = de+[b]d+[a]e+[ab]. The shared triples are enclosed in square brackets, indicating that the problem has been transformed into an addition problem. Participant B calculates the corresponding part and sends the result to Participant A. Participant A substitutes the x*y formula into this step to obtain the multiplication result.

[0090] To trace the source of the problematic data subset, the embodiment uses a technology for tracing erroneous data. This technology uses a filtered subset of problematic data and performs interval merging operations to form a subset describing the problematic data. This technology leverages the underlying technology of multi-party secure computation. During the entire execution process, the data of each participant does not leave the local server, and manual data review is not required, ensuring the security of the data and sensitive information of each participant.

[0091] Another embodiment is used to illustrate a vertical federated learning abnormal data debugging system based on data subsets. Referring to FIG8 , the system 800 includes:

[0092] The federated model training unit 810 is used by the initiator to model and train the federated model based on the vertical federated learning;

[0093] a problem data subset acquisition unit 820, configured to acquire a problem data subset from a data set using the trained federated model, wherein the prediction accuracy of the problem data subset in the federated model is lower than the prediction accuracy of other data subsets in the federated model;

[0094] The problem data subset screening unit 830 is used to screen the problem data subset based on the feature description combination to obtain the problem data subset with abnormal description;

[0095] The problem data subset correction unit 840 is used for the initiator to trace and correct the data based on the problem data subset with the abnormal description, and retrain the federated model after the correction.

[0096] The screening based on feature description combination in the problem data subset screening unit 830 also includes training a machine learning model for problem data subset screening. During the training process of the machine learning model, multiple public data sets are used to randomly generate rules to select data subsets for label destruction. The data subsets with real problems are marked as positive sample labels for model training, and the rest are negative sample labels. Feature data of different data subsets are collected through data subset discovery technology to form a data set for the machine learning model; the problem data subset is input into the machine learning model after training to obtain a problem data subset with abnormal description.

[0097] In addition to the upper module, the system 800 may also include other components. However, since these components are irrelevant to the content of the embodiment of the present disclosure, their illustration and description are omitted here.

[0098] For other specific working processes of the data subset-based vertical federated learning abnormal data debugging system 800, please refer to the description of the above-mentioned data subset-based vertical federated learning abnormal data debugging method embodiment, which will not be repeated here.

[0099] Another embodiment used to illustrate the system of the present invention can also be implemented using the computing device architecture shown in Figure 9 . Figure 9 illustrates the computing device architecture. As shown in Figure 9 , it includes a computer system 910, a system bus 930, one or more CPUs 940, input / output 920, and memory 950. Memory 950 can store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including those for the data subset-based longitudinal federated learning anomaly data debugging method. The architecture shown in Figure 9 is merely exemplary; when implementing different devices, one or more components in Figure 9 may be adjusted as needed. Memory 950, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data subset-based longitudinal federated learning anomaly data debugging method in the present invention (e.g., the federated model training unit 810, the problematic data subset acquisition unit 820, the problematic data subset screening unit 830, and the problematic data subset correction unit 840 in the data subset-based longitudinal federated learning anomaly data debugging system 800). One or more CPUs 940 execute various functional applications and data processing of the system of the present invention by running software programs, instructions, and modules stored in memory 950, that is, implementing the aforementioned data subset-based vertical federated learning abnormal data debugging method, which includes:

[0100] The initiator models and trains the federated model based on vertical federated learning;

[0101] Using the trained federated model to obtain a problem data subset in a data set, wherein the prediction accuracy of the problem data subset in the federated model is lower than the prediction accuracy of other data subsets in the federated model;

[0102] The problem data subset is screened based on a combination of feature descriptions to obtain a problem data subset with abnormal descriptions; the initiator performs data tracing and correction based on the problem data subset with abnormal descriptions, and retrains the federated model after correction.

[0103] Of course, the processor of the server provided in the embodiment of the present invention is not limited to executing the method operations described above, but can also execute relevant operations in the vertical federated learning abnormal data debugging method based on data subsets provided in any embodiment of the present invention.

[0104] The memory 950 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal, etc. In addition, the memory 950 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 950 may further include a memory remotely located relative to one or more CPUs 940, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0105] The input / output 920 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The input / output 920 may also include a display device such as a display screen.

[0106] Embodiments of the present invention also provide a non-transitory computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the data subset-based longitudinal federated learning anomaly data debugging method described in the above embodiments. The computer-readable storage medium of the embodiments of the present invention may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0107] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0108] The program code embodied on the storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0109] In addition, other specific working processes of a non-temporary computer-readable storage medium refer to the description of the above-mentioned embodiment of the vertical federated learning abnormal data debugging method based on data subsets, and are not repeated here.

[0110] In summary, the technical solutions provided by the above embodiments provide a method and system for debugging abnormal data in vertical federated learning based on data subsets, and an automated federated debugging application scenario for privacy protection. By designing a privacy-protected federated data subset evaluation technology, the relevant evaluation indicators of the federated data subset are correctly calculated while ensuring data privacy, forming a federated learning debugging framework based on the federated data subset, automatically locating abnormal data, and thus helping personnel solve the problem of abnormal performance of the federated learning model. During the debugging process, the initiator does not need to access the key privacy information of other participants, such as data and features, nor does the federated learning participants (or federated learning applications) need to leak or send each other the data, features, and other information held by the participants. Its beneficial effects include:

[0111] When a federated model performs poorly due to data errors, the federated model debugging process can automatically find the erroneous data without the data being exported locally. After deleting the erroneous data, the trained federated model can achieve normal performance.

[0112] Using privacy-preserving problem data subset search technology, without human intervention and without the data held by each participant being released locally, secure multi-party computing technology is used to identify the problem subset in the global dataset. For this problem subset, the prediction accuracy of the federated model on this data subset is much lower than the prediction accuracy of the model on other datasets other than the problem subset. The problem data subset found by this technology can be described by interpretable conditions that are easy for humans to understand. A screening method for problem data subsets is introduced. After the machine learning model has completed the screening of the problem data subset, the faulty data tracing technology is used to find the data in the dataset that caused the federated model to have label errors. Finally, the federated model is retrained by removing the traced faulty data to complete the debugging.

[0113] This machine learning-based problem subset filtering technology screens the many problem data subsets identified, eliminating those that don't meet the actual requirements and retaining those that negatively impact the performance of the federated model. The machine learning model used in this technology extracts features and labels from the existing problem data subsets, including their size and effect size. These features and labels are then used as training datasets to train the machine learning model.

[0114] This technology uses erroneous data tracing technology. This technology uses a filtered subset of problematic data and performs interval merging operations to form a subset of data describing the problem. Leveraging the underlying technology of multi-party secure computation, this technology ensures that all data from participants remains local throughout the entire execution process, eliminating the need for manual data review and ensuring the security of all participants' data and sensitive information.

[0115] In this document, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that comprises a series of elements includes not only those elements, but also includes other elements not expressly listed, or also includes elements inherent to such step or method.

[0116] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A method for debugging abnormal data in longitudinal federated learning based on data subsets, characterized in that: The method comprises: The initiator models and trains the federated model based on vertical federated learning; Using the trained federated model to obtain a problem data subset in a data set, wherein the prediction accuracy of the problem data subset in the federated model is lower than the prediction accuracy of other data subsets in the federated model; Screening the problem data subset based on the feature description combination to obtain the problem data subset with abnormal description; The initiator or participant performs data tracing and correction based on the problem data subset with the abnormal description, and retrains the federated model after the correction.

2. The method for debugging abnormal data in longitudinal federated learning based on data subsets according to claim 1 is characterized in that: Before screening the problem data subset based on the feature description combination, the discrete features are first classified according to the categories of discrete data, and the continuous features are segmented into data intervals.

3. The method for debugging abnormal data in longitudinal federated learning based on data subset according to claim 1 is characterized in that: The screening based on the combination of feature descriptions adopts a multi-party secure computing method to protect the data ID, and forms a data anomaly description by merging the data ID into intervals.

4. The method for debugging abnormal data in longitudinal federated learning based on data subsets according to claim 1 is characterized in that: The vertical federated learning abnormal data debugging method also includes a federal data subset privacy protection method based on a mask vector, specifically including: Align data IDs before modeling based on longitudinal federated learning; The mask vector is an array of full intersection data. The ID set corresponding to a feature in each data subset corresponds to a separate mask vector. The separate mask vector only exists in the data set to which the corresponding feature belongs. When a data subset contains multiple features, the mask vectors held by each feature owner are required to be multiplied using a multi-party secure computing method to determine the true mask vector of the current data subset. An evaluation metric for the current data subset is calculated based on the true mask vector.

5. The method for debugging abnormal data in longitudinal federated learning based on data subsets according to claim 1 is characterized in that: The screening based on feature description combination also includes training a machine learning model for screening problem data subsets. During the training process of the machine learning model, multiple public data sets are used to randomly generate rules to select data subsets for label destruction. The data subsets that actually have problems are marked as positive sample labels for model training, and the rest are negative sample labels. Feature data of different data subsets are collected through data subset discovery technology to form a data set for the machine learning model; the problem data subset is input into the machine learning model after training to obtain a problem data subset with anomaly descriptions.

6. The method for debugging abnormal data in longitudinal federated learning based on data subsets according to claim 3 or claim 4, characterized in that: The multi-party secure computing method is implemented using secret sharing technology as the underlying technology. The secret sharing technology includes: randomly splitting a number into two or more numbers, and the split numbers belong to different computing parties. Each computing party performs arithmetic calculations under privacy protection based on the assigned data.

7. A vertical federated learning abnormal data debugging system based on data subsets, characterized in that: The system comprises: The federated model training unit is used by the initiator to model and train the federated model based on vertical federated learning; A problem data subset acquisition unit, used to acquire a problem data subset in a data set by using the trained federated model, wherein the prediction accuracy of the problem data subset in the federated model is lower than the prediction accuracy of other data subsets in the federated model; A problem data subset screening unit, used to screen the problem data subset based on a combination of feature descriptions to obtain a problem data subset with an abnormal description; The problem data subset correction unit is used for the initiator or the participant to trace and correct the data based on the problem data subset with the abnormal description, and retrain the federated model after the correction.

8. The data subset-based longitudinal federated learning abnormal data debugging system according to claim 7 is characterized in that: The screening based on feature description combination in the problem data subset screening unit also includes training a machine learning model used for problem data subset screening. During the training process of the machine learning model, multiple public data sets are used to randomly generate rules to select data subsets for label destruction. The data subsets that actually have problems are marked as positive sample labels for model training, and the rest are negative sample labels. Feature data of different data subsets are collected through data subset discovery technology to form a data set for the machine learning model; the problem data subset is input into the machine learning model after training to obtain a problem data subset with abnormal descriptions.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for debugging abnormal data of longitudinal federated learning based on data subsets as described in any one of claims 1 to 6 is implemented.

10. A non-transitory computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by the processor, the vertical federated learning abnormal data debugging method based on data subsets as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Intrusion detection method, device, equipment and storage medium

    CN113434859A

  • Heterogeneous software defect prediction method based on federal prototype learning

    CN114896169A

  • Abnormal transaction risk early warning method and system based on network modeling

    CN115908022A

  • Federal learning-based sample selection method and system, electronic equipment and medium

    CN116468130A

  • Federal learning data detection and correction method and device

    CN116521658A

Cited By

  • Training method and system of image pre-training model and storage medium

    CN120913029A

  • Distributed equipment collaborative prediction maintenance method based on federated learning

    CN121217593A