High-dimensional data feature filtering method and system based on federated learning with adaptive label distribution drift
By constructing a federated learning-based label distribution drift adaptive high-dimensional data feature filtering method, the problem of model performance degradation caused by label drift in high-dimensional data is solved. It achieves efficient feature filtering and data integration in heterogeneous and noisy environments, and is applicable to data analysis in the medical and financial fields.
Patent Information
- Application Number
- CN202510299343.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing federated learning methods cannot effectively filter irrelevant features when faced with label distribution drift in high-dimensional data, leading to a decline in model performance. They are particularly difficult to apply in heterogeneous and noisy environments, and have high computational complexity.
We employ a high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift. By constructing a general framework, introducing a label offset robust federated feature filtering method (LR-FFS), and combining it with a distributed error detection rate control process, we ensure the accuracy of feature filtering and privacy protection.
In the presence of label drift and heterogeneity, it effectively filters irrelevant features, reduces computational complexity, improves model performance and fairness of data analysis, and is suitable for multi-source data integration in the medical and financial fields, thereby reducing data processing costs.
Smart Images

Figure CN120236157B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of federated learning technology, and in particular relates to a high-dimensional data feature filtering method and system based on adaptive label distribution drift of federated learning. Background Technology
[0002] Given the continuous advancements in science and technology, high-dimensional data classification is becoming increasingly prevalent in scientific research and industrial applications. The rapid expansion of data brings unprecedented opportunities, but also significant challenges:
[0003] Privacy breaches. In fields such as healthcare, data is typically collected and maintained by institutions in various locations. This data is highly sensitive, and its use is strictly regulated. Even if personal identifiers (such as names and dates of birth) are removed, privacy risks still exist. For example, a patient's facial information can be reconstructed from computed tomography (CT) or magnetic resonance imaging (MRI) data. Therefore, data sharing or centralized processing is generally prohibited.
[0004] Computational complexity. Processing large datasets presents significant computational challenges, such as easily exceeding storage capacity limits and consuming a large amount of computation time.
[0005] Data quality. Data collected by a single institution often faces challenges such as small volume and lack of diversity. For example, due to the low incidence of certain diseases, a single medical institution may struggle to collect sufficient data. In such cases, inter-institutional collaboration becomes crucial. Furthermore, data from a single institution is often susceptible to noise and outliers, which can degrade model performance. Therefore, robust and effective data analysis methods are particularly important.
[0006] Statistical heterogeneity. Heterogeneity refers to the differences in data distribution among institutions. Due to differences in equipment and geographical location, inherent data heterogeneity is very common in practical applications. A large body of literature has documented the impact of this heterogeneity on model training, including unstable convergence, suboptimal model performance, and even negative results.
[0007] To leverage data from each institution while ensuring security, distributed processing and federated learning have become increasingly popular approaches. In this framework, independent institutions, hospitals, or smartphones, referred to as clients, collaborate to train a global model. Clients communicate only with a central server (e.g., a service provider) and do not share raw data, thus ensuring data privacy while improving the efficiency of statistical inference. This approach has been widely adopted across multiple fields.
[0008] While federated learning helps integrate multi-source data, its implementation typically relies on platforms such as Spark. High-dimensional datasets not only incur significant computational costs but can also lead to overfitting and spurious correlations, known as the "curse of dimensionality." In such cases, it's generally accepted that only a subset of features significantly contribute to the classification task (the sparsity assumption). Therefore, pre-screening irrelevant features before formal analysis is crucial. This preprocessing is called feature filtering, and its core principle is to use statistics to measure the correlation (i.e., "utility") between each feature and the class response, and to eliminate features with low utility. Based on this idea, distributed and federated feature filtering methods have emerged, greatly improving computational efficiency. However, they neglect the impact of data heterogeneity across different clients, rendering the original methods ineffective in their feature filtering capabilities.
[0009] Among the various forms of heterogeneity, discrepancies in label distribution, often referred to as “label shift” or “label skewness,” are a widespread problem. These discrepancies are typically caused by geographical factors. In cases of label shift, it is generally assumed that characteristics within the same category are consistent across clients. For example, 2024 US cancer incidence data show that lung cancer rates in Kentucky, West Virginia, and Arkansas are three times higher than in Utah (75-84 cases per 100,000 vs. 25 cases per 100,000), reflecting historical differences in smoking rates. However, despite geographical differences, characteristics associated with specific types of cancer remain consistent.
[0010] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:
[0011] (1) Most methods fail when heterogeneity exists. Existing methods are usually based on the assumption that data from different clients are independent and identically distributed. However, in practical applications, data often exhibits heterogeneity (such as label drift), which leads to inconsistencies in parameter estimates (especially parameters related to class proportions) from different clients. This inconsistency causes a significant deviation between the aggregated parameters and those obtained through centralized estimation, severely damaging the statistical performance of the method and making it impossible to properly filter out irrelevant features. This bias is further exacerbated when class data is limited or even missing for some clients, making existing methods difficult to apply effectively in real-world scenarios.
[0012] (2) Methods robust to heterogeneity are sensitive to data noise. Some existing methods that are not sensitive to heterogeneity (such as PSIS and FAIR) are generally based on the statistical framework of mean test. However, these methods are highly sensitive to data noise and outliers. When there is noise in the data or the data exhibits a heavy-tailed distribution, the performance of these methods will degrade significantly. Even the presence of only one outlier in the data may cause the method to fail, severely limiting its application in real-world scenarios. In addition, when dealing with high-dimensional data, these methods often require additional regularization or dimensionality reduction processing, further increasing computational complexity and implementation difficulty. Summary of the Invention
[0013] To address the problems existing in the prior art, this invention provides a high-dimensional data feature filtering method and system based on federated learning and adaptive label distribution drift.
[0014] This invention is implemented as follows: A high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift includes:
[0015] A general framework for feature filtering: This framework incorporates various existing feature filtering methods and uses similar approaches for estimation and theoretical analysis.
[0016] Feature filtering methods for label drift scenarios: Based on a general framework, special weights are used to effectively address the impact of label drift;
[0017] Federated Feature Filtering Estimation Process: Estimating the LR-FFS utility function in a federated manner in environments where label offsets may exist, as well as other methods within a general framework;
[0018] Distributed Error Detection Rate (FDR) control process: Based on a given FDR control level, new feature utility is obtained by constructing pseudo-features, and the feature filtering threshold level is determined according to the FDR control relationship.
[0019] Furthermore, the specific steps of the federated feature filtering estimation process for the general framework include:
[0020] S1: The client iterates through each category and estimates the coefficient ζ. r :
[0021] S1.1: Iterate through each client and category r = 1, ..., R. The proportion of category r on the l-th client can be estimated as follows:
[0022] S1.2: The client uploads the category percentage and sample size to the central server, i.e.
[0023] S1.3: The central server passes through Estimate the overall category proportion, and through Estimate ζ r ;
[0024] S2: Expand each category into multiple terms using binomial expansion:
[0025]
[0026] S3: Among them The formula that needs to be estimated is d1 = 1, ..., d. Decompose it into two component functions and estimate them separately, i.e. in
[0027]
[0028] S4: Construction Estimation and The U statistic has a symmetric kernel, using and express;
[0029] S5: Iterate through each client and r = 1, ..., R, on the l-th client. and It can be estimated as:
[0030]
[0031]
[0032] in For data segment The combination of d+1 elements in;
[0033] S6: Each client will and sample size n l Upload to the central server;
[0034] S7: Central server via Aggregate parameters to obtain Estimate in
[0035] S8: Central server via Get ω j,r,d The estimate;
[0036] S9: Central server via To obtain the utility of each feature;
[0037] S10: For a given filtering threshold γ, we choose to preserve features.
[0038]
[0039] Furthermore, the specific steps of the federated feature filtering estimation process for LR-FFS include:
[0040] S1: Iterate through each client and its features in parallel;
[0041] S2: The client iterates through each category of data and calculates the corresponding... and aggregate weight α l,r ;
[0042] S3: The client will Send to the central server;
[0043] S4: The central server iterates through the data for each category, and... Weighting;
[0044] S5: Central server computing Will As a feature X j The utility;
[0045] S6: The central server receives... Features are preserved based on the selected threshold:
[0046]
[0047] Furthermore, the specific steps of the distributed error detection rate control process include:
[0048] S1: Each client will randomly shuffle its data to obtain pseudo-features;
[0049] S2: Apply the steps of the federated feature filtering estimation process for the general framework and the federated feature filtering estimation process for LR-FFS to the original features and pseudo-features respectively to obtain the utility values of the features. and and through Calculate new utility
[0050] S3: Central server via Determine the threshold
[0051] S4: The central server retains features based on selected thresholds.
[0052]
[0053] Another objective of this invention is to provide a high-dimensional data feature filtering system based on federated learning and adaptive label distribution drift, comprising:
[0054] A general feature filtering module for a general feature filtering framework: it incorporates various existing feature filtering methods into a general framework and uses a similar approach for estimation and theoretical analysis;
[0055] The label drift feature filtering module is designed for feature filtering methods in label drift scenarios: it adopts special weights on the basis of a general framework to effectively deal with the impact of label drift.
[0056] The federated feature filtering module is used for the federated feature filtering estimation process: estimating the LR-FFS utility function in a federated manner in environments where label offsets may exist, as well as other methods within a general framework;
[0057] The Error Detection Rate (FDR) control module is used for the distributed FDR control process: based on a given FDR control level, new feature utility is obtained by constructing pseudo-features, and the feature filtering threshold level is determined according to the FDR control relationship.
[0058] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift.
[0059] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift.
[0060] Another objective of this invention is to provide an information data processing terminal for implementing the high-dimensional data feature filtering system based on federated learning and adaptive label distribution drift.
[0061] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:
[0062] First, this invention provides a general framework for federated feature filtering, which integrates a large number of existing methods (such as CRU, MV-SIS, and CAVS) as special cases. This framework not only unifies the analysis and implementation of these methods but also allows for the simultaneous study of their large-sample properties. Specifically, it introduces a novel feature filtering method—label-shiftrobust federated feature screening (LR-FFS)—and proposes a corresponding distributed procedure for accurately quantifying the marginal importance of numerical features in classification problems.
[0063] The technical solution of this invention comprises four modules: First, this paper proposes a general framework to unify existing feature filtering methods and introduces a novel tool—Label-Offset Robust Federated Feature Filtering (LR-FFS)—along with its federated estimation process. This framework facilitates a systematic and unified analysis of existing methods and effectively mitigates the impact of label offset. LR-FFS handles label offset by utilizing conditional distribution functions and expectations without increasing computational burden. Simultaneously, it exhibits strong robustness against model misspecification and outliers. Furthermore, the federated process not only ensures computational efficiency and privacy protection but also maintains screening performance comparable to centralized processing methods while preserving privacy. Finally, we provide a false positive rate (FDR) control method for federated feature filtering to ensure the reliability and accuracy of results during the screening process. Experimental results and theoretical analysis show that LR-FFS performs excellently in different client environments, especially under challenges related to client class distribution, sample size, and missing class data, still providing efficient and accurate feature filtering.
[0064] Secondly, in the medical field, such as DNA and RNA sequencing, or in the financial field, such as credit assessment, LR-FFS can integrate multi-source data, breaking down data silos and performing feature filtering without sharing the original data, retaining effective features. This process can significantly reduce the complexity of data processing, computational resource consumption, data collection costs, and hardware and maintenance costs, enabling more small and medium-sized enterprises and research institutions to utilize high-dimensional data for analysis and decision-making. It also accelerates the application of technologies such as artificial intelligence across various industries, achieving cost reduction and efficiency improvement, and possesses high commercial value.
[0065] Currently, existing distributed or federated feature filtering methods rely on the assumption that data from different clients are independent and identically distributed, which fails to address the challenges posed by data heterogeneity, rendering their original advantages ineffective. This invention, LR-FFS, effectively utilizes the properties of conditional distribution and conditional expectation, identifying a statistic insensitive to class distribution. Furthermore, based on this statistic, a unified analysis and bias correction are performed on numerous existing feature filtering methods. In addition, we pioneeringly provide a process and theoretical guarantee for controlling the error detection rate in this scenario. This method fills these two technological gaps, providing a new technical path and solution for this field.
[0066] This invention demonstrates superior performance under label drift conditions, reducing model bias caused by distributional differences and improving the fairness and reliability of data analysis. More extreme examples show that the proposed LR-FFS remains effective even when client-side class data is missing or the sample size is small. It ensures the interests of minority groups, overcomes bias against disadvantaged groups, and guarantees algorithmic fairness. Attached Figure Description
[0067] Figure 1 This is a flowchart of a high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift provided in an embodiment of the present invention.
[0068] Figure 2 This is a block diagram of a high-dimensional data feature filtering system based on federated learning and adaptive label distribution drift, provided in an embodiment of the present invention.
[0069] Figure 3 These are simulation results from setting (B) provided in this embodiment of the invention, where the proportion of each category among different clients follows a Dirichlet distribution. The first row shows the success rate (SSR), and the second row shows the logarithmic ranking (log(wRank)) of the weakest correlated feature;
[0070] Figure 4 These are the simulation results of settings (E) and (F) provided in the embodiments of the present invention. The first column shows the success screening rate (SSR), and the second row shows the logarithmic ranking (log(wRank)) of the weakest correlation feature.
[0071] Figure 5 This refers to the classification accuracy (based on KNN classification) of different filtering methods in the TCGA example provided in this embodiment of the invention.
[0072] Figure 6 This is a schematic diagram of the traditional feature filtering process provided in an embodiment of the present invention;
[0073] Figure 7 This is a schematic diagram of the federal feature filtering process provided in an embodiment of the present invention;
[0074] Figure 8 This is a schematic diagram of the federated feature filtering process under the general framework of the present invention provided in the embodiments of the present invention;
[0075] Figure 9 This is a schematic diagram of the actual federated feature filtering process of the present invention provided in an embodiment of the present invention;
[0076] Figure 10 This is a schematic diagram of the distributed error detection rate control process provided in an embodiment of the present invention. Detailed Implementation
[0077] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0078] like Figure 1As shown in the figure, the high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift provided by this invention includes the following steps:
[0079] S101, a general framework for feature filtering: it incorporates various existing feature filtering methods into a general framework and uses a similar approach for estimation and theoretical analysis;
[0080] S102, Feature filtering method for label drift scenarios: Based on the general framework, special weights are used to effectively deal with the impact of label drift;
[0081] S103, Federated Feature Filtering Estimation Process: Estimating the LR-FFS utility function in a federated manner in environments where label offsets may exist, as well as other methods within a general framework;
[0082] S104, Distributed Error Detection Rate Control Process: Based on the given FDR control level, new feature utility is obtained by constructing pseudo-features, and the feature filtering threshold level is determined according to the FDR control relationship.
[0083] The specific steps of the federated feature filtering estimation process for a general framework provided in this embodiment of the invention include:
[0084] S1: The client iterates through each category and estimates the coefficient ζ. r :
[0085] S1.1: Iterate through each client and category r = 1, ..., R. The proportion of category r on the l-th client can be estimated as follows:
[0086] S1.2: The client uploads the category percentage and sample size to the central server, i.e.
[0087] S1.3: The central server passes through Estimate the overall category proportion, and through Estimate ζ r ;
[0088] S2: Expand each category into multiple terms using binomial expansion:
[0089]
[0090] S3: Among them The formula that needs to be estimated is d1 = 1, ..., d. Decompose it into two component functions and estimate them separately, i.e. in
[0091]
[0092] S4: Construction Estimation and The U statistic has a symmetric kernel, using and express;
[0093] S5: Iterate through each client and r = 1, ..., R, on the l-th client. and It can be estimated as:
[0094]
[0095] in For data segment The combination of d+1 elements in;
[0096] S6: Each client will and sample size n l Upload to the central server;
[0097] S7: Central server via Aggregate parameters to obtain Estimate in
[0098] S8: Central server via Get ω j,r,d The estimate;
[0099] S9: Central server via To obtain the utility of each feature;
[0100] S10: For a given filtering threshold γ, we choose to preserve features.
[0101]
[0102] The specific steps of the federated feature filtering estimation process for LR-FFS provided in this embodiment of the invention include:
[0103] S1: Iterate through each client and its features in parallel;
[0104] S2: The client iterates through each category of data and calculates the corresponding... and aggregate weight α l,r ;
[0105] S3: The client will Send to the central server;
[0106] S4: The central server iterates through the data for each category, and... Weighting;
[0107] S5: Central server computing Will As a feature X j The utility;
[0108] S6: The central server receives... Features are preserved based on the selected threshold:
[0109]
[0110] The specific steps of the distributed error detection rate control process provided in this embodiment of the invention include:
[0111] S1: Each client will randomly shuffle its data to obtain pseudo-features;
[0112] S2: Apply the steps of the federated feature filtering estimation process for the general framework and the federated feature filtering estimation process for LR-FFS to the original features and pseudo-features respectively to obtain the utility values of the features. and and through Calculate new utility
[0113] S3: Central server via Determine the threshold
[0114] S4: The central server retains features based on selected thresholds.
[0115]
[0116] like Figure 2 As shown in the figure, a high-dimensional data feature filtering system based on federated learning and adaptive label distribution drift provided by an embodiment of the present invention includes:
[0117] A general feature filtering module for a general feature filtering framework: it incorporates various existing feature filtering methods into a general framework and uses a similar approach for estimation and theoretical analysis;
[0118] The label drift feature filtering module is designed for feature filtering methods in label drift scenarios: it adopts special weights on the basis of a general framework to effectively deal with the impact of label drift.
[0119] The federated feature filtering module is used for the federated feature filtering estimation process: estimating the LR-FFS utility function in a federated manner in environments where label offsets may exist, as well as other methods within a general framework;
[0120] The Error Detection Rate (FDR) control module is used for the distributed FDR control process: based on a given FDR control level, new feature utility is obtained by constructing pseudo-features, and the feature filtering threshold level is determined according to the FDR control relationship.
[0121] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift.
[0122] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift.
[0123] Another objective of this invention is to provide an information data processing terminal for implementing the high-dimensional data feature filtering system based on federated learning and adaptive label distribution drift.
[0124] Specific implementation of the present invention:
[0125] First, this invention provides a general framework for federated feature filtering, which integrates a large number of existing methods (such as CRU, MV-SIS, and CAVS) as special cases. This framework not only unifies the analysis and implementation of these methods but also allows for the simultaneous study of their large-sample properties. In particular, a novel feature filtering method—label-shiftrobust federated feature screening (LR-FFS)—is introduced, and a corresponding distributed procedure is proposed for accurately quantifying the marginal importance of numerical features in classification problems.
[0126] Consider a classification problem, let Y∈{y1,…,y} R Let} be a categorical response variable with R categories, X = (X1, ..., X...). p ) T Let p be a variable containing p numerical features. Assume a complete dataset. Naturally divided into m data segments Each data segment Located in a client, containing n in (X,Y) l There are 100 observations, and each client processes the data independently. The total number of observations across all clients is 100. And n l<< p. Furthermore, let F(Y|X) denote the conditional cumulative distribution function of Y given X, and let P(Y|X) denote the conditional probability density function of Y given X.
[0127] A certain degree of heterogeneity in data distribution is allowed among clients, satisfying the label drift condition. That is, assuming the nth client... l The observations are from the joint distribution P. l (X,Y)=P l (X∣Y)P l (Y), where the subscript l denotes the distribution function of the l-th client. The distribution function of the response variable (label), P... l (Y) may differ between different clients, but the conditional probability density function, P l (X|Y) is the same across different clients, and is represented by P(X|Y).
[0128] The technical solution of this invention comprises four modules: a general feature filtering framework, a feature filtering method for label drift scenarios, a federated feature filtering estimation process, and a distributed error detection rate control process. Before introducing the four modules, it is necessary to provide an overview of the traditional feature filtering process, the specific implementation of which is as follows:
[0129] Traditional feature filtering process:
[0130] Feature filtering and feature selection methods are based on the sparsity assumption in high-dimensional problems, meaning that only a few predictors or features (i.e., X) will lead to the response (Y). Following this assumption, numerous feature selection methods have been proposed in recent years, including LASSO, SCAD, and MCP. While these methods have been successfully applied in many high-dimensional analysis fields, in fields such as genomics and protein analysis, with the exponential growth of data dimensionality due to modern technology, their application in ultra-high-dimensional problems faces significant challenges due to inherent computational complexity. Therefore, feature filtering methods—that is, removing irrelevant features before formal data analysis—have gradually become a focus of attention.
[0131] Feature filtering methods are mainly divided into model-based methods and model-free methods. Model-based methods include Feature Annealing Independence Rule (FAIR) and Paired Deterministic Independence Screening (PSIS). These methods rely on specific models, such as linear models, generalized linear models, and nonparametric regression models. Model-free methods include MV-SIS, Fusion Kolmogorov Filter (FKF), and Categorical Adaptive Variable Screening (CAVS). These methods are typically based on mean and independence hypothesis testing, measuring the marginal utility of features, such as the two-sample t-statistic, Kolmogorov-Smirnov statistic, and Cramér-von Mises distance. This utility (or test statistic) reflects the strength of the relationship between the feature and the response variable. The larger the utility value, the stronger the correlation between the feature and the response variable, and the easier it is to reject the null hypothesis, thus confirming the importance of the feature.
[0132] Therefore, a classic non-distributed feature filtering process involves the following steps:
[0133] (1): Define the feature filtering scenario and select the appropriate method, and define the set (active set) related to the response variable:
[0134]
[0135] (2): For a given random sample {(X) i ,Y i ):1≤i≤n}, the utility of each feature can be estimated.
[0136] (3): For a given filtering threshold γ, select the features to retain.
[0137]
[0138] Determining the threshold is often difficult, and typically the top d features with the highest utility values are retained. The threshold and the number of retained features d are usually related to the computing power of the computing device, the total number of features, and some error detection rate (FDR) control methods. However, these methods assume that all data is stored on a single computer, and therefore are not suitable for distributed scenarios with communication bottlenecks between clients.
[0139] To address this limitation, Li Xingxiang and Xu Chen (2020) pioneered a distributed feature filtering framework based on aggregated relevance screening, enabling clients to calculate utility without exchanging original data. In step (2), feature utility estimation is performed in a distributed manner, further broadening the application scenarios of feature filtering. Specifically, the process is as follows:
[0140] (1): Assume Y and feature X jThe correlation measure between them (i.e., X) j The utility can be expressed as:
[0141]
[0142] Where g is a predetermined function, θ j,1 ,…,θ j,s Let be a function of s component parameters that can be estimated relatively easily and unbiasedly. Assume... For θ j,h The unbiased estimate (kernel) is symmetric;
[0143] (2): For each data segment θ through local U statistics j,h Make an estimate:
[0144]
[0145] in For data segment The index set on;
[0146] (3): x is obtained through weighted aggregation. j Its utility:
[0147]
[0148] in
[0149] (4): For a given number of retained features d, select the model
[0150]
[0151] General feature filtering framework:
[0152] The utility function for each feature is defined as:
[0153]
[0154] in It is category Y = y t The utility value of the j-th feature. Here, d represents the order of difference, k is the exponent, and ζ is the value of the j-th feature. r It is usually related to Y=y r The proportion-dependent weighting parameter. Similar to the Kolmogorov-Smirnov distance, ω j,r,d Quantified from Y=y r And Y≠y r Whether the samples come from the same distribution. Its core idea is similar to FKF. Small differences indicate that X jIt has nothing to do with Y.
[0155] By adjusting parameters k, d, and weight ζ r A general framework can include many existing utility methods. For example, CAVS has weights of ζ. r =1-P(Y=y r This focuses on the utility of categories with lower category proportions. In contrast, CRU uses weights I(Y = y). r The square of the variance is more biased towards categories with a proportion close to 0.5.
[0156] Regarding the order of the parameter d, MV-SIS explores second-order differences (d=2), while CRU and CAVS focus on the distribution of first-order differences (d=1). When d=2, ω is estimated using the U statistic. j,r,2 The computational complexity is typically O(N). 3 p), which is much higher than the estimated d=1, and the complexity is O(N). 2 p) of ω j,r,1 For cases where d > 2, the computational burden increases further. Therefore, to improve computational efficiency, we focus on the utility method for d = 1. For ease of explanation, the following content will use ω... j and ω j,r It is expressed as the utility value using the first-order difference (d=1).
[0157] It is noted that existing methods are closely related to class proportions, and distributed or federated feature filtering methods assume that samples are independently and identically distributed across different clients. However, when label offsets cause different class proportions across clients, the impact of label offsets becomes significant when using the original distributed estimation method. This raises a key question: how to ensure that the objectives of different clients are consistent, i.e., that utility is insensitive to class proportions, in the case of label offsets? To address this issue, a novel feature filtering method is innovatively proposed.
[0158] Feature filtering methods for label drift scenarios:
[0159] To better mitigate the impact of label offset, a special weight was used. This is named Label-Shift Robust Federated Feature Selection (LR-FFS). Specifically, the utility of LR-FFS is formally defined as follows:
[0160]
[0161] The third equation is because...
[0162] The robustness of LR-FFS to label offset can be explained from two perspectives: the choice of statistics and coefficients.
[0163] First, the statistic ω j It is derived from the conditional expectation of the conditional distribution function, which helps to identify and mitigate the effects of label offset. Proposition 1 states that, under certain conditions, LR-FFS is effective for Y = y r The proportional change is robust. However, except for Y = y r In addition, changes in the proportions of the remaining R-1 categories can still affect the utility value. To address this issue, the remaining R-1 categories are merged into a single category Y≠y. r And pay attention to X j In Y = y r And Y≠y r Distribution differences under certain conditions.
[0164] Proposition 1:
[0165] When besides Y = y r When the proportions of other categories remain constant among the remaining R-1 categories, the expected... With Y = y r The proportion is irrelevant.
[0166] On the other hand, the coefficient weights of LR-FFS It is independent of class proportions, thus minimizing the impact of label shifts. As mentioned earlier, the weights in existing feature selection methods (such as CRU and CAVS) are functions of class proportions and are susceptible to label shifts. In contrast, LR-FFS naturally mitigates this vulnerability by ensuring that utility estimates remain consistent across clients, regardless of label distribution shifts.
[0167] It is important to emphasize that the goal is to mitigate, rather than completely eliminate, the impact of label shift. This approach aligns with the client drift mitigation principle in classic federated learning, which aims to align the local model more closely with the global model by adjusting the local objective. The scheme achieves a delicate balance between reducing the impact of label shift and maintaining estimation accuracy.
[0168] Federal Feature Filtering Estimation Process:
[0169] This section describes how to estimate the LR-FFS utility function using a federated approach, even when label offsets may exist. It reviews the proposed LR-FFS utility... It is the key statistic. To estimate γ... j,r We chose to use the One-Time Aggregation (OSA) method to obtain a robust global estimate with minimal inter-machine communication cost.
[0170] γ can j,r Decompose into two component functions and estimate them separately: in
[0171]
[0172] First estimate U j,r In the l-th client, U j,r It can be estimated using the bivariate U statistic:
[0173]
[0174] When each client will The data is transmitted to a central server, which then performs a weighted average. Aggregating uploads from various clients in This represents the effective sample size for the l-th client. It's important to note that due to label offset, It's U j,r Biased estimation, based on local client data to obtain U j,r An unbiased estimate is impossible. Fortunately, LR-FFS can correct this bias and design γ j,r A consistent estimator. Before delving into the details, some additional notation needs to be introduced. Let... The solution to the following equation:
[0175]
[0176] At the same time, set Note Through some algebraic transformations, we can obtain:
[0177]
[0178] Therefore, if a consistent estimate can be made It can be corrected The deviation, and obtain γ j,r A consistent estimator. Similar to an estimator. Using weighted U-statistic estimation Right now in
[0179] It can be easily checked. Finally, define γ j,r The estimator is
[0180] The general framework for federated feature filtering estimation process:
[0181] The previous section discussed how to estimate LR-FFS in a federated manner; other methods within the framework can be estimated in a similar way. The key lies in how to estimate ζ.r and ω j,r,d .
[0182] ζ r The estimate is direct; note that ζ r It is the category proportion π r The function, let's denote it as ζ r =g(π) r The following steps should be taken:
[0183] (1): Iterate through each client and r = 1, ..., R. The proportion of category r on the l-th client can be estimated as follows:
[0184] (2): The client uploads the category percentage and sample size to the central server, i.e.
[0185] (3): The central server passes through Estimate the overall category proportion, and through Estimate ζ r .
[0186] For ω j,r,d ω only when d = 1 j,r,d It can be broken down into and This avoids the introduction of interactive items and reduces estimation complexity. Among other things... The estimation process is the same as that of the federated feature filtering estimation process. When d>1, ω j,r,d The estimation becomes relatively complex, involving the following steps:
[0187] (1): By binomial expansion, ω j,r,d To split:
[0188]
[0189] in This is a formula that needs to be estimated. When d1 = 0, it can be easily derived.
[0190] (2): Traverse d1 = 1, ..., d, and Decompose it into two component functions and estimate them separately, i.e. in
[0191]
[0192] (3): Construction estimation and The U statistic has a symmetric kernel, using and express;
[0193] (4): Iterate through each client and r = 1, ..., R, on the l-th client, and It can be estimated as:
[0194]
[0195] in For data segment The combination of d+1 elements in;
[0196] (5): Each client will and sample size n l Upload to the central server;
[0197] (6): The central server passes through Aggregate parameters to obtain Estimate in
[0198] (7): The central server passes through Get ω j,r,d The estimate;
[0199] (8): The central server passes through To obtain the utility of each feature.
[0200] Distributed error detection rate control process:
[0201] The False Discovery Rate (FDR) is the proportion of falsely rejected null hypotheses out of all rejected null hypotheses. It is a crucial metric for evaluating model performance, particularly valuable in fields like industrial production and biomedicine, where it can be used to assess factors such as control failure rate and false positive rate. This section will explore how to more precisely control the FDR. Specifically, for each feature X... j On each client, the data is independently shuffled (permuted) to generate a "pseudo" feature X'. j Then, using the federated feature filtering process, X is calculated respectively. j and X' j The utility value is denoted as ω. j and ω' j Therefore, a new marginal utility φ is defined. j =ω j -ω' j Used to characterize X j The relationship between the response variable Y and the response variable Y.
[0202] The generated "pseudo" feature X' j and original feature X jIt has the same distribution and is independent of Y, therefore its utility value ω' j It should be close to zero. If X j If it is a relevant feature, then φ j It should be significantly greater than zero; otherwise, its value should be close to zero, and δ j The probability of a result being positive or negative is approximately equal. This property is known as the marginal symmetry property.
[0203] Given a threshold γ>0, the set of features to be retained is defined as follows:
[0204]
[0205] but The FDP (False Discovery Proportion) can be obtained from the following formula:
[0206]
[0207] in pass estimate, a∨b=max{a,b}. Therefore... The FDR can be generated by get.
[0208] In fact, since the irrelevant sets are unknown, inspired by marginal symmetry, we have
[0209]
[0210] thereby It can be done estimate.
[0211] Therefore, for a pre-defined FDR control level α, the threshold γ can be selected as:
[0212]
[0213] This invention proposes a general framework and an estimation process for LR-FFS. The estimation processes are similar, and the algorithm steps are explained using LR-FFS as an example. Utilizing the γ... j,r The forms of the two components are estimated separately, according to the appendix. Figure 6 , Figure 7 and Figure 8 The algorithm steps include:
[0214]
[0215]
[0216] Through this process, γ can be obtained. j,r The algorithm estimates the quantity. However, it has some drawbacks. Specifically, the transmission... This could potentially leak client-specific category ratio information, raising privacy concerns. To enhance privacy protection, the estimated value will be... Modify to an equivalent form (see appendix) Figure 9 Algorithm 2) avoids transmitting sensitive client information. Proposition 2 proves the equivalence of the two algorithms.
[0217]
[0218]
[0219] Proposition 2:
[0220] In Algorithm 2, step S.(2) aggregation Estimating aggregation parameters separately and The molecules are equivalent, specifically, the following relationships exist:
[0221] and
[0222] This invention provides a feature filtering method for label drift scenarios. Combining the above-mentioned technical solutions and the technical problems solved, the advantages and positive effects of the solution provided by this invention are as follows:
[0223] Robust to label drift: The utility value of LR-FFS is composed of the conditional distribution function and conditional expectation, which can accurately identify the impact of label drift on utility estimation. This method is insensitive to class proportions, achieving a good balance between mitigating the impact of label drift and ensuring statistical estimation efficiency, thus exhibiting strong robustness.
[0224] High computational efficiency: In Algorithm 2, the computational complexity for each client is O(n log n). This complexity is comparable to the corresponding steps in existing literature. Therefore, handling label offsets does not introduce additional computational burden. Table 1 summarizes the comparison of existing methods in terms of computational complexity, transmission cost, and robustness. LR-FFS, with its simplicity and robustness, is particularly suitable for scenarios with large sample sizes N and high-dimensional data p.
[0225] Robustness to outliers and noise: LR-FFS does not impose special restrictions on the model, thus exhibiting robustness even in cases of model misspecification. Furthermore, thanks to the properties of the distribution function, this utility effectively absorbs the effects of outliers and noise, making it suitable for scenarios with heavy-tailed distributions.
[0226] Robustness to Malicious Client-Side Attacks (Byzantine Attacks): In federated learning environments, besides outliers and noisy data, an extreme case is malicious client-side attacks. While LR-FFS is not specifically designed to defend against such attacks, it is robust due to its computation of ω. j,r The process does not rely on category-based weighting, thus reducing the impact of errors from malicious clients. Furthermore, the aggregation properties of the estimators in Algorithm 2 support more robust aggregation methods, such as median of mean or mean of median, further enhancing resistance to malicious attacks.
[0227] Table 1 Comparison of Feature Filtering Methods for Classification Problems
[0228]
[0229]
[0230] The proposed method is theoretically guaranteed, which establishes the statistical efficiency relationship between distributed and centralized estimation, as well as the properties of feature filtering. Before introducing the theoretical properties, some notation and conditional assumptions need to be explained.
[0231] Define the heterogeneity coefficient as This coefficient effectively measures the degree of heterogeneity of category r. When there is no heterogeneity, i.e. When, θ r =1. When extreme heterogeneity exists, and category r exists only in a single client, θ r =0. As θ r The reduction in [something] increases the degree of heterogeneity among clients.
[0232] Condition 1: Assume there exist three positive constants, b1, b2, and b3, such that and
[0233] Condition 2: Assume there exist positive constants c > 0 and 0 ≤ κ < 1 / 2, such that Established.
[0234] Condition 3: Assume the number of categories satisfies R = O(N) ξ ), where ξ>0 satisfies
[0235] Condition 4: There exists a constant. make Established.
[0236] Condition 1 requires that the proportion of each category be neither too large nor too small, while imposing certain restrictions on the degree of label offset. This condition relaxes the requirement for independent and identically distributed (IID) data from client clients; even if some clients have little or no data in certain categories, the condition still holds as long as other clients have sufficient data. Conditions 2 and 3 are similar to the assumptions in traditional feature filtering literature. Condition 2 allows the minimum true signal to be on the order of N. -k Condition 3 allows the number of categories of the response variable to gradually increase as the sample size N grows. Condition 4 ensures that, at the population level, relevant and irrelevant features can be clearly distinguished.
[0237] Conclusion 1: Assuming conditions 1 and 3 hold, for any positive constant c1 and category r = 1, ..., R, there exists a positive constant c2 such that...
[0238]
[0239] Conclusion 2: and The variance can be expressed as:
[0240]
[0241] Specifically, when conditions 1 and 3 are true and m = O(N), The mean square error has the following order:
[0242]
[0243] Conclusions 1 and 2 demonstrate that even if the number of features grows exponentially with the sample size, when log(p) = O(N) is satisfied... α (where α∈(0,1-2κ-4ξ)), estimator It still maintains consistency. The constant c2 reflects the heterogeneity information of the class distribution, and is consistent with... Positive correlation. Its error bound matches the efficiency of traditional centralized feature filtering.
[0244] Conclusion 3 – LR-FFS Deterministic Filtering Property: Continuing with the notation of Conclusion 1, when the feature retention threshold γ = cN -η When, for any constant c3>0, there exists c4>0 such that the following equation holds:
[0245]
[0246] Specifically, when condition 2 is true, there exists a positive constant c5, and...
[0247]
[0248] in This is the actual model size.
[0249] Conclusion 4 – LR-FFS Ranking Consistency Property: Continuing with the notation of Conclusion 3, when condition 4 is true, there exists a constant c6 such that the following equation holds:
[0250]
[0251] In conclusions 3 and 4, ω j The minimum signal strength meets the feature identifiability conditions commonly found in the literature. The method does not impose restrictions on the moment conditions of features, thus exhibiting robustness to heavy-tailed distributions. Compared to CRU, LR-FFS can adapt to the heterogeneity of response distributions among different clients. When the total sample size... When the feature size is large enough, LR-FFS can remove most irrelevant features with high probability while retaining all relevant features, thus ensuring deterministic filtering properties. Its convergence rate matches that of centralized feature filtering methods, proving the efficiency of this distributed method.
[0252] When condition 4 is true, there is a significant gap between the utility values of relevant and irrelevant features. This proves a theoretical result stronger than deterministic filtering: when log(p) = o(N) 1-2η-4ξ When LR-FFS is in a state of near 1, it can rank all relevant features higher than irrelevant features (Conclusion 4), thus ensuring that there is an ideal threshold to distinguish between relevant and irrelevant features.
[0253] Conclusion 5 – General Framework Deterministic Filtering Property: Assuming the number of categories R is fixed, ζ r It is a continuous function of category proportion, when the feature retention threshold γ = cN -η Furthermore, when condition 1 holds, for any constant c8 > 0, there exists c9 > 0 such that the following equation holds:
[0254]
[0255] Specifically, when condition 2 is true, there exists a positive constant c. 10 ,have
[0256]
[0257] in This is the actual model size.
[0258] Conclusion 6 - General framework ranking consistency property: Continuing with the notations in Conclusion 5, when Condition 4 holds, there exists a constant c 11 , such that the following equation holds:
[0259]
[0260] As can be seen from Conclusion 5 and Conclusion 6, as d increases, the error bounds tend to become looser, and to achieve a similar convergence rate, more sample sizes are needed for support.
[0261] Algorithms 1 and 2 discussed how to estimate the utility function in a distributed or federated manner. In addition, the present invention also proposes an algorithm for distributed FDR control, according to Appendix Figure 10 , the algorithm steps include:
[0262]
[0263]
[0264] Using the permutation method to construct "pseudo" features that are independent of Y while maintaining the same distribution as the original features. A similar method is to construct knockoff features, which also ensures that the "pseudo" features are correlated with the original features and are exchangeable, which helps to better control FDR in feature filtering. However, this method cannot be directly applied to high-dimensional problems because it requires 2p < n. Liu et al. (2022) and Pang and Xia (2024) respectively extended this method to high-dimensional scenarios in non-distributed and distributed feature filtering. Through two-step filtering, first reduce the number of features to d through preliminary screening, ensuring that knockoff features are constructed under the condition of satisfying 2d < n l to overcome the dimensionality constraint. However, this method will increase additional computational costs. Another significant drawback is that when the sample sizes of some clients are small, to meet the condition of 2d < min n l , many relevant features may be excluded, which may be counterproductive in practice. Conclusion 7 provides the theoretical properties regarding the estimation of the set of relevant features .
[0265] Conclusion 7 - For any defined if there exists a sequence c n → ∞, when (n, p) → ∞, and c n / p → 0 holds. Then for any α ∈ (0, 1), the threshold selected in Algorithm 3 and the corresponding selected feature set satisfy:
[0266]
[0267] Under more relaxed conditions, this method can effectively control FDR at a given α level. The condition requires that the growth rate of p is faster than that of c. n In high-dimensional scenarios, this requirement can usually be easily met.
[0268] This invention provides a model-free distributed or federated feature filtering algorithm that effectively addresses label drift scenarios. The invention comprises four modules:
[0269] A general framework for feature filtering: By incorporating various existing feature filtering methods into a general framework, estimation and theoretical analysis can be performed in a similar manner.
[0270] Feature filtering method for label drift scenarios: Based on the general framework, special weights are adopted to effectively deal with the impact of label drift.
[0271] The process of federated feature filtering under the general framework:
[0272] (1): Traverse each client and feature in parallel to estimate the proportion of client categories;
[0273] (2): The client uploads the category percentage and sample size to the central server;
[0274] (3): The central server aggregates and estimates the overall category proportion, and obtains ζ. r The estimate.
[0275] (4): By binomial expansion, ω j,r,d To split:
[0276]
[0277] (5): Traverse d1 = 1, ..., d, and Split into two component functions And estimate using the U statistic respectively;
[0278] (6): Each client will and sample size n l Uploaded to the central server, where it aggregates the data. The estimate;
[0279] (7): The central server passes through Get ω j,r,d The estimate, combined with the ζ obtained in step (3), r The estimate yields the utility of each feature;
[0280] (8): The central server retains features based on the selected threshold.
[0281] Federated Feature Filtering Estimation Process: Estimating the LR-FFS utility function in environments where label bias may exist, the steps (see Algorithm 2) include:
[0282] (1): Iterate through each client and its features in parallel;
[0283] (2): The client iterates through each category of data and calculates the corresponding... and aggregate weight α l,r ;
[0284] (3): The client will Upload to the central server;
[0285] (4): The central server iterates through the data of each category, through... Weighting;
[0286] (5): Central server calculation Will As a feature X j The utility;
[0287] (6): The central server received Features are preserved based on the selected threshold:
[0288]
[0289] Distributed error detection rate control process. This includes the following steps:
[0290] (1): Each client will randomly shuffle its data to obtain pseudo-features;
[0291] (2): Apply Algorithm 2 to both the original features and the pseudo-features to obtain the utility values of the features. and and through Calculate new utility
[0292] (3): The central server passes through Determine the threshold
[0293] (4): The central server retains features based on the selected threshold:
[0294]
[0295] This invention provides an information data processing terminal, which is used to implement the feature-based filtering algorithm to reduce computational costs.
[0296] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it causes the storage device to perform the feature filtering algorithm to reduce the storage of redundant features.
[0297] This invention provides a set of high-dimensional tabular numerical feature data, such as SNPs, genomics, risk control data, etc. The feature filtering algorithm is applied to the data to ensure that subsequent steps can be executed.
[0298] In the financial industry, accurate credit risk assessment is crucial. Overestimating user credit risk hinders efficient capital turnover and weakens the economic impact on financial institutions and even society as a whole. Underestimating user credit risk can lead to default risk, and more seriously, trigger a chain reaction that could result in black swan events. To comprehensively characterize user credit features, collecting user data from different institutions (banks, enterprises, etc.) is particularly important. This data naturally possesses high dimensionality and non-independent identically distributed characteristics due to the institution's attributes and scale. To build a more comprehensive assessment model, such as a binary classification model for determining whether a loan has been granted, multiple banks or financial institutions act as clients, each possessing a large amount of customer credit records, financial status, and other data. Under the LR-FFS method, joint feature filtering is performed to remove redundant features useless for classification, and new assessment models are built independently or jointly based on the retained features. Compared to results without feature filtering, LR-FFS protects customer privacy, reduces computational complexity, removes redundant features, and improves the accuracy and interpretability of model assessment, providing strong support for financial institutions to formulate reasonable credit policies.
[0299] Medical data is highly sensitive. On the one hand, due to the high technical and acquisition costs, such as SNP and DNA sequencing data, the data often exhibits characteristics of small sample sizes and extremely high dimensionality. On the other hand, data from different hospitals often show significant non-independent identically distributed characteristics due to their respective specialties, hospital levels, and geographical locations. To analyze such data, collaboration among multiple hospitals and data filtering and preprocessing are crucial. Multiple hospitals, acting as clients, each store a large amount of patient medical records, imaging data, and examination reports. Taking the analysis of inducing factors of type 2 diabetes in genome-wide association studies as an example, hospitals upload the corresponding statistics of local single nucleotide polymorphism (SNP) data at the initiative of a specific project (central server). Through integration by the central server, the association between each SNP and type 2 diabetes is obtained, and irrelevant features are removed. The entire process utilizes a much larger amount of data, enabling a more comprehensive, accurate, and less biased assessment of the importance of each SNP compared to institutions relying solely on their own data. Ultimately, this ensures the accuracy of downstream analysis and improves the precision of disease diagnosis and drug development.
[0300] Example 1: Simulated Dataset
[0301] Different heterogeneous scenarios were constructed through numerical simulation, covering linear models and generalized linear models, and considering the influence of various aspects such as data distribution (e.g., whether it is a heavy-tailed distribution, whether there is noise).
[0302] The first example considers a scenario with mean differences. N (X, Y) samples are randomly and independently generated, where the category response variable (label) in the l-th client follows a distribution P(Y = r) = π. r l r = 1, ..., R. For the r-th category, p = 10000 features are generated in the following manner:
[0303] X = μ r +ε,
[0304] Where μ r =(μ r1 ,…,μ rp ) T It is the mean parameter, ε=(ε1,…,ε p ) T It is a random noise vector. If a feature X j There is no difference in the mean, i.e., μ 1j =…=μ rj This can be considered an irrelevant feature. Consider the following four different experimental settings:
[0305] Assume there are 30 clients in total, and each client has n... l =100 samples. Distribution of different client categories P(Y=r)=π r l Influenced by the distributional heterogeneity parameter α, the noise vector ε follows a standard normal distribution, N(0,1). Assume R = 7, μ 1j =0.34, 1≤j≤8. Then the set of relevant features is...
[0306] To simulate the heterogeneity of the category distribution, on the l-th client, for each category r, a uniformly distributed random number is generated from the interval (1, α). After normalization, the proportion of category r on client l is calculated. As α increases, the heterogeneity of category distribution among different clients also increases.
[0307] Assume there are 30 clients in total, and each client has n... l = 100 samples. The distribution of different client categories follows a Dirichlet distribution controlled by parameter α. The noise vector ε independently follows a t-distribution with 2 degrees of freedom. For R = 5, 6, 7, we have μ11 =…=μ 14 =μ 25 =…=μ 28 =0.45,0.47,0.50.
[0308] The set of relevant features is then:
[0309] Assume there are a total of 16 clients, with four groups each for sample sizes of 100, 200, 300, and 400, for a total sample size of N = 4000. Set R = 8, the number of missing classes for each client ranges from 0 to 4, while the remaining classes have the same proportion. The noise term ε independently follows a standard log-normal distribution (i.e., log(ε) ~ N(0,1)). Set μ... 1j =0.32, for 1≤j≤10, and μ 2j =0.08, for 1≤j≤10. The index set of relevant features is
[0310] This setting examines the effect of FDR control in Algorithm 3. The heterogeneity setting and sample size are the same as in setting (A), assuming R = 5, and for 1 ≤ j ≤ 8, μ 1j =0.4, considering the heterogeneity coefficient α=5.
[0311] In the above settings: (A) is the simplest feature filtering scenario. (B) assumes that the class distribution among different clients follows a Dirichlet distribution, and that heterogeneity increases as α decreases. (C) considers the cases of missing class labels and variations in sample size among clients. (D) reports the FDR control effect under threshold selection in the distributed error detection rate control process.
[0312] The second example considers a scenario where the probability density functions differ under a generalized linear model. Assume there are R classes, and N = 3000 samples are generated from a multinomial logistic regression model log(P(Y=1|X))∝Xβ+ι, where β=(β1,…,β1). p ) T Let β represent a vector containing p = 8000 regression coefficients, where ι is a constant. Set most elements in β to zero to ensure that only a subset of the non-zero coefficients affect the response variable. Consider the following two settings:
[0313] EX ~ N(0,∑1), where Σ1 is a p×p identity matrix. Let the set of indices for the relevant features be... And for Set β j =(-1) W ×1, where W~Bernoulli(0.5), ι=-0.25.
[0314] Furthermore, we set P(Y=2|X) / 1.2=P(Y=3|X)=…=P(Y=R|X), and replace the 30 samples with random noise that follows a uniform distribution in [0,100]. The class heterogeneity among clients follows a Dirichlet distribution.
[0315] FX ~ N(0,∑2), where ∑2=[σ j,h ] p×p Its element is defined as: σ j,j =1, when |jh|=1 When |jh|=2 When |jh|≥3, σ j,h =0. Set the index set of relevant features as... And for Set β j =(-1) W ×1.5, where W~Bernoulli(0.5), ι=-0.2.
[0316] Similar to setting (E), the class probabilities are adjusted such that P(Y=2|X) / 0.8=P(Y=3|X)=…=P(Y=R|X), and noise is introduced. The class heterogeneity setting follows setting (A).
[0317] For settings (E) and (F), 50 features were retained out of p = 8000 features.
[0318] For each setting, the LR-FFS method was applied for federated feature filtering. For comparison, existing classification-based utility methods—CRU, PSIS, FKF, MV-SIS, and CAVS—were also used. These distributed algorithms were used for feature selection, based on the method of Li et al. (2022). To simulate potential noise in the data, 50 samples were randomly selected from these clients, and all features were replaced with random numbers uniformly distributed between 0 and 100. To obtain a suitable threshold γ in settings (A) through (C) and ensure data privacy, the strategy of Li and Xu (2024) was followed. First, a set (Z1,...,Z2) containing q = 1000 auxiliary features was created by permuting the randomly selected features. q Since these auxiliary features are independent of Y, the threshold is set to... in It is Y and Z j OSA estimation of the screening utility between the two.
[0319] Screening accuracy was assessed using the success rate (SSR), positive selection rate (PSR), and free selection rate (FDR) across 200 replicate experiments at T=200.
[0320]
[0321] in, This represents the set of indices of the features retained in the t-th iteration. Additionally, the average number of features retained after filtering is also reported. This includes the feature size and the average rank (wRank) of the weakest correlated feature in each simulation. For each method, the average computation time (in seconds) required to perform the distributed screening on a local machine is also reported. Table 2 shows the results for all performance metrics, with additional charts highlighting the SSR and wRank results for visualization purposes.
[0322] Table 2 shows the results with R=7 in setting (A) (the results in parentheses indicate the results after adding noise).
[0323]
[0324]
[0325] The upward arrows indicate that higher values are better, while the downward arrows indicate that lower values are better.
[0326] Figure 3 The simulation results in setting (B) show the proportions of each category among different clients following a Dirichlet distribution. The first row shows the success rate (SSR), and the second row shows the log-rank (log(wRank)) of the weakest correlated feature.
[0327] Table 3 shows the simulation results for setting (C), where each client contains only partial category data (the data in parentheses indicates the results after adding noise).
[0328]
[0329] Table 4 shows the simulation results for setting (D) (the results with noise added are shown in parentheses).
[0330]
[0331]
[0332]
[0333] Figure 4 The simulation results (E) and (F) are set up. The first column shows the success rate (SSR), and the second row shows the logarithmic ranking (log(wRank)) of the weakest correlated features.
[0334] As shown in Table 2, MV-SIS and FKF perform poorly in solving the distributed feature selection problem and consume a significant amount of time. Therefore, in other analyses, only the results of PSIS, CRU, and CAVS methods are presented.
[0335] In setting (A), without outliers, PSIS demonstrates robustness to label shifts and achieves effective feature filtering. However, in settings (B) and (C), when features exhibit heavy-tailed distributions or outliers are present, PSIS performs close to random guessing, significantly reducing its effectiveness. LR-FFS consistently delivers the best performance across all scenarios, particularly excelling in settings with moderate client heterogeneity. Although CAVS performs slightly worse than LR-FFS, it still exhibits certain advantages by using the maximum value as a special weight. For CRU, label shifts were observed to severely impact the filtering results, reducing their accuracy and reliability. Setting (D) shows the correct filtering rate for different features and the FDR results under different FDR control levels. It can be seen that, through Algorithm 3, all distributed feature filtering methods can effectively control FDR, achieving an ideal level.
[0336] The classification problems in settings (E) and (F) are more complex, involving multiple classes, different class distributions, label shifts, and the presence of noise. As expected, the correlation between features makes accurate feature filtering more challenging in these two configurations. Despite the significant differences in the overall class distribution, the label shift has a relatively smaller impact on filtering compared to settings (A) to (C). Nevertheless, LR-FFS still demonstrates superior accuracy compared to other methods in all these challenging scenarios.
[0337] Example 2: TCGA Dataset
[0338] The proposed method was applied to a dataset of invasive breast cancer, containing comprehensive data from 981 patients across 38 institutions. This dataset, part of the PanCancer Atlas project, covers data on mutated genes, patient demographics, and tumor type. The dataset can be downloaded from the official website (pancanatlas).
[0339] The primary objective is to develop a classifier for identifying breast cancer gene subtypes as a 5-class classification task. Although the dataset contains 20,531 features from mRNA expression data, the limited sample size presents a significant challenge for accurate classification. Furthermore, the heterogeneity in subtype proportions across institutions adds further difficulty to model training. Due to ethical and privacy concerns, many institutions cannot share raw data, necessitating the use of federated feature filtering methods in a medical context to ensure data privacy and compliance while fully utilizing the complete dataset across institutions.
[0340] During model training, institutions with a minimum sample size of 32 were designated as clients in the training set, while the other institutions served as the test set. Ultimately, the training set contained 13 clients and 829 samples, and the test set contained 152 samples. Detailed sample sizes for each client are shown in Table 2.
[0341] Table 5 Sample size for each institution
[0342]
[0343] In addition to analyzing the original dataset, experiments were conducted on noise contamination and attacks. In the noise test, all features of 30 samples in the training set were replaced with random numbers drawn from a uniform distribution from 0 to 30. In the attack test, the label of one client sample in the training set was randomly shuffled. Each test was repeated 50 times. To ensure the robustness of the results, the number of features whose utility value appeared more than 45 times among the top 100 features was reported.
[0344] The LR-FFS method was applied and compared with CRU, PSIS, and CAVS methods for feature selection, ultimately retaining key features. Using the selected features, a K=40 K-nearest neighbor (KNN) classifier was trained. The average accuracy of distributed estimation and summary data estimation on the test set is also reported across 50 repeated experiments. Results are presented in... Figure 5 As shown in Table 6, in addition to high average accuracy, LR-FFS can also retain similar features in multiple repeated filtering processes, demonstrating the stability of the method.
[0345] Table 6 shows that in 50 repeated trials, the number of instances where the top 100 features in terms of importance exceeded 45.
[0346]
[0347] PSIS performed poorly due to the inherent outliers and noise in medical data. Even with the addition of extra noise and attacks, LR-FFS consistently maintained superior feature selection and demonstrated stability. In comparisons of aggregated data and federated screening results, PSIS remained consistent, while the other three methods showed some variation, with LR-FFS exhibiting the least change.
[0348] Comparative analysis of LR-FFS with other feature filtering methods in Examples 1 and 2 verifies the effectiveness and superiority of the label drift robust federated feature filtering method proposed in this invention. First, regarding the impact of label drift on utility estimation, this invention can accurately identify and effectively handle the statistical bias caused by label shift, avoiding the performance degradation that may occur in traditional methods when facing label drift. Compared with existing methods, this invention successfully alleviates the challenges caused by client data heterogeneity while ensuring statistical estimation efficiency, and demonstrates excellent feature selection performance in different application scenarios. Second, the feature filtering method proposed in this invention has significant computational advantages. Compared with other distributed feature filtering methods, this method effectively solves data heterogeneity without introducing additional computational burden; at the same time, compared with traditional centralized processing methods, this method adopts a "divide and conquer" strategy, significantly improving computational efficiency, and its statistical efficiency is similar to that of centralized estimation results, making it more competitive in practical applications. Finally, the proposed distributed FDR control flow provides solid theoretical support and practical basis for various distributed feature filtering methods. This process effectively controls the false discovery rate (FDR) in heterogeneous multi-client environments, thereby ensuring high reliability and accuracy of statistical inference during the screening process.
[0349] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0350] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift, characterized by comprising the following steps: Step 1, a general framework for feature filtering: Incorporate various existing feature filtering methods into a general framework, and use similar methods for estimation and theoretical analysis; Step 2, Feature filtering method for label drift scenarios: Based on the general framework, special weights are used to effectively deal with the impact of label drift; Step 3, Federated Feature Filtering Estimation Process: Estimating the LR-FFS utility function in a federated manner in environments where label offsets may exist, as well as other methods within a general framework; Step 4, Distributed Error Detection Rate Control Process: Based on the given FDR control level, new feature utility is obtained by constructing pseudo-features, and the feature filtering threshold level is determined according to the FDR control relationship; The specific steps of the federated feature filtering estimation process for the general framework include: (1): The client iterates through each category and estimates the coefficient ζ. r : 1.1: Iterate through each client and category r = 1, ..., R. The proportion of category r on the l-th client can be estimated as follows: 1.2: The client uploads the category percentage and sample size to the central server, i.e. 1.3: The central server passes through Estimate the overall category proportion, and through Estimate ζ r ; (2): Expand each category into multiple terms using binomial expansion: (3): Among them The formula that needs to be estimated is d1 = 1, ..., d. Decompose it into two component functions and estimate them separately, i.e. in (4): Construction estimation and The U statistic has a symmetric kernel, using and express; (5): Iterate through each client and r = 1, ..., R, on the l-th client, and It can be estimated as: in For data segment The combination of d+1 elements in; (6): Each client will and sample size n l Upload to the central server; (7): The central server passes through Aggregate parameters to obtain Estimate in (8): The central server passes through We obtain estimates of ωj,r,d; (9): The central server passes through To obtain the utility of each feature; (1)0: For a given filtering threshold γ, select the feature to retain.
2. The high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift as described in claim 1, characterized in that, The specific steps of the federated feature filtering estimation process for LR-FFS include: (1): Iterate through each client and its features in parallel; (2): The client iterates through each category of data and calculates the corresponding... and aggregate weight α l,r ; (3): The client will Upload to the central server; (4): The central server iterates through the data of each category, through... Weighting; (5): Central server calculation Will As a feature X j The utility; (6): The central server received Features are preserved based on the selected threshold:
3. The high-dimensional data feature filtering method based on federated learning and adaptive label distribution drift as described in claim 2, characterized in that, The specific steps of the distributed error detection rate control process include: (1): Each client will randomly shuffle its data to obtain pseudo-features; (2): Apply the steps of the federated feature filtering estimation process for the general framework and the federated feature filtering estimation process for LR-FFS to the original features and pseudo features respectively to obtain the utility value of the features. and and through Calculate new utility (3): The central server passes through Determine the threshold (4): The central server retains features based on the selected threshold:
4. A high-dimensional data feature filtering system based on federated learning and adaptive label distribution shift, implementing the high-dimensional data feature filtering method based on federated learning and adaptive label distribution shift as described in any one of claims 1-3, characterized in that, The high-dimensional data feature filtering system based on federated learning and adaptive label distribution shift includes: A general feature filtering module for a general feature filtering framework: it incorporates various existing feature filtering methods into a general framework and uses a similar approach for estimation and theoretical analysis; The label drift feature filtering module is designed for feature filtering methods in label drift scenarios: it adopts special weights on the basis of a general framework to effectively deal with the impact of label drift. The federated feature filtering module is used for the federated feature filtering estimation process: estimating the LR-FFS utility function in a federated manner in environments where label offsets may exist, as well as other methods within a general framework; The Error Detection Rate (FDR) control module is used for the distributed FDR control process: based on a given FDR control level, new feature utility is obtained by constructing pseudo-features, and the feature filtering threshold level is determined according to the FDR control relationship.
5. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the high-dimensional data feature filtering method based on label distribution drift adaptive as described in any one of claims 1-3.
6. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the high-dimensional data feature filtering method based on label distribution drift adaptive as described in any one of claims 1-3.
7. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the high-dimensional data feature filtering system based on federated learning and adaptive label distribution drift as described in claim 4.
Citation Information
Patent Citations
Federal learning privacy evaluation method under cross-domain heterogeneous scene
CN115952507A
Privacy-preserving asynchronous federated learning of vertical partition data
CN116034382A