Tag distribution drift adaptive high-dimensional data feature filtering method and system based on federated learning

By introducing a feature filtering method that adaptively adaptively in label distribution drift in federated learning framework, the problems of data heterogeneity and label distribution drift in high-dimensional data are solved, and efficient and accurate feature filtering effect is achieved.

CN120236157AActive Publication Date: 2025-07-01RENMIN UNIVERSITY OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510299343.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with data heterogeneity, especially the problem of label distribution drift, when processing high-dimensional data, resulting in failure of feature filtering capabilities.

Method used

A high-dimensional data feature filtering method based on federated learning is proposed to effectively deal with label drift and data heterogeneity by introducing a common feature filtering framework, feature filtering method for label drift scenarios, federal feature filtering estimation process and distributed error discovery rate control process.

Benefits of technology

This method can accurately quantify the marginal importance of numerical features in classification problems in the presence of label drift and data heterogeneity, improve the accuracy and robustness of feature filtering, reduce the computational complexity, and break the data island phenomenon in multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236157A_ABST
    Figure CN120236157A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of federated learning, and discloses a tag distribution drift adaptive high-dimensional data feature filtering method and system based on federated learning, and the method comprises the steps: firstly providing a general framework for unifying an existing feature filtering method, and introducing a novel tool-tag offset robust federated feature filtering (LR-FFS); and a federated estimation process thereof. The framework is beneficial to systematic unified analysis of the existing method and effectively alleviates the influence caused by label offset. LR-FFS processes label offsets by utilizing conditional distribution functions and expectations without increasing computational burden. At the same time, the method has strong robustness for model missetting and abnormal values. In addition, the federation process not only ensures the calculation efficiency and privacy protection, but also can maintain the screening effect equivalent to that of a centralized processing method while maintaining privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and particularly relates to a high-dimensional data feature filtering method and system for adapting to label distribution drift based on federated learning. Background Art

[0002] Given the continuous progress of science and technology, high-dimensional data classification has become increasingly common in scientific research and industrial applications. The rapid expansion of data has brought unprecedented opportunities, as well as significant challenges:

[0003] Privacy leakage. In fields such as healthcare, data is usually collected and maintained by local institutions. This data is highly sensitive and its use is strictly regulated. Even after removing personal identifiers (such as names and dates of birth), privacy risks still exist. For example, a patient's facial information can be reconstructed from computed tomography (CT) or magnetic resonance imaging (MRI) data. Therefore, data sharing or centralized processing is usually prohibited.

[0004] Computational complexity. Processing large datasets poses significant computational challenges, such as easily exceeding storage capacity limits and consuming a large amount of computing time.

[0005] Data quality. Data collected by a single institution usually faces the problems of small data volume and lack of diversity. For example, due to the low incidence of certain diseases, it is difficult for a single medical institution to collect sufficient data. In this case, cross-institutional cooperation becomes crucial. In addition, data from a single institution is often vulnerable to noise and outliers, which may reduce model performance. Therefore, powerful and robust data analysis methods become particularly important.

[0006] Statistical heterogeneity. Heterogeneity refers to the differences in data distribution between institutions. Due to differences in equipment and geographical locations, the inherent heterogeneity of data is very common in practical applications. A large number of literatures have documented the impact of this heterogeneity on model training, including unstable convergence, suboptimal model performance, and even negative results.

[0007] To utilize the data of each institution while ensuring security, distributed processing and federated learning have become increasingly popular methods. In this framework, independent institutions, hospitals, or smartphones, etc., are called clients, and they cooperate to train a global model. The clients only communicate with a central server (such as a service provider) without sharing the original data, thus ensuring data privacy while improving the efficiency of statistical inference. This method has been widely applied in many fields.

[0008] Although federated learning helps to integrate multi-source data, its implementation usually relies on platforms such as Spark. High-dimensional datasets not only bring huge computational costs, but even lead to overfitting and spurious correlations, namely the "curse of dimensionality". In this case, it is generally believed that only some features contribute significantly to the classification task (sparsity assumption). Therefore, it is particularly important to pre-screen irrelevant features before formal analysis. This preprocessing process is called feature filtering, and its core lies in using statistics to measure the correlation (i.e., "utility") between each feature and the class response, and removing features with low utility. Based on this idea, distributed and federated feature filtering methods have emerged, greatly improving the computational efficiency, but ignoring the impact of data distribution heterogeneity among different clients, rendering the feature filtering ability of the original method ineffective.

[0009] Among various heterogeneities, label distribution differences, usually referred to as "label shift" or "label distribution skew", are widespread problems. These differences are usually caused by geographical factors. In the case of label shift, it is usually assumed that the features within the same class are consistent among clients. For example, data on the incidence of cancer in the United States in 2024 shows that the incidence of lung cancer in Kentucky, West Virginia, and Arkansas is three times that in Utah (75-84 cases per 100,000 people compared to 25 cases), which reflects historical differences in smoking rates. However, despite geographical differences, the features related to a specific type of cancer remain consistent.

[0010] Through the above analysis, the problems and defects of the existing technologies are as follows:

[0011] (1) Most methods fail when there is heterogeneity. Existing methods usually build on the assumption of independent and identically distributed data among different clients. However, in practical applications, data often has heterogeneity (such as label drift), which leads to inconsistencies in parameter estimation (especially parameters related to class proportions) among different clients. This inconsistency causes a significant deviation between the aggregated parameters and the parameters obtained by centralized estimation, seriously undermining the statistical performance of the method and resulting in the inability to correctly filter out irrelevant features. Especially in cases where the class data of some clients is limited or even missing, this deviation will be further exacerbated, making it difficult for existing methods to be effectively applied in practical scenarios.

[0012] (2) Methods robust to heterogeneity are sensitive to data noise. Some existing methods that are insensitive to heterogeneity (such as PSIS and FAIR) are usually based on the statistical framework of mean tests. However, these methods are highly sensitive to data noise and outliers. When there is noise in the data or it exhibits a heavy-tailed distribution, the performance of these methods will significantly decline. Even if there is only one outlier in the data, it may cause the method to fail, severely limiting its application effect in practical scenarios. In addition, when dealing with high-dimensional data, these methods often require additional regularization or dimensionality reduction processing, further increasing the computational complexity and implementation difficulty. Summary of the Invention

[0013] In view of the problems existing in the prior art, the present invention provides a high-dimensional data feature filtering method and system based on federated learning for adaptive label distribution drift.

[0014] The present invention is implemented as follows. A high-dimensional data feature filtering method based on federated learning for adaptive label distribution drift includes:

[0015] General feature filtering framework: Incorporate various existing feature filtering methods into the general framework and perform estimation and theoretical analysis in a similar manner;

[0016] Feature filtering method for label drift scenarios: Adopt special weights on the basis of the general framework to effectively cope with the impact of label drift;

[0017] Federated feature filtering estimation process: Estimate the LR-FFS utility function and other methods within the general framework in a federated manner in an environment where label offset may exist;

[0018] Distributed false discovery rate control process: Based on a given FDR control level, obtain the new feature utility by constructing pseudo-features, and determine the feature filtering threshold level according to the FDR control relationship.

[0019] Furthermore, the specific steps of the federated feature filtering estimation process for the general framework include:

[0020] S1: The client traverses each category and estimates the coefficient ζ r :

[0021] S1.1: Traverse each client and category r = 1,..., R. On the l-th client, the proportion of category r can be estimated as

[0022] S1.2: The client uploads the category proportion and the sample size to the central server, that is,

[0023] S1.3: The central server passes through Estimate the overall class proportion, and through estimate ζ r ;

[0024] S2: By means of binomial expansion, expand each category into multiple terms:

[0025]

[0026] S3: Among them is the expression to be estimated. Traverse d1 = 1, …, d, and split into two component functions and estimate them separately, that is Among them

[0027]

[0028] S4: Construct the U-statistic symmetric kernels for estimating and , denoted by and ;

[0029] S5: Traverse each client with r = 1, …, R. On the l-th client, and can be estimated as:

[0030]

[0031] Among them is the combination of d + 1 elements in the data segment ;

[0032] S6: Each client uploads and the sample size n l to the central server;

[0033] S7: The central server aggregates the parameters through to obtain the estimate of Among them Among them

[0034] S8: The central server obtains the estimate of ω through j,r,d ;

[0035] S9: The central server obtains the utility of each feature through ;

[0036] S10: For a given filtering threshold γ, we choose to retain the features

[0037]

[0038] Furthermore, the specific steps of the federated feature filtering and estimation process for LR-FFS include:

[0039] S1: Traverse each client and feature in parallel;

[0040] S2: Each client traverses each category of data and calculates the corresponding and the aggregation weight α l,r ;

[0041] S3: The client transmits to the central server;

[0042] S4: The central server traverses each category of data and performs weighting through ;

[0043] S5: The central server calculates and uses as the utility of feature X j ;

[0044] S6: The central server obtains and retains features based on the selected threshold:

[0045]

[0046] Furthermore, the specific steps of the distributed false discovery rate control process include:

[0047] S1: Each client randomly shuffles the data it owns to obtain pseudo-features;

[0048] S2: Perform the steps described in claim 2 or claim 3 on the original features and pseudo-features respectively to obtain the utility values of the features and and calculate the new utility through ;

[0049] S3: The central server determines the threshold through ;

[0050] S4: The central server retains features based on the selected threshold:

[0051]

[0052] Another object of the present invention is to provide a high-dimensional data feature filtering system based on federated learning for label distribution drift adaptation, including:

[0053] A general feature filtering module for a general feature filtering framework: incorporating multiple existing feature filtering methods into a general framework and performing estimation and theoretical analysis in a similar manner;

[0054] A label drift feature filtering module for feature filtering methods in the label drift scenario: using special weights on the basis of the general framework to effectively cope with the influence of label drift;

[0055] A federated feature filtering module for the federated feature filtering estimation process: estimating the LR-FFS utility function and other methods within the general framework in a federated manner in an environment where label offset may exist;

[0056] An FDR control module for the distributed false discovery rate control process: based on a given FDR control level, obtaining new feature utilities by constructing pseudo-features and determining the feature filtering threshold level according to the FDR control relationship.

[0057] Another object of the present invention is to provide a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the high-dimensional data feature filtering method based on federated learning for label distribution drift adaptation.

[0058] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the high-dimensional data feature filtering method based on federated learning for label distribution drift adaptation.

[0059] Another object of the present invention is to provide an information data processing terminal for implementing the high-dimensional data feature filtering system based on federated learning for label distribution drift adaptation.

[0060] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are as follows:

[0061] First, the present invention provides a general federated feature filtering framework that integrates a large number of existing methods (such as CRU, MV-SIS, and CAVS) as special cases. It can not only uniformly analyze and implement these methods but also study their large-sample properties simultaneously. In particular, a new type of feature filtering method, label-shift robust federated feature screening (LR-FFS), is introduced, and a corresponding distributed program is proposed to accurately quantify the marginal importance of numerical features in classification problems.

[0062] The technical solution of the present invention includes four modules: First, this paper proposes a general framework to unify existing feature filtering methods, and introduces a new tool, label shift-robust federated feature filtering (LR-FFS), and its federated estimation process. This framework helps to systematically and uniformly analyze existing methods and effectively alleviate the impact of label shift. LR-FFS handles label shift by utilizing conditional distribution functions and expectations without increasing the computational burden. At the same time, it is robust to model misspecification and outliers. In addition, the federated process not only ensures computational efficiency and privacy protection, but also maintains a screening effect comparable to that of centralized processing methods while maintaining privacy. Finally, we provide a false positive rate (FDR) control method for federated feature filtering to ensure the reliability and accuracy of the results in the screening process. Experimental results and theoretical analysis show that LR-FFS performs well in different client environments, especially under challenges such as client category distribution, sample size, and missing category data, and can still provide efficient and accurate feature filtering.

[0063] Second, in the medical field, such as DNA and RNA sequencing, or in the financial field, such as credit assessment, LR-FFS can integrate multi-source data, break the data island phenomenon, complete feature filtering without sharing the original data, and retain effective features. This process can greatly reduce the complexity of data processing, computing resource consumption, data collection costs, hardware costs and maintenance costs, so that more small and medium-sized enterprises and research institutions can use high-dimensional data for analysis and decision-making, accelerate the implementation of technologies such as artificial intelligence in various industries, achieve cost reduction and efficiency improvement, and have high commercial value.

[0064] At present, the existing distributed or federated feature filtering is based on the assumption that different client data are independent and identically distributed, which cannot cope with the challenges brought by data heterogeneity, and the original excellent properties become invalid. The LR-FFS of the present invention effectively uses the properties of conditional distribution and conditional expectation, finds a statistic that is insensitive to category distribution, and based on this statistic, uniformly analyzes and corrects the deviations of a large number of existing feature filtering methods. In addition, we have also pioneered the process and theoretical guarantee for false discovery rate control in this scenario. The method fills the technical gaps in these two aspects and provides a new technical path and solution for this field.

[0065] The present invention performs well in the case of label drift, reduces the model bias caused by distribution differences, and improves the fairness and reliability of data analysis. More extreme, the LR-FFS proposed by the present invention is still effective when there is missing category data on the client or the sample size is small. It can ensure the interests of minority groups, overcome prejudice against disadvantaged groups, and ensure algorithmic fairness. Brief Description of the Drawings

[0066] Figure 1 is a flowchart of a high-dimensional data feature filtering method based on federated learning with adaptive label distribution drift provided by an embodiment of the present invention;

[0067] Figure 2 is a block diagram of a high-dimensional data feature filtering system based on federated learning with adaptive label distribution drift provided by an embodiment of the present invention;

[0068] Figure 3 is the simulation result in setting (B) provided by an embodiment of the present invention, where the proportion of each category between different clients follows a Dirichlet distribution. The first row shows the successful screening rate (SSR), and the second row shows the logarithmic ranking of the weakest relevant feature (log(wRank));

[0069] Figure 4 is the simulation results of settings (E) and (F) provided by an embodiment of the present invention. The first column shows the successful screening rate (SSR), and the second row shows the logarithmic ranking of the weakest relevant feature (log(wRank));

[0070] Figure 5 is the classification accuracy of different filtering methods in the TCGA example provided by an embodiment of the present invention (based on KNN classification);

[0071] Figure 6 is a schematic flow diagram of traditional feature filtering provided by an embodiment of the present invention;

[0072] Figure 7 is a schematic flow diagram of the federated feature filtering of the present invention provided by an embodiment of the present invention;

[0073] Figure 8 is a schematic flow diagram of the federated feature filtering under the general framework of the present invention provided by an embodiment of the present invention;

[0074] Figure 9 is a schematic flow diagram of the actual federated feature filtering of the present invention provided by an embodiment of the present invention;

[0075] Figure 10 is a schematic flow diagram of the distributed false discovery rate control of the present invention provided by an embodiment of the present invention. Detailed Description of the Invention

[0076] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0077] As Figure 1As shown in the figure, a high-dimensional data feature filtering method based on federated learning for adapting to label distribution drift provided by an embodiment of the present invention includes the following steps:

[0078] S101, General feature filtering framework: Incorporate various existing feature filtering methods into a general framework and perform estimation and theoretical analysis in a similar manner.

[0079] S102, Feature filtering method for label drift scenario: Adopt special weights on the basis of the general framework to effectively cope with the impact of label drift.

[0080] S103, Federated feature filtering estimation process: Estimate the LR-FFS utility function and other methods within the general framework in a federated manner in an environment where label offset may exist.

[0081] S104, Distributed false discovery rate control process: Based on a given FDR control level, obtain the new feature utility by constructing pseudo-features, and determine the feature filtering threshold level according to the FDR control relationship.

[0082] The specific steps of the federated feature filtering estimation process for the general framework provided by the embodiment of the present invention include:

[0083] S1: The client traverses each category and estimates the coefficient ζ r :

[0084] S1.1: Traverse each client and category r = 1, …, R. On the l-th client, the proportion of category r can be estimated as

[0085] S1.2: The client uploads the category proportion and the sample size to the central server, that is

[0086] S1.3: The central server estimates the overall category proportion through and estimates ζ through r ;

[0087] S2: Expand each category into multiple terms by means of binomial expansion:

[0088]

[0089] S3: Among them is the expression to be estimated. Traverse d1 = 1, …, d and split into two component functions and estimate them separately, that is Among them

[0090] ​

[0091] S4: Construction estimation and the symmetric kernel of the U - statistic, denoted by and ;

[0092] S5: Traverse each client for r = 1, …, R. On the l - th client, and can be estimated as:

[0093]

[0094] where is the combination of d + 1 elements in the data segment ;

[0095] S6: Each client uploads and the sample size n l to the central server;

[0096] S7: The central server aggregates the parameters through to obtain the estimate of where where

[0097] S8: The central server obtains the estimate of ω through j,r,d ;

[0098] S9: The central server obtains the utility of each feature through ;

[0099] S10: For a given filtering threshold γ, we choose to retain the features

[0100]

[0101] The specific steps of the federated feature filtering estimation process for LR - FFS provided by the embodiments of the present invention include:

[0102] S1: Traverse each client and feature in parallel;

[0103] S2: The client traverses each category of data and calculates the corresponding and the aggregation weight α l,r ;

[0104] S3: The client transmits to the central server;

[0105] S4: The central server traverses each category of data and weights through ;

[0106] S5: Central server calculation Take as feature X j utility;

[0107] S6: The central server obtains Retain features based on the selected threshold:

[0108]

[0109] The specific steps of the distributed error discovery rate control process provided by the embodiments of the present invention include:

[0110] S1: Each client randomly shuffles the data it owns to obtain pseudo-features;

[0111] S2: Perform the steps described in claim 2 or claim 3 on the original features and pseudo-features respectively to obtain the utility values of the features and and calculate the new utility through calculate new utility

[0112] S3: The central server determines the threshold through determine the threshold

[0113] S4: The central server retains features based on the selected threshold:

[0114]

[0115] As Figure 2 shown, a high-dimensional data feature filtering system based on federated learning for label distribution drift adaptation provided by the embodiments of the present invention includes:

[0116] A general feature filtering module for the framework of general feature filtering: Incorporate various existing feature filtering methods into the general framework and perform estimation and theoretical analysis in a similar manner;

[0117] A label drift feature filtering module for feature filtering methods in the label drift scenario: Adopt special weights on the basis of the general framework to effectively cope with the influence of label drift;

[0118] A federated feature filtering module for the federated feature filtering estimation process: Estimate the LR-FFS utility function and other methods within the general framework in a federated manner in an environment where label offset may exist;

[0119] The false discovery rate control module is used for the distributed false discovery rate control process: Based on the given FDR control level, new feature utilities are obtained by constructing pseudo-features, and the feature filtering threshold level is determined according to the FDR control relationship.

[0120] Another object of the present invention is to provide a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the high-dimensional data feature filtering method based on federated learning for label distribution drift adaptation.

[0121] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the high-dimensional data feature filtering method based on federated learning for label distribution drift adaptation.

[0122] Another object of the present invention is to provide an information data processing terminal, which is used to implement the high-dimensional data feature filtering system based on federated learning for label distribution drift adaptation.

[0123] The specific implementation of the present invention:

[0124] First of all, the present invention provides a general federated feature filtering framework, which integrates a large number of existing methods (such as CRU, MV-SIS, and CAVS) as special cases. It can not only uniformly analyze and implement these methods, but also study their large-sample properties at the same time. In particular, a new type of feature filtering method - label-shift robust federated feature screening (LR-FFS) is introduced, and a corresponding distributed program is proposed to accurately quantify the marginal importance of numerical features in classification problems.

[0125] Consider a classification problem. Let Y ∈ {y1, …, y R} be a classification response variable with R categories, and x = (X1, …, X p ) T be a variable containing p numerical features. Assume that the complete data set is naturally divided into m data segments Each data segment is located on a client and contains n l observations in (X, Y). These clients independently process the data respectively, and the total number of observations of all clients is and n l<<p. In addition, let F(Y|x) denote the conditional cumulative distribution function of Y given X, and let P(Y|X) denote the conditional probability density function of Y given X.

[0126] Allow for a certain degree of heterogeneity in the data distribution among clients, satisfying the label drift condition. That is, assume that the n l observations of the l-th client are from the joint distribution P l (X,Y) = P l (X|Y)P l (Y), where the subscript l represents the distribution function of the l-th client. Among them, the distribution function of the response variable (label), P l (Y), can vary among different clients, but the conditional probability density function, P l (X|Y), is the same among different clients and is denoted by P(X|Y).

[0127] Implementing the technical solution of the present invention includes four modules: a general feature filtering framework, a feature filtering method for the label drift scenario, a federated feature filtering estimation process, and a distributed false discovery rate control process. Before introducing the four modules, it is necessary to overview the traditional feature filtering process, and the specific implementation is as follows:

[0128] Traditional feature filtering process:

[0129] Feature filtering and feature selection methods are based on the sparsity assumption in high-dimensional problems, that is, only a few predictors or features (i.e., X) will cause the response (Y). Following this assumption, a large number of feature selection methods have been proposed in recent years, including LASSO, SCAD, MCP, etc. Although these methods have been successfully applied to many high-dimensional analysis fields, in fields such as genomics and protein analysis, with the development of modern technologies, the data dimension has increased exponentially. Limited by the inherent computational complexity properties, the application of these methods in ultra-high-dimensional problems faces great challenges. In view of this, feature filtering methods - that is, removing irrelevant features before formal data analysis - have gradually become the focus of attention.

[0130] Feature filtering methods are mainly divided into model-based methods and model-free methods. Model-based methods include Feature Annealing Independence Rule (FAIR) and Pairwise Sure Independence Screening (PSIS), which rely on specific models such as linear models, generalized linear models, non-parametric regression models, etc. Model-free methods include MV-SIS, Fused Kolmogorov Filter (FKF), and Class Adaptive Variable Screening (CAVS), which are usually based on mean and independence hypothesis tests to measure the marginal utility of features, such as two-sample t-statistic, Kolmogorov-Smirnov statistic, Cramér-von Mises distance, etc. This utility (or test statistic) reflects the strength of the relationship between the feature and the response variable. The larger the utility value, the stronger the correlation between the feature and the response variable, and the easier it is to reject the null hypothesis, thus confirming the importance of the feature.

[0131] Therefore, a classic non-distributed feature filtering process is as follows:

[0132] (1): Define the feature filtering scenario and select the corresponding method, and define the set (active set) related to the response variable:

[0133]

[0134] (2): For a given random sample {(X i , Y i ): 1 ≤ i ≤ n}, the utility of each feature can be estimated

[0135] (3): For a given filtering threshold γ, select the retained features

[0136]

[0137] Determining the threshold is usually difficult. Usually, the d features with the largest utility values are retained. The threshold and the number of retained features d are usually related to the computing power of the computing device, the total number of features, and some False Discovery Rate (FDR) control methods. However, these methods assume that all data is stored on a single computer, so they are not applicable to distributed scenarios with communication bottlenecks between clients.

[0138] To address this limitation, Li Xingxiang and Xu Chen (2020) pioneered a distributed feature filtering framework based on Aggregated Correlation Screening, enabling clients to calculate the utility without exchanging the original data. In step (2), the estimation of feature utility is carried out in a distributed manner, which further broadens the application scenarios of feature filtering. Specifically, the process is as follows:

[0139] (1): Assume Y and feature X jThe correlation measure between (i.e., X j ) can be expressed as:

[0140] ω j = g(θ j,1 , …, θ j,s ),

[0141] where g is some pre-determined function, and θ j,1 , …, θ j,s are s component parameter functions and can be unbiasedly estimated relatively easily. Assume is an unbiased estimator (kernel) of θ j,h and is symmetric;

[0142] (2): For each data segment estimate θ j,h through local U-statistics:

[0143]

[0144] where is the index set on the data segment ;

[0145] (3): Obtain the utility of X j through weighted aggregation:

[0146]

[0147] where

[0148] (4): For a given number of retained features d, select the model

[0149]

[0150] General feature filtering framework:

[0151] Define the utility function for each feature as:

[0152]

[0153] where is the utility value of the j-th feature in the class Y = y r . Here, d represents the order of the difference, k is the exponent, and ζ r is the weight parameter usually related to the proportion of Y = y r . Similar to the Kolmogorov-Smirnov distance, ω j,r,d quantifies the difference between Y = y r and Y ≠ y rWhether the samples are from the same distribution. Its core idea is similar to that of FKF. The small difference indicates that X j is independent of Y.

[0154] By adjusting the parameters k, d, and the weight ζ r , the general framework can include many existing utility methods. For example, the weight of CAVS is ζ r = 1 - P(Y = y r ), which focuses on the utility of the category with a relatively low proportion. In contrast, the weight used by CRU is the square of the variance of I(Y = y r ), which is more biased towards the category with a class proportion close to 0.5.

[0155] Regarding the order of the parameter d, MV - SIS explores the second - order difference (d = 2), while CRU and CAVS focus on the distribution of the first - order difference (d = 1). When d = 2, the U - statistic is used to estimate ω j,r,2 The computational complexity is usually O(N 3 p), which is much higher than estimating ω with a complexity of O(N 2 p) when d = 1. j,r,1 . For the case of d > 2, the computational burden will further increase. Therefore, to improve computational efficiency, the utility methods with d = 1 are focused on. For ease of explanation, the following content will denote ω j and ω j,r as the utility values using the first - order difference (d = 1).

[0156] Note that the existing methods are closely related to the class proportion. The distributed or federated feature filtering methods assume that the samples are independently and identically distributed among different clients. However, when the class proportions of different clients are different due to label shift, the impact of label shift becomes significant when using the original distributed estimation methods. This leads to a key question: in the case of label shift, how to ensure that the goals of different clients are consistent, that is, the utility is insensitive to the class proportion? To address this issue, a new feature filtering method is innovatively proposed.

[0157] Feature filtering method for the label drift scenario:

[0158] To better mitigate the impact of label shift, a special weight named label - shift - robust federated feature screening (LR - FFS) is adopted. Specifically, the utility formal definition of LR - FFS is as follows:

[0159]

[0160] The third equation is because

[0161] The robustness of LR-FFS to label shift can be explained from two perspectives: the choice of statistic and coefficient.

[0162] First, the statistic ω j is derived from the conditional expectation of the conditional distribution function, which helps to identify and mitigate the impact of label shift. Proposition 1 shows that under certain conditions, LR-FFS is robust to proportional changes in Y = y r . However, except for Y = y r , proportional changes in the remaining R - 1 classes may still affect the utility value. To address this issue, the remaining R - 1 classes are combined into a single class Y ≠ y r , and the distribution difference of X j under the conditions of Y = y r and Y ≠ y r is concerned.

[0163] Proposition 1:

[0164] When the proportions of other classes except Y = y r remain fixed among the remaining R - 1 classes, the expectation is independent of the proportion of Y = y r .

[0165] On the other hand, the coefficient weights of LR-FFS are independent of class proportions, thus minimizing the impact of label shift. As mentioned before, the weights in existing feature screening methods (such as CRU and CAVS) are functions of class proportions and are vulnerable to label shift. In contrast, LR-FFS naturally mitigates this vulnerability by ensuring that utility estimates are consistent across clients, regardless of whether the label distribution shifts.

[0166] It should be emphasized that the goal is to mitigate rather than completely eliminate the impact of label shift. This approach is consistent with the principle of client drift mitigation in classical federated learning, that is, by adjusting the local objective to make the local model more closely aligned with the global model. The scheme achieves a delicate balance between reducing the impact of label shift and maintaining estimation accuracy.

[0167] Federated Feature Filtering Estimation Process:

[0168] This section describes how to estimate the LR-FFS utility function federally in the presence of possible label shift. Recall the proposed LR-FFS utility, which is the key statistic. To estimate γ j,r , the One-Shot Aggregation (OSA) method is chosen to obtain a robust global estimate with minimal inter-machine communication cost.

[0169] γ can be decomposed into two component functions and estimated separately: j,r where

[0170]

[0171] First, estimate U j,r , at the l-th client, U j,r can be estimated by a binary U-statistic:

[0172]

[0173] When each client transmits to the central server, the server aggregates the uploaded by each client through weighted average where is the effective sample size of the l-th client. Note that due to label shift, is a biased estimate of U j,r , and it is impossible to obtain an unbiased estimate of U j,r based on local client data. Fortunately, LR-FFS can correct this bias and design a consistent estimator of γ j,r . Before discussing the details in depth, some additional notations need to be introduced. Let be the solution of the following equation:

[0174]

[0175] Meanwhile, let Note that Through some algebraic transformations, we can obtain:

[0176]

[0177] Therefore, if we can consistently estimate , we can correct the bias of and obtain a consistent estimator of γ ,r . Similar to the estimator , use a weighted U-statistic to estimate , that is where

[0178] It can be easily checked that Finally, define the estimator of γ j,r as

[0179] The estimation process of the general framework of federated feature filtering:

[0180] ​The previous section discussed how to federally estimate LR-FFS. For other methods within the framework, a similar approach can be taken for estimation. The key lies in how to estimate ζ r and ω j,r,d .

[0181] ζ r The estimation of ζ is straightforward. Note that ζ r is a function of the class proportion π r . Let's denote it as ζ r = g(π r ). The following steps are taken:

[0182] (1): For each client with r = 1, …, R, on the l-th client, the proportion of class r can be estimated as

[0183] (2): The client uploads the class proportion and the sample size to the central server, i.e.,

[0184] (3): The central server estimates the overall class proportion through and estimates ζ r ..

[0185] For ω j,r,d , only when d = 1, ω j,r,d can be split into and which avoids the introduction of interaction terms and reduces the estimation complexity. Among them, the estimation of is the same as the federated feature filtering estimation process. When d > 1, the estimation of ω j,r,d becomes relatively complex and involves the following steps:

[0186] (1): Through binomial expansion, ω j,r,d is split as follows:

[0187]

[0188] where is the expression to be estimated. When d1 = 0, it can be simply derived that

[0189] (2): For d1 = 1, …, d, is split into two component functions and estimated separately, i.e., where

[0190]

[0191] (3): Construct the estimates and The symmetric kernel of the U - statistic is denoted by and .

[0192] (4): For each client with \(r = 1,\ldots,R\), on the \(l\) - th client, and can be estimated as:

[0193]

[0194] where is the data segment the \(d\) + combinations of 1 element in;

[0195] (5): Each client uploads and the sample size \(n\) l to the central server;

[0196] (6): The central server aggregates the parameters through to obtain the estimate of where where

[0197] (7): The central server obtains the estimate of \(\omega\) through j,r,d ;

[0198] (8): The central server obtains the utility of each feature through .

[0199] Distributed false discovery rate control process:

[0200] The false discovery rate (FDR) is the proportion of the number of falsely rejected null hypotheses to the total number of rejected null hypotheses. It is one of the important indicators for evaluating model performance, especially valuable in fields such as industrial production and biomedicine, such as controlling the defective rate and false positive rate, etc. This subsection will explore how to more precisely control the FDR. Specifically, for each feature \(X\) j , the data is independently randomly shuffled (permuted) on each client to generate a "pseudo" feature \(X'\) j . Then, using the federated feature filtering process, the utility values of \(X\) j and \(X'\) j are calculated respectively, denoted as \(\omega\) j and \(\omega'\) j . Based on this, a new marginal utility \(\varphi\) j =\(\omega\) j -\(\omega\) j ' is defined to characterize \(X\) jThe relationship with the response variable Y.

[0201] The generated "pseudo" feature X' j and the original feature X j have the same distribution and are independent of Y. Therefore, its utility value ω' j should be close to zero. If X j is a relevant feature, then φ j should be significantly greater than zero; otherwise, its value should be close to zero, and the probability that φ j is positive or negative is approximately equal. This property is called the marginal symmetry property.

[0202] Given a threshold γ > 0, the set of retained features is defined as:

[0203]

[0204] Then The False Discovery Proportion (FDP) of can be obtained by the following formula:

[0205]

[0206] where is estimated through , a ∨ b = max{a, b}. Thus The FDR of can be obtained from .

[0207] In fact, since the irrelevant set is unknown, inspired by the marginal symmetry, there is

[0208]

[0209] Thus can be estimated through .

[0210] Therefore, for a pre-given FDR control level α, the threshold γ can be selected as

[0211]

[0212] The present invention proposes a general framework and an estimation process for LR-FFS. Their estimation processes are similar. Taking LR-FFS as an example, the algorithm steps are described. Using the form of separately estimating the two components of γj,r , according to Appendix Figure 6 , Figure 7 and Figure 8 , the algorithm steps include:

[0213]

[0214]

[0215] Through this process, an estimator of γ j,r can be obtained. However, this algorithm has some drawbacks. Specifically, the transmission may disclose client-specific class proportion information, thus raising privacy concerns. To enhance privacy protection, the estimator is modified to an equivalent form (see Appendix Figure 9 and Algorithm 2) to avoid transmitting sensitive information of the client. Proposition 2 proves the equivalence of these two algorithms.

[0216]

[0217] Proposition 2:

[0218] The aggregation in step S.(2) of Algorithm 2 is equivalent to separately estimating the numerators of the aggregation parameters and Specifically, the following relationship holds:

[0219] and

[0220] The present invention provides a feature filtering method in the label drift scenario. Combining the above technical solutions and the technical problems to be solved, the advantages and positive effects of the solution given by the present invention are as follows:

[0221] Robust to label drift: The utility value of LR-FFS is composed of the conditional distribution function and the conditional expectation, and can accurately identify the impact of label drift on utility estimation. This method is insensitive to the class proportion, and achieves a good balance between alleviating the impact of label drift and ensuring the statistical estimation efficiency, and has strong robustness.

[0222] Efficient computing power: In Algorithm 2, the computational complexity of each client is This complexity is comparable to the corresponding steps in the existing literature. Therefore, dealing with label offset does not introduce additional computational burden. Table 1 summarizes the comparison of the existing methods in terms of computational complexity, transmission cost, and robustness. Due to its simplicity and robustness, LR-FFS is particularly suitable for scenarios with large sample size N and high-dimensional data p.

[0223] Robust to outliers and noise: LR-FFS has no special restrictions on the model, so it is robust in the case of model misspecification. In addition, due to the properties of the distribution function, this utility can effectively absorb the impact of outliers and noise, and is suitable for scenarios with heavy-tailed distributions.

[0224] Robustness to malicious attacks on clients (Byzantine attacks): In a federated learning environment, in addition to outliers and noisy data, an extreme case is malicious client attacks. Although LR-FFS is not specifically designed to defend against such attacks, since its calculation of ω j,r does not weight based on class proportions during the process, thereby reducing the error impact brought by malicious clients. In addition, the aggregation characteristics of the estimators in Algorithm 2 support the adoption of more robust aggregation methods, such as the median of means or the mean of medians, etc., further enhancing the resistance to malicious attacks.

[0225] Table 1 Comparison of feature filtering methods for classification problems

[0226]

[0227] Provides theoretical guarantees for the proposed method. The theoretical guarantees determine the statistical efficiency relationship between distributed and centralized estimation, and also determine the nature of feature filtering. Before introducing the theoretical properties, some notations and conditional assumptions need to be given.

[0228] Define the heterogeneity degree coefficient as This coefficient effectively measures the heterogeneity degree of class r. When there is no heterogeneity, that is, When, When there is extreme heterogeneity and class r only exists in a single client, As decreases, the heterogeneity degree among clients increases.

[0229] Condition 1: Assume there exist three positive constants, b1, b2, and b3, such that and

[0230] Condition 2: Assume there exist a positive constant c > 0 and 0 ≤ κ < 1 / 2, such that holds.

[0231] Condition 3: Assume the number of classes satisfies R = O(N ξ ), where ξ > 0 satisfies

[0232] Condition 4: There exists a constant such that holds.

[0233] Condition 1 requires that the proportion of each category is neither too large nor too small, and at the same time imposes certain restrictions on the degree of label shift. This condition relaxes the independent and identically distributed (IID) requirement for client data. Even if some clients have less or even missing data in some categories, as long as the data of other clients is sufficient, this condition still holds. Conditions 2 and 3 are similar to the assumptions in traditional feature filtering literature. Condition 2 allows the magnitude of the minimum true signal to be N -κ . Condition 3 allows the number of response variable categories to gradually increase as the sample size N grows. Condition 4 ensures that at the overall level, relevant features and irrelevant features can be clearly distinguished.

[0234] Conclusion 1: Assume that Conditions 1 and 3 hold. For any positive constant c1 and category r = 1, …, R, there exists a positive constant c2 such that

[0235]

[0236] Conclusion 2: and The variance of can be expressed as:

[0237]

[0238] In particular, when Conditions 1 and 3 hold and m = O ( N ) , The mean squared error of has the following order:

[0239]

[0240] Conclusions 1 and 2 prove that even if the number of features grows exponentially with the sample size, when log(p) = O(N α )(where α ∈ (0, 1 - 2κ - 4ξ)), the estimator still has consistency. The constant c2 reflects the heterogeneity information of the category distribution and is positively correlated with . Its error bound matches the efficiency of traditional centralized feature filtering.

[0241] Conclusion 3 - LR-FFS Sure screening property: Continuing the notation of Conclusion 1, when the retained feature threshold γ = cN -η , for any constant c3 > 0, there exists c4 > 0 such that the following holds:

[0242]

[0243] In particular, when Condition 2 holds, there exists a positive constant c5 such that

[0244]

[0245] where is the true model size.

[0246] Conclusion 4 - LR-FFS Ranking consistency property: Continuing the notation of Conclusion 3, when Condition 4 holds, there exists a constant c6 such that the following holds:

[0247]

[0248] In Conclusions 3 and 4, ω j 's minimum signal strength satisfies the characteristic identifiability condition common in the literature. The method does not impose restrictions on the moment conditions of the features, so it is robust to heavy-tailed distributions. Compared with CRU, LR-FFS can adapt to the heterogeneity of response distributions among different clients. When the total sample size is large enough, LR-FFS can eliminate most irrelevant features with high probability and retain all relevant features, thus ensuring the sure screening property. Its convergence rate matches that of the centralized feature filtering method, demonstrating the efficiency of this distributed method.

[0249] When Condition 4 holds, there is a clear gap between the utility values of relevant and irrelevant features. A theoretical result stronger than the sure screening property is proven: when log(p) = o(N 1-2η-4ζ ), LR-FFS can rank all relevant features higher than irrelevant features with probability approaching 1 (Conclusion 4), thus ensuring the existence of an ideal threshold to distinguish relevant and irrelevant features.

[0250] Conclusion 5 - Sure screening property of the general framework: Assume the number of classes R is fixed, ζ r is a continuous function of the class proportions. When the retained feature threshold γ = cN -η and Condition 1 holds, for any constant c8 > 0, there exists c9 > 0 such that the following holds:

[0251]

[0252] In particular, when Condition 2 holds, there exists a positive constant c 10 , such that

[0253]

[0254] where is the true model size.

[0255] Conclusion 6 - General framework ranking consistency property: Continuing with the notations in Conclusion 5, when Condition 4 holds, there exists a constant c 11 , such that the following equation holds:

[0256]

[0257] As can be seen from Conclusion 5 and Conclusion 6, as d increases, the error bounds tend to become looser, and more sample sizes are needed to achieve a similar convergence rate.

[0258] Algorithms 1 and 2 discuss how to estimate the utility function in a distributed or federated manner. In addition, the present invention also proposes an algorithm for distributed FDR control. According to Appendix Figure 10 , the algorithm steps include:

[0259]

[0260] Use the permutation method to construct "pseudo" features that are independent of Y while maintaining the same distribution as the original features. A similar method is to construct knockoff features, which also ensures that the "pseudo" features are correlated with the original features and are exchangeable, which helps to better control FDR in feature filtering. However, this method cannot be directly applied to high-dimensional problems because it requires 2p < n. Liu et al. (2022) and Pang and Xia (2024) respectively generalize this method to high-dimensional scenarios in non-distributed and distributed feature filtering. Through two-step filtering, first reduce the number of features to d through preliminary screening, ensuring that knockoff features are constructed under the condition of satisfying 2d < n l to overcome the dimensionality constraint. However, this method will increase additional computational costs. Another significant drawback is that when the sample sizes of some clients are small, in order to satisfy the condition of 2d < min n l , many relevant features may be excluded, which may be counterproductive in practice. Conclusion 7 provides the theoretical properties regarding the estimation of the set of relevant features .

[0261] Conclusion 7 - For any defined if there exists a sequence c n → ∞, when (n, p) → ∞, and c n / p → 0 holds. Then for any α ∈ (0, 1), the threshold selected in Algorithm 3 and the corresponding selected feature set satisfy:

[0262]

[0263] Under relatively loose conditions, this method can effectively control the FDR at a given α level. The condition requires that the growth rate of p is faster than c n , and this requirement can usually be easily met in high-dimensional scenarios.

[0264] The present invention provides a model-free distributed or federated feature filtering algorithm, which effectively addresses the scenario of label drift. The invention includes four modules:

[0265] General feature filtering framework: Incorporating various existing feature filtering methods into a general framework allows for estimation and theoretical analysis in a similar manner.

[0266] Feature filtering method for label drift scenarios: Special weights are adopted on the basis of the general framework to effectively address the impact of label drift.

[0267] Process of federated feature filtering under the general framework:

[0268] (1): Traverse each client and feature in parallel to estimate the proportion of client categories;

[0269] (2): The client uploads the category proportion and sample size to the central server;

[0270] (3): The central server aggregates to estimate the overall category proportion and obtains an estimate of ζ r estimation.

[0271] (4): Through binomial expansion, split ω j,r,d :

[0272]

[0273] (5): Traverse d1 = 1,..., d, and split into two component functions and estimate them separately using U statistics;

[0274] (6): Each client uploads and the sample size n l to the central server, and the central server aggregates to obtain an estimate of ;

[0275] (7): The central server obtains an estimate of ω through j,r,d , and combines the estimate of ζ r obtained in step (3) to obtain the utility of each feature;

[0276] (8): The central server retains features based on the selected threshold.

[0277] Federal Feature Filtering Estimation Process: Estimating the LR-FFS utility function in an environment where label shift may exist. The steps (see Algorithm 2) include:

[0278] (1): Traverse each client and feature in parallel;

[0279] (2): Each client traverses each category of data and calculates the corresponding and aggregation weight α l,r ;

[0280] (3): The client uploads to the central server;

[0281] (4): The central server traverses each category of data and weights it through ;

[0282] (5): The central server calculates and takes as the utility of feature X j ;

[0283] (6): The central server obtains and retains features based on the selected threshold:

[0284]

[0285] Distributed False Discovery Rate Control Process. It includes the following processes:

[0286] (1): Each client randomly shuffles the data it owns to obtain pseudo-features;

[0287] (2): Apply Algorithm 2 to the original features and pseudo-features respectively to obtain the utility values of the features and and calculate the new utility through ;

[0288] (3): The central server determines the threshold through ;

[0289] (4): The central server retains features based on the selected threshold:

[0290]

[0291] An embodiment of the present invention provides an information data processing terminal, which is used to implement the feature filtering algorithm to reduce the calculation cost.

[0292] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the storage device to execute the feature filtering algorithm to reduce the preservation of redundant features.

[0293] An embodiment of the present invention provides a set of high-dimensional tabular numerical feature data, such as SNP, genomics, risk control data, etc., and the data executes the feature filtering algorithm to ensure that subsequent steps can be executed.

[0294] In the financial industry, accurately assessing credit risk is crucial. Overestimating the credit risk of users is not conducive to the good turnover of funds and weakens the economic effects of financial institutions and even the entire society. Underestimating the credit risk of users may lead to default risks, and more seriously, the chain reaction brings black swan events. In order to comprehensively characterize the credit characteristics of users, it is particularly important to collect user data from different institutions (banks, enterprises, etc.). The data of these institutions will naturally have the characteristics of high-dimensional and non-independent and identically distributed due to institutional attributes and scales. In order to establish a more comprehensive evaluation model, such as a binary classification model for whether to grant loans. Multiple banks or financial institutions act as clients, each having a large amount of data on the credit records, financial status, etc. of their customers. Under the LR-FFS method, feature filtering is jointly performed to eliminate redundant features that are useless for classification, and a new evaluation model is independently or jointly established based on the retained features. Compared with the results without feature filtering, LR-FFS protects the privacy of customers, reduces the computational complexity, eliminates redundant features, improves the accuracy and interpretability of model evaluation, and provides strong support for financial institutions to formulate reasonable credit policies.

[0295] Medical data is highly sensitive. On the one hand, usually due to its expensive technical costs and acquisition costs, such as SNP and DNA sequencing data, the data often exhibits the characteristics of small sample size and ultra-high dimension. On the other hand, the data of different hospitals often shows significant non-independent and identically distributed characteristics due to their respective specialties, hospital levels, geographical locations, etc. In order to analyze such data, it is crucial to jointly filter and preprocess the data by multiple hospitals. Multiple hospitals act as clients, each storing a large amount of medical records, imaging data, inspection reports, etc. of patients. Taking the analysis of the inducing factors of type 2 diabetes in genome-wide association studies as an example, hospitals upload the corresponding statistics of local single nucleotide polymorphism (SNP) data under the initiative of a specific project (central server), and through the integration of the central server, the correlation between each SNP and type 2 diabetes is obtained, and irrelevant features are eliminated. The whole process utilizes a larger amount of data, and compared with institutions only using their own data, it can obtain the importance of each SNP more comprehensively, accurately, and with less bias. Finally, it ensures the accuracy of downstream analysis and improves the accuracy of disease diagnosis and the precision of drug research and development.

[0296] Example 1: Simulated dataset

[0297] Different heterogeneous scenarios were constructed through numerical simulation, covering linear models, generalized linear models, and considering various influences such as data distribution (such as whether it is a heavy-tailed distribution and whether there is noise).

[0298] The first example considers a scenario with a mean difference. N (X, Y) samples were randomly and independently generated, where the categorical response variable (label) follows the distribution P(Y = r) = π r l , r = 1, …, R. For the r-th category, p = 10000 features were generated in the following way:

[0299] X = μ r + ε,

[0300] where μ r =(μ r1 , …, μ rp ) T is the mean parameter, and ε = (ε1, …, ε p ) T is a random noise vector. If a feature X j has no difference in mean, i.e., μ 1j = … = μ rj , it can be regarded as an irrelevant feature. Consider the following 4 different experimental settings:

[0301] Suppose there are 30 clients in total, and each client has a sample size of n l = 100. The different client category distributions P(Y = r) = π r l are affected by the distribution heterogeneity parameter α, and the noise vector ε follows the standard normal distribution, N(0, 1). Suppose R = 7, μ 1j = 0.34, 1 ≤ j ≤ 8. Then the set of relevant features is

[0302] where, to simulate the heterogeneity of the category distribution, on the l-th client, for each category r, a random number uniformly distributed in the interval (1, α) is generated By normalizing, the proportion of category r on client l is calculated As α increases, the degree of category distribution heterogeneity among different clients also increases.

[0303] Suppose there are 30 clients in total, and each client has a sample size of n l = 100. The different client category distributions follow the Dirichlet distribution controlled by the parameter α. The noise vector ε independently follows the t-distribution with 2 degrees of freedom. For R = 5, 6, 7, there is μ11 = … = μ 14 = μ 25 = … = μ 28 = 0.45, 0.47, 0.50. Then the set of relevant features is

[0304] Suppose there are a total of 16 clients, with four groups each having a sample size of 100, 200, 300, and 400, and the total sample size N = 4000. Set R = 8, the number of missing classes for each client ranges from 0 to 4, and the remaining classes have the same proportion. The noise term ε independently follows a standard lognormal distribution (i.e., log(ε) ~ N(0,1)). Set μ 1j = 0.32 for 1 ≤ j ≤ 10, and μ 2j = 0.08 for 1 ≤ j ≤ 10. The index set of relevant features is

[0305] This setting examines the effect of FDR control in Algorithm 3. The heterogeneity setting and sample size are the same as in Setting (A), assume R = 5, and for 1 ≤ j ≤ 8, μ 1j = 0.4, considering the heterogeneity coefficient α = 5.

[0306] In the above settings: (A) is the simplest feature filtering scenario. (B) assumes that the class distribution between different clients follows a Dirichlet distribution, and as α decreases, the heterogeneity increases. (C) considers the case of missing class labels and varying sample sizes between clients. (D) reports the FDR control effect under the threshold selection in the distributed false discovery rate control process.

[0307] The second example considers a scenario where there are differences in the probability density functions under a generalized linear model. Suppose there are R classes and N = 3000 samples are generated from a multinomial logistic regression model log(P(Y = 1|X)) ∝ Xβ + ι. Where β = (β1, …, β p ) T represents a vector containing p = 8000 regression coefficients, and ι is a constant. Set most of the elements in β to zero to ensure that only features with some non - zero coefficients affect the response variable. Consider the following two settings:

[0308] E.X ~ N(0, ∑1), where ∑1 is a p×p identity matrix. Set the index set of relevant features to be And for Set β j = (-1) W × 1, where W ~ Bernoulli(0.5), ι = -0.25.

[0309] In addition, set P(Y = 2|X) / 1.2 = P(Y = 3|X) = … = P(Y = R|X), and replace 30 samples with random noise that follows a uniform distribution on [0, 100]. The inter-client class heterogeneity follows a Dirichlet distribution.

[0310] F.X ∼ N(0, ∑2), where Σ2 = [σ j,h p×p , and its elements are defined as: σ j,j = 1 when |j - h| = 1 when |j - h| = 2 when |j - h| ≥ 3, σ j,h = 0. Set the index set of relevant features as and for set β j = (-1) W ×1.5, where W ∼ Bernoulli(0.5), ι = -0.2.

[0311] Similar to setting (E), adjust the class probabilities such that P(Y = 2|X) / 0.8 = P(Y = 3|X) = … = P(Y = R|X), and introduce noise. The setting of class heterogeneity follows setting (A).

[0312] For settings (E) and (F), among p = 8000 features, 50 features were retained.

[0313] For each setting, the LR-FFS method is applied for federated feature filtering. For comparison, existing classification-based utility methods are also used: CRU, PSIS, FKF, MV-SIS, and CAVS. These distributed algorithms are used for feature screening, based on the method of Li et al. (2022). To simulate the possible noise in the data, 50 samples are randomly selected from these clients, and all features are replaced with random numbers that follow a uniform distribution between 0 and 100. To obtain an appropriate threshold γ and ensure data privacy in settings (A) to (C), the strategy of Li and Xu (2024) is followed. First, a set of q = 1000 auxiliary features (Z1,..., Z q ) is created by permuting randomly selected features. Since these auxiliary features are independent of Y, set the threshold as where is the OSA estimate of the screening utility between Y and Z j .

[0314] The screening accuracy is evaluated by the success screening rate (SSR), positive selection rate (PSR), and FDR for T = 200 repeated experiments:

[0315] ​

[0316] Among them, represents the index set of the retained features in the t-th iteration. In addition, the average number of retained features after screening is reported i.e., the size of the feature, and the average rank (wRank) of the weakest relevant feature in each simulation. For each method, the average computation time (in seconds) required for the local machine to perform distributed screening is also reported. Table 2 shows the results of all performance metrics. For visualization, other charts focus on presenting the SSR and wRank results.

[0317] Table 2 Results of R = 7 in setting (A) (results after adding noise are shown in parentheses)

[0318]

[0319]

[0320] Among them, the upward arrow indicates that the higher the value, the better, and the downward arrow indicates that the lower the value, the better.

[0321] Figure 3 Simulation results in setting (B), where the proportion of each category among different clients follows the Dirichlet distribution. The first row shows the success screening rate (SSR), and the second row shows the logarithmic rank of the weakest relevant feature (log(wRank))

[0322] Table 3 Simulation results of setting (C), where each client only contains partial category data (results after adding noise are shown in parentheses)

[0323]

[0324]

[0325] Table 4 Simulation results of setting (D) (results with added noise are in parentheses)

[0326]

[0327]

[0328] Figure 4 Simulation results of settings (E) and (F), the first column shows the success screening rate (SSR), and the second row shows the logarithmic rank of the weakest relevant feature (log(wRank))

[0329] As can be seen from Table 2, MV-SIS and FKF perform poorly in solving the distributed feature screening problem and consume more time. Therefore, in other analyses, only the results of the PSIS, CRU, and CAVS methods are presented.

[0330] In setting (A), in the absence of outliers, PSIS shows robustness to label shift and achieves effective feature filtering. However, in settings (B) and (C), when the features exhibit heavy-tailed distributions or there are outliers, the performance of PSIS is close to random guessing, significantly reducing its effectiveness. LR-FFS consistently provides the best performance in all scenarios, especially performing excellently in settings with moderate client heterogeneity. Although CAVS has slightly worse performance than LR-FFS, it still shows certain advantages by using the maximum value as a special weight. For CRU, it is observed that label shift seriously affects the screening results, reducing its accuracy and reliability. Setting (D) shows the correct screening rates of different features and the FDR results at different FDR control levels. It can be seen that through Algorithm 3, all distributed feature filtering methods can effectively control FDR and reach an ideal level.

[0331] The classification problems in settings (E) and (F) are more complex, involving multiple classes, different class distributions, label shift, and the presence of noise. As expected, the correlation between features makes these two configurations face greater challenges in accurate feature filtering. Although the overall class distributions are significantly different, compared with settings (A) to (C), the impact of label shift on screening is relatively small. Nevertheless, in all these challenging scenarios, LR-FFS still shows better accuracy than other methods.

[0332] Example 2: TCGA dataset

[0333] The proposed method is applied to the Breast Invasive Carcinoma dataset, which contains comprehensive data of 981 patients from 38 institutions. This dataset is part of the PanCancer Atlas project and covers data on mutated genes, patient demographics, and tumor types. The dataset can be downloaded from the official website (pancanatlas).

[0334] The main goal is to develop a classifier for identifying breast cancer gene subtypes, which is regarded as a five-class classification task. Although the mRNA expression data of the dataset contains 20,531 features, due to the limited number of samples, this poses significant challenges to accurate classification. Additionally, due to the heterogeneity of the contribution of each institution to the subtype ratio, model training faces additional difficulties. Due to ethical and privacy issues, many institutions are unable to share the original data, which requires the use of a federated feature filtering method in a medical context to ensure data privacy and compliance while making full use of the complete cross-institutional dataset.

[0335] In model training, institutions with a minimum sample size of 32 were designated as clients in the training set, while other institutions served as the test set. Eventually, the training set contained 13 clients and 829 samples, and the test set contained 152 samples. The detailed sample sizes of each client are shown in Table 2.

[0336] Table 5 Sample sizes of each institution

[0337]

[0338]

[0339] In addition to analyzing the original dataset, experiments on noise pollution and attacks were also conducted. In the noise test, all features of 30 samples in the training set were replaced with random numbers drawn from a uniform distribution from 0 to 30. In the attack test, the labels of the samples of one client in the training set were randomly shuffled. Each test was repeated 50 times. To ensure the robustness of the results, the number of features with a utility value appearing more than 45 times among the top 100 features was reported.

[0340] The LR-FFS method was applied and compared with the CRU, PSIS, and CAVS methods for feature screening, and finally the key features were retained. Using the selected features, a K-nearest neighbor (KNN) classifier with K = 40 was trained. Meanwhile, the average accuracies of the distributed estimation and the aggregated data estimation on the test set in 50 repeated experiments were also reported. The results are shown in Figure 5 and Table 6. It can be seen that in addition to the relatively high average accuracy, LR-FFS can also retain similar features in multiple repeated filtrations, demonstrating the stability of the method.

[0341] Table 6 Number of instances where the number of features ranked among the top 100 in terms of importance exceeded 45 in 50 repeated trials

[0342]

[0343] Due to the outliers and noise inherent in medical data, PSIS performs poorly. Even when additional noise and attacks are added, LR-FFS always maintains superior feature screening performance and exhibits stability. In the comparison of aggregated data and federated screening results, PSIS always maintains consistency, while the other three methods show certain differences, with the smallest change in LR-FFS.

[0344] Through the comparative analysis of LR-FFS and other feature filtering methods in Example 1 and Example 2, the effectiveness and superiority of the proposed label drift robust federated feature filtering method of the present invention are verified. First, regarding the impact of label drift on utility estimation, the present invention can accurately identify the statistical bias caused by label shift and effectively handle it, avoiding the possible performance degradation of traditional methods when facing label drift. Compared with existing methods, the present invention successfully alleviates the challenges brought by client data heterogeneity while ensuring statistical estimation efficiency, and demonstrates excellent feature screening performance in different application scenarios. Second, the feature filtering method proposed by the present invention has significant computational advantages. Compared with other distributed feature filtering methods, the method of the present invention can effectively solve data heterogeneity without introducing additional computational burdens; at the same time, compared with traditional centralized processing methods, the method of the present invention adopts a "divide and conquer" strategy, significantly improving computational efficiency, and having a statistical efficiency similar to that of centralized estimation results, with higher practical application competitiveness. Finally, the proposed distributed FDR control process provides a solid theoretical support and practical basis for various distributed feature filtering methods. Through this process, the false discovery rate (FDR) can be effectively controlled in a multi-client heterogeneous environment, thus ensuring high reliability and accuracy of statistical inferences during the screening process.

[0345] It should be noted that the embodiments of the present invention can be implemented through hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.

[0346] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A high-dimensional data feature filtering method based on federated learning with adaptive label distribution drift, characterized in that: The following steps are involved: Step 1: General feature filtering framework: Incorporate multiple existing feature filtering methods into a general framework and use a similar approach for estimation and theoretical analysis; Step 2: Feature filtering method for label drift scenarios: special weights are used based on the general framework to effectively deal with the impact of label drift; Step 3, federated feature filter estimation process: estimate the LR-FFS utility function in a federated manner in an environment where label bias may exist, as well as other methods within the general framework; Step 4, distributed false discovery rate control process: Based on the given FDR control level, the new feature utility is obtained by constructing pseudo features, and the feature filtering threshold level is determined according to the FDR control relationship.

2. The high-dimensional data feature filtering method based on federated learning with adaptive label distribution drift as claimed in claim 1, characterized in that: The specific steps of the federated feature filtering estimation process for the general framework include: (1): The client traverses each category and estimates the coefficient ζ r : 1.1: Traverse each client and category r = 1, ..., R. On the lth client, the proportion of category r can be estimated as 1.2: The client uploads the category proportion and sample size to the central server, that is 1.3: Central server through Estimate the overall category share and pass Estimation r ; (2): Expand each category into multiple terms by binomial expansion: (3): Among them is the formula to be estimated, traversing d1=1,…,d, Split into two component functions and estimate them separately, namely in (4): Construction estimate and The U statistic of the symmetric kernel is and express; (5): Traverse each client and r=1,…,R, on the lth client, and can be estimated as: in For data segment The combination of d+1 elements in ; (6): Each client will and sample size n l Upload to the central server; (7): The central server passes Aggregate parameters to get Estimates in (8): The central server passes Get ω j,r,d estimates; (9): The central server passes Get the utility of each feature; (1) 0: For a given filtering threshold γ, select the retained features 3. The high-dimensional data feature filtering method based on federated learning with adaptive label distribution drift as claimed in claim 1, characterized in that: The specific steps of the federated feature filtering estimation process for LR-FFS include: (1): Traverse each client and feature in parallel; (2): The client traverses each category of data and calculates the corresponding and the aggregation weight α l,r ; (3): The client will Upload to the central server; (4): The central server traverses each category of data through weighting; (5): Central server computing Will As feature X j The utility of (6): The central server obtains Retain features based on a chosen threshold:

4. The high-dimensional data feature filtering method based on federated learning with adaptive label distribution drift as claimed in claim 1, characterized in that: The specific steps of the distributed error discovery rate control process include: (1): Each client randomly shuffles the data it has to obtain pseudo features; (2): Implement the steps of claim 2 or claim 3 on the original feature and the pseudo feature respectively to obtain the utility value of the feature and and through Calculating new utility (3): The central server passes Determine the threshold (4): The central server retains features based on the selected threshold:

5. A high-dimensional data feature filtering system with label distribution drift adaptation based on federated learning, which implements the high-dimensional data feature filtering method with label distribution drift adaptation based on federated learning as claimed in any one of claims 1 to 4, characterized in that: The label distribution drift adaptive high-dimensional data feature filtering system based on federated learning includes: A general feature filtering module, a framework for general feature filtering: multiple existing feature filtering methods are incorporated into the general framework, and similar methods are used for estimation and theoretical analysis; The label drift feature filtering module is used for feature filtering methods in label drift scenarios: special weights are used based on the general framework to effectively deal with the impact of label drift; The federated feature filtering module is used for the federated feature filtering estimation process: estimating the LR-FFS utility function in a federated manner in an environment where label drift may exist, as well as other methods within a general framework; The false discovery rate control module is used in the distributed false discovery rate control process: based on the given FDR control level, the new feature utility is obtained by constructing pseudo features, and the feature filtering threshold level is determined according to the FDR control relationship.

6. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the high-dimensional data feature filtering method based on federated learning with adaptive label distribution drift as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the label distribution drift adaptive high-dimensional data feature filtering method based on federated learning as described in any one of claims 1 to 4.

8. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the high-dimensional data feature filtering system with adaptive label distribution drift based on federated learning as described in claim 5.

Citation Information

Patent Citations

  • Federal learning privacy evaluation method under cross-domain heterogeneous scene

    CN115952507A

  • Privacy-preserving asynchronous federated learning of vertical partition data

    CN116034382A

  • Federal learning method and system based on difference perception collaboration

    CN118036766A

  • Reliable federated learning method and system based on AC-GAN and dynamic probability scheduling

    CN119167226A

  • High-robustness heterogeneous software defect prediction algorithm based on federated learning

    CN119312155A