Adaptive screening method and device for input features and medium
By using an adaptive screening method, the minimum subset of covariates in high-dimensional data is determined, the center selection type is selected, and the marginal utility of covariate components and response variables is quantified. This solves the problems of accuracy and robustness in feature screening of high-dimensional data and optimizes the training effect and prediction performance of artificial intelligence models.
Patent Information
- Application Number
- CN202511309256.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for processing high-dimensional data suffer from low computational efficiency, low statistical accuracy, and limited ability to handle nonlinear and non-Gaussian distributions, which affects the training effect and prediction performance of artificial intelligence models.
An adaptive screening method is adopted, which determines the minimum subset of covariates, selects the center selection type, quantifies the marginal utility of covariate components and response variables, uses test statistics to rank and significance levels to determine the screening threshold, and dynamically adjusts the screening criteria to optimize the subset of covariates.
It improves the accuracy and robustness of feature screening for high-dimensional data, adapts to complex and ever-changing data scenarios, balances the retention and removal of effective and ineffective covariates, and ensures the scientific nature and reliability of the screening process.
Smart Images

Figure CN120822007A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, and medium for adaptively screening input features. Background Art
[0002] With the rapid development of artificial intelligence and big data technology, high-dimensional data has become increasingly common in various application scenarios. In practical applications, the number of covariates often far exceeds the sample size, which makes it difficult to establish the response variable. Y ∈ R With high-dimensional covariates x The relationship model between high-dimensional data and feature selection poses a huge challenge. Specifically, feature selection for high-dimensional data is not only related to computational efficiency, statistical accuracy, and algorithm consistency, but also directly affects the training results, predictive performance, and interpretability of downstream large models. In industry and academia, high-dimensional data processing and feature selection technology have become key links in optimizing the performance of artificial intelligence models. Among existing technologies, the earliest methods based on the Pearson correlation coefficient are only applicable to linear models and have very limited processing capabilities for nonlinear, non-Gaussian, or long-tail data. Subsequently, feature selection methods targeting conditional means and conditional distributions have continued to develop, such as generalized linear models, nonparametric additive models, and marginal utility measures based on Martingale Difference Correlation (MDC), Distance Correlation (DC), and Projection Correlation (PC). These methods have improved selection accuracy and robustness to a certain extent, but still have many shortcomings. Summary of the Invention
[0003] In order to solve the above problems, the present application proposes an adaptive screening method for input features, including: determining a minimum covariate subset that is informative about a response variable under multiple pre-set conditions in a high-dimensional scenario, and determining a center selection type based on the minimum covariate subset, the center selection type including center distribution selection, center mean selection and center quantile selection; determining multiple covariate components, quantifying the multiple covariate components to determine the marginal utilities between the multiple covariate components and the response variable, and screening the marginal utilities according to the center selection type; determining test statistics for the multiple covariate components, sorting the test statistics, determining a screening threshold based on the sorted test statistics, determining a minimum index based on a pre-set significance level to determine the proportion of invalid covariate components under the screening threshold, and determining a center selection index set based on the screening threshold.
[0004] In one example, the method further includes: for central distribution selection, using a distribution dependence measure to determine the degree of influence of the covariate component on the distribution shape of the response variable by comparing the conditional variance with the unconditional variance; for central mean selection, using a mean dependence measure to determine the explanatory power of the covariate component on the mean of the response variable by comparing the conditional mean square error with the overall variance; for central quantile selection, using a quantile dependence measure to determine the information contribution of the covariate component at the quantile level by comparing the conditional and unconditional variances of the quantile residual function.
[0005] In one example, the method further includes: determining the proportion of invalid covariate components selected by the test statistic under the screening threshold based on the marginal symmetry of the statistics corresponding to the invalid covariate components and the sorted test statistic; a pre-set significance level, determining a minimum index based on the significance level to determine the proportion of invalid covariate components after the statistics corresponding to the minimum index, thereby achieving adaptive adjustment corresponding to the screening threshold.
[0006] In one example, the method also includes: when determining that the covariate support set is open and convex, deriving a screening subset for central distribution selection from a variable subset selected based on central mean selection through power transform and Fourier transform; traversing multiple quantile levels within a pre-set interval, deriving a screening subset for central quantile selection at different quantile levels based on a variable subset selected based on central distribution selection, so as to realize information integration between different types of center selection.
[0007] In one example, the method further includes: determining a preset random sample, sorting the random sample according to the value of the covariate component, and dividing the sorted random sample into multiple slices of equal size; calculating sample estimation values corresponding to the multiple dependency metrics based on the slices, and calculating the test statistic corresponding to the covariate component based on the sample estimation value, so as to perform screening and threshold determination based on the test statistic.
[0008] In one example, the method further includes: determining a false discovery rate during the screening process, monitoring changes in the false discovery rate, and when it is detected that the false discovery rate is greater than a preset threshold, adjusting the screening threshold to determine a balance between retaining valid covariate components and eliminating invalid covariate components.
[0009] In one example, the method further includes: determining a derivation result when screening the subset, and performing a consistency check on the derivation result so that the derived subset is statistically consistent with a result of the original center selection type.
[0010] In one example, the method further includes: training the screened minimum covariate subset through a preset machine learning algorithm to determine the performance index of the minimum covariate subset; dynamically adjusting the screening criteria of the covariate component according to the performance index, and determining the center selection type to optimize the composition of the covariate subset.
[0011] On the other hand, the present application also proposes an adaptive screening device for input features, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the adaptive screening device for the input features can perform: determining a minimum covariate subset that is informative for the response variable under multiple pre-set conditions in a high-dimensional scenario, and determining a center selection type based on the minimum covariate subset, the center selection type including center distribution selection, center mean selection and center quantile selection; determining multiple covariate components, quantifying the multiple covariate components to determine the marginal utility between the multiple covariate components and the response variable, and screening the marginal utility according to the center selection type; determining the test statistics of the multiple covariate components, sorting the test statistics, determining a screening threshold based on the sorted test statistics, determining a minimum index based on a pre-set significance level to determine the proportion of invalid covariate components under the screening threshold, and determining a center selection index set based on the screening threshold.
[0012] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, which are configured to: determine a minimum covariate subset that is informative about a response variable under multiple pre-set conditions in a high-dimensional scenario, and determine a center selection type based on the minimum covariate subset, wherein the center selection type includes center distribution selection, center mean selection, and center quantile selection; determine multiple covariate components, quantify the multiple covariate components to determine the marginal utility between the multiple covariate components and the response variable, and screen the marginal utility based on the center selection type; determine the test statistics of the multiple covariate components, sort the test statistics, determine a screening threshold based on the sorted test statistics, determine a minimum index based on a pre-set significance level, determine the proportion of invalid covariate components under the screening threshold, and determine a center selection index set based on the screening threshold.
[0013] This application can accurately capture information at different levels by determining a variety of center selection types and using corresponding dependency metrics, effectively identifying the minimum covariate subset that is informative for the response variable, and improving screening accuracy. The center selection type can be flexibly switched according to different needs, and information integration between different types can be achieved through power transform, Fourier transform and quantile level traversal to adapt to complex and changing data scenarios. The screening threshold is determined by using the test statistic sorting and marginal symmetry to achieve adaptive adjustment. While controlling the false discovery rate, it balances the retention and elimination of effective and invalid covariate components. The screening process and results are guaranteed in multiple dimensions, such as monitoring the false discovery rate and verifying the consistency of the derivation results, to ensure the scientificity and reliability of the screening. This application dynamically adjusts the screening criteria, optimizes the composition of the covariate subset, and improves the effect of subsequent analysis or modeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 Schematic diagram of a flow chart of an adaptive screening method of input features in an embodiment of the present application; Figure 2 Schematic diagram of an adaptive screening device for input features in an embodiment of the present application. DETAILED DESCRIPTION
[0015] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0016] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0017] like Figure 1 As shown, in order to solve the above problems, an embodiment of the present application provides an adaptive screening method of input features, the method comprising: S101. Determine a minimum subset of covariates that is informative for a response variable under multiple pre-set conditions in a high-dimensional scenario, and determine a center selection type based on the minimum subset of covariates, wherein the center selection type includes center distribution selection, center mean selection, and center quantile selection.
[0018] First, we introduce the concept of "center selection" to formally define which predictors are informative in high- and ultra-high-dimensional settings. Center selection aims to find a subset containing the fewest covariates that captures a specific distribution characteristic of the response variable given all covariates. The study focuses on three aspects: the entire conditional distribution, the conditional mean, and the conditional quantile. For each aspect, we define and explore the existence and uniqueness of the corresponding center selection, and investigate the relationship between different types of center selection.
[0019] Specifically, center selection concepts include: conditional distribution screening (CDS), which screens for the smallest subset of covariates that influences the entire conditional distribution of the response variable Y under a given covariate x; conditional mean screening (CMS), which screens for the smallest subset of covariates that influences the conditional mean E(Y|x); and conditional quantile screening (CQS), which screens for the smallest subset of covariates that influences the conditional quantile. When the support of the covariate x is open and convex, CDS, CMS, and CQS all guarantee the existence and uniqueness of the final screening set. Regarding cross-type relations, assuming the support of x is open and convex, the CDS screening subset can be obtained from the CMS screening subset through power and Fourier transforms. Similarly, assuming the support of x is open and convex, the CQS screening subset can be obtained from the CDS screening subset by traversing all quantile levels within (0,1).
[0020] S102 : Determine a plurality of covariate components, quantify the plurality of covariate components to determine marginal utilities between the plurality of covariate components and a response variable, and screen the marginal utilities according to the center selection type.
[0021] Assumptions p The first dimension of the covariate x j The component is Xj ( j =1,2,…, p ). In the context of high-dimensional or ultra-high-dimensional data, for each covariate component Xj With the response variable Y quantify the marginal utility between them, thus laying the foundation for subsequent screening. Xj , calculate its Y The marginal utility under the condition is denoted as . The marginal utility can be a mean dependence measure, a quantile dependence measure, or a dependence measure covering the entire conditional distribution, which needs to be determined according to the type of center selection. Through this quantitative method, the relationship between each covariate component and the response variable can be converted into a comparable numerical value, thereby providing an operational indicator for screening under high-dimensional data. In addition, for the subsequent application of the false discovery rate (FDR) control method, each covariate component is also defined Xj The corresponding test statistic ,in γ For the time Y and Xj In independent case The convergence rate of j =1,2,…, p .
[0022] S103. Determine test statistics for the multiple covariate components, sort the test statistics, determine a screening threshold based on the sorted test statistics, determine a minimum index based on a preset significance level to determine the proportion of invalid covariate components under the screening threshold, and determine a center selection index set based on the screening threshold.
[0023] For the calculated p Test statistics , j =1,2,…, p Sort the values in ascending order , thus forming an ordered sequence. This facilitates subsequent threshold search for FDR control; on the other hand, it directly reflects the importance ranking of each covariate component. In a high-dimensional data environment, this sorting operation can map extremely high-dimensional data into a one-dimensional ordered space, providing an intuitive and effective basis for data-driven threshold control while ensuring that the marginal utility of each covariate component forms a highly comparable reference standard within the entire analytical system.
[0024] In one embodiment, after the sorting is completed, the screening threshold is determined by the FDR control mechanism. Specifically, for each sorted statistic , the false discovery rate (FDR) is approximately calculated as follows:
[0025] Among them, | | represents the absolute value, #{} represents the counting of set elements, Indicates taking the smaller one. By taking advantage of the marginal symmetry of the statistic corresponding to the invalid covariate component, the number of covariate components in the left tail and the right tail are ratio-calculated to approximate the proportion of selecting invalid covariate components under a given threshold.
[0026] Then, at the pre-specified significance level α Under the conditions, find the minimum index that meets the conditions j ∗, so that This process enables adaptive adjustment of the threshold, that is, while effectively controlling the false discovery rate, it retains as many effective covariate components that have a practical effect on the model as possible, while accurately eliminating those ineffective covariate components. This successfully solves the problem of difficult to accurately select thresholds in high-dimensional data scenarios, ensuring that the algorithm can still maintain robust and efficient operation when the number of samples is limited.
[0027] In one embodiment, the threshold value obtained above is Determine the following set of center selection indices: .
[0028] This index set represents the minimum subset of covariates that is informative about the response variable under given data conditions. Specifically, for central mean selection and central quantile selection, this index set can be output based on conditional mean correlation or conditional quantile correlation, respectively; while for central distribution selection, different index sets can be combined by multiple traversal operations based on mean selection or quantile selection, such as using power transforms, Fourier transforms, or traversing different quantile levels, to obtain information about the entire conditional distribution. With this approach, a unified treatment of different variable / feature selection types is achieved, while ensuring that the output results are theoretically provable and operationally feasible.
[0029] In one embodiment, in the proposed data-driven feature screening framework, to effectively identify informative covariate components in high-dimensional and ultra-high-dimensional data, a dependency measure is introduced as the first step in the framework's implementation. Next, three metrics that meet the assumptions of the screening framework are presented to facilitate practical application. These metrics, without requiring model assumptions, can capture distributional dependence, mean dependence, and quantile dependence, respectively, and are naturally compatible with the CDS, CMS, and CQS frameworks proposed in this invention.
[0030] Unlike methods that employ complex modeling assumptions, the core goal in designing dependency metrics is to construct statistical indicators that are theoretically provable, free of distributional assumptions, and estimable in slices, while also satisfying marginal symmetry and consistency requirements. To this end, the slicing method has been systematically improved to not only characterize dependencies at the distributional level but also to test for dependencies under mean and quantile conditions, thereby forming a unified measurement system.
[0031] Theoretically, the three proposed dependency measures are shown to satisfy marginal symmetry and consistency conditions, ensuring their statistical validity in high-dimensional feature screening and FDR control. Specifically, for uninformative covariate components, these statistics exhibit asymptotic symmetry around zero and converge steadily to their theoretical expectations as the sample size increases, effectively avoiding bias and overfitting risks.
[0032] In one embodiment, the following indicators are used in the distribution dependency test and measurement: :
[0033] Among them, μ is the probability measure, I is the indicator function, E and Var represent the expectation and variance of the random variable respectively, and Var represents the conditional variance. This measure measures the conditional variance by comparing the conditional variance with the unconditional variance. The impact on the Y distribution shape. If the two are equal, it means Independent of Y at the distribution level.
[0034] For mean dependence, the following measure is used :
[0035] Among them, this metric compares the conditional mean square error and the overall variance Var(Y), reflecting The ability to explain the mean of Y. When it is independent of Y at the mean level, the conditional mean square error is equal to the population variance and the metric value approaches zero.
[0036] In order to characterize the conditional dependence at different quantile levels, the following indicators are introduced :
[0037] Among them, the function , For Y Quantile. This measure accurately captures the variance of the quantile residual function under conditions and without conditions. The information contribution at a particular quantile level is particularly sensitive to heteroskedasticity and differences in the tail distribution.
[0038] In one embodiment, in actual calculation, the slice estimation strategy in statistics is used to obtain the sample estimation values of the above three metrics: , and thus calculate the statistic (j=1,2,…p). The specific steps of the slice estimation strategy are as follows: given a set of random samples of size n and covariate dimension p (j=1,2,…p). First, for any j, according to ,(i=1,2,…n) values for random samples Then, the sorted samples are divided into H slices of equal size, each slice of size c. For the sake of generality, assume that c ≥ 2 and n = Hc, where H∈N. For j = 1, 2, … p, the observation value in the hth slice is recorded as , where k=1,2,...,c and h=1,2,...,H.
[0039] like Figure 2 As shown, the embodiment of the present application further provides an adaptive screening device for input features, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the adaptive screening device of the input feature to perform: determining, in a high-dimensional scenario, a minimum covariate subset that is informative about a response variable under a plurality of pre-set conditions, and determining a center selection type according to the minimum covariate subset, wherein the center selection type includes center distribution selection, center mean selection, and center quantile selection; determining a plurality of covariate components, quantifying the plurality of covariate components to determine marginal utilities between the plurality of covariate components and a response variable, and screening the marginal utilities according to the center selection type; Determine test statistics for the multiple covariate components, sort the test statistics, determine a screening threshold based on the sorted test statistics, determine a minimum index based on a preset significance level to determine the proportion of invalid covariate components at the screening threshold, and determine a center selection index set based on the screening threshold.
[0040] The embodiment of the present application further provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to: determining, in a high-dimensional scenario, a minimum covariate subset that is informative about a response variable under a plurality of pre-set conditions, and determining a center selection type according to the minimum covariate subset, wherein the center selection type includes center distribution selection, center mean selection, and center quantile selection; determining a plurality of covariate components, quantifying the plurality of covariate components to determine marginal utilities between the plurality of covariate components and a response variable, and screening the marginal utilities according to the center selection type; Determine test statistics for the multiple covariate components, sort the test statistics, determine a screening threshold based on the sorted test statistics, determine a minimum index based on a preset significance level to determine the proportion of invalid covariate components at the screening threshold, and determine a center selection index set based on the screening threshold.
[0041] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0042] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0043] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0044] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0045] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0046] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0047] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0048] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0049] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0050] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0051] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. An adaptive screening method for input features, characterized in that: include: determining, in a high-dimensional scenario, a minimum covariate subset that is informative about a response variable under a plurality of pre-set conditions, and determining a center selection type according to the minimum covariate subset, wherein the center selection type includes center distribution selection, center mean selection, and center quantile selection; determining a plurality of covariate components, quantifying the plurality of covariate components to determine marginal utilities between the plurality of covariate components and a response variable, and screening the marginal utilities according to the center selection type; Determine test statistics for the multiple covariate components, sort the test statistics, determine a screening threshold based on the sorted test statistics, determine a minimum index based on a preset significance level to determine the proportion of invalid covariate components at the screening threshold, and determine a center selection index set based on the screening threshold.
2. The method according to claim 1, characterized in that The method further comprises: For central distribution selection, a distribution dependence measure is used to determine the degree of influence of the covariate component on the distribution shape of the response variable by comparing the conditional variance with the unconditional variance; For the central mean selection, the mean dependence measure was used to determine the explanatory power of the covariate components on the mean of the response variable by comparing the conditional mean square error with the overall variance; For central quantile selection, the quantile dependence measure is used to determine the information contribution of the covariate components at the quantile level by comparing the variance of the quantile residual function under condition and without condition.
3. The method according to claim 1, characterized in that The method further comprises: determining, based on marginal symmetry of statistics corresponding to invalid covariate components and the sorted test statistics, a proportion of the test statistics selecting invalid covariate components under the screening threshold; A pre-set significance level is used to determine a minimum index according to the significance level to determine the proportion of invalid covariate components after the statistic corresponding to the minimum index, thereby achieving adaptive adjustment corresponding to the screening threshold.
4. The method according to claim 1, wherein The method further comprises: When the covariate support set is determined to be open and convex, the selected subset of variables selected by the central mean is derived through power transform and Fourier transform; Traversing multiple quantile levels within a preset interval, the variable subset selected based on the center distribution is used to derive the screening subset of center quantile selection at different quantile levels to achieve information integration between different types of center selection.
5. The method according to claim 2, characterized in that The method further comprises: Determining a preset random sample, sorting the random sample according to the value of the covariate component, and dividing the sorted random sample into a plurality of slices of equal size; Sample estimation values corresponding to the plurality of dependency metrics are calculated according to the slices, and test statistics corresponding to the covariate components are calculated according to the sample estimation values, so as to perform screening and threshold determination according to the test statistics.
6. The method according to claim 1, characterized in that The method further comprises: During the screening process, the false discovery rate is determined, and changes in the false discovery rate are monitored. When it is detected that the false discovery rate is greater than a preset threshold, the screening threshold is adjusted to determine a balance between retaining valid covariate components and eliminating invalid covariate components.
7. The method according to claim 1, characterized in that The method further comprises: Determine the derivation result when screening the subset, and perform consistency check on the derivation result to ensure that the derived subset is statistically consistent with the result of the original center selection type.
8. The method according to claim 1, characterized in that The method further comprises: Training the screened minimum covariate subset using a preset machine learning algorithm to determine a performance indicator of the minimum covariate subset; The screening criteria of the covariate components are dynamically adjusted according to the performance indicators, and the center selection type is determined to optimize the composition of the covariate subset.
9. An adaptive screening device for input features, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the adaptive screening device of the input feature to perform: determining, in a high-dimensional scenario, a minimum covariate subset that is informative about a response variable under a plurality of pre-set conditions, and determining a center selection type according to the minimum covariate subset, wherein the center selection type includes center distribution selection, center mean selection, and center quantile selection; determining a plurality of covariate components, quantifying the plurality of covariate components to determine marginal utilities between the plurality of covariate components and a response variable, and screening the marginal utilities according to the center selection type; Determine test statistics for the multiple covariate components, sort the test statistics, determine a screening threshold based on the sorted test statistics, determine a minimum index based on a preset significance level to determine the proportion of invalid covariate components at the screening threshold, and determine a center selection index set based on the screening threshold.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: determining, in a high-dimensional scenario, a minimum covariate subset that is informative about a response variable under a plurality of pre-set conditions, and determining a center selection type according to the minimum covariate subset, wherein the center selection type includes center distribution selection, center mean selection, and center quantile selection; determining a plurality of covariate components, quantifying the plurality of covariate components to determine marginal utilities between the plurality of covariate components and a response variable, and screening the marginal utilities according to the center selection type; Determine test statistics for the multiple covariate components, sort the test statistics, determine a screening threshold based on the sorted test statistics, determine a minimum index based on a preset significance level to determine the proportion of invalid covariate components at the screening threshold, and determine a center selection index set based on the screening threshold.
Citation Information
Patent Citations
Feature screening method and device, electronic equipment and storage medium
CN117932287A
Feature screening method and device, storage medium and electronic equipment
CN120372237A
Distribution Wide Estimated Risk Scoring to Decrease the Probability of Covariate Imbalances Adversely Affecting Randomized Trial Outcomes
US20130211805A1
Data model processing in machine learning using a reduced set of features
US20210342735A1