Explanatable high-throughput multi-omics marker feature screening method

By adopting the random feature discarding mechanism and attention mechanism in multi-omics medical data analysis, combined with the exponential moving average coefficient, the "black box" characteristics and data missing problems of the deep learning model are solved, the interpretability and robustness of the model are improved, and more accurate predictions are achieved.

CN120808867APending Publication Date: 2025-10-17ZHEJIANG UNIV CITY COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510919852.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing deep learning models have "black box" characteristics and data missing problems in multi-omics medical data analysis, resulting in insufficient interpretability and robustness, making them difficult to be widely used in clinical practice.

Method used

A random feature dropout mechanism is combined with an attention mechanism and an exponential moving average coefficient to retain key biomarkers and enhance the interpretability and robustness of the model through dynamic nonlinear feature screening and weighted random masking.

Benefits of technology

It improves the interpretability and robustness of the model, reduces the risk of overfitting, can better adapt to incomplete data in the real world, and provide more accurate prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808867A_ABST
    Figure CN120808867A_ABST
Patent Text Reader

Abstract

The invention relates to an interpretable high-throughput multi-omics marker feature screening method, which comprises the following steps: acquiring a medical multi-omics data set, and acquiring an input tensor according to the medical multi-omics data set; performing random feature discarding on the input tensor; carrying out dynamic nonlinear feature screening based on attention on the processed tensor; and performing result prediction on the screened features through a decision model. The method has the beneficial effects that a random feature discarding mechanism for keeping marking features is adopted, so that the model is helped to adapt to common incomplete data in the real world, and the over-fitting risk can be reduced. In addition, the interpretability of the prediction result can be enhanced through attention-based dynamic nonlinear feature screening.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to an interpretable high-throughput multi-omic marker feature screening method. BACKGROUND

[0002] With the development of high-throughput biological technology and artificial intelligence, the role of multi-omic data in disease diagnosis, prognosis evaluation and personalized treatment is increasingly prominent. Multi-omics integrates multiple biological hierarchical information such as genomics, transcriptomics, proteomics, metabolomics and imagingomics, providing a comprehensive perspective for revealing the complex mechanism of diseases. Especially in the clinical research of major diseases such as cancer, diabetes and pneumonia, multi-omic modeling has been widely used in biomarker discovery, drug target identification, efficacy prediction and individualized intervention strategy development.

[0003] In recent years, deep learning (DL) methods have shown significant advantages in multi-omic data analysis due to their powerful non-linear modeling capabilities. However, DL models are often considered as "black boxes" and lack transparency and interpretability, which limits their credible deployment in the medical field. In particular, in decision-making scenarios involving patient health, the interpretability of the model is a key factor in gaining the trust of doctors and achieving clinical transformation. Therefore, developing a modeling framework with high performance and strong interpretability has become one of the core challenges in current multi-omic medical artificial intelligence research. To this end, a variety of explainable AI (XAI) methods have been introduced into the multi-omic field, such as SHAP, DeepSHAP, DeepLIFT, PermFIT and other post-hoc explanation techniques. Although these methods have certain universality and practicality, they often rely on complex computational processes or simplified approximation assumptions, making it difficult to accurately reflect the true behavior of deep models. In addition, some modeling methods based on attention mechanisms have built-in interpretability, but they mostly act on the intermediate hidden layers of the model, lacking direct correspondence with the original biological space, thus limiting their practical value in multi-omic applications.

[0004] On the other hand, the multi-omics data in the real world generally has a large number of missing values, which seriously affects the stability and generalization ability of the model. Traditionally, researchers use imputation methods (such as classic interpolation, GANs, diffusion models, etc.) to fill in the missing data. However, these methods perform poorly in high missing rate or small sample cases, easily introducing noise and exacerbating the risk of overfitting. Dropout, as a classic regularization method, simulates diverse missing patterns by randomly masking features, and performs well in improving model robustness. However, the standard Dropout mechanism lacks selective retention of key biomarkers during training, which may lead to the loss of important features and affect the learning effect and interpretability of the model. Therefore, how to design a new Dropout mechanism that can enhance model robustness and improve interpretability has become a technical problem that needs to be solved.

[0005] In summary, the existing technology still has many limitations in processing multi-omics medical data: on the one hand, the "black box" characteristics of deep learning models hinder their widespread application in clinical practice; on the other hand, traditional imputation and Dropout mechanisms cannot effectively deal with the contradiction between data missing and key feature retention. Therefore, it is urgent to propose a new modeling framework that can ensure model performance while considering interpretability, robustness and biological relevance to promote the application of multi-omics artificial intelligence in precision medicine. SUMMARY

[0006] The purpose of the present application is to overcome the deficiencies in the prior art and provide an interpretable high-throughput multi-omics marker feature screening method.

[0007] In a first aspect, an interpretable high-throughput multi-omics marker feature screening method is provided, comprising:

[0008] Step 1, obtaining a medical multi-omics dataset, and obtaining an input tensor according to the medical multi-omics dataset;

[0009] Step 2, randomly discarding features of the input tensor;

[0010] Step 3, performing dynamic nonlinear feature screening based on attention on the tensor processed in step 2;

[0011] Step 4, predicting the results of the features screened by the decision model.

[0012] As a preferred embodiment, step 2 comprises:

[0013] Step 2.1, determining the dropout probability of each feature in the input tensor;

[0014] Step 2.2, explicitly retaining part of the features with low dropout probability;

[0015] Step 2.3, weighted random masking is performed on the remaining features.

[0016] As preferred, step 3 comprises:

[0017] Step 3.1, a feedforward network is inputted with the tensor processed in step 2 to obtain a correlation coefficient tensor reflecting the importance of features;

[0018] Step 3.2, the tensor processed in step 2 is multiplied element-wise with the correlation coefficient tensor to obtain target features.

[0019] As preferred, in step 3.1, the feedforward network is a multilayer perceptron composed of multiple fully connected layers, each followed by a ReLU activation function, and finally ended by an output layer using a Tanh activation function.

[0020] As preferred, step 3 further comprises:

[0021] Step 3.3, an exponential moving average coefficient is calculated; the exponential moving average coefficient is used to reflect the importance of features changing over time.

[0022] As preferred, in step 2.1, the dropout probability of each feature in the input tensor is calculated according to the exponential moving average coefficient.

[0023] As preferred, in step 4, the decision model is constructed by a Transformer, ResNets or KANs.

[0024] In a second aspect, an interpretable high-throughput multi-omics marker feature screening system is provided for executing the method of any one of the first aspect, comprising:

[0025] An acquisition module is configured to acquire a medical multi-omics dataset and obtain an input tensor according to the medical multi-omics dataset;

[0026] A dropout module is configured to perform random feature dropout on the input tensor;

[0027] A screening module is configured to perform attention-based dynamic nonlinear feature screening on the tensor processed by the dropout module;

[0028] A prediction module is configured to perform result prediction on the features screened by the decision model.

[0029] In a third aspect, a computer storage medium is provided, and the computer storage medium stores a computer program; when the computer program runs on a computer, the computer program causes the computer to execute the method of any one of the first aspect.

[0030] In a fourth aspect, an electronic device is provided, comprising:

[0031] a memory for storing a computer program;

[0032] a processor for executing the computer program to implement the method according to any one of the first aspect.

[0033] The present application has the following advantages:

[0034] 1. The present application adopts a random feature dropping mechanism with label feature preservation, which not only helps the model adapt to incomplete data commonly seen in the real world, but also reduces the risk of overfitting.

[0035] 2. The present application can enhance the interpretability of the prediction results through attention-based dynamic nonlinear feature screening.

[0036] 3. The present application forms a bidirectional optimization mechanism through the cooperation of the exponential moving average coefficient and the random feature dropping mechanism. The exponential moving average coefficient is fed back to the random feature dropping mechanism to guide the dropping of irrelevant features during training; on the other hand, since the random feature dropping mechanism retains important features and drops irrelevant features, it speeds up the convergence speed in learning meaningful feature weights. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 is a flowchart of an interpretable high-throughput multi-omics label feature screening method provided by the present application. DETAILED DESCRIPTION

[0038] The present application will be further described below in conjunction with examples. The following examples are only used to help understand the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

[0039] Example 1

[0040] To solve the problems of the prior art, the present application provides an interpretable high-throughput multi-omics label feature screening method, as shown in Figure 1 , which comprises:

[0041] Step 1, obtaining a medical multi-omics data set, and obtaining an input tensor according to the medical multi-omics data set.

[0042] Step 2, randomly dropping features of the input tensor.

[0043] The application provides a specially designed Dropout mechanism in step 2, which aims to simulate missing values in biomedical high-dimensional feature data while retaining key biomarkers.

[0044] Specifically, step 2 includes:

[0045] Step 2.1, determining the dropout probability of each feature in the input tensor.

[0046] For example, for an input tensor x ∈ R B×D , where B is the batch size and D is the number of features, the application dynamically calculates the dropout probability p = [p1, …, p D ] ∈ [0, 1] D for each feature in step 2.1.

[0047] Step 2.2, explicitly retaining partial features with low dropout probability.

[0048] Specifically, the top features with the lowest dropout probability will be explicitly retained to ensure that they are not masked.

[0049] Step 2.3, weighted random masking for remaining features.

[0050] During forward propagation, each sample undergoes a weighted random masking process to simulate missing values. Specifically, for each sample i, a mask ratio r i is uniformly selected from the interval [0, 1-γ], where γ ∈ [0, 1] controls the proportion of biomarkers that need to be preferentially retained. The number of features that need to be masked is Then, based on the sampling probability p, a multinomial distribution is constructed, and n mask positions are randomly selected based on the weight to perform masking operations. Finally, the masked positions are replaced with a predefined fill value v fill (default -1).

[0051] In addition, in order to alleviate the problem of too large numerical range difference between different features, the feature values will be normalized to the interval [0, 1] before step 2.

[0052] Step 3, performing attention-based dynamic nonlinear feature selection on the tensor processed in step 2.

[0053] Step 4, predicting the results of the selected features by a decision model.

[0054] Embodiment 2:

[0055] Based on embodiment 1, the application provides a more specific and interpretable high-throughput multi-omics marker feature selection method, which includes:

[0056] Step 1, obtaining a medical multi-omics dataset, and obtaining an input tensor according to the medical multi-omics dataset.

[0057] Step 2, performing random feature dropout on the input tensor.

[0058] In step 2, the dropout probability of each feature in the input tensor is calculated according to the exponential moving average coefficient obtained in subsequent step 3. For example, if the EMA feedback w t = [w1, …, w D ] ∈ R D , the dropout probability is calculated as follows:

[0059]

[0060] This calculation ensures that the more important the feature, the lower the probability of being discarded.

[0061] Step 3, performing attention-based dynamic nonlinear feature screening on the tensor processed in step 2.

[0062] Step 3 is a learnable and lightweight feature selection mechanism, which aims to adaptively assign importance weights to input features according to their relevance to downstream tasks.

[0063] Specifically, step 3 includes:

[0064] Step 3.1, inputting the tensor processed in step 2 into a feedforward network to obtain a correlation coefficient tensor reflecting feature importance.

[0065] For example, input the tensor to step 3, which generates a correlation coefficient tensor c ∈ R B×D , where each row is a vector whose values represent the relative importance of each feature in the corresponding sample. This calculation can be formalized as follows:

[0066]

[0067] where f θ represents a feedforward network, which is implemented as a multi-layer perceptron (MLP) controlled by parameters θ, consisting of multiple fully connected layers, each followed by a ReLU activation function, and finally ending with an output layer using a Tanh activation function. This mapping converts the input into a normalized weight matrix c ∈ R B×D reflecting the relative importance of each feature.

[0068] Step 3.2, element-wise multiplication of the tensor processed in step 2 and the correlation coefficient tensor to obtain the target feature.

[0069] Specifically, the selected feature is calculated by the following method:

[0070]

[0071] Where ⊙ represents element-wise multiplication operation. Through this operation, non-important features are suppressed, and significant features are emphasized. Then it is sent to the downstream decision step for result prediction.

[0072] Step 3.3, calculate the exponential moving average coefficient; the exponential moving average coefficient is used to reflect the importance of the feature changing over time.

[0073] Specifically, the calculation formula of the exponential moving average (EMA) coefficient is:

[0074] w t =α·w t-1 +(1-α)·E(c)

[0075] Where E(c) represents the batch mean of the coefficient. Alpha is an artificially given weight coefficient (value between 0-1, default 0.99), which controls the EMA update speed. This EMA provides a smooth estimate of the importance of the feature changing over time.

[0076] Step 4, the selected features are predicted by the decision model.

[0077] In step 4, the decision model is constructed by Transformer, ResNets or KANs.

[0078] It should be noted that the same or similar parts in this embodiment as in embodiment 1 can be mutually referenced, and will not be repeated in this application.

[0079] Embodiment 3:

[0080] Based on embodiment 2, the application embodiment 3 provides an interpretable high-throughput multi-omics marker feature screening system, comprising:

[0081] An acquisition module is configured to acquire a medical multi-omics dataset and acquire an input tensor based on the medical multi-omics dataset.

[0082] A discard module is configured to perform random feature discarding on the input tensor.

[0083] A screening module is configured to perform attention-based dynamic nonlinear feature screening on the tensor processed by the discard module.

[0084] a prediction module configured to predict the result of the screened features by the decision model.

[0085] During the training process, the dropout module randomly discards part of the feature values of each sample. Then, the screening module processes these samples to generate weight coefficients for screening features, and then inputs the screened features into the prediction module for prediction. These weight coefficients are recorded by exponential moving average (EMA) and fed back to the dropout module to guide the dropout module to discard unimportant features and retain marker features. Specifically, in order to retain key biomarkers during feature dropout, the model sorts the features in descending order of importance, and retains the top k features according to a ratio threshold γ (default 20%), and the remaining features are randomly discarded. This mechanism not only helps the model adapt to incomplete data commonly seen in the real world, but also reduces the risk of overfitting.

[0086] In the test phase, the dropout module will be discarded directly, and the EMA coefficient is fixed. By sorting the EMA coefficient, features with statistical significance can be identified. At the same time, the individualized coefficient generated by the screening module can also explain each prediction result and provide interpretability beyond traditional statistical significance. For example, by obtaining the feature weight coefficient of each individual through the feature screening module, and sorting these coefficients from large to small, it can be known which features play a key role in the decision of the individual.

[0087] In addition, the screening module can dynamically learn important features while suppressing irrelevant features without any modification to the downstream model structure. It serves as a differentiable, end-to-end trainable feature selector suitable for high-dimensional multi-omics data modeling. More importantly, it works with the dropout module to form a bidirectional optimization mechanism. Specifically, EMA is fed back to the dropout module to guide the discarding of irrelevant features during training; on the other hand, since the dropout module retains important features and discards irrelevant features, it speeds up the convergence speed of the screening module in learning meaningful feature weights.

[0088] It should be noted that the system provided in this embodiment corresponds to the method provided in Embodiment 2, so in this embodiment, the same or similar parts as Embodiment 2 can be mutually referenced, and will not be repeated in this application.

Claims

1. An interpretable high-throughput multi-omics signature screening method, characterized in that: include: Step 1: Obtain a medical multi-omics dataset, and obtain an input tensor based on the medical multi-omics dataset; Step 2: Perform random feature discarding on the input tensor; Step 3: Perform attention-based dynamic nonlinear feature screening on the tensor processed in step 2; Step 4: Use the decision model to predict the results of the screened features.

2. The interpretable high-throughput multi-omics marker feature screening method according to claim 1, characterized in that: Step 2 includes: Step 2.

1. Determine the drop probability for each feature in the input tensor. Step 2.2, explicitly retain some features with low discard probability; Step 2.3: Perform weighted random masking on the remaining features.

3. The interpretable high-throughput multi-omics marker feature screening method according to claim 2, characterized in that: Step 3 includes: Step 3.1: Input the tensor processed in step 2 into the feedforward network to obtain the correlation coefficient tensor reflecting the importance of the feature; Step 3.2: Multiply the tensor processed in step 2 by the correlation coefficient tensor element by element to obtain the target feature.

4. The interpretable high-throughput multi-omics signature screening method according to claim 3, characterized in that: In step 3.1, the feedforward network is a multi-layer perceptron, which is composed of multiple fully connected layers, each of which is followed by a ReLU activation function and finally ends with an output layer using a Tanh activation function.

5. The interpretable high-throughput multi-omics signature screening method according to claim 4, characterized in that: Step 3 also includes: Step 3.3: Calculate the exponential moving average coefficient; the exponential moving average coefficient is used to reflect the feature importance that changes over time.

6. The interpretable high-throughput multi-omics signature screening method according to claim 5, characterized in that: In step 2.1, the drop probability of each feature in the input tensor is calculated based on the exponential moving average coefficient.

7. The interpretable high-throughput multi-omics signature screening method according to claim 6, characterized in that: In step 4, the decision model is constructed by Transformer, ResNets or KANs.

8. An interpretable high-throughput multi-omics signature screening system, characterized in that: Used to perform the method according to any one of claims 1 to 7, comprising: an acquisition module, configured to acquire a medical multi-omics dataset and obtain an input tensor based on the medical multi-omics dataset; A drop module, configured to perform random feature drop on the input tensor; The filtering module is used to perform attention-based dynamic nonlinear feature filtering on the tensor processed by the discard module; The prediction module is used to predict the results of the screened features through the decision model.

9. A computer storage medium, characterized in that The computer storage medium stores a computer program; when the computer program is run on a computer, the computer executes the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 7.