Decision rule algorithm with biological interpretability based on different types of association relationships among molecules

By constructing a decision rule algorithm based on different types of intermolecular associations, this system analyzes the associations between molecular features in breast cancer, solving the problem of integrating genomics and metabolomics data in existing technologies. This improves the accuracy of early diagnosis and personalized treatment of breast cancer and provides easily interpretable decision rules and biomarker screening.

CN120977378APending Publication Date: 2025-11-18ANSHAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511015592.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate genomics and metabolomics data to identify accurate, simple, and easily interpretable decision rules for the early diagnosis and personalized treatment of breast cancer, and fail to fully understand the changes in molecular relationships during cancer development.

Method used

We construct a decision rule algorithm based on different types of intermolecular associations. Through systematic analysis of positive linear, negative linear and nonlinear associations, we use the joint probability density function to screen out key molecular feature pairs with biological interpretability, and construct decision rules to reflect the development process of breast cancer, eliminating individual differences and overfitting problems.

Benefits of technology

It improves the accuracy of early diagnosis and treatment of breast cancer, provides clinical application value that is easily interpreted biologically, screens out biomarkers with strong predictive ability, and helps in the study of the pathogenesis of breast cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977378A_ABST
    Figure CN120977378A_ABST
Patent Text Reader

Abstract

The invention relates to a decision rule algorithm with biological interpretability based on different types of association relationships among molecules, which comprises the following steps of: inputting omics data of breast cancer, and comprehensively analyzing the omics data by fully utilizing positive linear association relationships, negative linear association relationships and non-linear association relationships among the molecules; molecular feature pairs (fi, fj) with strong predictive ability are screened by utilizing a joint probability density function delta ij to serve as clinical management markers of the breast cancer, and accurate, simple and easy-to-explain decision rules are constructed based on the selected markers to guide early diagnosis and personalized treatment of the breast cancer; according to the method, the influence of individual differences on data analysis is effectively eliminated, and the problem of overfitting caused by a small number of samples or high complexity of molecular expression data can also be effectively solved, so that the constructed decision rule can reflect the change of an intermolecular association relationship more truly and effectively; therefore, clinical early diagnosis of the breast cancer and improvement of personalized treatment effects are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biological data analysis, and particularly relates to a decision rule algorithm with biological interpretability based on different types of intermolecular association relationships. BACKGROUND

[0002] Breast cancer is one of the most common malignant tumors, which seriously threatens the life and health of female siblings. In 86% of countries worldwide, breast cancer ranks first in the incidence rate among all female tumor patients; and in 173 countries, it ranks among the top two in the mortality rate of all female tumor patients. When diagnosed with breast cancer, 32% of patients are in the advanced stage, which has missed the best treatment time, thus is not conducive to achieving good clinical treatment effect, and even seriously affects the quality of life of patients. According to statistics, by 2040, it is estimated that there will be 3 million new breast cancer patients and 1 million deaths worldwide each year. Although some traditional treatment methods are widely used in clinical practice of breast cancer, the recurrence rate and prognosis effect after treatment are still not ideal, which brings great mental pressure and economic burden to patients. Near-infrared technology NIR has the advantages of low cost, non-invasive and high efficiency, and is a new means of clinical breast cancer treatment. However, different patients have different metabolic responses to different near-infrared technology treatment schemes, and researchers are not very clear about the different metabolic activity changes of patients receiving different near-infrared technology. Therefore, there are still great challenges in the research of personalized treatment of breast cancer in clinical application.

[0003] The clinical application management of breast cancer, including early accurate diagnosis and personalized treatment, involves complex interactions among multiple factors, such as genes, metabolism and environmental factors, etc. In addition, due to the existence of individual differences, biochemical reaction activities such as gene regulation and metabolic phenotype are not the same in different patients, and this individual difference will directly affect the accuracy of breast cancer clinical data analysis research, and then affect the clinical effect of early diagnosis and drug treatment of breast cancer. In the life organism, the changes in the concentration and types of molecules caused by the disorder of key metabolic pathways such as non-essential amino acid metabolism, proline metabolism and glucose metabolism can induce tumor cells to use metabolic reprogramming biochemical reactions to meet the uncontrolled proliferation demand, thereby leading to the occurrence of cancer in the organism, in addition, the metabolic stress activities and other related metabolic mechanisms in tumor infiltrating immune cells directly affect the functional activity of immune cells and anti-tumor immune response, so that tumor cells can escape the monitoring of the immune system. Therefore, in-depth analysis of the differences in metabolic reaction activities during the occurrence and development of breast cancer and before and after treatment under different schemes can help better understand the pathogenesis of breast cancer and improve the effect of clinical early diagnosis, monitoring and personalized treatment.

[0004] During the development of cancer, the relevant metabolic reaction activities are realized by the interaction of different levels of molecules such as genes and metabolites. For example, abnormal changes in genes can cause disturbances in metabolic pathways closely related to cancer, thereby affecting the normal expression of the body's metabolic level and metabolites; and the disorder of metabolites will affect the epigenetic and interfere with a series of biochemical reaction activities such as transcriptional regulatory factors. Therefore, effectively integrating genomic data and metabolomic data and systematically and comprehensively mining biological information at different levels can help better promote clinical research on cancer risk prediction, patient diagnosis and precision medicine and other related issues. With the rapid development of high-resolution separation analysis technology, more and more high-dimensional genomic and metabolomic data provide a basis for establishing a scientific and advanced diagnostic platform. However, it is a great challenge to identify a small number of key molecules from these complex data to deduce accurate, efficient and easily interpretable decision rules for early diagnosis and personalized treatment, and these decision rules are crucial for clinical application management. Although some algorithms including support vector machine, deep learning and Transformer have high classification performance, they often evaluate the results of disease classification based on complex decision boundaries, and this decision boundary is a complex function constructed based on hundreds or even thousands of molecules, which is difficult to explain biologically, thus limiting their further application in early diagnosis and personalized treatment of diseases. Therefore, it is of great significance for clinical translational medicine application research to construct accurate, simple and easily biologically interpretable decision rules for accurate disease classification and early prediction of treatment effect.

[0005] In addition, the correlation between molecules in the living body is very complex, there are linear correlation and nonlinear correlation, and the linear correlation includes positive linear correlation and negative linear correlation. Therefore, it is not conducive to comprehensively understand the mechanism of cancer development and pathogenesis to explore the changes of the interaction between molecules in the process of cancer development from the perspective of single linear or nonlinear correlation. SUMMARY

[0006] The present application provides a decision rule algorithm with biological interpretability based on different types of intermolecular correlation, which systematically explores the dynamic changes of the interaction between molecular features in the development of breast cancer from the perspectives of positive linear correlation, negative linear correlation and nonlinear correlation between molecular features, and effectively eliminates the influence of individual differences on data analysis and the overfitting problem caused by the small number of samples or the high complexity of molecular expression data, so that the constructed decision rule can more truly and effectively reflect the changes of the correlation between molecular features in the development of breast cancer, and screen molecular feature pairs with strong prediction ability. i ,f jAs a clinical management marker of breast cancer, in addition, the decision rule proposed by the present application is easy to be explained in clinical biology, so that the screened marker has strong clinical application value.

[0007] In order to achieve the above purpose, the present application adopts the following technical solutions:

[0008] A decision rule algorithm with biological interpretability based on different types of correlation between molecules, comprising the following steps:

[0009] 1) input the omics data of different stages of breast cancer occurrence and development;

[0010] 2) use the function y(1)=ax to represent the positive linear correlation between molecular features f i and f j , wherein f i is equivalent to x, f j is equivalent to y(1), and a is a positive number;

[0011] Use the function y(2)=-b j x / a i +b j to represent the negative linear correlation between molecular features f i and f j , wherein f i is equivalent to x, and f j is equivalent to y(2);

[0012] Wherein, the coefficient a i is set to u i +δσ i , b j is set to u j +δσ j , the parameter u i is the mean of the molecular feature f i , the parameter σ i is the standard deviation of the molecular feature f i , u j is the mean of the molecular feature f j , the parameter σ j is the standard deviation of the molecular feature f j , and δ is a balance factor;

[0013] 3) the functions y(1)=ax and y(2)=-b j x / a i +b j divide the data space into four regions: the first region is y(1)≥ax and y(2)<-b j x / a i +b j; The second region is y(1) ≥ ax and y(2) ≥ -b j x / a i +b j ; The third region is y(1) < ax and y(2) ≥ -b j x / a i +b j ; The fourth region is y(1) < ax and y(2) < -b j x / a i +b j , Therefore, the combined functions y(1) and y(2) are used to represent both the positive and negative linear correlation relationships and the non - linear correlation relationships between molecules;

[0014] 4) Statistically count the frequencies of adjacent two classes c q and c q+1 samples in the four different regions constructed based on the molecular feature pairs (f i , f j );

[0015] 5) Calculate the differential information contained in the adjacent two classes c i , f j ) in the sample data of c q and c q+1 . This differential information is measured using the joint probability density function Δ ij , and all the molecular feature pairs are sorted in descending order according to the joint probability density function value Δ;

[0016] 6) Screen the top k non - overlapping molecular feature pairs according to the joint probability density function value Δ to form a feature pair set PS for constructing a decision rule. The screening method is: if the molecular feature pair (f i , f j ) ∈ PS, then all subsequent molecular feature pairs containing f i or f j are not in the set PS;

[0017] 7) Construct a decision rule RI - MFR: For any molecular feature pair (f i , f j ) belonging to the set PS and a given unknown sample S new , RI - MFR first calculates which region the unknown sample S new belongs to according to the functions y(1) and y(2). When it belongs to the t region, it predicts the class label of the unknown sample S ij according to the size relationship between p q (c ij |t) and p q+1 (c new |t) in the training set, where pij (c q |t) for c q ij |t) for c q+1 q+1 |t) for c ij q |t) for c ij q+1 |t) for c new q |t) for c q+1

[0018] 8) k pairs of molecular features in the set PS will construct k decision rules RI-MFR, each of which predicts the class label of unknown sample S new , and the final class label of S new is determined by majority voting.

[0019] Further, the difference information is measured by joint probability density function Δ ij , and the specific formula is as follows:

[0020]

[0021] Where p ij (t) is the frequency of all samples in the tth region, is the joint probability of molecular feature pair (f i , f j ) in the tth region of c q class samples, is the joint probability of molecular feature pair (f i , f j ) in the tth region of c q+1 class samples; Δ ij value is normalized using formula (2) and (3):

[0022]

[0023] N-Δ ij = (Δ ij,max - Δ ij ) / Δ ij,max (3)

[0024] Where Δ ij,max is the maximum joint probability density function, N-Δ ij is the normalized joint probability density function, p ij (c q ) and p​​​​​​ij (c q+1 ) are the number of samples of class c q and c q+1 , respectively, and n is the total number of samples.

[0025] Compared with the prior art, the present application has the following beneficial effects:

[0026] 1) The present application measures the discriminant ability of (f i ,f j ) pairs for different classes of samples using the joint probability density function based on the frequency of occurrence of (f i ,f j ) pairs in each region, and since the correlation between the expression values of a pair of molecular features (f i ,f j ) is calculated in the same sample, the present application can effectively eliminate the influence of sample fluctuations or individual differences on data analysis, so that the constructed decision rule can more accurately use the real changes in the correlation between molecules to screen for early diagnosis markers of cancer and make accurate predictions of diseases, thereby improving the early diagnosis and treatment effect of breast cancer;

[0027] 2) The present application uses sample mean and standard deviation and analyzes the size relationship between the expression values of a pair of molecular features (f i ,f j ) to construct a decision rule, which can effectively solve the overfitting problem caused by small sample size or high complexity of molecular expression data;

[0028] 3) In view of the complex and diverse correlation between different omics molecules in the biochemical reaction process, the present application analyzes the changes of the interaction between different omics molecules in the process of canceration of the body from three aspects of positive linear correlation, negative linear correlation and nonlinear correlation to construct a decision rule, so that the changes of the relevant metabolic mechanisms in the process of occurrence and development of breast cancer can be more comprehensively reflected, which can provide help for the in-depth study of the pathogenesis of breast cancer and provide important target for the early diagnosis and personalized treatment of breast cancer;

[0029] 4) The design idea and method principle of the present application have good universality and can also be applied to the pathogenesis, clinical early diagnosis and precision medicine research of other cancers or complex diseases. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is a schematic diagram of the positive linear correlation between the molecular features (f i ,f j ) of the present application.

[0031] Figure 2 is a schematic diagram of the negative linear correlation between the molecular features (f i ,fj ) negative linear correlation relationship between the schematic diagram.

[0032] Figure 3 is the molecular characteristics (f i ,f j ) nonlinear correlation relationship between the schematic diagram.

[0033] Figure 4 is the partial least squares discriminant analysis model based on the marker set screened out by RI-MFR in the embodiment of the application.

[0034] Figure 5 is the partial least squares discriminant analysis model based on the marker set screened out by the highest score pair algorithm TSP in the embodiment of the application.

[0035] Figure 6 is the partial least squares discriminant analysis model based on the marker set screened out by the improved highest score pair algorithm M-TSP in the embodiment of the application.

[0036] Figure 7 is the partial least squares discriminant analysis model based on the marker set screened out by the weighted highest score pair algorithm weight-TSP in the embodiment of the application.

[0037] Figure 8 is the partial least squares discriminant analysis model based on the marker set screened out by the fusion different marker mode algorithm CDBP in the embodiment of the application. DETAILED DESCRIPTION

[0038] The specific embodiments of the application will be further described below in conjunction with the accompanying drawings:

[0039] This invention presents a biologically interpretable decision rule algorithm based on different types of intermolecular associations. Using breast cancer genomics and metabolomics data as research objects, it aims to discover biomarkers for early diagnosis and personalized treatment assessment of breast cancer, exploring changes in metabolic activity mechanisms during the development of breast cancer, thereby contributing to the advancement of clinical research and application in breast cancer. In living organisms, the intermolecular associations between different omics molecules are complex and diverse, and different types of associations contribute differently to clinical research on breast cancer under various circumstances. Systematic and comprehensive exploration of the differential changes in intermolecular associations between different omics molecules during the development of breast cancer helps to screen for prospective warning signals and adaptive assessment biomarkers that can effectively predict the occurrence of breast cancer, thereby improving the clinical diagnosis and treatment outcomes of breast cancer. This invention proposes a biologically interpretable decision rule algorithm, RI-MFR, based on different types of intermolecular associations. It utilizes a joint probability density function to deeply compare and analyze the differential changes in positive linear, negative linear, and nonlinear associations between molecular features during the process of carcinogenesis, screening key molecular features with strong predictive capabilities to construct precise and simple decision rules for effective clinical management of breast cancer.

[0040] Let F = {f1, f2, ..., f m} is defined as the set of molecular features, where m represents the number of molecular features; X = {x1, x2, ..., x...} n} is defined as a sample set, where n represents the number of samples; C = {c1, c2, ..., c z} is defined as a set of class labels, where z represents the number of class labels.

[0041] 1) Input omics data on different stages of breast cancer development.

[0042] 2) such as Figure 1 As shown, the function y(1) = ax can represent the molecular characteristic f i and f j A positive linear correlation exists between two points, where f is the positive linear correlation. i Equivalent to x, f j Equivalent to y(1), where a takes the value 1, i.e., the function is y(1) = x, as shown below. Figure 2 As shown, the function y(2)=-b j x / a i +b j It can represent molecular characteristics f i and f j A negative linear correlation exists between two points, where f is the negative linear correlation. i Equivalent to x, f j Equivalent to y(2), in order to eliminate the overfitting problem caused by the small sample size or high complexity of molecular expression data, the coefficient ai Set to u i +δσ i , b j Set to u j +δσ j , the parameter u i is the mean of the molecular feature f i , the parameter σ i is the standard deviation of the molecular feature f i , u j is the mean of the molecular feature f j , the parameter σ j is the standard deviation of the molecular feature f j is the standard deviation of the molecular feature f, and δ is the balance factor.

[0043] 3) As Figure 3 shown, the functions y(1) = x and y(2) = -b j x / a i +b j divide the data space into four regions: The first region is y(1) ≥ x and y(2) < -b j x / a i +b j ; The second region is y(1) ≥ x and y(2) ≥ -b j x / a i +b j ; The third region is y(1) < x and y(2) ≥ -b j x / a i +b j ; The fourth region is y(1) < x and y(2) < -b j x / a i +b j , therefore, the joint functions y(1) and y(2) can be used to represent both the positive and negative linear correlation relationships and the non - linear correlation relationships between molecules.

[0044] 4) Statistically count the frequencies of adjacent two classes c q and c q+1 samples in the four different regions constructed based on the molecular feature pair (f i , f j ).

[0045] 5) Calculate the differential information contained in the adjacent two classes c i , f j ) in the sample data of c q and c q+1 . This differential information is measured by the joint probability density function Δ ij , and all the molecular feature pairs are sorted in descending order according to the joint probability density function value Δ;

[0046] The difference information adopts joint probability density function Δ ij The measurement is carried out, and the specific formula is as follows:

[0047]

[0048] Wherein, p ij (t) is the frequency of all samples in the tth region, is the joint probability of the molecular feature pair (f i , f j ) in the tth region of the c q+1 class sample, i j is the joint probability of the molecular feature pair (f i , f j ) in the tth region of the c q+1 class sample; the Δ ij value is normalized using formula (2) and (3):

[0049]

[0050] N-Δ ij = (Δ ij,max - Δ ij ) / Δ ij,max (3)

[0051] Wherein, Δ ij,max is the maximum joint probability density function, N-Δ ij is the normalized joint probability density function, p ij (c q |t) is the frequency of the c q class sample in the tth region, p ij (c q+1 |t) is the frequency of the c q+1 class sample in the tth region, p ij (c q ) and p ij (c q+1 ) are the c q and c q+1 class sample numbers respectively, and n is the number of all samples.

[0052] 6) According to the joint probability density function value Δ, the first k disjoint molecular feature pairs are screened to form a feature pair set PS for constructing a decision rule, and the screening method is: if the molecular feature pair (f i , f j ) ∈ PS, then all subsequent molecular feature pairs (f i , f j ) or (f i , f v ) containing f w or f jf j ) are in the set PS, i.e. or where f v is the vth molecular feature, and f w is the wth molecular feature.

[0053] 7) Constructing the decision rule RI-MFR: for any molecular feature pair (f i ,f j ) belonging to the set PS and a given unknown sample S new , the RI-MFR first calculates which region the unknown sample S new belongs to according to the functions y(1) and y(2), when it belongs to the t region, the class label of the unknown sample S ij is predicted according to the size relationship of p q (c ij |t) and p q+1 (c new |t); when p ij (c q |t) > p ij (c q+1 |t), S new is predicted as the c q class, and vice versa, then it is predicted as the c q+1 class.

[0054] 8) The k disjoint molecular feature pairs in the set PS will construct k decision rules RI-MFR, each decision rule RI-MFR predicts the class label of the unknown sample S new , and the final class label of S new is determined by majority voting.

[0055] The present application fully explores the changes of different forms of correlation between molecules in different pathological states of the body, finds important potential biomarkers for early warning of early breast cancer and precise treatment, and constructs accurate, simple and easy-to-biologically-explain decision rules based on the screened biomarkers, so as to effectively improve the clinical early diagnosis and treatment effect of breast cancer, provide help for further understanding of the pathogenesis of breast cancer, and provide theoretical guidance for research in the fields of disease omics data analysis and translational medicine.

[0056] The following examples are implemented on the premise of the technical solutions of the present application, and detailed implementation modes and specific operation processes are given, but the protection scope of the present application is not limited to the following examples. The methods used in the following examples are all conventional methods unless otherwise specified.

[0057] Example:

[0058] Screening of clinical markers of breast cancer based on omics data and metabolomics data related to body metabolism.

[0059] (1) Collection of breast cancer omics data related to body metabolism;

[0060] The training set of breast cancer omics data in this experiment was derived from the Cancer Genome Atlas database, containing 113 control group healthy samples N and 1109 model group samples M, which included breast cancer stage I samples, breast cancer stage II samples, breast cancer stage III samples, breast cancer stage IV samples and breast cancer stage V samples. Among them, breast cancer stage I samples and breast cancer stage II samples constitute early stage breast cancer samples E.S. Based on functional enrichment analysis, genes closely related to body metabolism were determined and analyzed subsequently. The validation set of breast cancer omics data in this experiment was derived from the Gene Expression Omnibus database, consisting of data sets GSE162228, GSE161533, GSE29431 and GSE42568, which were used to further verify the clinical early diagnosis effect of potential markers of breast cancer screened based on the training set.

[0061] (2) Collection of breast cancer metabolomics data related to body metabolism;

[0062] The breast cancer metabolomics data in this experiment consisted of 15 control group rat samples Ctrl, 6 rat samples treated with near-infrared NIR and 6 rat samples treated with near-infrared combined with ferritin-bound cytophaga capulata Cy@AFT+NIR.

[0063] (3) To discover potential markers of different stages of breast cancer and individualized treatment suitability evaluation markers, this study divided the above two data sets into four two-class sub-problems: N vs. E.S., N vs. M, Ctrl vs. NIR and Ctrl vs. Cy@AFT+NIR.

[0064] (4) Related parameter settings: 20 times 5-fold cross-validation, k value set to 9; the mean and standard deviation of cross-validation classification accuracy were used to measure the effectiveness of the algorithm.

[0065] (5) Table 1 shows the mean and standard deviation of classification accuracy of the marker set screened based on the present application on the early diagnosis sub-problems; for the sub-problems N vs. M and N vs. E.S., the classification accuracy of the present application is the highest among all the comparison methods, which are 97.10±0.35 and 98.09±0.53, respectively; to further illustrate the effectiveness of the present application, a binary logistic regression model is constructed based on the marker set selected by the present application to classify the validation set, and Table 2 shows the classification results of different methods on the validation set. The experiment shows that the AUC value and standard error of the present application are the best among all the comparison methods.

[0066] (6) Figures 4-8 Table 1 shows the mean and standard deviation of classification accuracy of the marker set screened based on the present application on the early diagnosis sub-problems; for the sub-problems N vs. M and N vs. E.S., the classification accuracy of the present application is the highest among all the comparison methods, which are 97.10±0.35 and 98.09±0.53, respectively; to further illustrate the effectiveness of the present application, a binary logistic regression model is constructed based on the marker set selected by the present application to classify the validation set, and Table 2 shows the classification results of different methods on the validation set. The experiment shows that the AUC value and standard error of the present application are the best among all the comparison methods.

[0067] Table 1. Performance of different algorithms on the genomic discovery set data

[0068]

[0069] Table 2. Performance of different algorithms on the genomic validation set data

[0070]

Claims

1. A biologically interpretable decision rule algorithm based on different types of intermolecular relationships, characterized in that, Includes the following steps: 1) Input omics data on different stages of breast cancer development; 2) Use the function y(1) = ax to represent the molecular characteristic f i and f j A positive linear correlation exists between two points, where f is the positive linear correlation. i Equivalent to x, f j Equivalent to y(1), where a is a positive number; Using the function y(2) = -b j x / a i +b j Indicates molecular characteristics f i and f j A negative linear correlation exists between two points, where f is the negative linear correlation. i Equivalent to x, f j Equivalent to y(2); Wherein, coefficient a i Set to u i +δσ i b j Set to u j +δσ j , parameter u i molecular characteristics f i The mean, parameter σ i molecular characteristics f i Standard deviation, u j molecular characteristics f j The mean, parameter σ j molecular characteristics f j The standard deviation of , where δ is the balance factor; 3) The functions y(1) = ax and y(2) = -b j x / a i +b j Divide the data space into four regions: The first region is y(1) ≥ ax and y(2) < -b j x / a i +b j ; The second region is y(1) ≥ ax and y(2) ≥ -b j x / a i +b j ; The third region is y(1) < ax and y(2) ≥ -b j x / a i +b j ; The fourth region is y(1) < ax and y(2) < -b j x / a i +b j , therefore, the combined functions y(1) and y(2) are used to represent both the positive and negative linear correlation relationships and the non-linear correlation relationship between molecules; 4) Count the number of adjacent c-classes q and c q+1 Samples based on molecular feature pairs (f i ,f j The frequency in the four different regions constructed; 5) Calculate molecular feature pairs (f i ,f j In two adjacent classes c q and c q+1 The discrepancy information contained in the sample data is expressed using the joint probability density function Δ. i j is used to measure and all molecular feature pairs are sorted in descending order according to the joint probability density function value Δ; 6) Select the top k disjoint molecular feature pairs based on the joint probability density function value Δ, forming a feature pair set PS, which is used to construct decision rules. The selection method is: if the molecular feature pair (f i ,f j If f ∈ PS, then all subsequent cases containing f i or f j None of the molecular feature pairs are in the set PS; 7) Constructing the RI-MFR decision rule: For any pair of molecular features (f) belonging to the set PS i ,f j ) and the given unknown sample S new RI-MFR first calculates the unknown sample S based on functions y(1) and y(2). new Which region does it belong to? When it belongs to region t, it depends on p in the training set. ij (c q |t) and p ij (c q+1 Predicting the relationship between the size of |t) for unknown sample S new The class label, where p ij (c q |t) is c q The frequency of class sample in region t, p ij (c q+1 |t) is c q+1 The frequency of class samples in the t-th region; when p ij (c q |t)>p ij (c q+1 When |t), S new The prediction is c q If the class is different, then the prediction is c. q+1 kind; 8) The k disjoint molecular feature pairs in the set PS will be used to construct k decision rules RI-MFR. Each decision rule RI-MFR applies to the unknown sample S. new Predicting the class labels and determining S using majority voting. new The final class label.

2. The decision rule algorithm with biological interpretability based on different types of intermolecular associations according to claim 1, characterized in that, The difference information is expressed using the joint probability density function Δ. ij The specific formula for measurement is as follows: Where, p ij (t) represents the frequency of all samples in the t-th region. For molecular characteristic pairs (f i ,f j ) in c q The joint probability of the t-th region in the class sample. For molecular characteristic pairs (f i ,f j ) in c q+1 The joint probability of the t-th region in the class sample; using formulas (2) and (3) to calculate Δ ij Normalize the values: N-W ij =(D ij,max -D ij ) / D ij,max (3) Where, Δ ij,max For the maximum joint probability density function, N-Δ ij p is the normalized joint probability density function. ij (c q ) and p ij (c q+1 ) are respectively c q and c q+1 The number of samples in each class, where n is the total number of samples.