High-specificity gene / site combination screening method based on Bayesian optimization

By applying Bayesian optimization algorithm and Gaussian process model in gene screening, combined with two-factor vertical crosslinking kernel function, the problem of insufficient specificity and accuracy of existing gene screening methods is solved, and a more efficient and specific gene/locus combination screening effect is achieved.

CN120015128APending Publication Date: 2025-05-16SUZHOU HAIMIAO BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510044523.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-11
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing gene screening methods have the disadvantages of high false positive rate, low specificity, sensitivity to parameters or sample size, easy to overfit, and poor ability to capture nonlinear relationships, making it difficult to effectively screen out gene/locus combinations with high specificity and high accuracy.

Method used

Using a computer algorithm based on Bayesian optimization, the screening process of gene/locus combination is optimized by constructing a Gaussian process model and using two-factor vertical crosslinking kernel function to improve the efficiency, specificity and accuracy of computational gene/locus combination.

Benefits of technology

It significantly improves the efficiency, specificity and accuracy of gene/locus combination screening, increases the upper limit of gene/locus number screening, and can more scientifically screen out highly specific gene/locus combinations related to diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015128A_ABST
    Figure CN120015128A_ABST
Patent Text Reader

Abstract

The invention discloses a gene / site combination screening method based on Bayesian optimization, a gene / site combination with the highest sample type specificity is screened by using a Gaussian process model, the kernel function of the Gaussian process model is a two-factor vertical cross-linked kernel, and comprises two cores of basic variation and discrete variation; on the basis of differences in the same sample type, the dispersion degree between different sample types is used as an auxiliary item, and the relation between two gene / site combinations is calculated. According to the method disclosed by the invention, the correlation of each gene / site in different gene / site combinations in the same sample type and in different sample types can be evaluated, and the gene / site characteristics required in biology are fully fused into the gene / site screening process; the high efficiency, specificity and accuracy of computational gene / site screening are remarkably improved, and meanwhile the upper limit of the number of gene / site screening is remarkably increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and specifically relates to a computer algorithm based on Bayesian optimization that can be applied to biological gene / site screening logic, which is suitable for screening gene / site combinations based on a Gaussian process model and ultimately applied to disease screening and classification. Background Art

[0002] Gene screening is the process of selecting genes related to specific biological processes, diseases or phenotypes from a large number of genes through experimental or computational analysis methods. The purpose of gene screening is to find genes that play a key role in disease occurrence, cell function regulation or other biological phenomena, and to provide candidate targets for subsequent functional verification or clinical application.

[0003] There are two main types of genetic screening:

[0004] Experimental screening: Direct intervention in gene function through experimental means, such as RNA interference (RNAi) or CRISPR-Cas9 gene editing technology, to conduct functional loss or enhancement experiments on thousands of genes. By observing phenotypic changes in cells or organisms, researchers can determine which genes are important under specific conditions. This type of method is often used to discover potential therapeutic targets related to cancer, neurodegenerative diseases, etc.

[0005] Computational screening: Based on gene expression data, biomarkers or other high-throughput omics data, statistical and machine learning methods are used for analysis. Through differential gene expression analysis, association analysis, cluster analysis and other techniques, genes significantly associated with diseases or specific phenotypes can be screened out from thousands of genes. This type of screening is often used to extract information from large-scale data sets, such as screening key driver genes for cancer through RNA sequencing data.

[0006] In the process of gene screening, the importance of combinatorial gene markers has become increasingly prominent. This type of marker is not simply a combination of several genes related to a certain disease, but through precise screening and optimization, the combination of these genes can reflect the occurrence, development and individual differences of the disease with higher specificity and accuracy. For example, compared with single gene markers, combinatorial gene markers can significantly reduce false positive and false negative rates, and improve the diagnosis and prediction capabilities of complex diseases. This method is particularly suitable for multi-factor driven diseases, such as cancer and autoimmune diseases, and can more comprehensively capture their potential molecular mechanisms, thus laying a solid foundation for personalized treatment and precision medicine.

[0007] Common methods for computational gene screening include differential expression analysis, cluster analysis, machine learning, and gene network analysis. For example, differential expression analysis uses statistical tools (such as DESeq2) to screen out genes that change significantly in disease samples, but is easily limited by sample size and false positive rate; cluster analysis (such as K-means) is used to discover the co-regulatory pattern of genes, but is sensitive to parameters and susceptible to noise; machine learning methods such as random forests can screen key genes based on feature importance, but are prone to overfitting in small sample data; gene network analysis screens core genes by constructing co-expression networks, but relies on prior knowledge of network topology. Although these gene screening methods can screen genes related to diseases, they generally have disadvantages such as high false positive rate, low specificity, sensitivity to parameters or sample size, limited number of sample type distinctions, easy overfitting, and poor ability to capture nonlinear relationships.

[0008] Bayesian optimization is a powerful tool for global optimization, especially for expensive and non-differentiable black box functions and experimental costs. This method characterizes the distribution of the objective function by constructing a probabilistic model (usually a Gaussian process) and combines exploration and exploitation strategies to select the input point for the next evaluation. The core idea is to maximize the "expected improvement" (EI) or other acquisition functions to balance the exploration of unsearched areas and the use of local information near the current best point.

[0009] The main steps of the Bayesian optimization model include:

[0010] Initialization: Randomly select several points and calculate the objective function value to build an initial Gaussian process model.

[0011] Build a surrogate model: Based on known data points, the Gaussian process generates an approximate model of the target function, providing the predicted mean and uncertainty for each input point.

[0012] Select next evaluation point: Determine the next point by optimizing an acquisition function (such as EI) that takes uncertainty into account while potentially improving the optimal solution.

[0013] Update the model: Calculate the objective function value at the new evaluation point, add the result to the existing dataset, and retrain the Gaussian process model.

[0014] Iteration: Repeat steps 3 and 4 until a termination condition is met (such as the number of iterations, computational budget, or improvement in the objective function is less than a certain threshold).

[0015] Bayesian optimization has a wide range of applications in areas such as hyperparameter tuning and experimental design in machine learning, and performs particularly well in scenarios with high computational and experimental costs.

[0016] In the biomedical field, there is often a deviation between theoretical expectations and actual experimental results. The same gene may be associated with multiple diseases, and there are individual differences among patients. Therefore, how to design a reasonable computational method to scientifically solve the above problems is a difficult problem in current research. Summary of the invention

[0017] In view of the shortcomings of the prior art, the present invention provides a computational gene / site combination screening method, which realizes a more reasonable computational gene / site combination screening process by modifying and optimizing the Bayesian optimization method. Different from the Bayesian optimization kernel function that simply describes the size relationship between two input variables, the two-factor vertical cross-link kernel described in the present invention can evaluate the correlation of each gene / site in different gene / site combinations in the same sample type and different sample types, fully integrating the biologically required gene / site characteristics into the gene / site screening process, significantly improving the efficiency, specificity and accuracy of computational gene / site screening, and significantly increasing the upper limit of the number of gene / site screening.

[0018] The specific technical solutions of the present invention are as follows:

[0019] A screening method for gene / site combinations based on Bayesian optimization uses a Gaussian process model to screen the gene / site combination with the highest sample type specificity. The Gaussian process model includes: a kernel (covariance) function, a predicted mean function, a predicted variance function, and an expected improvement function. The kernel function is a two-factor vertical cross-link kernel, which includes two cores, basic variation and discrete variation. The two-factor vertical cross-link kernel is based on the difference in the same sample type and uses the discrete degree between different sample types as an auxiliary item to calculate the relationship between two gene / site combinations. The kernel function is

[0020] k is the covariance of the two gene / site combinations, and its value is the correlation between the two gene / site combinations. The higher the correlation between the two gene / site combinations, the higher k(x1,x2) will be. That is, the value of k(x1,x2) is positively correlated with the correlation between the two gene / site combinations.

[0021] Theoretically, there is no limit to the number of genes / sites constituting a gene / site combination, which is a positive integer greater than 1, preferably a positive integer of 2-10.

[0022] In a specific example of the present invention, a gene combination formed by 5 genes is screened using the comprehensive recognition accuracy of the sample type prediction model as the evaluation criterion. The gene combination with the highest comprehensive recognition accuracy can be used as a specific marker for a specific sample type (such as disease, race, age, gender or other biological or environmental factors that affect genetic characteristics).

[0023] x1, x2 represent the averaged feature data combination variables corresponding to the two groups of gene / site combinations, n is the number of sample types, and the value range of n is any integer greater than 1; α, β, γ, λ are hyperparameters defined by prior knowledge;

[0024] BV is the abbreviation of the original basic variation function (Basal variation) in the method of the present invention: The meaning of this basic variation is the correlation between two different gene / locus combinations in the same sample type.

[0025] EXP{} is an exponential function with the natural constant e as the base, and its purpose is to limit the value range of BV to a real number range less than 1, and at the same time smoothly adjust the change of BV value caused by the difference between two groups of gene / site combinations in the same sample type.

[0026] Π is the ratio of a circle to a circle, and its purpose is to normalize the calculated radian value, which is the angle between the spatial multidimensional vectors of the averaged feature data combination corresponding to the two sets of gene / site combinations.

[0027] arccos() is the inverse cosine function, which specifically means calculating the radian value of the angle between the spatial multidimensional vectors of the averaged feature data combination corresponding to two sets of gene / site combinations.

[0028] x1i and x2i are the spatial multidimensional vectors of the averaged feature data combination corresponding to the two groups of gene / site combinations under the i-th sample type, and ||x1i|| and ||x2i|| are the moduli of the spatial multidimensional vectors.

[0029] γ is a hyperparameter whose value is defined by prior knowledge. Its theoretical value range is the real number R. Its purpose is to coordinate the importance of the angle change and length change between the spatial vectors projected by the averaged feature data combinations corresponding to two different gene / site combinations in the same sample type.

[0030] max() is a maximum value function, which aims to normalize the difference in the modulus length of the spatial vector projected by the mean feature data corresponding to two different gene / site combinations in the same sample type.

[0031] β is a hyperparameter, which is defined by prior knowledge and has a theoretical value range of real number R. Its purpose is to coordinate the influence of the discrete degree of BV values ​​of the averaged feature data combination corresponding to two different gene / site combinations in multiple sample types on the BV average value. It is a coordination parameter of the penalty function. σB.V. is the standard deviation between BVs in different sample types. Its purpose is to evaluate the discrete degree of basic variation between the averaged feature data combinations corresponding to two gene / site combinations in different sample types, and try to ensure that in each sample type, the difference between the averaged feature data combinations corresponding to the two gene / site combinations is as close to the average value of BV as possible, so that the final calculated result can better represent the specific situation in each sample type.

[0032] α is a hyperparameter, which is defined by prior knowledge. Its theoretical value range is the real number R. Its purpose is to coordinate the influence of the minimum value of the basic variation of the averaged feature data combination corresponding to two different gene / site combinations in each sample type on the covariance value. It is also the coordination parameter of the penalty function.

[0033] log() is a logarithmic function with base 10, and its function value range is less than 0. It is used as a penalty function to smooth the influence of the minimum value that the basic variation of the averaged feature data combination corresponding to the two groups of gene / site combinations in each sample type can reach on the final covariance.

[0034] q is the minimum threshold of the basic variation function value set artificially, ranging from 0 to 1, and its maximum value is 1. When the value is 1, it means that the difference between the averaged feature data combinations corresponding to the two groups of gene / site combinations in the same sample type is 0. The preferred q value is set between 0.9 and 0.99. The smaller the q value is, the more inclined to explore new space, and the larger the q value is, the more inclined to stably explore known areas.

[0035] Ⅱ() is an indicator function, and its return value is 0 or 1. It serves as a penalty function to coordinate the influence of the number of basic variations of the averaged feature data combination corresponding to the two groups of gene / site combinations in all sample types on the final covariance.

[0036] m is the number of genes / locus in the gene / locus combination, and its value range is any positive integer greater than 1.

[0037] DV is the abbreviation of the discrete variation function invented in the method of the present invention: The meaning of this discrete variation is the difference in the degree of dispersion of the mean-valued characteristic data of each gene / site in all sample types in the two gene / site combinations.

[0038] sin{} is a sine function, which aims to smoothly adjust the difference in the discrete degree of the averaged characteristic data of each gene / site in all sample types in the two gene / site combinations.

[0039] Delta[] is a difference calculation function, which aims to calculate the difference in the degree of dispersion of the mean characteristic data of each gene / site in all sample types in two gene / site combinations.

[0040] Max() and Min() are maximum and minimum value functions respectively, and their purpose is to obtain the maximum and minimum values ​​of the averaged characteristic data of the jth gene / site in all sample types.

[0041] x is the combination of averaged feature data of different sample types corresponding to the jth gene / site in the gene / site combination, ri represents the averaged feature data of the i-th sample type corresponding to the jth gene / site in the gene / site combination, and rμ represents the mean of ri of the jth gene / site in all sample types.

[0042] The method of the present invention comprises:

[0043] (1) Building a dataset

[0044] Select a gene information database with known sample types, select n sample types, select several independent samples for each sample type, each independent sample contains characteristic data of several genes / sites, normalize the characteristic data of each gene / site, arrange them in a matrix according to sample type and characteristic data, and obtain a standardized data set;

[0045] The number of independent samples of each sample type can be the same or different, and the number of independent samples of the same sample type corresponding to all genes / sites in the gene / site combination is the same. Generally, in order to meet the needs of neural network model training, the number of independent samples of each sample type should be no less than 100.

[0046] The feature data in the standardized data set are averaged according to the feature data of all independent samples corresponding to the same gene under the same sample type to obtain the average feature data of the gene / site, and the average feature data of m genes / sites are set to form a combination, and s combinations are randomly selected from the standardized data set. The average feature data combination is used as input, and the analysis result of the sample type prediction model corresponding to the combination is used as output. The average feature data combination matrix X and the analysis result matrix Y of the sample type prediction model corresponding to the average feature data combination matrix are constructed, and these two matrices are combined to define the initial data set Ds, where s is the number of randomly selected average feature data combinations;

[0047] The sample type prediction model is composed of a sample type accuracy calculation model and a calculation formula. The sample type accuracy calculation model is selected from a neural network model, a linear regression model, a support vector machine, a decision tree, a random forest, a k-nearest neighbor algorithm or a logistic regression model. The sample type accuracy calculation model takes the characteristic data of all independent samples of all sample types for each gene / site in the gene / site combination as input, and takes the sample type recognition accuracy of each gene / site combination for all sample types as output. The calculation formula calculates the comprehensive accuracy of sample type recognition for each gene / site combination, and the calculation formula is: f(x)=α2Q1+β2Q2+γ2Q3、f(x)=median(accuracy)、 One or more combinations of, wherein f(x) is the function value of the comprehensive accuracy of sample type identification, n is the number of sample types, accuracy is the recognition accuracy of each sample type obtained by the sample type accuracy calculation model, σ' is the standard deviation between the recognition accuracy of different sample types, log() is a logarithmic function with base 10, II() is an indicator function, p is the defined minimum accuracy threshold of each sample type, max() and min() are the maximum and minimum calculation formulas, Q1, Q2, Q3 are the 25th, 50th and 75th percentiles, median() is the median calculation formula, Ωi is the weight of the accuracy of the i-th sample type, β1, α1, α2, β2, γ2, Q1, Q2, Q3 are hyperparameters, which are defined by prior knowledge or manually. A specific example of the present invention, the calculation formula is

[0048] (2) Using the Gaussian process model to perform Gaussian process cycles

[0049] a. Construct the covariance matrix: Use the two-factor vertical cross-link kernel function to calculate the covariance between all the averaged feature data combinations in the initial data set, and construct the covariance matrix K(X,X);

[0050] b. Reselect feature data from the standardized data set and process it by averaging to form a new average feature data combination x * As input, the new averaged feature data combination does not belong to the initial data set, and the kernel function value (covariance) of the new averaged feature data combination and the averaged feature data combination in the initial data set, the kernel function value of the new averaged feature data combination itself, the predicted variance function value, the predicted mean function value and the expected improvement function value are calculated;

[0051] The predicted mean function is μ(x * )=k(x * ,X)×K(X,X)-1 ×y; the prediction variance function is σ 2 (x * )=k(x * ,x * )-k(x * ,X)×K(X,X) -1 ×k(X,x * ), the expected improvement function is EI(x * )=E[max(0,μ(x * )-y * )]=(μ(x * )-y * )×Φ(Z)+σ(x * )×φ(Z),

[0052] where μ(x * ) is the predicted mean function value, k(x * ,X) is x * The covariance matrix between the matrix X of all the averaged feature data combinations in the initial data set, K(X,X) is the autocovariance matrix between all the averaged feature data combinations in the initial data set, -1 is the inverse operation of the matrix, and y is the comprehensive accuracy matrix of sample type recognition in the initial data set; σ 2 (x * ) is the predicted variance function value, k(x * ,x * ) is x * The autocovariance of k(X,x * ) is the averaged feature data combination of all gene / site combinations in the initial data set and x * The covariance matrix between EI(x * ) is the expected improvement function value, y * is the value with the highest comprehensive accuracy in sample type recognition in the initial data set, Z is the standardized improvement, and the calculation formula is Z = (μ (x * )-y * ) / σ(x * ), Φ(Z) is the cumulative distribution function, which means μ(x * ) exceeds y * The probability of x is φ(Z), which is the probability density function. * The distribution density around the predicted mean, σ(x * ) is the square root of the forecast variance, i.e. the forecast standard deviation, which reflects the confidence of the current model in the new combination;

[0053] c. Evaluate the expected improvement, select the gene / site combination corresponding to the averaged feature data combination with the largest expected improvement as the target combination for gene / site screening, calculate the comprehensive accuracy of sample type identification of the target combination using the sample type prediction model, and update the averaged feature data combination of the target combination and the corresponding comprehensive accuracy of sample type identification (i.e., the analysis result of the sample type prediction model) to the initial data set;

[0054] d. Iteration: Repeat steps a, b, and c until the termination condition is met, and screen out the gene / site combination with the highest sample type specificity.

[0055] The hyperparameters in each function in the above method can be defined manually or determined after optimization using the initial data set.

[0056] The method of the present invention, the termination condition of step d is that the comprehensive accuracy of the identification of the sample type prediction model does not improve for several consecutive cycles, reaches a threshold set by humans, the number of cycles set by humans, the improvement of the objective function value is lower than the threshold, the maximum improvement of the acquisition function value is lower than the threshold, the model training performance reaches the expectation, the optimization time reaches the upper limit, etc. Preferably, the comprehensive accuracy of the identification of the sample type prediction model does not improve for 5-10 consecutive cycles, reaches 90-98% accuracy, or the number of cycles is 10-200 times.

[0057] Preferably, the neural network model is a multi-layer fully connected neural network model.

[0058] A specific example of the present invention, the multi-layer fully connected neural network model includes an input layer, 0 to several hidden layers and an output layer, wherein the input feature of the input layer is the feature data of the gene / site combination in all sample types, and the activation function used is Relu; the first hidden layer is a fully connected layer, comprising several neurons, and the activation function used is Relu; the second hidden layer is a fully connected layer, comprising several neurons, and the activation function used is Relu; the output layer comprises n (the number of sample types) neurons, and the softmax activation function is used; the learning rate is set; the loss function is Sparse Categorical Crossentropy, L2 regularization is used in each fully connected layer to avoid overfitting, the L2 parameter value is set, and the early stopping function is added through the Keras callback function (Callbacks), and the model early stopping condition is that the loss function value does not decrease for several consecutive cycles, the number of training rounds is set, and the number of feature data combinations of genes / sites for each parameter update in training is set. The neural network model is developed and run using Python, and the high-level framework Keras of the model is provided by Python's TensorFlow. The model development environment is Python 3.11.5, TensorFlow 2.16.1, and it is developed and run through PyCharmCommunity Edition2023.2.1IDE. The platform for model development and operation is a computer equipped with Windows 11Pro23H2 system, the processor is 13th Gen Intel(R)Core(TM)i7-13700KF 3.40GHz, and the memory is 64GB dual-channel DDR4Hynix memory.

[0059] In the method of the present invention, the gene information database with known sample types can be GEO, ENA, SRA, dbSNP, TCGA, gnomAD, GTEx, ICGC, PRIDE, miRBase, NONCODE, TarBase, miRGator, etc. The characteristic data of genes are biological characteristics obtained from DNA sequencing data, RNA sequencing data, DNA methylation sequencing data, miRNA sequencing data, protein expression spectrum, DNA copy number variation spectrum, metabolomics data, genomics data, proteomics data, transcriptomics data, epigenomics data, immunoomics data, pharmacomics data, systemomics data, phenotypic data, lipidomics data, glycomics data, signalomics data, toxicomics data or evolutionomics data.

[0060] The biological characteristics include methylation rate, gene expression level, mutation status, gene copy number, miRNA expression abundance, protein expression level, single nucleotide polymorphism (SNP) frequency, structural variation copy number, epigenetic modification intensity, non-coding RNA expression abundance, phosphorylated protein ratio, protein interaction intensity, metabolite concentration, metabolic pathway activity, immune cell infiltration ratio, cytokine concentration, bacterial flora abundance, metabolite contribution, target inhibition efficiency, sugar metabolism concentration or lipid metabolism concentration, etc.

[0061] The method of the present invention can be used to perform calculations and comparisons on the same or different types of gene feature data, preferably on the same type of gene feature data.

[0062] In the method of the present invention, the sample type is one of disease, race, age, gender or other biological or environmental factors that affect gene characteristics. The different sample types in the method of the present invention are further subdivisions of the same type of factors. For example, when the sample type is disease, the different sample types described in the present invention are different diseases (such as lung cancer, liver cancer, gastric cancer, intestinal cancer, etc.); when the sample type is race, the different sample types described in the present invention are different races (such as yellow, white, black, brown).

[0063] Another object of the present invention is to provide a Gaussian process model for high-specificity gene / site combination screening, including a kernel function, a prediction mean function, a prediction variance function and an expected improvement function, wherein the kernel function is a two-factor vertical cross-link kernel for calculating the relationship between two gene / site combinations, and the kernel function is The symbols in the above functions are defined the same as before.

[0064] The prediction mean calculation formula of Gaussian process is: μ(x * )=k(x * ,X)×K(X,X) -1 ×y is used to calculate the predicted mean of the averaged feature data combination of the target gene / site combination;

[0065] The prediction variance calculation formula of Gaussian process is: 2 (x * )=k(x * ,x * )-k(x * ,X)×K(X,X) -1 ×k(X,x * ) is used to calculate the predicted variance of the averaged feature data combination of the target gene / site combination, the expected improvement function of the Gaussian process: EI(x * )=E[max(0,μ(x * )-y *)]=(μ(x * )-y * )×Φ(Z)+σ(x * )×φ(Z) is used to calculate the expected improvement of the averaged feature data combination of the target gene / site combination. The calculation method of the prediction mean, prediction variance, and expected improvement function of the above Gaussian process is the prior art.

[0066] The symbols in the above functions are defined the same as before.

[0067] The purpose of the present invention is to provide a mathematical method for computational gene / site combination screening, the core of which is to use the two-factor vertical cross-linking kernel to calculate the relationship between two groups of gene / site combinations. The realization of the functions of the present invention depends on computers, computer programs such as Python, and operators with corresponding computer programming capabilities and data processing capabilities.

[0068] Another object of the present invention is to provide an application of the Gaussian process model in gene / site combination screening, including differentially expressed gene screening, differentially methylated gene screening, differentially mutated gene screening, differentially copied number variation gene screening, differentially expressed protein coding gene screening, differentially expressed miRNA screening, differentially methylated site screening, differentially mutated site screening, differentially hydroxymethylated modification site screening, differentially acetylated modification site screening, differentially frequent single nucleotide polymorphism site screening, differentially copied number structural variation gene screening, differentially epigenetic modification intensity gene screening, differentially expressed abundance non-coding RNA screening, differentially phosphorylated ratio protein coding gene screening, differentially interacting intensity protein coding gene screening, differentially concentrated metabolite related gene screening, differentially active metabolic pathway gene screening, differentially infiltrating ratio immune cell coding gene screening, differentially concentrated cytokine coding gene screening, differentially abundant bacterial flora coding gene screening, differentially contributing metabolic substance coding gene screening, and differentially inhibiting efficiency target corresponding gene / site screening.

[0069] The Gaussian process model described in the present invention can be used to screen gene marker combinations related to tumors, genetic abnormalities, metabolic abnormalities, diseases caused by infection with drug-resistant pathogens, autoimmune diseases, neurodegenerative diseases, cardiovascular diseases, infectious diseases, inflammatory diseases, rare diseases, endocrine diseases, blood system diseases, developmental disorders, mental illnesses, skin diseases, degenerative bone and joint diseases, or diseases caused by environmental exposure.

[0070] Advantages of the present invention:

[0071] The present invention establishes a highly specific, highly sensitive and low-cost screening method for calculating gene / site combinations for predicting sample types. The Gaussian process model based on the dual-factor vertical cross-linking kernel can be used to normalize the optimization process of the Bayesian optimization model, so that the optimization process always remains within a reasonable and correct range of variation, making the predicted mean of the model optimization process more accurate and the final prediction result more accurate. At the same time, the basic variation and discrete variation calculation formulas described in the present invention can force the changes in the model optimization process to be more in line with biological needs, thereby making the model optimization results more specific and sensitive. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the embodiments of the present invention are briefly introduced below.

[0073] Obviously, the drawings described below are only drawings of some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work, but these other drawings also belong to the drawings required for use in the embodiments of the present invention.

[0074] Figure 1 It is a flowchart of the gene / site combination screening method of the present invention. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solution, beneficial effects and significant improvements of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings provided in the embodiments of the present invention.

[0076] Obviously, all the described embodiments are only partial embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of the present invention.

[0077] For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0078] It should also be noted that the following specific embodiments may be combined with each other, and the same or similar concepts or processes therein may not be repeated in some embodiments.

[0079] The technical solution of the present invention is described in detail below with specific embodiments.

[0080] Example 1: Cost comparison between the exhaustive method and the gene / site screening method of the present invention

[0081] 1.1 Data Preparation

[0082] The purpose of this example is to demonstrate that the gene screening method described in the present invention can significantly reduce the time cost and experimental cost of gene screening. Specifically, this example will use randomly generated data as simulated gene expression level data, and use the exhaustive method and the gene screening method described in the present invention to analyze the randomly generated data, respectively, to find the most suitable gene combination for disease judgment and screening.

[0083] A data matrix with 1000 rows and 5000 columns was randomly generated by computer as the standardized data set. The value of each element of the data matrix ranged from 0 to 1. Each row of the data matrix was defined as the normalized expression level of a gene in all independent samples of different disease types. The 1st to 1000th columns of the data matrix were defined as the normalized expression levels of different genes in 1000 independent samples of disease A, the 1001st to 2000th columns were the normalized expression levels of different genes in 1000 independent samples of disease B, the 2001st to 3000th columns were the normalized expression levels of different genes in 1000 independent samples of disease C, the 3001st to 4000th columns were the normalized expression levels of different genes in independent samples of disease D, and the 4001st to 5000th columns were the normalized expression levels of different genes in 1000 independent samples of disease E.

[0084] The genome combination consisting of 5 genes was taken as the research object, and the gene combination was screened based on the sample type prediction accuracy as the evaluation standard. Finally, the gene combination with the highest prediction accuracy can be used as a diagnostic marker for specific diseases.

[0085] Specifically, the ultimate goal of this embodiment is to screen the gene combination (consisting of 5 genes) with the highest prediction accuracy for a specific disease (sample type) from 1000 genes. Each gene screened has gene expression data of 1000 independent samples of 5 diseases (a total of 5000). First, the newly selected gene combination is cyclically compared by exhaustive enumeration or the gene screening method described in the present invention, and then the expression data of all independent samples of the screened gene combination are grouped, wherein 800 expression data are randomly selected from each gene of each disease as the training data set of the neural network model, and the remaining 200 gene expression data are used as the test data set. A multi-layer fully connected neural network model with fixed hyperparameters is used as a sample type accuracy calculation model to train and verify the expression data of the combination formed by the screened 5 genes. The method described in the present invention further uses a calculation formula to calculate the comprehensive accuracy of sample type identification as an evaluation criterion. The calculation formula is Where f(x) is the function value of the comprehensive accuracy of sample type recognition, n is the number of sample types, accuracy is the recognition accuracy of each sample type obtained by the sample type accuracy calculation model, σ' is the standard deviation between the recognition accuracy of different sample types, log() is the logarithmic function with base 10, II() is the indicator function, p is the defined minimum accuracy threshold of each sample type, β1 and α1 are hyperparameters defined by prior knowledge, in this embodiment, p = 0.95, β1 = 0.45, α1 = -0.37. The exhaustive method uses a multi-layer fully connected neural network model with fixed hyperparameters to predict the sample type accuracy of each combination.

[0086] The multi-layer fully connected neural network model includes an input layer, two hidden layers and an output layer, wherein the input feature of the input layer is the expression level of 5 genes in all independent samples of 5 diseases; the first hidden layer is a fully connected layer, comprising 64 neurons, and the activation function used is Relu; the second hidden layer is a fully connected layer, comprising 32 neurons, and the activation function used is Relu; the output layer comprises 5 neurons, and a softmax activation function is used; the learning rate is set to 0.001; the loss function is Sparse Categorical Crossentropy, L2 regularization is used in each fully connected layer to avoid overfitting, and the L2 parameter value is set to 0.01, and an early stopping function is added through a Keras callback function (Callbacks), and the model early stopping condition is that the loss function value does not decrease for five consecutive cycles, the number of training rounds is set to 1000 rounds, and the number of batches for updating parameters each time during training is set to 50. The label in the model training process is the true sample type (disease type) of the sample, and the sample type with the highest probability in the output layer is used as the final conclusion of the model output, and the conclusion output by the model is compared with the true sample type to obtain the accuracy of the model in each sample type. The neural network model is developed and run using Python, and the high-level framework Keras of the model is provided by Python's TensorFlow. The model development environment is Python 3.11.5, TensorFlow2.16.1, and it is developed and run through PyCharm Community Edition 2023.2.1IDE. The platform for model development and operation is a computer equipped with Windows 11Pro 23H2 system, the processor is 13th Gen Intel(R)Core(TM)i7-13700KF 3.40GHz, and the memory is 64GB dual-channel DDR4 Hynix memory.

[0087] 1.2 Exhaustive gene screening

[0088] The principle of exhaustive gene screening is as follows: First, all possible combinations of 5 genes are randomly selected from 1,000 genes and listed. Then, all expression data of each gene combination are predicted by the neural network model one by one, and finally the group with the best prediction result is selected as the final selected gene combination. The whole process uses a computer program to perform the corresponding exhaustive operation and the operation of the neural network model. Each complete program operation includes: selecting one of all gene combinations, running the neural network model, and statistically obtaining the disease prediction accuracy results.

[0089] 1.3 Gene screening based on dual-factor vertical cross-linking core of the present invention

[0090] The gene screening principle described in the present invention is as follows: first, 1000 combinations are randomly selected from all gene combinations, each combination consists of 5 genes, and the gene expression data of these 1000 gene combinations are respectively subjected to calculation and analysis of the sample type prediction model to obtain the corresponding comprehensive accuracy of sample type identification, and the expression level average of each gene in each gene combination in 5 sample types is calculated as the averaged feature data combination of the gene combination, and the averaged feature data combination of the gene combination and its corresponding comprehensive accuracy of sample type identification are integrated to construct an initial data set, and the initial data set is used to optimize the hyperparameters of the Gaussian process model.

[0091] Use the Gaussian process model to perform a Gaussian process loop:

[0092] a. Construct the covariance matrix: Use the kernel function of the two-factor vertical cross-link kernel to calculate the covariance between all gene combinations (meaning feature data combinations) in the initial data set, and construct the covariance matrix K(X,X);

[0093] The kernel function is β=0.4, α=0.35, q=0.95, λ=1.19

[0094]

[0095] b. Randomly select 100,000 gene combinations from the standardized data set, average the feature data of all gene combinations to form new averaged feature data combinations, which do not belong to the initial data set, and calculate the kernel function value of the new averaged feature data combination and the averaged feature data combination in the initial data set, the kernel function value of the new combination itself, the predicted variance function value, the predicted mean function value and the expected improvement function value.

[0096] The predicted mean function is μ(x * )=k(x * ,X)×K(X,X) -1×y; the prediction variance function is σ 2 (x * )=k(x * ,x * )-k(x * ,X)×K(X,X) -1 ×k(X,x * ), the expected improvement function is EI(x * )=E[max(0,μ(x * )-y * )]=(μ(x * )-y * )×Φ(Z)+σ(x * )×φ(Z).

[0097] c. Evaluate the expected improvement, compare and select the gene combination corresponding to the averaged feature data combination with the largest expected improvement function value as the final selection of the current cycle. If the expected improvements of all currently selected combinations are 0, reselect the averaged feature data combination corresponding to the new unselected gene combination for calculation, use the sample type prediction model to calculate the comprehensive accuracy of sample type recognition of all sample types for the gene combination selected in the current cycle, and update the averaged feature data combination of the current gene combination and its corresponding comprehensive accuracy of sample type recognition to the initial data set.

[0098] d. Iteration: Repeat steps a, b, and c until all gene combinations have been selected.

[0099] 1.4 Cost comparison

[0100] This example is a theoretical experiment, mainly for the following reasons. First, in actual operation, the gene screening is often far more than 1,000 genes, which may be tens of thousands of genes or hundreds of thousands of methylation sites. This example uses 1,000 genes for theoretical experimental verification, which is sufficient to prove the efficiency and practicality of the method described in the present invention. Second, when screening 1,000 genes, if 5 genes are selected as a gene combination, there will be 8.3×10 9 There are possible combinations. Even with a high-performance computer, the exhaustive method will not be able to be experimentally verified in practice. Therefore, this embodiment calculates and analyzes the time cost and computing cost of the exhaustive method and the gene screening method described in the present invention from a theoretical perspective.

[0101] In this theoretical experiment, the computer used is a computer with Windows 11Pro 23H2 system, a 13th Gen Intel(R) Core(TM) i7-13700KF 3.40GHz processor, and a 64GB dual-channel DDR4 Hynix memory. Using this computer, it takes about 5 seconds to run the neural network model once per single thread, and about 60 seconds to run a complete Gaussian process cycle. When considering the extreme case of calculating all gene combinations to find the best gene combination, the exhaustive method requires a total of 8.3×10 9 The gene screening method of the present invention requires a total of 83,000 Gaussian cycles and 84,000 neural network model operations.

[0102] Table 1: Comparison of theoretical experimental results using the exhaustive method and the gene screening method of the present invention

[0103] Exhaustive genetic screening Gene screening method of the present invention Theoretical number of neural network runs <![CDATA[8.3×10 9 times]]> 84000 times The number of times the Gaussian process optimization model is run / 83000 times Theoretical running time of neural network ca. 1315 About 5 days The time required to run the Gaussian process optimization model / About 57 days Total time ca. 1315 About 62 days

[0104] As shown in Table 1, when the computer described in this embodiment is used to perform the exhaustive method or the gene screening method described in the present invention to select 5 genes in 1000 genes, the method described in the present invention takes about 62 days in extreme cases, while the exhaustive method takes about 1315 years. Obviously, the work that cannot be completed using the exhaustive method will be solved in the method described in the present invention, and the time cost of the gene screening process using the exhaustive method will reach 7741 times compared to the method described in the present invention. Even if a higher-performance computer, more core parallel computing, and GPU acceleration and other technologies are used to shorten the running time, the exhaustive method still cannot be completed in a suitable time when facing a real larger big data sample. For the method described in the present invention, even in the most extreme case, Gaussian process analysis is performed on all gene combinations, which can also be completed in a reasonable time using higher-performance computers or GPU acceleration and other technical means, and in the actual Gaussian process model operation process, it is not necessary to perform Gaussian process analysis on all combinations, and it is often only necessary to find the optimal solution in dozens to hundreds of cycles. Therefore, the gene screening method described in the present invention is compared with the exhaustive method. It has the characteristics of high efficiency and high feasibility.

[0105] Example 2: Example application of the gene site screening process described in the present invention

[0106] 1. Data Preparation

[0107] The purpose of this example is to prove that the Gaussian process model using the covariance calculation method of the present invention can screen out a better gene / site combination. Specifically, this example uses part of the methylation sequencing data in a real disease-related database (TCGA database) as the original data, and performs corresponding data processing, and then uses 6 different Gaussian process optimization models for screening. The kernel functions of these 6 optimization models are different, including:

[0108] Square exponential kernel: k(x,x′)=exp(-(x-x')^2 / 2σ^2), k is the value of the covariance function; x and x′ are two sets of averaged feature data combinations; exp() is an exponential function with the natural constant e as the base; σ is the scale parameter. σ=1

[0109] Linear kernel: k(x,x′)=x·x′+c, k is the value of the covariance function; x and x′ are two sets of averaged feature data combinations; c is a constant bias term. c=1

[0110] Polynomial kernel: k(x,x′)=(x·x′+c)^d, k is the value of the covariance function; x and x′ are two sets of averaged feature data combinations; c is a constant bias term, and d is the order of the polynomial. c=1,d=2

[0111] Exponential kernel: k(x,x′)=exp(-abs(x-x') / 2σ), k is the value of the covariance function; x and x′ are two sets of averaged feature data combinations; exp() is an exponential function with the natural constant e as the base; σ is the scale parameter; abs() function is the absolute value function.

[0112] σ=1

[0113] Periodic kernel: k(x,x′)=exp((2(sin(πabs(xx′) / p))^2) / -σ^2), k is the value of the covariance function; x and x′ are two sets of averaged feature data combinations; exp() is an exponential function with the natural constant e as the base; σ is the amplitude control parameter; abs() function is an absolute value function. sin() is a sine function; p is a period parameter. σ=1, p=2.

[0114] And the dual-factor vertical cross-linked core of the present invention: β=0.4, α=0.3, q=0.95, λ=1.12

[0115]

[0116] Specifically, the methylation profiles of genomic DNA of bladder urothelial carcinoma tissue (TCGA-BLCA), colorectal cancer tissue (TCGA-COAD & TCGA-READ) and normal tissue were obtained from the TCGA database, and all selected data were sequencing data (IDAT format data) of the Illumina 450k methylation chip, and the Illumina 450k methylation chip contained methylation detection of 485,577 CpG sites. In this embodiment, a total of 412 bladder urothelial carcinoma tissue samples, 393 colorectal cancer tissue samples, and 371 normal tissue samples were collected, and a total of 374,689 CpG site data were obtained after quality control screening of the data as a standardized data set for analysis in this embodiment. In this example, the beta value (methylation rate) after bioinformatics analysis is used as the characteristic data of the CpG site, with the purpose of finding 3 CpG sites (the 3 CpG sites of the optimal result may reflect the relationship between 1 to 3 genes and the disease, which may be 3 CpG sites of a gene, or 3 CpG sites shared by 2 to 3 genes), wherein 80% of the independent sample methylation data are randomly selected from each CpG site of each sample type as the training data set of the neural network model, and the remaining 20% ​​of the methylation sequencing data are used as the test data set. Referring to the method in Example 1, the sample type accuracy calculation model of the multi-layer fully connected neural network model with fixed hyperparameters was used to train and verify the methylation sequencing data of the three CpG sites. The neural network model includes an input layer, a hidden layer and an output layer, wherein the input feature of the input layer is the methylation rate of the three CpG sites, 16 neurons are fully connected, and the activation function used is Relu; the hidden layer is a fully connected layer, containing 8 neurons, and the activation function used is Relu; the output layer contains 3 neurons and uses a softmax activation function; the learning rate is set to 0.001; the loss function is SparseCategorical Crossentropy, L2 regularization is used in each fully connected layer to avoid overfitting, and the L2 parameter value is set to 0.001. The early stopping function is added through the Keras callback function (Callbacks). The model early stopping condition is that the loss function value does not decrease for 5 consecutive cycles. The number of training rounds is set to 100 rounds, and the number of batches for each parameter update is set to 50. The label during the model training process is the real sample type (cancer type) of the sample. The sample type with the highest probability in the output layer is used as the output conclusion of the model. By comparing the final conclusion of the model with the real sample type, the accuracy of each sample type is obtained. The neural network model is developed and run using Python, and the high-level framework of the model, Keras, is provided by Python's TensorFlow.The model development environment is Python 3.11.5, TensorFlow 2.16.1, and it is developed and run through PyCharm Community Edition 2023.2.1 IDE. The platform for model development and operation is a computer equipped with Windows 11Pro 23H2 system, the processor is 13th Gen Intel(R) Core(TM) i7-13700KF 3.40GHz, and the memory is 64GB dual-channel DDR4 Hynix memory.

[0117] 2.2 CpG site screening process

[0118] In this embodiment, 80,000 combinations are randomly selected from all CpG combinations, each combination consists of 3 CpGs, and the neural network model is run on the methylation sequencing data of these 80,000 CpG combinations in three sample types, and the corresponding sample type identification comprehensive accuracy is calculated in combination with the calculation formula, the mean of the methylation sequencing data of all CpG combinations is calculated according to the sample type as the averaged feature data combination of the CpG combination, the averaged feature data combination of the CpG combination and its corresponding sample type identification comprehensive accuracy are integrated to construct an initial data set, and the initial data set is used to optimize the hyperparameters of the Gaussian process model.

[0119] Use the Gaussian process model to perform a Gaussian process loop:

[0120] a. Constructing the covariance matrix: Using the above six kernel functions, respectively, we calculate the covariance between the averaged feature data combinations of all CpG combinations in the initial data set, and construct the covariance matrix K(X,X) corresponding to each of the six kernel functions;

[0121] b. Select methylation sequencing data of 100,000 new CpG combinations from the standardized data set, and calculate the mean according to the sample type to obtain the corresponding averaged feature data combination. The averaged feature data combination of the new CpG site combination does not belong to the averaged feature data combination matrix of the initial data set (in the first 40 cycles, 80% are randomly selected new combinations, and 20% are similar combinations of combinations with greater expected improvements in previous cycles; in the subsequent 60 cycles, 50% are randomly explored new combinations, and 50% are similar combinations of combinations with greater expected improvements in previous cycles). Perform multi-threaded parallel calculations on the feature data combinations of these 100,000 CpG combinations for the corresponding kernel function values, predicted mean function values, predicted variance function values, and expected improvement function values ​​(the predicted mean function, predicted variance function, and expected improvement function are the same as in Example 1);

[0122] c. Compare and select the averaged feature data combination of the CpG site combination with the largest expected improvement function value as the final selection of the current cycle, and perform neural network model analysis on the methylation sequencing data of all sample types of the CpG site combination selected in the current cycle to obtain the corresponding sample type recognition accuracy, using the calculation formula β1=0.45,α1=-0.35,p=0.95Calculate the comprehensive recognition accuracy of the neural network model in all sample types under the optimization scheme of applying the two-factor vertical cross-linking kernel. The calculation method of the comprehensive recognition accuracy under the optimization scheme of applying other kernel functions is to obtain the average of the recognition accuracy of each sample type, and update the current CpG combination of averaged feature data and its corresponding sample type recognition comprehensive accuracy data to the initial data set;

[0123] d. Iteration: Repeat steps a, b, and c until 100 cycles are completed.

[0124] 2.3 Experimental Results

[0125] The results of CpG site screening using different optimization models are shown in Table 2 below, where the dual-factor vertical cross-linking kernel described in the present invention can greatly improve the comprehensive cancer prediction accuracy and minimum accuracy of the feature data of the selected optimal CpG site combination in the test data set of the neural network model under the conditions of comparable time cost and computational cost.

[0126] Table 2: Statistics of CpG screening results using different optimization models

[0127]

Claims

1. A screening method for gene / locus combinations based on Bayesian optimization, characterized in that The Gaussian process model is used to screen the gene / site combination with the highest sample type specificity. The Gaussian process model includes: a kernel function, a predicted mean function, a predicted variance function, and an expected improvement function. The kernel function is a two-factor vertical cross-link kernel. The two-factor vertical cross-link kernel contains two cores, basic variation and discrete variation. It is based on the difference in the same sample type and uses the discrete degree between different sample types as an auxiliary item to calculate the relationship between two gene / site combinations. The kernel function is k is the covariance of the two gene / site combinations, x1 and x2 represent the averaged feature data combination variables corresponding to the two gene / site combinations, n is the number of sample types, and the value range of n is any integer greater than 1; α, β, γ, and λ are hyperparameters defined by prior knowledge; BV is the basic variogram: EXP{} is an exponential function with the natural constant e as the base, Π is pi, arccos() is the inverse cosine function, x1i and x2i are the spatial multidimensional vectors of the averaged feature data combination corresponding to the two groups of gene / site combinations under the i-th sample type, ||x1i|| and ||x2i|| are the moduli of the spatial multidimensional vector, and max() is the maximum value function; σB.V. is the standard deviation between BVs in different sample types; log() is a logarithmic function with base 10; q is the minimum threshold of the basic variation function value set artificially; Ⅱ() is an indicator function, and its return value is 0 or 1; m is the number of genes / locus in the gene / locus combination, and its value range is a positive integer greater than 1; DV is the discrete variogram: sin{} is a sine function, Delta[] is a difference calculation function, Max() and Min() are maximum and minimum value functions respectively, x is a combination of averaged feature data of different sample types corresponding to the jth gene / site in the gene / site combination, ri represents the averaged feature data of the i-th sample type corresponding to the jth gene / site in the gene / site combination, and rμ represents the mean of ri of the jth gene / site in all sample types; The method comprises: (1) Building a dataset Select a gene information database with known sample types, select n sample types, select several independent samples for each sample type, each independent sample contains characteristic data of several genes / sites, normalize the characteristic data of each gene / site, arrange them in a matrix according to sample type and characteristic data, and obtain a standardized data set; The feature data in the standardized data set are averaged according to the feature data of all independent samples corresponding to the same gene under the same sample type to obtain the average feature data of the gene / site, and the average feature data of m genes / sites are set to form a combination, and s combinations are randomly selected from the standardized data set. The average feature data combination is used as input, and the analysis result of the sample type prediction model corresponding to the combination is used as output. The average feature data combination matrix X and the analysis result matrix Y of the sample type prediction model corresponding to the average feature data combination matrix are constructed, and these two matrices are combined to define the initial data set Ds, where s is the number of randomly selected average feature data combinations; The sample type prediction model is composed of a sample type accuracy calculation model and a calculation formula. The sample type accuracy calculation model is selected from a neural network model, a linear regression model, a support vector machine, a decision tree, a random forest, a k-nearest neighbor algorithm or a logistic regression model. The sample type accuracy calculation model takes the characteristic data of all independent samples of all sample types for each gene / site in the gene / site combination as input, and takes the sample type recognition accuracy of each gene / site combination for all sample types as output. The calculation formula calculates the comprehensive accuracy of sample type recognition for each characteristic data combination, and the calculation formula is: f(x)=α2Q1+β2Q2+γ2Q3、f(x)=median(accuracy)、 One or more combinations of, where f(x) is the function value of the comprehensive accuracy of sample type recognition, n is the number of sample types, accuracy is the recognition accuracy of each sample type obtained by the sample type accuracy calculation model, σ' is the standard deviation between the recognition accuracy of different sample types, log() is the logarithmic function with base 10, II() is the indicator function, p is the defined minimum accuracy threshold of each sample type, max() and min() are the maximum and minimum calculation formulas, Q1, Q2, Q3 are the 25th, 50th and 75th percentiles, median() is the median calculation formula, Ωi is the weight of the accuracy of the i-th sample type, β1, α1, α2, β2, γ2, Q1, Q2, Q3 are hyperparameters, which are defined by prior knowledge or manually; (2) Using the Gaussian process model to perform Gaussian process cycles a. Construct the covariance matrix: Use the two-factor vertical cross-link kernel function to calculate the covariance between all the averaged feature data combinations in the initial data set, and construct the covariance matrix K(X,X); b. Reselect feature data from the standardized data set and process it by averaging to form a new average feature data combination x * As input, the new averaged feature data combination does not belong to the initial data set, and the kernel function value of the new averaged feature data combination and the averaged feature data combination in the initial data set, the kernel function value of the new averaged feature data combination itself, the predicted variance function value, the predicted mean function value and the expected improvement function value are calculated; The predicted mean function is μ(x * )=k(x * ,X)×K(X,X) -1 ×y; the prediction variance function is σ 2 (x * )=k(x * ,x * )-k(x * ,X)×K(X,X) -1 ×k(X,x * ), the expected improvement function is EI(x * )=E[max(0,μ(x * )-y * )]=(μ(x * )-y * )×Φ(Z)+σ(x * )×φ(Z), where μ(x * ) is the predicted mean function value, k(x * ,X) is x * The covariance matrix between the matrix X of all the averaged feature data combinations in the initial data set, K(X,X) is the autocovariance matrix between all the averaged feature data combinations in the initial data set, -1 is the inverse operation of the matrix, y is the sample type recognition comprehensive accuracy matrix of the sample type prediction model corresponding to the averaged feature data combination of all gene / site combinations in the initial data set; σ 2 (x * ) is the predicted variance function value, k(x * ,x * ) is x * The autocovariance of k(X,x * ) is the averaged feature data combination of all gene / site combinations in the initial data set and x * The covariance matrix between EI(x * ) is the expected improvement function value, y * is the value with the highest comprehensive accuracy of sample type recognition of the sample type prediction model in the initial data set, Z is the standardized improvement margin, and the calculation formula is Z = (μ (x * )-y * ) / σ(x * ), Φ(Z) is the cumulative distribution function, which means μ(x * ) exceeds y * The probability of x is φ(Z), which is the probability density function. * The distribution density around the predicted mean, σ(x * ) is the square root of the forecast variance, i.e. the forecast standard deviation, which reflects the confidence of the current model in the new combination; c. Evaluate the expected improvement, select the gene / site combination corresponding to the averaged feature data combination with the largest expected improvement as the target combination for gene / site screening, calculate the comprehensive accuracy of sample type identification of the target combination using the sample type prediction model, and update the averaged feature data combination of the target combination and the corresponding comprehensive accuracy of sample type identification to the initial data set; d. Iteration: Repeat steps a, b, and c until the termination condition is met, and screen out the gene / site combination with the highest sample type specificity.

2. The method according to claim 1, characterized in that The calculation formula is 3. The method according to claim 1, characterized in that The termination conditions of step d are that the comprehensive accuracy of the sample type prediction model does not improve for several consecutive cycles, reaches a manually set threshold, the number of cycles is manually specified, the improvement of the objective function value is lower than the threshold, the maximum improvement of the acquisition function value is lower than the threshold, the model training performance reaches the expectation, or the optimization time reaches the upper limit.

4. The method according to claim 3, characterized in that The termination condition of step d is that the comprehensive recognition accuracy of the sample type prediction model does not increase for 5-10 consecutive cycles, reaches 90-98% accuracy, or the number of cycles is 10-200 times.

5. The method according to claim 1, characterized in that The n is 2-100.

6. The method according to claim 1, characterized in that The neural network model is a multi-layer fully connected neural network model.

7. The method according to claim 1, characterized in that The characteristic data are biological characteristics obtained based on DNA sequencing data, RNA sequencing data, DNA methylation sequencing data, miRNA sequencing data, protein expression profile, DNA copy number variation profile, metabolomics data, genomics data, proteomics data, transcriptomics data, epigenomics data, immunoomics data, pharmacomics data, systems omics data, phenotypic omics data, lipidomics data, glycomics data, signalomics data, toxicomics data or evolutionomics data.

8. The method according to claim 7, characterized in that The biological feature is selected from methylation rate, gene expression level, mutation status, gene copy number, miRNA expression abundance, protein expression level, single nucleotide polymorphism frequency, structural variation copy number, epigenetic modification intensity, non-coding RNA expression abundance, phosphorylated protein ratio, protein interaction intensity, metabolite concentration, metabolic pathway activity, immune cell infiltration ratio, cytokine concentration, bacterial flora abundance, metabolite contribution, target inhibition efficiency, sugar metabolism concentration or lipid metabolism concentration.

9. The method according to claim 1, characterized in that The sample type is disease, race, age, gender or other biological or environmental factors that affect genetic characteristics.

10. A Gaussian process model for high-specificity gene / locus combination screening, characterized in that It includes kernel function, prediction mean function, prediction variance function and expected improvement function. The kernel function is a two-factor vertical cross-link kernel, which is used to calculate the relationship between two gene / site combinations. The kernel function is k is the covariance of the two gene / site combinations, x1 and x2 represent the averaged feature data combination variables corresponding to the two gene / site combinations, n is the number of sample types, and the value range of n is any integer greater than 1; α, β, γ, and λ are hyperparameters defined by prior knowledge; BV is the basic variogram: EXP{} is an exponential function with the natural constant e as the base, Π is pi, arccos() is the inverse cosine function, x1i and x2i are the spatial multidimensional vectors of the averaged feature data combination corresponding to the two groups of gene / site combinations under the i-th sample type, ||x1i|| and ||x2i|| are the moduli of the spatial multidimensional vector, and max() is the maximum value function; σB.V. is the standard deviation between BVs in different sample types; log() is a logarithmic function with base 10; q is the minimum threshold of the basic variation function value set artificially; Ⅱ() is an indicator function, and its return value is 0 or 1; m is the number of genes / locus in the gene / locus combination, and its value range is any positive integer greater than 1; DV is the discrete variogram: sin{} is the sine function, Delta[] is the difference calculation function, Max() and Min() are the maximum and minimum value functions respectively, x is the combination of the averaged feature data of the jth gene / site in the gene / site combination corresponding to different sample types, ri represents the averaged feature data of the jth gene / site in the gene / site combination corresponding to the ith sample type, and rμ represents the mean of ri of the jth gene / site in all sample types.

11. The application of the Gaussian process model in gene / locus combination screening as claimed in claim 10, characterized in that The screening includes screening of differentially expressed genes, differentially methylated genes, differentially mutated genes, differential copy number variation genes, coding genes of differentially expressed proteins, differentially expressed miRNAs, differentially methylated sites, differentially mutated sites, differentially hydroxymethylated modification sites, differentially acetylated modification sites, differentially frequent single nucleotide polymorphism sites, structural variation genes of differential copy numbers, genes of differential epigenetic modification intensity, non-coding RNA of differential expression abundance, coding genes of differential phosphorylation ratio proteins, coding genes of differential interaction intensity proteins, genes related to differential concentration metabolites, genes of differentially active metabolic pathways, coding genes of differentially infiltrating ratio immune cells, coding genes of differentially concentrated cytokines, coding genes of differentially abundant bacterial communities, coding genes of differentially contributing metabolites, and corresponding gene / site screening of differentially inhibited efficiency targets.

12. The use as claimed in claim 11, characterized in that the sample type for gene feature screening is selected from tumors, genetic abnormalities, metabolic abnormalities, diseases caused by infection with drug-resistant pathogens, autoimmune diseases, neurodegenerative diseases, cardiovascular diseases, infectious diseases, inflammatory diseases, rare diseases, endocrine diseases, blood system diseases, developmental disorders, mental illnesses, skin diseases, degenerative bone and joint diseases or diseases caused by environmental exposure.