High-specificity gene / site combination screening method based on two-factor vertical cross-linked nucleus
By using the two-factor vertical crosslinking nucleus method in gene screening, the correlation between gene/locus combination was evaluated, and problems such as high false positive rate and low specificity of gene screening in the prior art were solved, and more efficient and accurate gene/locus combination screening was achieved.
Patent Information
- Application Number
- CN202510044524.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-11
- Publication Date
- 2025-05-16
AI Technical Summary
The existing gene screening methods have the disadvantages of high false positive rate, low specificity, sensitivity to parameters or sample size, easy to overfit, and poor ability to capture nonlinear relationships, and it is difficult to scientifically solve the complex relationship between genes and diseases.
The computational gene/site combination screening method based on two-factor vertical crosslinking nuclei was used to construct a two-factor vertical crosslinking kernel to evaluate the correlation between gene/site combinations in the same and different sample types, and combined with biological characteristics, the efficiency and accuracy of screening were significantly improved.
It significantly improves the efficiency and accuracy of gene/locus screening, increases the upper limit of screening, and provides a high specificity and high sensitivity gene/locus combination screening model, which can more accurately predict sample types.
Smart Images

Figure CN120015129A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bioinformatics, and specifically relates to a computer algorithm based on a dual-factor vertical cross-linking core that can be applied to biological gene / site screening logic, suitable for screening gene / site combinations and ultimately applied to disease screening and classification. Background Art
[0002] Gene screening is the process of selecting genes related to specific biological processes, diseases or phenotypes from a large number of genes through experimental or computational analysis methods. The purpose of gene screening is to find genes that play a key role in disease occurrence, cell function regulation or other biological phenomena, and to provide candidate targets for subsequent functional verification or clinical application.
[0003] There are two main types of genetic screening:
[0004] Experimental screening: Direct intervention in gene function through experimental means, such as RNA interference (RNAi) or CRISPR-Cas9 gene editing technology, to conduct functional loss or enhancement experiments on thousands of genes. By observing phenotypic changes in cells or organisms, researchers can determine which genes are important under specific conditions. This type of method is often used to discover potential therapeutic targets related to cancer, neurodegenerative diseases, etc.
[0005] Computational screening: Based on gene expression data, biomarkers or other high-throughput omics data, statistical and machine learning methods are used for analysis. Through differential gene expression analysis, association analysis, cluster analysis and other techniques, genes significantly associated with diseases or specific phenotypes can be screened out from thousands of genes. This type of screening is often used to extract information from large-scale data sets, such as screening key driver genes for cancer through RNA sequencing data.
[0006] In the process of gene screening, the importance of combinatorial gene markers has become increasingly prominent. This type of marker is not simply a combination of several genes related to a certain disease, but through precise screening and optimization, these genes can reflect the occurrence, development and individual differences of the disease with higher specificity and accuracy. For example, compared with single gene markers, combinatorial gene markers can significantly reduce false positive and false negative rates, and improve the diagnosis and prediction capabilities of complex diseases. This method is particularly suitable for multi-factor driven diseases, such as cancer and autoimmune diseases, and can more comprehensively capture their potential molecular mechanisms, thus laying a solid foundation for personalized treatment and precision medicine.
[0007] Common methods for computational gene screening include differential expression analysis, cluster analysis, machine learning, and gene network analysis. For example, differential expression analysis uses statistical tools (such as DESeq2) to screen out genes that change significantly in disease samples, but is easily limited by sample size and false positive rate; cluster analysis (such as K-means) is used to discover the co-regulatory pattern of genes, but is sensitive to parameters and susceptible to noise; machine learning methods such as random forests can screen key genes based on feature importance, but are prone to overfitting in small sample data; gene network analysis screens core genes by constructing co-expression networks, but relies on prior knowledge of network topology. Although these gene screening methods can screen genes related to diseases, they generally have disadvantages such as high false positive rate, low specificity, sensitivity to parameters or sample size, limited number of sample type distinctions, easy overfitting, and poor ability to capture nonlinear relationships.
[0008] In the biomedical field, there is often a deviation between theoretical expectations and actual experimental results. The same gene may be associated with multiple diseases, and there are individual differences among patients. Therefore, how to design a reasonable computational method to scientifically solve the above problems is a difficult problem in current research. Summary of the invention
[0009] In view of the deficiencies in the prior art, the present invention provides a computational gene / site combination screening method, which evaluates the correlation of each gene / site in different gene / site combinations in the same sample type and different sample types through a constructed two-factor vertical cross-linking core, fully integrating the biologically required gene / site characteristics into the gene / site screening process, significantly improving the efficiency and accuracy of computational gene / site screening, and significantly increasing the upper limit of the number of genes / sites screened.
[0010] The specific technical solutions of the present invention are as follows:
[0011] A gene / site combination screening method based on a two-factor vertical crosslinking kernel is used to screen the gene / site combination with the highest sample type specificity using a two-factor vertical crosslinking kernel function. The two-factor vertical crosslinking kernel contains two cores, basic variation and discrete variation. It is based on the difference in the same sample type and uses the discrete degree between different sample types as an auxiliary item to calculate the possibility that the target gene / site combination exceeds the current optimal gene / site combination. The two-factor vertical crosslinking kernel function is k
[0012]
[0013] k is the function value of the two-factor vertical cross-linking kernel, and its value is the possibility that the target gene / site combination exceeds the current optimal gene / site combination. When the possibility that the target gene / site combination exceeds the current optimal gene / site combination is higher, k(x1,x2) will also be higher, that is, the value of k(x1,x2) is positively correlated with the possibility that the target gene / site combination exceeds the current optimal gene / site combination.
[0014] Theoretically, there is no limit to the number of genes / sites constituting a gene / site combination, which is a positive integer greater than 1, preferably a positive integer of 2-10.
[0015] In a specific example of the present invention, a gene combination formed by 5 genes is screened using the comprehensive recognition accuracy of the sample type prediction model as the evaluation criterion. The gene combination with the highest comprehensive recognition accuracy can be used as a specific marker for a specific sample type (such as disease, race, age, gender or other biological or environmental factors that affect genetic characteristics).
[0016] x1, x2 represent the averaged feature data combination variables corresponding to the two sets of gene / site combinations (target gene / site combination and current optimal gene / site combination), n is the number of sample types, and the value range of n is any integer greater than 1; α, β, γ, λ are hyperparameters defined by prior knowledge;
[0017] BV is the abbreviation of the original basic variation function (Basal variation) in the method of the present invention: The meaning of this basic variation is the correlation between two different gene / locus combinations in the same sample type.
[0018] EXP{} is an exponential function with the natural constant e as the base, and its purpose is to limit the value range of BV to a real number range less than 1, and at the same time smoothly adjust the change of BV value caused by the difference between two groups of gene / site combinations in the same sample type.
[0019] Π is the ratio of a circle to a circle, and its purpose is to normalize the calculated radian value, which is the angle between the spatial multidimensional vectors of the averaged feature data combination corresponding to the two sets of gene / site combinations.
[0020] arccos() is the inverse cosine function, which specifically means calculating the radian value of the angle between the spatial multidimensional vectors of the averaged feature data combination corresponding to two sets of gene / site combinations.
[0021] x1i and x2i are the spatial multidimensional vectors of the averaged feature data combination corresponding to the two groups of gene / site combinations under the i-th sample type, and ||x1i|| and ||x2i|| are the moduli of the spatial multidimensional vectors.
[0022] γ is a hyperparameter whose value is defined by prior knowledge. Its theoretical value range is the real number R. Its purpose is to coordinate the importance of the angle change and length change between the spatial vectors projected by the mean feature data combinations corresponding to two different gene / site combinations in the same sample type.
[0023] max() is a maximum value function, which aims to normalize the difference in the modulus length of the spatial vector projected by the mean feature data combination of two different gene / site combinations in the same sample type.
[0024] β is a hyperparameter, which is defined by prior knowledge and has a theoretical value range of real number R. Its purpose is to coordinate the influence of the discrete degree of the BV value of the averaged feature data combination corresponding to two different gene / site combinations in multiple sample types on the BV average value. It is a coordination parameter of the penalty function. σB.V. is the standard deviation between BVs in different sample types. Its purpose is to evaluate the discrete degree of the basic variation between the averaged feature data combinations corresponding to two gene / site combinations in different sample types, and try to ensure that in each sample type, the difference between the averaged feature data combinations corresponding to the two gene / site combinations is as close to the average value of BV as possible, so that the final calculated result can better represent the specific situation in each sample type.
[0025] α is a hyperparameter, which is defined by prior knowledge and has a theoretical value range of real number R. Its purpose is to coordinate the influence of the minimum value of the basic variation of the mean feature data combination corresponding to two different gene / site combinations in each sample type on the two-factor vertical cross-linking kernel function. It is also the coordination parameter of the penalty function.
[0026] log() is a logarithmic function with base 10, and its function value range is less than 0. It is used as a penalty function to smooth the influence of the minimum value of the basic variation of the averaged feature data combination corresponding to the two groups of gene / site combinations in each sample type on the final two-factor vertical cross-linking kernel function value.
[0027] q is the minimum threshold of the basic variation function value set artificially, ranging from 0 to 1, and its maximum value is 1. When the value is 1, it means that the difference between the averaged feature data combinations corresponding to the two groups of gene / site combinations in the same sample type is 0. The preferred q value is set between 0.9 and 0.99. The smaller the q value is, the more inclined to explore new space, and the larger the q value is, the more inclined to stably explore known areas.
[0028] Ⅱ() is an indicator function, and its return value is 0 or 1. It serves as a penalty function to coordinate the influence of the number of basic variations that can reach the minimum value of the averaged feature data combination corresponding to the two groups of gene / site combinations in all sample types on the final two-factor vertical cross-linking kernel function value.
[0029] m is the number of genes / locus in the gene / locus combination, and its value range is any positive integer greater than 1.
[0030] DV is the abbreviation of the discrete variation function invented in the method of the present invention: The meaning of this discrete variation is the difference in the degree of dispersion of the mean-valued characteristic data of each gene / site in all sample types in the two gene / site combinations.
[0031] sin{} is a sine function, which aims to smoothly adjust the difference in the discrete degree of the averaged characteristic data of each gene / site in all sample types in the two gene / site combinations.
[0032] Delta[] is a difference calculation function, which aims to calculate the difference in the degree of dispersion of the mean characteristic data of each gene / site in all sample types in two gene / site combinations.
[0033] Max() and Min() are maximum and minimum value functions respectively, and their purpose is to obtain the maximum and minimum values of the averaged characteristic data of the jth gene / site in all sample types.
[0034] x is the combination of averaged feature data of different sample types corresponding to the jth gene / site in the gene / site combination, ri represents the averaged feature data of the i-th sample type corresponding to the jth gene / site in the gene / site combination, and rμ represents the mean of ri of the jth gene / site in all sample types.
[0035] The method of the present invention comprises:
[0036] (1) Building a dataset
[0037] Select a gene information database with known sample types, select n sample types, select several independent samples for each sample type, each independent sample contains characteristic data of several genes / sites, normalize the characteristic data of each gene / site, arrange them in a matrix according to sample type and characteristic data, and obtain a standardized data set;
[0038] The number of independent samples of each sample type can be the same or different, and the number of independent samples of the same sample type corresponding to all genes / sites in the gene / site combination is the same. Generally, in order to meet the needs of neural network model training, the number of independent samples of each sample type should be no less than 100.
[0039] The feature data in the standardized data set are averaged according to the feature data of all independent samples corresponding to the same gene under the same sample type to obtain the average feature data of the gene / site, and the average feature data of m genes / sites are set to form a combination, and s combinations are randomly selected from the standardized data set. The average feature data combination is used as input, and the analysis result of the sample type prediction model corresponding to the combination is used as output. The average feature data combination matrix X and the analysis result matrix Y of the sample type prediction model corresponding to the average feature data combination matrix are constructed, and these two matrices are combined to define the initial data set Ds, where s is the number of randomly selected average feature data combinations;
[0040] The sample type prediction model is composed of a sample type accuracy calculation model and a calculation formula. The sample type accuracy calculation model is selected from a neural network model, a linear regression model, a support vector machine, a decision tree, a random forest, a k-nearest neighbor algorithm or a logistic regression model. The sample type accuracy calculation model takes the characteristic data of all independent samples of all sample types for each gene / site in the gene / site combination as input, and takes the sample type recognition accuracy of each gene / site combination for all sample types as output. The calculation formula calculates the comprehensive accuracy of sample type recognition for each gene / site combination, and the calculation formula is: f(x)=α2Q1+β2Q2+γ2Q3、f(x)=median(accuracy)、 One or more combinations of, wherein f(x) is the function value of the comprehensive accuracy of sample type identification, n is the number of sample types, accuracy is the recognition accuracy of each sample type obtained by the sample type accuracy calculation model, σ' is the standard deviation between the recognition accuracy of different sample types, log() is a logarithmic function with base 10, II() is an indicator function, p is the defined minimum accuracy threshold of each sample type, max() and min() are the maximum and minimum calculation formulas, Q1, Q2, Q3 are the 25th, 50th and 75th percentiles, median() is the median calculation formula, Ωi is the weight of the accuracy of the i-th sample type, β1, α1, α2, β2, γ2, Q1, Q2, Q3 are hyperparameters, which are defined by prior knowledge or manually. A specific example of the present invention, the calculation formula is
[0041] (2) Gene / site combinatorial screening cycle using a two-factor vertical cross-linking core
[0042] a. Select the averaged feature data combination with the highest comprehensive accuracy of sample type identification in the initial data set (the averaged feature data combination of the current optimal gene / site) as the benchmark and as the comparison object for the gene / site combination to be tested;
[0043] b. Reselect feature data from the standardized data set and form a new averaged feature data combination through averaging. The new averaged feature data combination does not belong to the initial data set. The selection method of the new averaged feature data combination is a heuristic search strategy combined with random selection or a completely random selection. The heuristic search strategy is selected from one or more of genetic algorithms, simulated annealing, particle swarm optimization, taboo search, ant colony optimization, differential evolution, stochastic gradient descent, adaptive search algorithm, and bee colony optimization.
[0044] Preferably, the selection method is a genetic algorithm, including a crossover operation: forming a new gene / site combination to be tested by recombining some genes of a better gene / site combination in the initial data set or a mutation operation: randomly modifying some genes in a better gene / site combination in the initial data set to form a new gene / site combination to be tested.
[0045] c. Calculate the two-factor vertical cross-link kernel function value between the averaged feature data combination of the new gene / site combination and the averaged feature data combination of the current optimal gene / site combination, select the gene / site combination corresponding to the averaged feature data combination with the largest function value as the target gene / site combination, use the sample type prediction model to calculate the comprehensive accuracy of sample type identification of the target gene / site combination, and update the averaged feature data combination of the target gene / site combination and its corresponding comprehensive accuracy of sample type identification (the result of the sample type prediction model) to the initial data set.
[0046] d. Iteration: Repeat steps a, b, and c until the termination condition is met, and screen out the gene / site combination with the highest sample type specificity.
[0047] The method of the present invention, the termination condition of step d is that the comprehensive accuracy of the identification of the sample type prediction model does not improve for several consecutive cycles, reaches a threshold set by humans, the number of cycles set by humans, the improvement of the objective function value is lower than the threshold, the maximum improvement of the acquisition function value is lower than the threshold, the model training performance reaches the expectation, the optimization time reaches the upper limit, etc. Preferably, the comprehensive accuracy of the identification of the sample type prediction model does not improve for 5-10 consecutive cycles, reaches 90-98% accuracy, or the number of cycles is 10-200 times.
[0048] Preferably, the neural network model is a multi-layer fully connected neural network model.
[0049] A specific example of the present invention, the multi-layer fully connected neural network model includes an input layer, 0 to several hidden layers and an output layer, wherein the input feature of the input layer is the feature data of the gene / site combination in all sample types; the first hidden layer is a fully connected layer, comprising several neurons, and the activation function used is Relu; the second hidden layer is a fully connected layer, comprising several neurons, and the activation function used is Relu; the output layer comprises n (the number of sample types) neurons, and the softmax activation function is used; the learning rate is set; the loss function is Sparse CategoricalCrossentropy, L2 regularization is used in each fully connected layer to avoid overfitting, the L2 parameter value is set, and the early stopping function is added through the Keras callback function (Callbacks), and the model early stopping condition is that the loss function value does not decrease for several consecutive cycles, the number of training rounds is set, and the number of feature data combinations of genes / sites for each parameter update in the training is set. The neural network model is developed and run using Python, and the high-level framework Keras of the model is provided by Python's TensorFlow. The model development environment is Python 3.11.5, TensorFlow 2.16.1, and it is developed and run through PyCharm Community Edition 2023.2.1IDE. The platform for model development and operation is a computer equipped with Windows 11Pro 23H2 system, the processor is 13thGen Intel(R)Core(TM)i7-13700KF 3.40GHz, and the memory is 64GB dual-channel DDR4 Hynix memory.
[0050] In the method of the present invention, the gene database with known sample types can be GEO, ENA, SRA, dbSNP, TCGA, gnomAD, GTEx, ICGC, PRIDE, miRBase, NONCODE, TarBase, miRGator, etc. The characteristic data of genes are biological characteristics obtained from DNA sequencing data, RNA sequencing data, DNA methylation sequencing data, miRNA sequencing data, protein expression spectrum, DNA copy number variation spectrum, metabolomics data, genomics data, proteomics data, transcriptomics data, epigenomics data, immunoomics data, pharmacomics data, phylogenetics data, phenotypic data, lipidomics data, glycomics data, signalomics data, toxicomics data or evolutionomics data.
[0051] The biological characteristics include methylation rate, gene expression level, mutation status, gene copy number, miRNA expression abundance, protein expression level, single nucleotide polymorphism (SNP) frequency, structural variation copy number, epigenetic modification intensity, non-coding RNA expression abundance, phosphorylated protein ratio, protein interaction intensity, metabolite concentration, metabolic pathway activity, immune cell infiltration ratio, cytokine concentration, bacterial flora abundance, metabolite contribution, target inhibition efficiency, sugar metabolism concentration or lipid metabolism concentration, etc.
[0052] The method of the present invention can be used to perform calculations and comparisons on the same or different types of gene feature data, preferably on the same type of gene feature data.
[0053] In the method of the present invention, the sample type is one of disease, race, age, gender or other biological or environmental factors that affect gene characteristics. The different sample types in the method of the present invention are further subdivisions of the same type of factors. For example, when the sample type is disease, the different sample types described in the present invention are different diseases (such as lung cancer, liver cancer, gastric cancer, intestinal cancer, etc.); when the sample type is race, the different sample types described in the present invention are different races (such as yellow, white, black, brown).
[0054] Another object of the present invention is to provide a high-specificity gene / site combination screening model, the core of which is a two-factor vertical cross-linking kernel function that evaluates the possibility that the target gene / site combination exceeds the current optimal gene / site combination. The symbols in the above functions are defined as above. The purpose of the present invention is to provide a mathematical method for computational gene / site screening, the core of which is to use the two-factor vertical cross-linking kernel model to calculate the possibility that the target gene / site combination exceeds the current optimal gene / site combination. The realization of the functions of the present invention depends on computers, Python and other computer programs, and operators with corresponding computer programming and data processing capabilities.
[0055] Another object of the present invention is to provide an application of the dual-factor vertical cross-linking nuclear model in gene / site combination screening, including differentially expressed gene screening, differentially methylated gene screening, differentially mutated gene screening, differentially copied number variation gene screening, differentially expressed protein coding gene screening, differentially expressed miRNA screening, differentially methylated site screening, differentially mutated site screening, differentially hydroxymethylated modification site screening, differentially acetylated modification site screening, differentially frequent single nucleotide polymorphism site screening, differentially copied number structural variation gene screening, differentially epigenetic modification intensity gene screening, differentially expressed abundance non-coding RNA screening, differentially phosphorylated ratio protein coding gene screening, differentially interacting intensity protein coding gene screening, differentially concentrated metabolite related gene screening, differentially active metabolic pathway gene screening, differentially infiltrating ratio immune cell coding gene screening, differentially concentrated cytokine coding gene screening, differentially abundant bacterial flora coding gene screening, differentially contributing metabolic substance coding gene screening, and differentially inhibiting efficiency target corresponding gene / site screening.
[0056] The two-factor vertical cross-linked nuclear model of the present invention can be used to screen gene marker combinations related to tumors, genetic abnormalities, metabolic abnormalities, diseases caused by infection with drug-resistant pathogens, autoimmune diseases, neurodegenerative diseases, cardiovascular diseases, infectious diseases, inflammatory diseases, rare diseases, endocrine diseases, blood system diseases, developmental disorders, mental illnesses, skin diseases, degenerative bone and joint diseases, or diseases caused by environmental exposure.
[0057] Advantages of the present invention:
[0058] The present invention establishes a computational gene / site combination screening method for predicting sample types with high specificity, high sensitivity and low cost. The gene / site combination screening process can be standardized by constructing a two-factor vertical cross-linked core model, so that the optimization process always maintains reasonable and correct changes, and the model optimization process has directionality, so that the final prediction process is more accurate and efficient, and the number of sample type predictions is significantly increased. The basic variation and discrete variation calculation formulas described in the present invention can force the changes in the model optimization process to be more in line with biological needs, thereby making the model optimization results more specific and sensitive. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the embodiments of the present invention are briefly introduced below.
[0060] Obviously, the drawings described below are only drawings of some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work, but these other drawings also belong to the drawings required for use in the embodiments of the present invention.
[0061] Figure 1 It is a flowchart of the gene / site combination screening method of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solution, beneficial effects and significant improvements of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings provided in the embodiments of the present invention.
[0063] Obviously, all the described embodiments are only partial embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of the present invention.
[0064] For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0065] It should also be noted that the following specific embodiments may be combined with each other, and the same or similar concepts or processes therein may not be repeated in some embodiments.
[0066] The technical solution of the present invention is described in detail below with specific embodiments.
[0067] Example 1: Cost comparison between the exhaustive method and the gene / site screening method of the present invention
[0068] 1.1 Data Preparation
[0069] The purpose of this example is to demonstrate that the gene screening method described in the present invention can significantly reduce the time cost and experimental cost of gene screening. Specifically, this example will use randomly generated data as simulated gene expression level data, and use the exhaustive method and the gene screening method described in the present invention to analyze the randomly generated data, respectively, to find the most suitable gene combination for disease judgment and screening.
[0070] A data matrix with 1000 rows and 5000 columns was randomly generated by computer as the standardized data set. The value of each element of the data matrix ranged from 0 to 1. Each row of the data matrix was defined as the normalized expression level of a gene in all independent samples of different disease types. The 1st to 1000th columns of the data matrix were defined as the normalized expression levels of different genes in 1000 independent samples of disease A, the 1001st to 2000th columns were the normalized expression levels of different genes in 1000 independent samples of disease B, the 2001st to 3000th columns were the normalized expression levels of different genes in 1000 independent samples of disease C, the 3001st to 4000th columns were the normalized expression levels of different genes in independent samples of disease D, and the 4001st to 5000th columns were the normalized expression levels of different genes in 1000 independent samples of disease E.
[0071] The genome combination consisting of 5 genes was taken as the research object, and the gene combination was screened based on the sample type prediction accuracy as the evaluation standard. Finally, the gene combination with the highest prediction accuracy can be used as a diagnostic marker for specific diseases.
[0072] Specifically, the ultimate goal of this embodiment is to screen the gene combination (consisting of 5 genes) with the highest prediction accuracy for a specific disease (sample type) from 1000 genes. Each gene screened has gene expression data of 1000 independent samples of 5 diseases (a total of 5000). First, the newly selected gene combination is cyclically compared by exhaustive enumeration or the gene screening method described in the present invention, and then the expression data of all independent samples of the screened gene combination are grouped, wherein 800 expression data are randomly selected from each gene of each disease as the training data set of the neural network model, and the remaining 200 gene expression data are used as the test data set. A multi-layer fully connected neural network model with fixed hyperparameters is used as a sample type accuracy calculation model to train and verify the expression data of the combination formed by the screened 5 genes. The method described in the present invention further uses a calculation formula to calculate the comprehensive accuracy of sample type identification as an evaluation criterion. The calculation formula is Where f(x) is the function value of the comprehensive accuracy of sample type recognition, n is the number of sample types, accuracy is the recognition accuracy of each sample type obtained by the sample type accuracy calculation model, σ' is the standard deviation between the recognition accuracy of different sample types, log() is the logarithmic function with base 10, II() is the indicator function, p is the defined minimum accuracy threshold of each sample type, β1 and α1 are hyperparameters defined by prior knowledge, in this embodiment, p = 0.95, β1 = 0.45, α1 = -0.37. The exhaustive method uses a multi-layer fully connected neural network model with fixed hyperparameters to predict the sample type accuracy of each combination.
[0073] The multi-layer fully connected neural network model includes an input layer, two hidden layers and an output layer, wherein the input feature of the input layer is the expression level of 5 genes in all independent samples of 5 diseases; the first hidden layer is a fully connected layer, comprising 64 neurons, and the activation function used is Relu; the second hidden layer is a fully connected layer, comprising 32 neurons, and the activation function used is Relu; the output layer comprises 5 neurons, and a softmax activation function is used; the learning rate is set to 0.001; the loss function is Sparse Categorical Crossentropy, L2 regularization is used in each fully connected layer to avoid overfitting, and the L2 parameter value is set to 0.01, and an early stopping function is added through a Keras callback function (Callbacks), and the model early stopping condition is that the loss function value does not decrease for five consecutive cycles, the number of training rounds is set to 1000 rounds, and the number of batches for updating parameters each time during training is set to 50. The label in the model training process is the true sample type (disease type) of the sample, and the sample type with the highest probability in the output layer is used as the final conclusion of the model output, and the conclusion output by the model is compared with the true sample type to obtain the accuracy of the model in each sample type. The neural network model is developed and run using Python, and the high-level framework Keras of the model is provided by Python's TensorFlow. The model development environment is Python 3.11.5, TensorFlow2.16.1, and it is developed and run through PyCharm Community Edition 2023.2.1IDE. The platform for model development and operation is a computer equipped with Windows 11Pro 23H2 system, the processor is 13th Gen Intel(R)Core(TM)i7-13700KF 3.40GHz, and the memory is 64GB dual-channel DDR4 Hynix memory.
[0074] 1.2 Exhaustive gene screening
[0075] The principle of exhaustive gene screening is as follows: First, all possible combinations of 5 genes are randomly selected from 1,000 genes and listed. Then, all expression data of each gene combination are predicted by the neural network model one by one, and finally the group with the best prediction result is selected as the final selected gene combination. The whole process uses a computer program to perform the corresponding exhaustive operation and the operation of the neural network model. Each complete program operation includes: selecting one of all gene combinations, running the neural network model, and statistically obtaining the disease prediction accuracy results.
[0076] 1.3 Gene screening based on two-factor vertical cross-linking core
[0077] The gene screening principle described in the present invention is as follows: first, 1000 combinations are randomly selected from all gene combinations, each combination consists of 5 genes, and the gene expression data of these 1000 gene combinations are respectively subjected to neural network model and comprehensive accuracy calculation and analysis to obtain the corresponding comprehensive accuracy of sample type identification, and the expression level of each gene in each gene combination in 5 sample types is calculated as the averaged feature data combination of the gene combination, and the averaged feature data combination of the gene combination and its corresponding comprehensive accuracy of sample type identification are integrated to construct an initial data set, and the initial data set is used to optimize the hyperparameters of the two-factor vertical cross-linking kernel model.
[0078] Genetic screening cycles using a two-factor vertical cross-linking nuclear model:
[0079] a. Select the current optimal gene combination: Find the gene combination with the largest sample type prediction model function value from the initial data set as the current optimal gene combination.
[0080] b. Select target gene combinations: Randomly select 100,000 gene combinations from the standardized data set, and average the feature data of all gene combinations to form new averaged feature data combinations. The new averaged feature data combinations do not belong to the initial data set.
[0081] c. Calculate the two-factor vertical cross-link kernel function value between the averaged feature data combination corresponding to the new gene combination and the averaged feature data corresponding to the current optimal gene combination for each of the 100,000 gene combinations. The two-factor vertical cross-link kernel function is β=0.4, α=0.35, q=0.95, λ=1.19
[0082]
[0083] The gene set corresponding to the averaged feature data combination with the largest function value is selected as the target gene set, the sample type prediction model is used to calculate the comprehensive accuracy of sample type identification of the target gene set, and the averaged feature data combination of the target gene set and the corresponding comprehensive accuracy of sample type identification are updated to the initial data set;
[0084] d. Iteration: Repeat steps a, b, and c until all gene combinations have been selected.
[0085] 1.4 Cost Comparison
[0086] This example is a theoretical experiment, mainly for the following reasons. First, in actual operation, the gene screening is often far more than 1,000 genes, which may be tens of thousands of genes or hundreds of thousands of methylation sites. This example uses 1,000 genes for theoretical experimental verification, which is sufficient to prove the efficiency and practicality of the method described in the present invention. Second, when screening 1,000 genes, if 5 genes are selected as a gene combination, there will be 8.3×10 9 There are a lot of possible combinations. Even with a high-performance computer, the exhaustive method will not be able to be actually verified experimentally. Therefore, this embodiment calculates and analyzes the time cost and computational cost of the exhaustive method and the gene screening method described in the present invention from a theoretical perspective.
[0087] In this theoretical experiment, the computer used is a computer with Windows 11Pro 23H2 system, a 13th Gen Intel(R) Core(TM) i7-13700KF 3.40GHz processor, and a 64GB dual-channel DDR4 Hynix memory. Using this computer, it takes about 5 seconds to run the neural network model once, and about 1 second to run a complete two-factor vertical cross-link kernel calculation. When considering the extreme case of calculating all gene combinations to find the best gene combination, the exhaustive method requires a total of 8.3×10 9 The gene screening method of the present invention requires 83,000 double-factor vertical cross-linking nuclear cycles and 84,000 neural network model operations.
[0088] Table 1: Comparison of theoretical experimental results using the exhaustive method and the gene screening method of the present invention
[0089]
[0090]
[0091] As shown in Table 1, when the computer described in this embodiment is used to perform the exhaustive method or the gene screening method described in the present invention to perform a theoretical experiment of selecting 5 genes out of 1000 genes, the method described in the present invention takes about 6 days in extreme cases, while the exhaustive method takes about 1315 years. Obviously, the work that cannot be completed using the exhaustive method will be solved in the method described in the present invention, and the time cost of the gene screening process performed using the exhaustive method will reach 79995 times that of the method described in the present invention. Even if a higher-performance computer is used, more cores are used for parallel computing, and GPU acceleration and other technologies are used to shorten the running time, the exhaustive method still cannot be completed in a suitable time when facing a real larger large data sample. As for the method described in the present invention, even in the most extreme case, the two-factor vertical cross-linking kernel model analysis of all gene combinations can be completed in a very short time by using higher performance computers or GPU acceleration and other technical means. In the actual operation process of the two-factor vertical cross-linking kernel model, it is not necessary to perform two-factor vertical cross-linking kernel model analysis and neural network model operation on all combinations. It often only takes dozens to hundreds of cycles to find the optimal solution. Therefore, compared with the exhaustive method, the gene screening method described in the present invention is highly efficient and feasible.
[0092] Example 2: Example application of the gene / site screening process described in the present invention
[0093] 1. Data Preparation
[0094] The purpose of this example is to prove that the two-factor vertical cross-linking kernel described in the present invention can screen out a better feature combination. Specifically, this example uses partial methylation sequencing data in a real disease-related database (TCGA database) as raw data, and performs corresponding data processing, and then uses the two-factor vertical cross-linking kernel model method described in the present invention to find the best cancer-specific CpG site combination.
[0095] Specifically, the methylation profiles of genomic DNA of bladder urothelial carcinoma tissue (TCGA-BLCA), colorectal cancer tissue (TCGA-COAD & TCGA-READ) and normal tissue were obtained from the TCGA database, and all selected data were sequencing data (IDAT format data) of the Illumina 450k methylation chip, and the Illumina 450k methylation chip contained methylation detection of 485,577 CpG sites. In this embodiment, a total of 412 bladder urothelial carcinoma tissue samples, 393 colorectal cancer tissue samples, and 371 normal tissue samples were collected, and a total of 374,689 CpG site data were obtained after quality control screening of the data as a standardized data set for analysis in this embodiment. In this example, the beta value (methylation rate) after bioinformatics analysis is used as the characteristic data of the CpG site, with the purpose of finding 3 CpG sites (the 3 CpG sites of the optimal result may reflect the relationship between 1 to 3 genes and the disease, which may be 3 CpG sites of a gene, or 3 CpG sites shared by 2 to 3 genes), wherein 80% of the independent sample methylation data are randomly selected from each CpG site of each sample type as the training data set of the neural network model, and the remaining 20% of the methylation sequencing data are used as the test data set. Referring to the method in Example 1, the sample type accuracy calculation model of the multi-layer fully connected neural network model with fixed hyperparameters was used to train and verify the methylation sequencing data of the three CpG sites. The neural network model includes an input layer, a hidden layer and an output layer, wherein the input feature of the input layer is the methylation rate of the three CpG sites, 16 neurons are fully connected, and the activation function used is Relu; the hidden layer is a fully connected layer, containing 8 neurons, and the activation function used is Relu; the output layer contains 3 neurons and uses a softmax activation function; the learning rate is set to 0.001; the loss function is SparseCategorical Crossentropy, L2 regularization is used in each fully connected layer to avoid overfitting, and the L2 parameter value is set to 0.001. The early stopping function is added through the Keras callback function (Callbacks). The model early stopping condition is that the loss function value does not decrease for 5 consecutive cycles. The number of training rounds is set to 100 rounds, and the number of batches for each parameter update is set to 50. The label during the model training process is the real sample type (cancer type) of the sample. The sample type with the highest probability in the output layer is used as the output conclusion of the model. By comparing the final conclusion of the model with the real sample type, the accuracy of each sample type is obtained. The neural network model is developed and run using Python, and the high-level framework of the model, Keras, is provided by Python's TensorFlow.The model development environment is Python 3.11.5, TensorFlow 2.16.1, and it is developed and run through PyCharm Community Edition 2023.2.1 IDE. The platform for model development and operation is a computer equipped with Windows 11Pro 23H2 system, the processor is 13th Gen Intel(R) Core(TM) i7-13700KF 3.40GHz, and the memory is 64GB dual-channel DDR4 Hynix memory.
[0096] 2.2 CpG site screening process
[0097] In this embodiment, 80,000 combinations are randomly selected from all CpG combinations, each combination consists of 3 CpGs, and the neural network model is run on the methylation sequencing data of these 80,000 CpG combinations in three sample types respectively, and the corresponding sample type identification comprehensive accuracy is calculated in combination with the calculation formula, and the mean of the methylation sequencing data (methylation rate) of all CpG combinations is calculated according to the sample type as the averaged feature data combination of the CpG combination, and the averaged feature data combination of the CpG combination and its corresponding sample type identification comprehensive accuracy are integrated to construct an initial data set, and the hyperparameters of the two-factor vertical cross-linking kernel model are optimized using the initial data set.
[0098] CpG site screening cycle using a two-factor vertical cross-linking nuclear model:
[0099] a. Select the current optimal CpG site combination: Find the CpG site combination with the largest function value of the sample type prediction model from the initial data set as the current optimal CpG site combination.
[0100] b. Select target CpG site combinations: reselect 100,000 CpG site combinations from the standardized data set. The new CpG site combinations do not belong to the initial data set. 70% are randomly selected new CpG site combinations, and the remaining 30% are obtained from the top 20 CpG site combinations of the sample type prediction model function value selected from the initial data set using a genetic algorithm (crossover operation). Calculate the mean methylation rate of these 100,000 CpG site combinations according to sample type as the averaged feature data combination corresponding to the CpG site combination.
[0101] c. Calculate the two-factor vertical cross-link kernel function value between the averaged feature data combination of the new CpG site combination and the averaged feature data combination of the current optimal CpG site combination. The two-factor vertical cross-link kernel formula is: β=0.4, α=0.3, q=0.95, λ=1.12.
[0102]
[0103] Evaluate the two-factor vertical cross-linking kernel function value, compare and select the CpG site combination corresponding to the averaged feature data combination with the largest two-factor vertical cross-linking kernel function value as the final selection of the current cycle, perform neural network model analysis on the methylation sequencing data of the final CpG site combination selected in the current cycle in the three sample types, and use the calculation formula described in Example 1 (β1=0.45, α1=-0.3, p=0.95) to calculate the comprehensive recognition accuracy of each sample type, and update the averaged feature data combination of the current CpG site combination and its corresponding sample type recognition comprehensive accuracy to the initial data set.
[0104] d. Iteration: Repeat steps a, b, and c until 100 cycles are completed.
[0105] 2.3 Experimental Results
[0106] The results of CpG screening using the dual-factor vertical cross-linking core optimization model of the present invention are shown in Table 2 below. Using the dual-factor vertical cross-linking core optimization model, it only took about 6 hours to screen a CpG site combination with a minimum accuracy of 92.63% and an overall recognition accuracy of 95.47% in the test data set of the neural network model.
[0107] Table 2: Statistical table of results of CpG site screening using the two-factor vertical cross-linking nuclear model
[0108] Neural Network Model Multi-layer fully connected neural network model Number of neural network model runs 80100 times Total running time of the neural network model About 6 hours Number of runs of the two-factor vertical cross-linking kernel optimization model 100 times Running time of dual-factor vertical cross-linking kernel optimization model About 3 minutes Total time cost About 6 hours The comprehensive accuracy of identifying the optimal CpG combination in the neural network model 95.47% Minimum accuracy of optimal CpG combination in different sample types in neural network model 92.63%
Claims
1. A gene / site combination screening method based on a dual-factor vertical cross-linking core, characterized in that The two-factor vertical cross-linking kernel contains two cores: basic variation and discrete variation. It is based on the differences in the same sample type and uses the discrete degree between different sample types as an auxiliary item. The two-factor vertical cross-linking kernel is used to evaluate the possibility that the target gene / site combination exceeds the current optimal gene / site combination. The two-factor vertical cross-linking kernel function is: k is the function value of the two-factor vertical cross-linking kernel, and its value indicates the possibility that the target gene / site combination exceeds the current optimal gene / site combination. x1 and x2 represent the averaged feature data combination variables corresponding to the two groups of gene / site combinations, respectively. n is the number of sample types, and the value range of n is any integer greater than 1. α, β, γ, and λ are hyperparameters defined by prior knowledge. BV is the basic variogram: EXP{} is an exponential function with the natural constant e as the base, Π is pi, arccos() is the inverse cosine function, x1i and x2i are the spatial multidimensional vectors of the averaged feature data combination corresponding to the two groups of gene / site combinations under the i-th sample type, ||x1i|| and ||x2i|| are the moduli of the spatial multidimensional vector, and max() is the maximum value function; σB.V. is the standard deviation between BVs in different sample types; log() is a logarithmic function with base 10; q is the minimum threshold of the basic variation function value set artificially; Ⅱ() is an indicator function, and its return value is 0 or 1. m is the number of genes / locus in the gene / locus combination, and its value range is a positive integer greater than 1. DV is the discrete variogram: sin{} is a sine function, Delta[] is a difference calculation function, Max() and Min() are maximum and minimum value functions respectively, x is a combination of averaged feature data of different sample types corresponding to the jth gene / site in the gene / site combination, ri represents the averaged feature data of the i-th sample type corresponding to the jth gene / site in the gene / site combination, and rμ represents the mean of ri of the jth gene / site in all sample types; The method comprises: (1) Building a dataset Select a gene information database with known sample types, select n sample types, select several independent samples for each sample type, each independent sample contains characteristic data of several genes / sites, normalize the characteristic data of each gene / site, arrange them in a matrix according to sample type and characteristic data, and obtain a standardized data set; The feature data in the standardized data set are averaged according to the feature data of all independent samples corresponding to the same gene under the same sample type to obtain the average feature data of the gene / site, and the average feature data of m genes / sites are set to form a combination, and s combinations are randomly selected from the standardized data set. The average feature data combination is used as input, and the analysis result of the sample type prediction model corresponding to the combination is used as output. The average feature data combination matrix X and the analysis result matrix Y of the sample type prediction model corresponding to the average feature data combination matrix are constructed, and these two matrices are combined to define the initial data set Ds, where s is the number of randomly selected average feature data combinations; The sample type prediction model is composed of a sample type accuracy calculation model and a calculation formula. The sample type accuracy calculation model is selected from a neural network model, a linear regression model, a support vector machine, a decision tree, a random forest, a k-nearest neighbor algorithm or a logistic regression model. The sample type accuracy calculation model takes the characteristic data of all independent samples of all sample types for each gene / site in the gene / site combination as input, and takes the sample type recognition accuracy of each gene / site combination for all sample types as output. The calculation formula calculates the comprehensive accuracy of sample type recognition for each characteristic data combination, and the calculation formula is: f(x)=α2Q1+β2Q2+γ2Q3、f(x)=median(accuracy)、 One or more combinations of, where f(x) is the function value of the comprehensive accuracy of sample type recognition, n is the number of sample types, accuracy is the recognition accuracy of each sample type obtained by the sample type accuracy calculation model, σ' is the standard deviation between the recognition accuracy of different sample types, log() is the logarithmic function with base 10, II() is the indicator function, p is the defined minimum accuracy threshold of each sample type, max() and min() are the maximum and minimum calculation formulas, Q1, Q2, Q3 are the 25th, 50th and 75th percentiles, median() is the median calculation formula, Ωi is the weight of the accuracy of the i-th sample type, β1, α1, α2, β2, γ2, Q1, Q2, Q3 are hyperparameters, which are defined by prior knowledge or manually (2) Using a two-factor vertical cross-linking core to perform feature combination screening cycles a. Select the averaged feature data combination with the highest comprehensive accuracy of sample type identification in the initial data set as the benchmark and the comparison object of the gene / site combination to be tested, that is, the current optimal gene / site combination; b. Reselecting feature data from the standardized data set and performing averaging to form a new averaged feature data combination, the new averaged feature data combination does not belong to the initial data set, and the selection method of the new averaged feature data combination is a heuristic search strategy combined with random selection or a completely random selection, the heuristic search strategy is selected from one or more of genetic algorithm, simulated annealing, particle swarm optimization, taboo search, ant colony optimization, differential evolution, stochastic gradient descent, adaptive search algorithm, and bee colony optimization; c. Calculate the two-factor vertical cross-link kernel function value between the averaged feature data combination of the new gene / site combination and the averaged feature data combination of the current optimal gene / site combination, select the gene / site combination corresponding to the averaged feature data combination with the largest function value as the target gene / site combination, calculate the comprehensive accuracy of sample type identification of the target gene / site combination using the sample type prediction model, and update the averaged feature data combination of the target gene / site combination and the corresponding comprehensive accuracy of sample type identification to the initial data set; d. Iteration: Repeat steps a, b, and c until the termination condition is met, and screen out the gene / site combination with the highest sample type specificity.
2. The method according to claim 1, characterized in that The calculation formula is 3. The method according to claim 1, characterized in that The selection method for the new averaged feature data combination is a genetic algorithm combined with random selection. The genetic algorithm includes a crossover operation: forming a new gene / site combination to be tested by recombining some genes of a better gene / site combination in the initial data set or a mutation operation: randomly modifying some genes in a better gene / site combination in the initial data set to form a new gene / site combination to be tested).
4. The method according to claim 1, characterized in that The termination conditions of step d are that the comprehensive recognition accuracy of the sample type prediction model does not improve for several consecutive cycles, reaches a manually set threshold, the number of cycles is manually specified, the improvement of the objective function value is lower than the threshold, the maximum improvement of the acquisition function value is lower than the threshold, the model training performance reaches the expectation, or the optimization time reaches the upper limit.
5. The method according to claim 3, characterized in that The termination condition of step d is that the comprehensive recognition accuracy of the sample type prediction model does not increase for 5-10 consecutive cycles, reaches 90-98% accuracy, or the number of cycles is 10-200 times.
6. The method according to claim 1, characterized in that The n is 2-100.
7. The method according to claim 1, characterized in that The neural network model is a multi-layer fully connected neural network model.
8. The method according to claim 1, characterized in that The characteristic data are biological characteristics obtained based on DNA sequencing data, RNA sequencing data, DNA methylation sequencing data, miRNA sequencing data, protein expression profile, DNA copy number variation profile, metabolomics data, genomics data, proteomics data, transcriptomics data, epigenomics data, immunoomics data, pharmacomics data, systems omics data, phenotypic omics data, lipidomics data, glycomics data, signalomics data, toxicomics data or evolutionomics data.
9. The method according to claim 8, characterized in that The biological feature is selected from methylation rate, gene expression level, mutation status, gene copy number, miRNA expression abundance, protein expression level, single nucleotide polymorphism frequency, structural variation copy number, epigenetic modification intensity, non-coding RNA expression abundance, phosphorylated protein ratio, protein interaction intensity, metabolite concentration, metabolic pathway activity, immune cell infiltration ratio, cytokine concentration, bacterial flora abundance, metabolite contribution, target inhibition efficiency, sugar metabolism concentration or lipid metabolism concentration.
10. The method according to claim 1, characterized in that The sample type is disease, race, age, gender or other biological or environmental factors that affect genetic characteristics.
11. A two-factor vertical cross-linking nuclear model for high-specificity gene / site combination screening, characterized in that The two-factor vertical cross-link kernel function is k is the value of the two-factor vertical cross-linking kernel function of the two groups of gene / site combinations, x1, x2 are the averaged feature data combination variables corresponding to the two groups of gene / site combinations, n is the number of sample types, and the value range of n is any integer greater than 1; α, β, γ, λ are hyperparameters defined by prior knowledge; BV is the basic variogram: EXP{} is an exponential function with the natural constant e as the base, Π is pi, arccos() is the inverse cosine function, x1i and x2i are the spatial multidimensional vectors of the averaged feature data combination corresponding to the two groups of gene / site combinations under the i-th sample type, ||x1i|| and ||x2i|| are the moduli of the spatial multidimensional vector, and max() is the maximum value function; σB.V. is the standard deviation between BVs in different sample types; log() is a logarithmic function with base 10; q is the minimum threshold of the basic variation function value set artificially; Ⅱ() is an indicator function, and its return value is 0 or 1. m is the number of genes / locus in the gene / locus combination, and its value range is a positive integer greater than 1. DV is the discrete variogram: sin{} is the sine function, Delta[] is the difference calculation function, Max() and Min() are the maximum and minimum value functions respectively, x is the combination of the averaged feature data of the jth gene / site in the gene / site combination corresponding to different sample types, ri represents the averaged feature data of the jth gene / site in the gene / site combination corresponding to the ith sample type, and rμ represents the mean of ri of the jth gene / site in all sample types.
12. Application of the two-factor vertical cross-linking nuclear model in gene / site combination screening as claimed in claim 11, characterized in that The screening includes screening of differentially expressed genes, differentially methylated genes, differentially mutated genes, differential copy number variation genes, coding genes of differentially expressed proteins, differentially expressed miRNAs, differentially methylated sites, differentially mutated sites, differentially hydroxymethylated modification sites, differentially acetylated modification sites, differentially frequent single nucleotide polymorphism sites, structural variation genes of differential copy numbers, genes of differential epigenetic modification intensity, non-coding RNA of differential expression abundance, coding genes of differential phosphorylation ratio proteins, coding genes of differential interaction intensity proteins, genes related to differential concentration metabolites, genes of differentially active metabolic pathways, coding genes of differentially infiltrating ratio immune cells, coding genes of differentially concentrated cytokines, coding genes of differentially abundant bacterial communities, coding genes of differentially contributing metabolites, and corresponding gene / site screening of differentially inhibited efficiency targets.
13. The use as claimed in claim 12, characterized in that the sample type for gene feature screening is selected from tumors, genetic abnormalities, metabolic abnormalities, diseases caused by infection with drug-resistant pathogens, autoimmune diseases, neurodegenerative diseases, cardiovascular diseases, infectious diseases, inflammatory diseases, rare diseases, endocrine diseases, blood system diseases, developmental disorders, mental illnesses, skin diseases, degenerative bone and joint diseases or diseases caused by environmental exposure.