Gene expression marker combination screening method and device for judging tuberculosis infection

By screening and analyzing the gene expression data set, a tri-classification model is constructed and gene expression markers are gradually added to form a combination of gene expression markers, which solves the problem of difficult to accurately distinguish active tuberculosis, latent tuberculosis infection and non-tuberculosis in the prior art, and achieves a high-accurate tuberculosis diagnosis.

CN120183500APending Publication Date: 2025-06-20鲲鹏基因(北京)科学仪器有限公司 +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510167963.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately distinguish between active tuberculosis, latent tuberculosis infection and non-tuberculosis, and a single marker is difficult to achieve ideal diagnostic results.

Method used

By screening the gene expression data sets related to tuberculosis infection from the gene expression database, it is divided into active tuberculosis, latency tuberculosis and non-tuberculosis infection sample groups, pretreatment and differential analysis are performed, gene expression data sets that meet the preset degree differences, a three-class model is constructed, and gene expression markers are gradually added to form a combination of gene expression markers.

Benefits of technology

Accurate identification of active tuberculosis, latent tuberculosis and non-tuberculosis has been achieved, the accuracy of tuberculosis diagnosis has been improved, and clinical treatment and prevention have been guided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183500A_ABST
    Figure CN120183500A_ABST
Patent Text Reader

Abstract

The invention discloses a gene expression marker combination screening method and device for judging tuberculosis infection. The method comprises the following steps: dividing a gene expression data set related to tuberculosis infection into a plurality of sample groups; each sample group is preprocessed; performing difference analysis on every two sample groups meeting the preset requirements, and screening out candidate gene expression marker sets meeting preset degree differences of all the sample groups; constructing a three-classification model for each candidate gene expression marker in the candidate gene expression marker set, and selecting the candidate gene expression marker corresponding to the three-classification model with the highest accuracy rate as an initial candidate gene expression marker; after candidate gene expression markers are added one by one, a three-classification model is reconstructed, and newly added gene expression markers are selected according to the accuracy of the reconstructed three-classification model to form a gene expression marker combination. The gene expression marker combination capable of accurately identifying different tuberculosis types can be screened out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technologies, and in particular, to a method and device for screening a combination of gene expression markers for judging tuberculosis infection. Background Art

[0002] This section aims to provide background or context for the embodiments of the present invention described in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] Timely and accurate diagnosis of tuberculosis is crucial for disease control and patient treatment. Existing tuberculosis diagnoses have many limitations. For example, the sensitivity of sputum smear microscopy is low, the culture method is time-consuming, and professional laboratory conditions are required; the specificity of imaging methods such as chest X-ray is not high, making it difficult to distinguish tuberculosis timely and accurately. In addition, it is currently difficult to effectively distinguish active tuberculosis, latent tuberculosis infection, and non-tuberculosis patients clinically.

[0004] In recent years, the research on biomarkers has provided new ideas for tuberculosis diagnosis. However, a single biomarker often fails to achieve an ideal diagnostic effect. Therefore, developing a multi-marker combination that can accurately identify active tuberculosis, latent tuberculosis, and non-tuberculosis to improve the accuracy of tuberculosis diagnosis and guide clinical treatment and prevention is an urgent problem to be solved. Summary of the Invention

[0005] An embodiment of the present invention provides a method for screening a combination of gene expression markers for judging tuberculosis infection, which is used to screen out a combination of gene expression markers that can accurately identify active tuberculosis, latent tuberculosis infection, and non-tuberculosis. The method includes:

[0006] Screening a gene expression data set related to tuberculosis infection from a gene expression database;

[0007] Dividing the gene expression data set into multiple sample groups, where the multiple sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group;

[0008] Performing preprocessing on each sample group to obtain multiple sample groups that meet preset requirements;

[0009] Performing pairwise differential analysis on multiple sample groups that meet preset requirements, and screening out gene expression data sets that meet the preset degree of difference in all sample groups to form a gene expression data set, which is used as a candidate gene expression marker set;

[0010] For each candidate gene expression marker in the candidate gene expression marker set, constructing a three-classification model, and selecting the candidate gene expression marker corresponding to the three-classification model with the highest accuracy as the initial candidate gene expression marker;

[0011] After adding candidate gene expression markers one by one, a three-classification model is reconstructed, and the newly added gene expression markers are selected based on the accuracy of the reconstructed three-classification model.

[0012] An initial candidate gene expression marker and all newly added gene expression markers are combined to form a gene expression marker combination.

[0013] An embodiment of the present invention also provides a screening device for a gene expression marker combination for judging tuberculosis infection, which is used to screen out a gene expression marker combination that can accurately identify active tuberculosis, latent tuberculosis, and non-tuberculosis. The device includes:

[0014] A gene expression dataset screening module for screening gene expression datasets related to tuberculosis infection from a gene expression database;

[0015] A sample group division module for dividing the gene expression dataset into multiple sample groups, where the multiple sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group;

[0016] A preprocessing module for preprocessing each sample group to obtain multiple sample groups that meet preset requirements;

[0017] A differential analysis module for performing pairwise differential analysis on multiple sample groups that meet preset requirements, screening out gene expression datasets that meet the preset degree of difference in all sample groups to form a gene expression dataset set, and using it as a candidate gene expression marker set;

[0018] A combination screening module for constructing a three-classification model for each candidate gene expression marker in the candidate gene expression marker set, selecting the candidate gene expression marker corresponding to the three-classification model with the highest accuracy as the initial candidate gene expression marker; after adding candidate gene expression markers one by one, reconstructing the three-classification model, and selecting the newly added gene expression markers based on the accuracy of the reconstructed three-classification model; combining the initial candidate gene expression markers and all newly added gene expression markers to form a gene expression marker combination.

[0019] An embodiment of the present invention also provides a gene expression marker combination for judging tuberculosis infection, which can accurately identify active tuberculosis, latent tuberculosis, and non-tuberculosis. The gene expression marker combination is obtained by using the method for screening a gene expression marker combination for judging tuberculosis infection as described above, and the gene expression marker combination includes at least one of GBP5, PSMD13, KNTC1, ZC3H3, DUSP3, ATG12, GNAT2, SNRPD2, PCM1, and EHD4.

[0020] An embodiment of the present invention further provides a kit for judging tuberculosis infection, which can accurately identify active tuberculosis, latent tuberculosis and non-tuberculosis. The kit includes reagents for judging tuberculosis infection and primer pairs corresponding to the foregoing gene expression marker combinations.

[0021] An embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method for screening gene expression marker combinations for judging tuberculosis infection is implemented.

[0022] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method for screening gene expression marker combinations for judging tuberculosis infection is implemented.

[0023] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned method for screening gene expression marker combinations for judging tuberculosis infection is implemented.

[0024] In the embodiment of the present invention, a gene expression data set related to tuberculosis infection is screened from a gene expression database; the gene expression data set is divided into multiple sample groups, and the multiple sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group; each sample group is preprocessed to obtain multiple sample groups meeting preset requirements; pairwise differential analysis is performed on the multiple sample groups meeting the preset requirements, and a gene expression data set that meets the preset degree of difference in all sample groups is screened out to form a gene expression data set, which is used as a candidate gene expression marker set; for each candidate gene expression marker in the candidate gene expression marker set, a three-classification model is constructed, and the candidate gene expression marker corresponding to the three-classification model with the highest accuracy is selected as the initial candidate gene expression marker; after adding candidate gene expression markers one by one, a three-classification model is reconstructed, and the newly added gene expression markers are selected through the accuracy of the reconstructed three-classification model; the initial candidate gene expression markers and all the newly added gene expression markers are formed into a gene expression marker combination. In the above process, since three different sample groups, namely an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group, are used and pairwise differential analysis is performed, the screened gene expression data set meets the degree of difference and has high accuracy; by adding candidate gene expression markers one by one and reconstructing the three-classification model, the accuracy of the newly added gene expression markers selected can be very high, so as to achieve the purpose of accurately identifying active tuberculosis, latent tuberculosis and non-tuberculosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0026] Figure 1 It is a flowchart of a method for screening a gene expression marker combination for judging tuberculosis infection in an embodiment of the present invention;

[0027] Figure 2 It is a curve of the change in accuracy rate during the forward feature selection process of the test set of a three-classification model in an embodiment of the present invention;

[0028] Figure 3 It is a schematic structural diagram of a device for screening a gene expression marker combination for judging tuberculosis infection in an embodiment of the present invention;

[0029] Figure 4 It is another schematic structural diagram of a device for screening a gene expression marker combination for judging tuberculosis infection in an embodiment of the present invention;

[0030] Figure 5 It is a schematic diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further details the embodiments of the present invention with reference to the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.

[0032] Figure 1 It is a flowchart of a method for screening a gene expression marker combination for judging tuberculosis infection in an embodiment of the present invention, including:

[0033] Step 101: Screen a gene expression data set related to tuberculosis infection from a gene expression database;

[0034] Step 102: Divide the gene expression data set into multiple sample groups, and the multiple sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group;

[0035] Step 103: Preprocess each sample group to obtain multiple sample groups that meet the preset requirements;

[0036] Step 104: Perform pairwise difference analysis on multiple sample groups that meet the preset requirements, and screen out the gene expression data sets that meet the preset degree of difference among all sample groups to form a gene expression data set, which is used as the candidate gene expression marker set.

[0037] Step 105: For each candidate gene expression marker in the candidate gene expression marker set, construct a three-classification model, and select the candidate gene expression marker corresponding to the three-classification model with the highest accuracy as the initial candidate gene expression marker.

[0038] Step 106: After adding candidate gene expression markers one by one, reconstruct the three-classification model, and select the newly added gene expression markers through the accuracy of the reconstructed three-classification model.

[0039] Step 107: Form a gene expression marker combination with the initial candidate gene expression markers and all the newly added gene expression markers.

[0040] In the embodiment of the present invention, since three different sample groups, namely the active tuberculosis infection sample group, the latent tuberculosis infection sample group, and the non-tuberculosis infection sample group, are used and pairwise difference analysis is performed, the screened gene expression data set meets the degree of difference and has high accuracy; by adding candidate gene expression markers one by one and reconstructing the three-classification model, the accuracy of the newly added gene expression markers selected can be very high, so as to achieve the purpose of accurately identifying active tuberculosis, latent tuberculosis, and non-tuberculosis. Each step is introduced in detail below.

[0041] In step 101, screen out the gene expression data sets related to tuberculosis infection from the gene expression database.

[0042] Specifically, the screened gene expression data sets should ensure the representativeness and diversity of the data. The gene expression database can be searched with the keyword "tuberculosis" to obtain multiple gene expression data sets related to tuberculosis infection. These gene expression data sets cover different regions, human characteristics, age stages, and detection platforms, increasing the representativeness and diversity of the research.

[0043] In step 102, divide the gene expression data set into multiple sample groups, and the multiple sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group.

[0044] In the embodiment of the present invention, the division of the 3 sample groups is very scientific and in line with the actual situation. Note that the consistency among the samples in each sample group should be ensured.

[0045] In step 103, preprocess each sample group to obtain multiple sample groups that meet the preset requirements.

[0046] The purpose of preprocessing is to ensure the consistency and comparability among samples. Specifically, by filtering aspects such as the geographical distribution diversity of sample sources, the consistency evaluation of different detection platforms, and the representativeness of clinical subtypes of tuberculosis, a sample group meeting the preset requirements is finally selected.

[0047] In one embodiment, preprocessing is performed on each sample group to obtain multiple sample groups meeting the preset requirements, including:

[0048] Perform unified naming processing on the gene expression data sets within each sample group;

[0049] Obtain the gene expression data sets within each sample group that meet the missing rate condition. If the gene expression data set is related to key pathways, an imputation strategy is used to fill the gene expression data set, otherwise, the K-nearest neighbor algorithm is used to fill the gene expression data set;

[0050] Eliminate the technical batch differences and biological batch differences among all sample groups.

[0051] Specifically, when performing unified naming processing, a dedicated package renaming software, or a conversion tool or correspondence table provided by bioinformatics data, is used to standardize the names of gene expression data sets, avoiding the inconsistency of gene names among different gene expression data sets. Among them, the names of genes related to tuberculosis immunity are manually reviewed to ensure their accuracy.

[0052] Since the gene expression data set is related to key pathways, the imputation strategy is used to fill the gene expression data set to consider the expression characteristics and correlations of tuberculosis genes.

[0053] It should be noted that the two steps of dividing into multiple sample groups and unified naming processing can be swapped in sequence, but both need to be before batch effect processing.

[0054] The K-nearest neighbor algorithm is a method for dealing with missing data values. Its main functions are: filling the missing gene expression value data, avoiding losing information by directly deleting genes with missing values; estimating the missing values by finding the expression values of the K most similar samples, and being able to better maintain the distribution characteristics of the data.

[0055] Technical batch differences refer to the systematic differences brought by different detection platforms, the differences in laboratory conditions and operation procedures, the differences caused by different reagent batches, and the differences in instrument equipment. Biological batch differences refer to the differences in sample sources (characteristics of different regions and people), the differences in sample processing and preservation conditions, and the differences in sample collection times.

[0056] In step 104, pairwise difference analysis is performed on multiple sample groups that meet the preset requirements, and gene expression data sets that meet the preset degree of difference in all sample groups are screened out to form a gene expression data set, which is used as a candidate gene expression marker set;

[0057] In one embodiment, pairwise difference analysis is performed on multiple sample groups that meet the preset requirements, and gene expression data sets that meet the preset degree of difference in all sample groups are screened out to form a gene expression data set, including:

[0058] Divide multiple sample groups that meet the preset requirements into a training set and a test set;

[0059] Any two sample groups in the training set are formed into a pair of analysis difference groups;

[0060] For each pair of analysis difference groups, the training set of the pair of analysis difference groups is divided into subsets of a preset number of parts, and the same number of different subsets are taken each time for difference analysis to obtain multiple gene expression data sets that meet the preset degree of difference. After repeating the preset number of times, the multiple gene expression data sets found by each subset are intersected to form the gene expression data set of the pair of analysis difference groups;

[0061] The gene expression data sets of the three pairs of analysis difference groups are unioned to obtain the gene expression data set that meets the preset degree of difference in all sample groups.

[0062] Specifically, there are three sample groups, both divided into a training set and a test set. For example, the training set ratio can be 70% and the test set ratio can be 30%. Stratified sampling is used to ensure that the sample ratios of each group (active tuberculosis, latent tuberculosis, non-tuberculosis infection) are roughly the same.

[0063] Any two sample groups in the training set are formed into a pair of analysis difference groups. Then there are three pairs of analysis difference groups, namely:

[0064] Active tuberculosis analysis difference group vs latent tuberculosis analysis difference group;

[0065] Active tuberculosis analysis difference group vs non-tuberculosis infection analysis difference group;

[0066] Latent tuberculosis analysis difference group vs non-tuberculosis infection analysis difference group.

[0067] The preset number of copies can be determined according to the actual situation. For example, taking the pair of differential groups to be analyzed, namely the active tuberculosis differential group to be analyzed vs the latent tuberculosis differential group to be analyzed, as an example, the active tuberculosis differential group to be analyzed is divided into 10 copies, and the latent tuberculosis differential group to be analyzed is divided into 10 copies. The preset number of repetitions is 10 times, and the subsets taken each time are different. For the first time, 9 subsets, namely 1, 2, 3, 4, 5, 6, 7, 8, 9, are taken; for the second time, 9 subsets, namely 1, 2, 3, 4, 5, 6, 7, 8, 10, are taken, and so on. For the first time, differential analysis is performed on the 9 subsets of the active tuberculosis differential group to be analyzed and the 9 subsets of the latent tuberculosis differential group to be analyzed. In each pairwise differential analysis, the Wilcoxon two-sample rank sum test is used. The FDR < 0.05 is set as the significance threshold. Then, those with FDR < 0.05 in the differential analysis results are considered to have significant differences. Executing once in this way can obtain multiple gene expression data sets that meet the preset degree of difference. Repeating 10 times, multiple gene expression data sets that meet the preset degree of difference are obtained each time. Taking the intersection of the gene expression data sets of these 10 times, the gene expression data set of this pair of differential groups to be analyzed is obtained. Finally, taking the union of the gene expression data sets of the three pairs of differential groups to be analyzed, the gene expression data set that meets the preset degree of difference for all sample groups is obtained, that is, the candidate gene expression marker set.

[0068] In step 105, for each candidate gene expression marker in the candidate gene expression marker set, a three-classification model is constructed, and the candidate gene expression marker corresponding to the three-classification model with the highest accuracy is selected as the initial candidate gene expression marker;

[0069] In one embodiment, for each candidate gene expression marker in the candidate gene expression marker set, constructing a three-classification model and selecting the candidate gene expression marker corresponding to the three-classification model with the highest accuracy as the initial candidate gene expression marker includes:

[0070] For each candidate gene expression marker in the candidate gene expression marker set, based on the preset model parameters, a three-classification model is constructed. The three-classification model is trained using the training set, and the trained three-classification model is tested using the test set to obtain the accuracy. The candidate gene expression marker corresponding to the three-classification model with the highest accuracy is selected as the initial candidate gene expression marker.

[0071] Specifically, the three-classification model can select any one of random forest (RF), logistic regression (LR), support vector machine (SVM), or gradient boosting tree (XGBoost), or an experiment can be performed for each type, and the three-classification model with the highest accuracy is selected.

[0072] A three-classification model can be constructed for each candidate gene expression marker. In the process corresponding to step 105, based on the preset model parameters, since the purposes of step 105 and step 106 are to screen gene expression markers, these two steps can use fixed preset model parameters. During training, each candidate gene expression marker in the training set is used to train the three-classification model corresponding to this candidate gene expression marker. In the embodiments of the present invention, the accuracy rate refers to the ratio of the number of samples correctly classified by the model to the total number of samples.

[0073] In step 106, after adding candidate gene expression markers one by one, the three-classification model is reconstructed, and the newly added gene expression markers are selected through the accuracy rate of the reconstructed three-classification model;

[0074] In the embodiments of the present invention, the process of adding candidate gene expression markers one by one for training is called forward feature selection. Figure 2 This is the accuracy rate change curve of the forward feature selection process of the test set of the three-classification model in the embodiments of the present invention, and it can be understood the change process of the accuracy rate of the reconstructed three-classification model during the process of adding candidate gene expression markers one by one.

[0075] In one embodiment, after adding candidate gene expression markers one by one, the three-classification model is reconstructed, and the newly added gene expression markers are selected through the accuracy rate of the reconstructed three-classification model, including:

[0076] Taking the three-classification model constructed by the initial candidate gene expression markers as the current three-classification model;

[0077] Adding the initial candidate gene expression markers to the gene expression marker combination, and repeating the following steps until the improvement amount of the accuracy rate of the three-classification model is not greater than the improvement threshold:

[0078] Adding each unadded candidate gene expression marker to the gene expression marker combination respectively, reconstructing the three-classification model based on the preset model parameters, training the reconstructed three-classification model using the training set, testing the trained reconstructed three-classification model using the test set, and obtaining the accuracy rate corresponding to each candidate gene expression marker;

[0079] Calculating the improvement amount of the highest accuracy rate of the reconstructed three-classification models corresponding to all candidate gene expression markers compared to the accuracy rate of the current three-classification model;

[0080] If the improvement amount is greater than the improvement threshold, taking the reconstructed three-classification model with the highest accuracy rate as the current three-classification model, and taking the candidate gene expression marker corresponding to the current three-classification model as the newly added gene expression marker.

[0081] The termination condition of the above loop process is that the accuracy improvement of the three-classification model is not greater than the improvement threshold. Then, each time, the gene expression marker that can maximize the accuracy improvement of the three-classification model is used as the newly added gene expression marker, so as to finally screen out the gene expression marker combination.

[0082] In step 107, an initial candidate gene expression marker and all the newly added gene expression markers are combined to form a gene expression marker combination.

[0083] In the embodiment of the present invention, the gene expression marker combination is at least one of GBP5, PSMD13, KNTC1, ZC3H3, DUSP3, ATG12, GNAT2, SNRPD2, PCM1, and EHD4.

[0084] After obtaining the gene expression marker combination, a discrimination classification model can be constructed to determine the tuberculosis infection type of the sample according to the gene expression of the input sample.

[0085] In one embodiment, the method further includes:

[0086] Construct a discrimination classification model according to the gene expression marker combination and the corresponding tuberculosis infection type label, where the tuberculosis infection types include active tuberculosis infection, latent tuberculosis infection, and non-tuberculosis infection;

[0087] Take the initialized model parameters as the current model parameters, and repeat the following steps until the accuracy of the discrimination classification model reaches the preset accuracy requirement, and output the trained discrimination classification model, where the trained discrimination classification model is used to determine the tuberculosis infection type of the sample according to the gene expression of the input sample:

[0088] Use the training set to train the current model parameters of the discrimination classification model;

[0089] Use the test set to test the discrimination classification model to obtain the accuracy.

[0090] Specifically, the discrimination classification model can be constructed by using a random forest (RF), logistic regression (LR), support vector machine (SVM), or gradient boosting tree (XGBoost), and it also belongs to a three-classification model. The purpose of constructing the discrimination classification model this time is to train the optimal model parameters to achieve the best performance of the discrimination classification model. The above training set is the initial gene expression data set in step 101.

[0091] In addition, other independent validation sets (i.e., unfamiliar data sets) can be used to further evaluate the generalization ability of the trained discrimination classification model.

[0092] A single combination of gene expression markers cannot achieve the diagnosis and differentiation of patients with active tuberculosis, latent tuberculosis, and non-tuberculosis infections. The combination of gene expression markers (10 genes such as GBP5, PSMD13, KNTC1, etc.) is only a theoretical marker, which indicates "which genes" can be used to distinguish different types of tuberculosis infections. To detect the expression levels of these genes in actual clinical applications, specific primer-probes must be used; therefore, only the primer pairs designed based on this combination of gene expression markers can achieve the diagnosis and differentiation of patients with active tuberculosis, latent tuberculosis, and non-tuberculosis infections. Therefore, the steps to obtain the primer pairs for each gene expression marker are given below.

[0093] In one embodiment, the method further includes:

[0094] For each gene expression marker in the combination of gene expression markers, obtain the cDNA sequence of the gene expression marker, screen the primer candidate regions of the cDNA sequence, and obtain the first primer candidate regions;

[0095] Perform genome specificity verification on the screening of the first primer candidate regions to obtain the second primer candidate regions with a specificity score greater than the specificity threshold;

[0096] By comparing the human genome database, exclude the cross-amplification regions in the second primer candidate regions to obtain the primer pairs for each gene expression marker, and the primer pairs are used to identify the gene expression markers of the samples.

[0097] After designing the primer pairs for each gene expression marker, the expression levels of these gene expression markers can be detected by qPCR (quantitative real-time polymerase chain reaction) technology; therefore, the combination of gene expression markers is the theoretical basis, while the primer pairs are the practical application tools, and both are indispensable. Ultimately, it is the primer (also called primer-probe) pairs that achieve the diagnostic function.

[0098] Specifically, the primer design follows the following parameters: primer length 18 - 25bp, GC content 40 - 60%, the annealing temperature (Tm) values of the forward and reverse primers are between 55 - 65°C, and the Tm difference between the forward and reverse primers does not exceed 3°C. Calculate the complementarity, hairpin structure probability, and heterodimer formation risk of the primers through algorithms to ensure the efficiency and specificity of the primers.

[0099] Use special software to predict the secondary structure of the primers and evaluate their stability.

[0100] During the final screening, the specificity score (>0.9), amplification efficiency (>0.8), and structural stability.

[0101] The primer pairs for each gene expression marker in the aforementioned combination of gene expression markers are introduced below.

[0102] Primer pair for GBP5:

[0103] Forward primer: 5'-CTGTCTGCCATTACGCAACCTG-3'

[0104] Reverse primer: 5'-GTGTGAGACTGCACCGTAGATG-3'

[0105] Primer pair for PSMD13:

[0106] Forward primer: 5'-AGTTCTCGTTTCCAGTGAAACC-3'

[0107] Reverse primer: 5'-CTCAAAGGAAATGACCACTGG-3'

[0108] Primer pair for KNTC1:

[0109] Forward primer: 5'-GCCAGATATCTGAGTGTTCC-3'

[0110] Reverse primer: 5'-TCTCAGTACAAGGGGTAGAAGG-3'

[0111] Primer pair for ZC3H3:

[0112] Forward primer: 5'-CAGTTCTGTTACACTCAGTGG-3'

[0113] Reverse primer: 5'-ATCATGACTACCCTGGAAGG-3'

[0114] Primer pair for DUSP3:

[0115] Forward primer: 5'-TGCCGACTTCATTGACCAGGCT-3'

[0116] Reverse primer: 5'-CGTCCATCTTCTGCCGCATCAT-3'

[0117] Primer pair for ATG12:

[0118] Forward primer: 5'-GCTATCCTCAACCCTTTATTGC-3'

[0119] Reverse primer: 5'-CAAAGTGCTGGGATTACAGGA-3'

[0120] Primer pair for GNAT2:

[0121] Forward primer: 5'-ATCTGCAACACACCCTTCAC-3'

[0122] Reverse primer: 5'-AGCTAGGGAATATCTCAGAGG-3'

[0123] Primer pair for SNRPD2:

[0124] Forward primer: 5'-ACCAGAAAGATTCCAGGACAG-3'

[0125] Reverse primer: 5'-TCAACTCTGCATGTCTCAGG-3'

[0126] Primer pair for PCM1:

[0127] Forward primer: 5'-TTGTACTTCGATCTCCCTTTCC-3'

[0128] Reverse primer: 5'-TGATCTCCTGACCTCGTCAT-3'

[0129] Primer pair for EHD4:

[0130] Forward primer: 5'-ATACGCACGAGTCTCAAACTC-3'

[0131] Reverse primer: 5'-CTCCAGTCTGGGAGTTTATTTC-3'

[0132] In the embodiments of the present invention, a screening device for a gene expression marker combination for judging tuberculosis infection is also provided, as described in the following embodiments. Since the principle of the device for solving the problem is similar to that of the screening method for a gene expression marker combination for judging tuberculosis infection, the implementation of the device can refer to the implementation of the screening method for a gene expression marker combination for judging tuberculosis infection, and the repeated parts will not be described again.

[0133] Figure 3 It is a schematic structural diagram of the screening device for a gene expression marker combination for judging tuberculosis infection in the embodiments of the present invention, including:

[0134] A gene expression dataset screening module 301, configured to screen a gene expression dataset related to tuberculosis infection from a gene expression database;

[0135] A sample group division module 302, configured to divide the gene expression dataset into multiple sample groups, and the multiple sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group;

[0136] The preprocessing module 303 is used to preprocess each sample group to obtain multiple sample groups that meet the preset requirements;

[0137] The differential analysis module 304 is used to perform pairwise differential analysis on multiple sample groups that meet the preset requirements, screen out the gene expression data sets that meet the preset degree of difference in all sample groups to form a gene expression data set, and use it as a candidate gene expression marker set;

[0138] The combination screening module 305 is used to construct a three-classification model for each candidate gene expression marker in the candidate gene expression marker set, select the candidate gene expression marker corresponding to the three-classification model with the highest accuracy as the initial candidate gene expression marker; after adding candidate gene expression markers one by one, reconstruct the three-classification model, and select the newly added gene expression markers through the accuracy of the reconstructed three-classification model; form a gene expression marker combination with the initial candidate gene expression marker and all the newly added gene expression markers.

[0139] In one embodiment, the preprocessing module is used for:

[0140] Perform unified naming processing on the gene expression data sets within each sample group;

[0141] Obtain the gene expression data sets within each sample group that meet the missing rate condition. If the gene expression data set is a gene related to a key pathway, use an imputation strategy to fill the gene expression data set, otherwise, use the K-nearest neighbor algorithm to fill the gene expression data set;

[0142] Eliminate the technical batch differences and biological batch differences between all sample groups.

[0143] In one embodiment, the differential analysis module is used for:

[0144] Divide multiple sample groups that meet the preset requirements into a training set and a test set;

[0145] Form a pair of differential analysis groups with any two sample groups in the training set;

[0146] For each pair of differential analysis groups, divide the training set of the pair of differential analysis groups into subsets of a preset number of copies, take different subsets of the same number of copies for differential analysis each time, obtain multiple gene expression data sets that meet the preset degree of difference, and after repeating the preset number of times, take the intersection of the multiple gene expression data sets found by each subset as the gene expression data set of the pair of differential analysis groups;

[0147] Take the union of the gene expression data sets of the three pairs of differential analysis groups to obtain the gene expression data set that meets the preset degree of difference in all sample groups.

[0148] In one embodiment, the combined screening module is configured to:

[0149] For each candidate gene expression marker in the set of candidate gene expression markers, based on preset model parameters, construct a three-classification model, train the three-classification model using a training set, test the trained three-classification model using a test set, obtain an accuracy rate, and select the candidate gene expression marker corresponding to the three-classification model with the highest accuracy rate as the initial candidate gene expression marker.

[0150] In one embodiment, the combined screening module is configured to:

[0151] Take the three-classification model constructed by the initial candidate gene expression marker as the current three-classification model;

[0152] Add the initial candidate gene expression marker to the gene expression marker combination, and repeat the following steps until the increase in the accuracy rate of the three-classification model is not greater than the increase threshold:

[0153] Add each candidate gene expression marker that has not been added to the gene expression marker combination respectively, based on preset model parameters, reconstruct a three-classification model using the gene expression marker combination, train the reconstructed three-classification model using a training set, test the trained reconstructed three-classification model using a test set, and obtain the accuracy rate corresponding to each candidate gene expression marker;

[0154] Calculate the increase in the accuracy rate of the reconstructed three-classification model corresponding to all candidate gene expression markers compared to the accuracy rate of the current three-classification model;

[0155] If the increase is greater than the increase threshold, take the reconstructed three-classification model with the highest accuracy rate as the current three-classification model, and take the candidate gene expression marker corresponding to the current three-classification model as the newly added gene expression marker.

[0156] Figure 4 This is another structural schematic diagram of the gene expression marker combination screening device for judging tuberculosis infection in the embodiments of the present invention. In one embodiment, the device further includes a discrimination and classification model construction module 401, which is configured to:

[0157] Construct a discrimination and classification model according to the gene expression marker combination and the corresponding tuberculosis infection type label, where the tuberculosis infection type includes active tuberculosis infection, latent tuberculosis infection, and non-tuberculosis infection;

[0158] Take the initialized model parameters as the current model parameters, and repeat the following steps until the accuracy rate of the discrimination and classification model reaches the preset accuracy rate requirement, and output the trained discrimination and classification model, where the trained discrimination and classification model is used to judge the tuberculosis infection type of the sample according to the gene expression of the input sample:

[0159] Train the current model parameters of the discrimination classification model using the training set;

[0160] Test the discrimination classification model using the test set to obtain the accuracy rate.

[0161] In one embodiment, the device further includes a primer pair generation module 402, configured to:

[0162] For each gene expression marker in the gene expression marker combination, obtain the cDNA sequence of the gene expression marker, screen the primer candidate segments of the cDNA sequence, and obtain the first primer candidate segments;

[0163] Perform genome specificity verification on the screening of the first primer candidate segments to obtain second primer candidate segments with a specificity score greater than the specificity threshold;

[0164] By comparing with the human genome database, exclude the cross-amplification regions in the second primer candidate segments to obtain the primer pairs for each gene expression marker, and the primer pairs are used to identify the gene expression markers of the sample.

[0165] An embodiment of the present invention also provides a gene expression marker combination for judging tuberculosis infection, which is obtained by using the foregoing method for screening a gene expression marker combination for judging tuberculosis infection. The gene expression marker combination includes at least one of GBP5, PSMD13, KNTC1, ZC3H3, DUSP3, ATG12, GNAT2, SNRPD2, PCM1, and EHD4.

[0166] An embodiment of the present invention also provides a kit for judging tuberculosis infection, including reagents for judging tuberculosis infection and primer pairs corresponding to the foregoing gene expression marker combination.

[0167] In a gene expression detection kit, the kit may include: specific primers, amplification reagents, detection reagents, control reagents, and standards.

[0168] Specifically, the samples for which the kit is used for detection include one or any combination of cell lines, histological sections, tissue biopsies, paraffin-embedded tissues, body fluids, feces, colonic effluents, urine, plasma, serum, whole blood, isolated blood cells, and cells isolated from blood.

[0169] Here, it is explained that the reagent is a broad concept, including various chemical substances and substances used in biological experiments.

[0170] A primer specifically refers to a short single-stranded nucleic acid sequence used in DNA / RNA amplification.

[0171] A specific embodiment is given below to illustrate the effectiveness of the solution proposed in the embodiment of the present invention.

[0172] 1. Sample selection:

[0173] 100 samples of active tuberculosis infection confirmed by the clinical reference method, 80 samples of latent tuberculosis, and 50 samples without infection. The sample type is whole blood sample.

[0174] 2. RNA sample extraction:

[0175] RNA extraction was performed on the whole blood samples in a timely manner. The specific steps are as follows:

[0176] (1) Add the whole blood sample to the reagent, mix well, and let it stand at room temperature for 5 minutes.

[0177] (2) Add chloroform, shake vigorously and let it stand for 3 minutes, then centrifuge at 12000g for 10 minutes. After stratification, take the supernatant.

[0178] (3) Add an equal volume of isopropanol, mix gently, let it stand at room temperature for 10 minutes, and then centrifuge at 12000g for 10 minutes to precipitate RNA.

[0179] (4) Wash the RNA precipitate with 75% ethanol, dry it after centrifugation, and finally dissolve it in water of the preset type to obtain the required RNA sample.

[0180] 3. Detection of RNA expression markers:

[0181] The real-time quantitative PCR (qPCR) technique was used to detect RNA expression markers. Specifically, using the extracted RNA as a template, three pairs of primer pairs of the gene expression marker combination proposed in the embodiment of the present invention were added for qPCR experiments to synthesize cDNA. In this embodiment, the three pairs of primers are the primer pairs of the three gene expression markers GBP5, PSMD13, and DUSP3. Table 1 shows the reaction system for the qPCR experiment.

[0182] Table 1 Reaction system

[0183]

[0184] Table 2 shows an example of the reaction conditions and instruments in the embodiment of the present invention.

[0185] Table 2 Reaction conditions and instruments

[0186] Instrument ROCGENE Archimed X4 qPCR program Temperature Pre-denaturation 95℃ Denaturation 95℃ Annealing and extension 60℃

[0187] 4. Analysis of qPCR results:

[0188] The statistical results of the verification are shown in Table 3.

[0189] Table 3 Sample verification results

[0190]

[0191] As can be seen from Table 3, the detection accuracies of active tuberculosis samples, latent tuberculosis samples, and non-infected samples are as high as 85%, 81.2%, and 94% respectively. It shows that the combination of gene expression markers selected in the embodiments of the present invention has good sensitivity and specificity in detecting active tuberculosis.

[0192] In summary, in the method and device proposed in the embodiments of the present invention, a gene expression data set related to tuberculosis infection is screened from a gene expression database; the gene expression data set is divided into multiple sample groups, and the multiple sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group; each sample group is preprocessed to obtain multiple sample groups that meet preset requirements; pairwise differential analysis is performed on the multiple sample groups that meet the preset requirements, and a gene expression data set that meets the preset degree of difference in all sample groups is screened out to form a gene expression data set, which is used as a candidate gene expression marker set; for each candidate gene expression marker in the candidate gene expression marker set, a three-classification model is constructed, and the candidate gene expression marker corresponding to the three-classification model with the highest accuracy is selected as the initial candidate gene expression marker; after adding candidate gene expression markers one by one, a three-classification model is reconstructed, and the newly added gene expression markers are selected through the accuracy of the reconstructed three-classification model; the initial candidate gene expression markers and all the newly added gene expression markers are combined to form a gene expression marker combination. In the above process, since three different sample groups, namely an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group, are used and pairwise differential analysis is performed, the screened gene expression data set meets the degree of difference and has high accuracy; by adding candidate gene expression markers one by one and reconstructing the three-classification model, the accuracy of the newly added gene expression markers selected can be very high, so as to achieve the purpose of accurately identifying active tuberculosis, latent tuberculosis, and non-tuberculosis.

[0193] The embodiments of the present invention also provide a computer device, Figure 5 which is a schematic diagram of the computer device in the embodiments of the present invention. The computer device 500 includes a memory 510, a processor 520, and a computer program 530 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 530, the above method for screening a gene expression marker combination for judging tuberculosis infection is implemented.

[0194] The embodiments of the present invention also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method for screening a gene expression marker combination for judging tuberculosis infection is implemented.

[0195] An embodiment of the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for screening a gene expression marker combination for tuberculosis infection determination.

[0196] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0197] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0198] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0199] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0200] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for screening a combination of gene expression markers for determining tuberculosis infection, characterized in that: include: Screening gene expression datasets related to tuberculosis infection from gene expression databases; Dividing the gene expression data set into a plurality of sample groups, wherein the plurality of sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group; Preprocess each sample group to obtain multiple sample groups that meet preset requirements; Performing intergroup difference analysis on multiple sample groups that meet the preset requirements, screening out gene expression data sets that meet the preset degree of difference from all sample groups to form a gene expression data set, and using it as a candidate gene expression marker set; For each candidate gene expression marker in the candidate gene expression marker set, a three-classification model is constructed, and the candidate gene expression marker corresponding to the three-classification model with the highest accuracy is selected as the initial candidate gene expression marker; After adding candidate gene expression markers one by one, the three-classification model is reconstructed, and the newly added gene expression markers are selected according to the accuracy of the reconstructed three-classification model; The initial candidate gene expression markers and all the newly added gene expression markers are combined to form a gene expression marker combination.

2. The method according to claim 1, characterized in that Each sample group is preprocessed to obtain multiple sample groups that meet the preset requirements, including: Gene expression datasets in each sample group were uniformly named; Obtain a gene expression dataset that meets the missing rate conditions in each sample group. If the gene expression dataset is a key pathway-related gene, use the interpolation strategy to fill the gene expression dataset. Otherwise, use the K nearest neighbor algorithm to fill the gene expression dataset. Eliminate technical and biological batch variability between all sample groups.

3. The method according to claim 1, characterized in that Perform intergroup difference analysis on multiple sample groups that meet the preset requirements, and screen out gene expression data sets that meet the preset degree of difference from all sample groups to form a gene expression data set, including: Divide multiple sample groups that meet preset requirements into training sets and test sets; Any two sample groups in the training set form a pair of difference groups to be analyzed; For each difference group to be analyzed, the training set of the difference group to be analyzed is divided into a preset number of subsets, and the same number of different subsets are taken for difference analysis each time to obtain multiple gene expression data sets that meet the preset degree of difference. After repeating the preset number of times, the intersection of the multiple gene expression data sets found in each subset is taken as the gene expression data set of the difference group to be analyzed; The gene expression data sets of the three difference groups to be analyzed are combined to obtain the gene expression data sets of all sample groups that meet the preset degree of difference.

4. The method according to claim 1, characterized in that For each candidate gene expression marker in the candidate gene expression marker set, a three-classification model is constructed, and the candidate gene expression marker corresponding to the three-classification model with the highest accuracy is selected as the initial candidate gene expression marker, including: For each candidate gene expression marker in the candidate gene expression marker set, a three-classification model is constructed based on preset model parameters, the three-classification model is trained using a training set, and the trained three-classification model is tested using a test set to obtain the accuracy, and the candidate gene expression marker corresponding to the three-classification model with the highest accuracy is selected as the initial candidate gene expression marker.

5. The method according to claim 4, characterized in that After adding candidate gene expression markers one by one, the three-classification model was reconstructed, and the newly added gene expression markers were selected based on the accuracy of the reconstructed three-classification model, including: The three-classification model constructed by the initial candidate gene expression markers is used as the current three-classification model; The initial candidate gene expression markers are added to the gene expression marker combination, and the following steps are repeated until the accuracy improvement of the three-classification model is no greater than the improvement threshold: Each candidate gene expression marker that has not been added is added to the gene expression marker combination, and based on the preset model parameters, the three-classification model is reconstructed using the gene expression marker combination, the training set is used to train the reconstructed three-classification model, and the test set is used to test the trained reconstructed three-classification model to obtain the accuracy rate corresponding to each candidate gene expression marker; Calculate the improvement of the highest value of the accuracy of the reconstructed three-classification model corresponding to all candidate gene expression markers compared with the accuracy of the current three-classification model; If the improvement is greater than the improvement threshold, the reconstructed three-classification model with the highest accuracy is used as the current three-classification model, and the candidate gene expression markers corresponding to the current three-classification model are used as the newly added gene expression markers.

6. The method according to claim 1, characterized in that Also includes: constructing a discrimination and classification model based on the combination of gene expression markers and the corresponding tuberculosis infection type labels, wherein the tuberculosis infection types include active tuberculosis infection, latent tuberculosis infection and non-tuberculosis infection; The initialized model parameters are used as the current model parameters, and the following steps are repeatedly performed until the accuracy of the identification and classification model reaches the preset accuracy requirement, and the trained identification and classification model is output. The trained identification and classification model is used to determine the tuberculosis infection type of the sample according to the gene expression of the input sample: Using the training set to train the current model parameters of the discriminative classification model; The test set is used to test the identification and classification model to obtain the accuracy.

7. The method according to claim 1 or 6, characterized in that Also includes: For each gene expression marker in the gene expression marker combination, a cDNA sequence of the gene expression marker is obtained, and a primer candidate segment of the cDNA sequence is screened to obtain a first primer candidate segment; Performing genome-specific verification on the screening of the first primer candidate segment to obtain a second primer candidate segment having a specificity score greater than a specificity threshold; By comparing the human genome database and excluding the cross-amplification region in the second primer candidate segment, a primer pair for each gene expression marker is obtained, and the primer pair is used to identify the gene expression marker of the sample.

8. A gene expression marker combination screening device for judging tuberculosis infection, characterized in that: include: A gene expression data set screening module, used to screen gene expression data sets related to tuberculosis infection from a gene expression database; A sample group division module, used for dividing the gene expression data set into a plurality of sample groups, wherein the plurality of sample groups include an active tuberculosis infection sample group, a latent tuberculosis infection sample group, and a non-tuberculosis infection sample group; A preprocessing module is used to preprocess each sample group to obtain multiple sample groups that meet preset requirements; The difference analysis module is used to perform difference analysis between two groups of samples that meet the preset requirements, and screen out the gene expression data sets that meet the preset degree of difference of all sample groups to form a gene expression data set, which is used as a candidate gene expression marker set; The combined screening module is used to construct a three-classification model for each candidate gene expression marker in the candidate gene expression marker set, and select the candidate gene expression marker corresponding to the three-classification model with the highest accuracy as the initial candidate gene expression marker; after adding the candidate gene expression markers one by one, reconstruct the three-classification model, and select the newly added gene expression marker according to the accuracy of the reconstructed three-classification model; the initial candidate gene expression marker and all the newly added gene expression markers form a gene expression marker combination.

9. A combination of gene expression markers for determining tuberculosis infection, characterized in that: The method for screening a combination of gene expression markers for determining tuberculosis infection according to any one of claims 1 to 7 is used to obtain the gene expression marker combination, wherein the gene expression marker combination includes at least one of GBP5, PSMD13, KNTC1, ZC3H3, DUSP3, ATG12, GNAT2, SNRPD2, PCM1 and EHD4.

10. A kit for determining tuberculosis infection, characterized in that: It comprises a reagent for judging tuberculosis infection and a primer pair corresponding to the gene expression marker combination according to claim 9.

11. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

13. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.