Screening method and application of bovine tuberculosis diagnostic marker
Through machine learning, genes such as ADGRA3, PRELP and HBB were screened as diagnostic markers for bovine tuberculosis, and a multi-parameter joint diagnostic model was constructed, which solved the problem of accuracy in the diagnosis of bovine tuberculosis and realized an efficient and accurate diagnostic tool.
Patent Information
- Application Number
- CN202510763963.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-03
AI Technical Summary
In the existing technology, the diagnostic methods for bovine tuberculosis have poor accuracy. Traditional detection methods cannot meet the needs of efficient and accurate diagnosis. In addition, there are few markers for bovine tuberculosis, and the diagnostic efficacy of a single marker is limited, which cannot effectively respond to the continued occurrence of bovine tuberculosis.
By using machine learning technology combined with a comprehensive gene expression database, genes such as ADGRA3, PRELP and HBB were screened as diagnostic markers, and a multi-parameter joint diagnostic model was constructed. The expression levels of these genes were detected by primer pairs to construct a diagnostic tool.
It improves the diagnostic accuracy and precision of bovine tuberculosis, provides a multi-parameter joint diagnostic model, enhances the differentiation ability, realizes the accurate screening of bovine tuberculosis and reduces the risk of human-animal transmission.
Smart Images

Figure CN120738337A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bovine tuberculosis diagnosis, and in particular to a screening method and application of a bovine tuberculosis diagnosis marker. Background Art
[0002] Bovine tuberculosis (bTB) is a chronic zoonotic disease caused by Mycobacterium bovis, a bacterium that severely impacts livestock production and threatens public health worldwide. According to the World Organization for Animal Health (WOAH), bTB causes billions of dollars in economic losses worldwide. Furthermore, the frequent culling of infected cattle further exacerbates the economic burden and significantly hinders the sustainable development of the livestock industry.
[0003] Mycobacterium bovis not only affects animal health but can also be transmitted to humans through the respiratory tract or through inadequately sterilized dairy products. This is particularly serious in underdeveloped regions, posing a serious public health problem. Various in vivo and postmortem testing methods are available for diagnosing bovine tuberculosis. However, as an intracellular parasite, M. tuberculosis is primarily a cellular immune response, leading to inter-individual variability in the incidence of tuberculosis infection. Furthermore, host genetic factors influence the outcome of M. tuberculosis infection. This results in poor accuracy for traditional detection methods. Sequencing and bioinformatics analysis have further improved the accuracy of molecular diagnostics. Biomarkers are biological characteristics that can be used to indicate physiological or pathological processes or drug responses. Changes in gene expression reflect molecular changes in cells or tissues under specific conditions and can therefore serve as important biomarkers for disease diagnosis, prognosis assessment, or therapeutic efficacy monitoring. Therefore, methods for detecting gene expression fall under the category of biomarker detection. Machine learning (ML) methods have been widely used in the discovery and application of biomarkers for the diagnosis and management of bovine tuberculosis (bTB). Combining machine learning with biomarker research can significantly improve the accuracy, efficiency, and predictive power of diagnostic methods. For example, through ML-based screening, mRNA FCRL1 has been identified as a candidate biomarker for the early diagnosis of bovine tuberculosis. However, while markers for bovine tuberculosis detection have emerged, existing patents are still mostly focused on human tuberculosis. There are relatively few markers for bovine tuberculosis, which cannot cope with the increasing incidence of bovine tuberculosis. Furthermore, the diagnostic efficacy of a single marker is limited, and a combination of markers for bTB is not yet mature. Therefore, the development of diagnostic tools based on novel molecular markers is of urgent importance for achieving accurate screening for bovine tuberculosis in cattle herds and reducing the risk of human-animal transmission. Summary of the Invention
[0004] In response to the above-mentioned prior art, the present invention aims to provide a method and application for screening diagnostic markers for bovine tuberculosis. This study aims to utilize machine learning techniques, combined with RNA-seq data published in the Gene Expression Omnibus (GEO), to analyze gene expression profiles in blood samples from cattle infected with Mycobacterium bovis and healthy cattle, screening for and identifying biomarkers with diagnostic value. Furthermore, a machine learning diagnostic model will be constructed based on the selected marker genes, aiming to provide a new detection tool for bovine tuberculosis diagnosis.
[0005] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect of the present invention, a diagnostic marker set for bovine tuberculosis is provided, wherein the diagnostic marker set comprises at least two of ADGRA3, PRELP or HBB.
[0006] Furthermore, the Gene ID of ADGRA3 in NCBI is 166647, the Gene ID of PRELP in NCBI is 5549, and the Gene ID of HBB in NCBI is 3043.
[0007] The second aspect of the present invention provides the use of primer pairs for detecting the diagnostic marker set in a sample in the preparation of a diagnostic product for bovine tuberculosis.
[0008] Furthermore, the primer pair for detecting ADGRA3 is SeqID No: 1-2, the primer pair for detecting PRELP is SeqID No: 3-4, and the primer pair for detecting HBB is SeqID No: 5-6.
[0009] A third aspect of the present invention provides a product for diagnosing bovine tuberculosis, the product comprising a reagent for detecting the expression level of the diagnostic marker set in a sample.
[0010] A fourth aspect of the present invention provides a method for screening the diagnostic marker set for bovine tuberculosis, comprising the following steps: The genetic data of healthy cows and cows with bovine tuberculosis were compared and screened to obtain differentially expressed genes; the differentially expressed genes were verified by single gene clinical samples; and single genes verified by clinical samples were selected to form a diagnostic marker set for bovine tuberculosis.
[0011] In a fifth aspect, the present invention provides a pharmaceutical composition for treating bovine tuberculosis, comprising: an inhibitor of the functional expression of ADGRA3, PRELP or HBB.
[0012] Furthermore, it also includes pharmaceutically acceptable carriers or excipients.
[0013] Beneficial effects of the present invention: This study, using machine learning to systematically screen and validate with clinical samples, successfully identified three marker genes with specific diagnostic value for bTB infection as bovine tuberculosis diagnostic markers, namely ADGRA3, PRELP, and HBB, from among numerous previously unreported genes. The resulting marker genes were combined to construct a multi-parameter combined diagnostic model (training set AUC = 0.959, test set AUC = 0.847). This model's comprehensive discriminatory power improved compared to traditional single-gene detection methods and was also at a higher level than existing multi-gene detection methods, providing a new solution for the accurate diagnosis of bTB. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 are the chip data before and after normalization, where Figure 1 A in the middle is the chip data before normalization. Figure 1 B in the middle is the chip data after normalization.
[0015] Figure 2 Figure 2 is a heat map of differentially expressed genes (Top 20).
[0016] Figure 3 Volcano plot of differentially expressed genes.
[0017] Figure 4 This is a KEGG bubble chart.
[0018] Figure 5 Screening characteristic genes for support vector machine.
[0019] Figure 6 Screening of characteristic genes for LASSO.
[0020] Figure 7 Screening characteristic genes for random forest.
[0021] Figure 8 Screening characteristic genes for the intersection of three machine learning methods.
[0022] Figure 9 is the expression of core genes in the training set.
[0023] Figure 10 The expression of core genes in the validation set.
[0024] Figure 11 is the diagnostic ability of the core genes on the training set.
[0025] Figure 12 is the diagnostic ability of the core genes on the validation set.
[0026] Figure 13The mRNA expression of core genes in clinical samples, * indicates p < 0.05.
[0027] Figure 14 The AUC indicators of the training set and the validation set are ranked in the top 50 models.
[0028] Figure 15 is the AUC of the polygenic model, where Figure 15 A in the middle is the AUC of the multi-gene model on the training set, Figure 15 Figure B is the AUC of the multi-gene model on the validation set. DETAILED DESCRIPTION
[0029] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.
[0030] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the technical solution of the present application will be described in detail below with reference to specific embodiments.
[0031] The present invention first divides a data set from a public database (NCBI-GEO) into a training set and a validation set. The training set includes 49 tuberculosis-positive cattle (experimental group) and 49 tuberculosis-negative cattle (control group). The validation set includes 14 tuberculosis-positive cattle (experimental group) and 11 tuberculosis-negative cattle (control group). Using three machine learning algorithms, 14 potential core genes were screened in the training set, of which four genes were differentially expressed in the test set. Among them, three genes, ADGRA3, PRELP, and HBB, were differentially expressed in clinical samples. The single gene's ability to distinguish bTB in the training set and validation set was between 0.694 and 0.727, and in the validation set, between 0.738 and 0.846. A multi-gene diagnosis model was constructed using nine machine learning models and their parameter combinations. The optimal model was XGBoost ten-fold cross-validation (Cutoff: 0.5), which achieved a discrimination ability of 0.959 for Btb on the training set and a discrimination ability of 0.847 on the validation set.
[0032] The experimental materials used in the examples of the present invention but not specifically described are all conventional experimental materials in the art and can be purchased through commercial channels.
[0033] Example 1: Bioinformatics-based core gene screening (1) Chip data processing The limma package in R software was used to normalize the GSE255724 data from the public database (NCBI-GEO) to remove batch effects. Figure 1 As shown in A, the processed data is as follows Figure 1 As shown in B.
[0034] (2) Differential gene heat map and volcano map The normalized microarray data were used to construct a dataset, which was randomly divided into a training set and a validation set. The training set included 49 TB-positive cattle (Treat group) and 49 TB-negative cattle (Control group). The validation set included 14 TB-positive cattle (Treat group) and 11 TB-negative cattle (Control group).
[0035] Analyze the differentially expressed genes between the training set experimental group and the training set control group in the training set. The heat map of the top 20 differentially expressed genes is as follows Figure 2 The volcano plot of genes with significant differential expression between the experimental group and the control group is shown in Figure 3 A total of 22 genes with significantly differential expression were screened. Figure 3 Red dots indicate 20 significantly upregulated genes, green dots indicate 2 significantly upregulated genes, and black dots indicate genes without changes.
[0036] (3) KEGG enrichment function analysis KEGG enrichment function analysis was performed on 22 genes with significant differential expression. Figure 4 The results of gene enrichment analysis showed that the 22 differentially expressed genes were mainly involved in functions such as oxygen transport, hemoglobin complex, and extracellular matrix composition.
[0037] (4) Machine learning to screen potential core genes Through machine learning, 22 genes with significantly different expression were screened to obtain potential core genes. Figure 5 、 Figure 6 、 Figure 7 18, 14 and 20 potential core genes were screened using support vector machine recursive elimination (SVM-REF), lasso regression (LASSO) and random forest (RF) respectively. The intersection of the potential core genes screened by the three machine learning methods was taken to further screen out 14 potential core genes. The results are shown in Figure 8 .
[0038] (5) Expression of potential core genes Among the 14 potential core genes, 4 genes were differentially expressed in the training set and the validation set, and ADGRA3 (NCBI Gene ID: 166647), PRELP (NCBI Gene ID: 5549), HBB (NCBI Gene ID: 3043), and LOC107131333 (NCBI Gene ID: 107131333) were selected as core genes. The expression of core genes in the training set and the validation set is shown in Figure 9-10 .
[0039] (6) Single gene diagnostic capability The diagnostic capabilities of the core genes in the training set and validation set are shown in Figure 11-12 It can be seen that the AUC range of core genes (ADGRA3, PRELP, HBB, LOC107131333) in the training set is 0.694-0.784, and the AUC range in the validation set is 0.738-0.846.
[0040] In order to evaluate the actual effects of the above core genes, clinical samples were used for verification.
[0041] Example 2: Clinical sample verification Samples were collected from a farm in Henan Province. 5 mL of whole blood was collected with sodium heparin anticoagulation, stored at low temperatures, and delivered to the laboratory within 24 hours. Tuberculosis infection status was determined using a bTB interferon-γ ELISA kit (Wuhan Keqian Biological Co., Ltd.). Finally, three test-positive cattle (test group) and three test-negative cattle (control group) were selected for mRNA expression analysis of core genes (ADGRA3, PRELP, HBB, and LOC107131333). The main reagents and primer sequences used in the validation process are listed in Tables 1 and 2.
[0042] Table 1 Main reagents Table 2 Primer sequence list Depend on Figure 13 It can be seen that HBB, ADGRA3, and PRELP genes were significantly differentially expressed between the control group and the experimental group ( p <0.05). Therefore, HBB, ADGRA3, and PRELP were selected as marker genes, i.e., diagnostic markers for identifying bTB.
[0043] Example 3: Construction of a multi-gene joint diagnosis model based on machine learning To verify the robustness of the screened marker genes, the present invention uses a variety of models and their parameter combinations to construct a multi-gene diagnostic model for verification and comparison. The main models are as follows: logistic regression, linear discriminant analysis, quadratic discriminant analysis, k-nearest neighbor, XGBoost classification, ridge regression, lasso regression, support vector machine, stepwise logistic regression model, and naive Bayesian algorithm model. Figure 14 The models that rank in the top 50 in terms of AUC on both the training and validation sets are shown. The marker genes perform well in the above models, further supporting the generalization ability of these genes.
[0044] according to Figure 15 The XGBoost ten-fold cross validation (Cutoff: 0.5) algorithm performed best, and the AUC (Area Under Curve) of the multi-gene diagnosis model in the training set and test set were 0.959 and 0.847, respectively.
[0045] The XGBoost 10-fold cross validation method works best because it excels at processing data with high feature dimensions and nonlinear relationships between variables. It has the ability to handle complex interactions between variables and is particularly suitable for the multi-gene expression data involved in the present invention. Therefore, its excellent performance reflects more the fit between the model and the data characteristics rather than a high degree of dependence on the model. Therefore, the gene combination of the present invention exhibits strong stability across different modeling methods and has a certain degree of versatility, rather than relying solely on a specific model.
[0046] This study, using machine learning to systematically screen and validate with clinical samples, successfully identified three marker genes with specific diagnostic value for bTB infection as bovine tuberculosis diagnostic markers, namely ADGRA3, PRELP, and HBB, from among numerous previously unreported genes. The resulting marker genes were combined to construct a multi-parameter combined diagnostic model (training set AUC = 0.959, test set AUC = 0.847). This model's comprehensive discriminatory power improved compared to traditional single-gene detection methods and was also at a higher level than existing multi-gene detection methods, providing a new solution for the accurate diagnosis of bTB.
[0047] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A diagnostic marker set for bovine tuberculosis, characterized in that: The diagnostic marker set includes at least two of ADGRA3, PRELP or HBB.
2. The diagnostic marker set for bovine tuberculosis according to claim 1, characterized in that: The Gene ID of ADGRA3 in NCBI is 166647, the Gene ID of PRELP in NCBI is 5549, and the Gene ID of HBB in NCBI is 3043.
3. Use of a primer pair for detecting the diagnostic marker set according to claim 1 or 2 in a sample in the preparation of a diagnostic product for bovine tuberculosis.
4. The use according to claim 3, characterized in that The primer pair for detecting ADGRA3 is SeqID No: 1-2, the primer pair for detecting PRELP is SeqID No: 3-4, and the primer pair for detecting HBB is SeqID No: 5-6.
5. A product for diagnosing bovine tuberculosis, characterized in that: The product comprises a reagent for detecting the expression level of the diagnostic marker set according to claim 1 or 2 in a sample.
6. The method for screening a diagnostic marker set for bovine tuberculosis according to claim 1 or 2, characterized in that: The steps include: The genetic data of healthy cows and cows with bovine tuberculosis were compared and screened to obtain differentially expressed genes; the differentially expressed genes were verified by single gene clinical samples; and single genes verified by clinical samples were selected to form a diagnostic marker set for bovine tuberculosis.
7. A pharmaceutical composition for treating bovine tuberculosis, characterized in that: The pharmaceutical composition comprises: an ADGRA3, PRELP or HBB functional expression inhibitor.
8. The pharmaceutical composition for treating bovine tuberculosis according to claim 7, characterized in that It also includes pharmaceutically acceptable carriers or excipients.