Application of microbial marker in preparation of product for diagnosing hyperuricemia and device for diagnosing hyperuricemia
By detecting specific microbial biomarkers in feces and constructing a random forest model, the problem of the inability to identify the risk of hyperuricemia in the early stages of existing technologies has been solved, achieving highly accurate early screening and risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AIR FORCE MEDICAL CENT PLA
- Filing Date
- 2026-03-02
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies are insufficient to provide effective early warning tools to identify individuals at subclinical risk of hyperuricemia, and current research has not yet been translated into precise predictive tools that can be used in clinical practice or health management.
By detecting the abundance of Firmicutes bacterium CAG 124, Parapravotella clara, Clostridia bacterium, and Bacteroides intestinalis in feces, combined with gene sequencing and bioinformatics analysis, a diagnostic device based on a random forest model was constructed to assess the risk of hyperuricemia.
It provides a non-invasive, novel, and highly accurate tool for the early screening and risk stratification of hyperuricemia, and can accurately predict the tendency to develop hyperuricemia.
Smart Images

Figure CN121992124A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, and in particular to the application of a microbial biomarker in the preparation of products for diagnosing hyperuricemia, and an apparatus for diagnosing hyperuricemia. Background Technology
[0002] Hyperuricemia (HUA), as an independent risk factor for gout, chronic kidney disease, and cardiovascular and cerebrovascular diseases, poses a significant public health challenge. The core bottleneck in current prevention and control strategies lies in the lack of effective early warning tools. Clinical criteria relying on a single blood uric acid test can only provide "post-hoc diagnosis" after metabolic imbalance has occurred, failing to identify "subclinical" high-risk individuals with significantly elevated or drastically fluctuating uric acid levels who already exhibit a clear tendency towards hyperuricemia.
[0003] Recent studies have revealed that the gut microbiome plays a crucial role in human purine and uric acid metabolism. Gut microbiota not only directly participate in dietary purine uptake and intestinal breakdown of uric acid (accounting for approximately 30% of total uric acid excretion), but also indirectly regulate uric acid synthesis and renal excretion by influencing systemic inflammatory states, insulin sensitivity, and intestinal barrier function. Therefore, gut microbiome dysbiosis is considered a key environmental and physiological mediator in the occurrence and development of hyperuricemia. This provides a novel and highly promising source of biomarkers for early risk assessment of hyperuricemia.
[0004] However, most existing studies focus on describing the differences in overall gut microbiota structure between patients with hyperuricemia and healthy individuals, or on finding cross-sectional correlations between a few bacterial genera and uric acid levels. These findings have not yet been translated into precise predictive tools that can be used in clinical practice or health management.
[0005] In view of this, the present invention is hereby proposed. Summary of the Invention
[0006] The first objective of this invention is to provide the application of a reagent for detecting the abundance of microorganisms in feces in the preparation of products for diagnosing hyperuricemia, in order to solve the above-mentioned technical problems.
[0007] A second objective of the present invention is to provide a device for diagnosing hyperuricemia.
[0008] To achieve the above objectives, the following technical solution is adopted: In a first aspect, the present invention provides the application of a reagent for detecting the abundance of microorganisms in feces in the preparation of products for diagnosing hyperuricemia, wherein the microorganisms include Firmicutes bacterium CAG.124, Pararaprevotella clara, Clostridia bacterium, and Bacteroides intestinalis. It should be noted that Firmicutes bacterium CAG.124, Pararaprevotella clara, Clostridia bacterium, and Bacteroides intestinalis all refer to species of bacteria.
[0009] As a further technical solution, the abundance of Firmicutes.bacterium.CAG.124 was downregulated in the feces of patients with hyperuricemia compared with healthy subjects.
[0010] As a further technical solution, the abundance of Pararaprevotella clara was downregulated in the feces of patients with hyperuricemia compared with healthy subjects.
[0011] As a further technical solution, the abundance of Clostridia bacterium was downregulated in the feces of patients with hyperuricemia compared with healthy subjects.
[0012] As a further technical solution, the abundance of Bacteroides intestinalis was downregulated in the feces of patients with hyperuricemia compared with healthy subjects.
[0013] As a further technical solution, the abundance of microorganisms in feces is detected through gene sequencing and bioinformatics analysis.
[0014] Secondly, the present invention provides a device for diagnosing hyperuricemia, including a data acquisition module and a classification module; The data acquisition module is used to acquire microbial abundance data of the feces of the subjects to be tested; The classification module is used to input the microbial abundance data of the test subject's feces into a pre-trained classification model. The classification model classifies the microbial abundance data of the test subject's feces to predict whether the test subject has hyperuricemia or to assess whether the test subject is a susceptible population for hyperuricemia. The classification model was trained using the following method: a. Obtain microbial abundance data in the feces of healthy subjects and patients with hyperuricemia; b. Use the microbial abundance data to train a classification model and obtain a pre-trained classification model; The microorganisms include Firmicutes.bacterium.CAG.124, Parapravotella.clara, Clostridia.bacterium, and Bacteroides.intestinalis.
[0015] As a further technical solution, the abundance of microorganisms in feces is detected through gene sequencing and bioinformatics analysis.
[0016] As a further technical solution, the classification model includes a random forest model.
[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention utilizes high-throughput microbial sequencing technology to analyze the gut microbiota composition of a population cohort. Employing bioinformatics algorithms such as machine learning, it identifies and validates a set of core microbial biomarkers capable of reliably predicting the predisposition to hyperuricemia. Ultimately, based on these microbial biomarkers, a device for diagnosing hyperuricemia is developed. This device demonstrates high predictive accuracy and provides a novel, non-invasive tool with high clinical translational potential for early screening and risk stratification of hyperuricemia. Attached Figure Description
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 ROC curves for the test set. Detailed Implementation
[0020] The embodiments and examples of the present invention will be described in detail below. However, those skilled in the art will understand that the following embodiments and examples are for illustrative purposes only and should not be considered as limiting the scope of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Unless otherwise specified, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0021] In a first aspect, the present invention provides the application of reagents for detecting the abundance of microorganisms in feces in the preparation of products for diagnosing hyperuricemia, wherein the microorganisms include Firmicutes.bacterium.CAG.124, Parapravotella.clara, Clostridia.bacterium, and Bacteroides.intestinalis.
[0022] The inventors' research revealed that analysis of training set samples (including fecal samples from healthy subjects and patients with hyperuricemia) showed significant differences in the levels of Firmicutes bacterium CAG 124, Parapravotella clara, Clostridia bacterium, and Bacteroides intestinalis in the feces of healthy subjects and patients with hyperuricemia. By detecting the abundance characteristics of this biomarker combination, it is possible to effectively assess whether a population or individual has hyperuricemia or is likely to develop hyperuricemia. When an individual's gut microbiome exhibits a specific abundance pattern of this biomarker combination, it indicates that the individual has a higher risk of developing hyperuricemia in the future.
[0023] In the risk assessment and early prevention of hyperuricemia, gut microbiota testing can be performed on the individuals being consulted to analyze whether they carry the aforementioned high-risk microbial marker patterns. For individuals identified as having high-risk patterns, targeted and scientific early interventions can be implemented, such as suggesting adjustments to dietary structure and supplementation with specific prebiotics / probiotics. This is expected to intervene in the uric acid metabolism pathway at the microecological level and reduce the risk of developing hyperuricemia.
[0024] In some alternative implementations, the abundance of microorganisms in feces is detected by gene sequencing and bioinformatics analysis.
[0025] This invention does not impose specific limitations on the methods used for gene sequencing and bioinformatics analysis; any methods well-known to those skilled in the art can be used.
[0026] Secondly, the present invention provides a device for diagnosing hyperuricemia, including a data acquisition module and a classification module; The data acquisition module is used to acquire microbial abundance data of the feces of the subjects to be tested; The classification module is used to input the microbial abundance data of the test subject's feces into a pre-trained classification model. The classification model classifies the microbial abundance data of the test subject's feces to predict whether the test subject has hyperuricemia or to assess whether the test subject is a susceptible population for hyperuricemia. The classification model was trained using the following method: a. Obtain microbial abundance data in the feces of healthy subjects and patients with hyperuricemia; b. Use the microbial abundance data to train a classification model and obtain a pre-trained classification model; The microorganisms include Firmicutes.bacterium.CAG.124, Parapravotella.clara, Clostridia.bacterium, and Bacteroides.intestinalis.
[0027] The device for diagnosing hyperuricemia provided by this invention analyzes the abundance of Firmicutes.bacterium.CAG.124, Parapravotella.clara, Clostridia.bacterium, and Bacteroides.intestinalis in test subjects, which can accurately predict whether the subject has hyperuricemia or assess whether the subject is a susceptible population for hyperuricemia.
[0028] In some alternative implementations, the abundance of microorganisms in feces is detected by gene sequencing and bioinformatics analysis.
[0029] In some alternative implementations, the classification model includes, but is not limited to, the random forest model, with the random forest model being preferred.
[0030] The present invention will be further illustrated below with specific embodiments. However, it should be understood that these embodiments are merely for the purpose of more detailed illustration and should not be construed as limiting the present invention in any way.
[0031] Example 1: Microbial Identification and Model Construction for the Diagnosis of Hyperuricemia 1. Materials and Methods 1.1 Research Subjects To identify microbial biomarkers associated with hyperuricemia (HUA), the inventors recruited two cohorts at different times. All participants maintained stable lifestyles and metabolic states before arriving at the study site. All participants provided written informed consent after receiving a comprehensive briefing on the study details. Participants diagnosed with hyperuricemia according to internationally recognized diagnostic criteria (e.g., serum uric acid levels >420 μmol / L) were assigned to the case group, while participants with consistently normal serum uric acid levels and no related metabolic diseases were assigned to the control group. Thus, cohort 1 consisted of 26 cases and 33 controls, and cohort 2 consisted of 11 cases and 11 controls. Fecal samples were collected from all participants for subsequent testing.
[0032] 1.2 DNA Shearing DNA was extracted from feces, and the genomic DNA was randomly fragmented into segments of approximately 350 bp using a Covaris fragmenter.
[0033] 1.3 End Repair Response Fragmented DNA has 5' or 3' protrusions. An end-completion system is added to the purified DNA fragments. The exonuclease activity of T4 DNA polymerase digests the 3' single-strand protrusion, while the polymerase activity completes the 5' protrusion. At the same time, phosphokinase (PNK) adds a phosphate group necessary for subsequent ligation reactions to the 5' end. After purification with Agencourt AMPure XP magnetic beads, a library of blunt-ended DNA fragments with phosphate groups at the 5' end is finally obtained.
[0034] 1.4 Add an "A" suffix to the 3' end. Add a 3' end-added "A" buffer to the above system. Adding a single adenosine "A" to the 3' end of the modified double-stranded DNA prevents self-ligation between blunt ends of DNA fragments. It also complements the single "T" protrusion at the 5' end of the sequencing adapter in the next step, ensuring accurate ligation and effectively reducing tandem ligation between library fragments.
[0035] 1.5 Connecting Sequencing Adapters Add ligation buffer and double-stranded sequencing adapter to the above reaction system, and use T4 DNA ligase to ligate the Illumina sequencing adapter to both ends of the library DNA.
[0036] 1.6 Library Fragment Selection (Size Selection) For libraries with adapters, the Agencourt SPRIselect (Beckman Coulter, USA, Catalog # 2358413) nucleic acid fragment selection kit was used to perform fragment size selection simultaneously with library purification. A two-step selection method (Double Size Selection) was employed: first, small fragments to the left of the target domain were removed using SPRI magnetic beads (Left-side Size Selection); then, large fragments to the right of the target region were removed (Right-side Size Selection). This resulted in the selection of original libraries with appropriate fragment lengths for subsequent PCR amplification. The purified libraries removed excess sequencing adapters and adapter self-ligation products, avoiding invalid amplification during the PCR process and eliminating any impact on sequencing.
[0037] 1.7 PCR Amplification of DNA Libraries High-fidelity polymerase was used to amplify the original library to ensure a sufficient total library volume. Furthermore, this step effectively enriches only DNA fragments with adapters at both ends, as only such fragments can be amplified. While ensuring sufficient product, bias introduced by excessive amplification cycles was minimized; finally, Qubit 3.0 was used to accurately determine the concentration of each library.
[0038] 1.8 Library Quality Assessment After the library was constructed, the Agilent 5400 system (AATI) was used to detect the insert size of the library. If the insert size met the expectations, the effective concentration (1.5 nM) of the library was accurately quantified using qPCR to ensure the quality of the library.
[0039] 1.9 Bridge PCR After the library passes inspection, sequencing is performed on the Illumina Novaseq platform based on the effective concentration of the library and the data output. The next step is to seed the captured library onto a FlowCell chip for amplification. Two different DNA primers are seeded on the inner surface of the FlowCell channel. These primer sequences complement the adapter sequences at both ends of the DNA library and are covalently linked to the FlowCell. The specific process is as follows: a. Add the DNA library to the chip. Since the DNA sequences at both ends of the library are complementary to the primer sequences on the chip, complementary hybridization occurs. After hybridization, dNTPs and polymerase are added. The polymerase starts from the primer and synthesizes a DNA strand complementary to the original DNA sequence along the template; b. Add NaOH alkaline solution to dissociate the DNA double strand, washing away the original DNA strand that was not covalently linked to the chip, and retaining the newly synthesized DNA strand covalently linked to the chip; c. Add neutralizing solution to the fluidized bed to neutralize the alkaline solution. At this time, the other end of the DNA hybridizes with another primer on the chip. Add enzyme and dNTPs to synthesize a new DNA strand; Add alkaline solution again to separate the two DNA strands, add neutralizing solution again, and the DNA hybridizes with the new primer on the chip. Add enzyme and dNTPs to synthesize a new strand from the new primer. Repeat this process continuously, and the DNA strand grows exponentially.
[0040] 1.10 Sequencing on the Illumina PE150 platform PE150 stands for Pair-end 150bp, a high-throughput sequencing technology. In the constructed DNA fragment library, each insert fragment is sequenced from both ends, 150bp from each end. The specific process is as follows: After bridge PCR, the synthesized double strands are converted into sequenceable single strands; a specific group of one of the primers on the chip is cleaved, and the chip is washed with alkaline solution to dissociate the DNA double strands, washing away the cleaved DNA strands at the root, leaving the covalently linked strand; a neutral solution, sequencing primers, and fluorescently labeled dNTPs are added. Four different fluorescently labeled dNTPs have their 3' ends blocked by azido groups. Polymerase is then added to synthesize the dNTPs onto the new DNA strand. The ends are blocked by sodium azide groups, so each cycle can only extend one base. After completing one cycle, excess dNTPs and enzymes are washed away, and the sample is placed under a microscope for laser scanning. The fluorescence emitted indicates which base has been newly synthesized, and the template base can be deduced through the complementarity principle. After completing one cycle, chemical reagents are added to remove the sodium azide groups and fluorescent groups, exposing the 3' hydroxyl group. New dNTPs and new enzymes are added to extend another base. After the new base extension is complete, excess dNTPs and enzymes are washed away, and another round of microscopic laser scanning is performed to read the base. By continuously repeating this cycle, hundreds of bases can be read.
[0041] 1.11 Sequencing Result Preprocessing The raw data was preprocessed using FastP (https: / / github.com / OpenGene / fastp) to obtain clean data for subsequent analysis. The specific processing steps are as follows: a) If any sequencing read contains an adapter sequence, remove the paired read; b) If the number of low-quality (Q<=5) bases in any sequencing read exceeds 50% of the total bases in that read, remove the paired read; c) If the N content in any sequencing read exceeds 10% of the total bases in that read, remove the paired read.
[0042] If the sample contains host contamination, it needs to be compared with the host sequence to filter out reads that may originate from the host, using Bowtie2 software.
[0043] 1.12 Metagenome Assembly The clean data was assembled and analyzed using MEGAHIT software with the following assembly parameters: -- presets meta-large (--end to- end, --sensitive, -I 200, -X 400 (Karlsson FH et al., 2013; Nielsen HB et al., 2014). The assembled scaffolds were then broken at the N-connections to obtain scaffolds without N (Qin J et al., 2010; Li D et al., 2015).
[0044] 1.13 Gene prediction and abundance analysis MetaGeneMark (http: / / topaz.gatech.edu / GeneMark / ) was used to predict the ORF of scaftigs (>=500bp) for each sample, and information shorter than 100 nt in the prediction results was filtered out, using default parameters. The ORF prediction results were then deredundantd using CD-HIT software (http: / / www.bioinformatics.org / cd-hit / ) to obtain a non-redundant initial gene catalogue. Here, the non-redundant continuous gene-encoding nucleic acid sequences are referred to as genes, with the following parameter settings: -c 0.95, -G 0, -aS 0.9, -g 1, -d 0. Bowtie2 was used to align the clean data of each sample to the initial gene catalogue, calculating the number of aligned reads for each gene in each sample, with the following alignment parameters: --end-to-end, --sensitive, -I 200, -X 400. Genes with fewer than or equal to 2 reads in each sample were filtered out to obtain the final gene catalogue (unigenes) for subsequent analysis. Based on the number of aligned reads and gene length, the abundance information of each gene in each sample was calculated, as shown in the formula, where r is the number of aligned gene reads and L is the gene length. Based on the abundance information of each gene in each sample from the gene catalogue, basic statistical analysis, Core-Pan gene analysis, inter-sample correlation analysis, and Venn diagram analysis of gene number were performed.
[0045] 1.14 Species Annotation Using the DIAMOND software (https: / / github.com / bbuchfink / diamond / ), unigenes were compared with Micro_NR with the following parameters: blastp, -e 1e-5 (Karlsson FH et al., 2013). Micro_NR consists of bacterial, fungal, archaea, and viral sequences extracted from the NCBI NR database (https: / / www.ncbi.nlm.nih.gov / ).
[0046] For each sequence alignment result, select the one with evalue <= the smallest evalue. For the results in 10, since each sequence may have multiple alignment results, the LCA algorithm (used in the systematic classification of MEGAN software (https: / / en.wikipedia.org / wiki / Lowest_common_ancestor) was used to determine the species annotation information for that sequence. Starting from the LCA annotation results and gene abundance table, the abundance information and gene number table for each sample at each taxonomic level (kingdom, phylum, class, order, family, genus, species) were obtained. The abundance of a species in a sample is equal to the sum of the abundances of genes annotated as belonging to that species; the number of genes of a species in a sample is equal to the number of genes with a non-zero abundance among those annotated as belonging to that species.
[0047] Anosim analysis (R vegan package) was used to examine differences between groups; then LEfSe analysis was used to identify species with differences between groups; finally, cohort 1 was used to screen model biomarkers and a random forest model was constructed. The model was then validated using data from cohort 2.
[0048] 2. Results 2.1 Sample and Microbial Quantity To identify microbial biomarkers associated with hyperuricemia (HUA), the inventors performed metagenomic sequencing on fecal samples from two cohorts. Since both cohorts used the same sequencing platform and experimental procedures, the inventors performed standardized bioinformatics processing and quality control on the resulting microbiome data. Ultimately, cohort 1 contained 26 cases and 33 controls, with 7956 bacterial species. Cohort 2 contained 11 cases and 11 controls, with 6291 bacterial species.
[0049] 2.2 Results of the difference analysis The inventors performed a differential analysis on the bacterial species in the control and high uric acid groups in cohort 1, and ultimately found 12 bacterial species that showed significant differences between the two groups (P<0.05, |LDA|>2). These included Vescimonas sp., Firmicutes bacterium CAG:124, Clostridia bacterium, Parapravotella clara, Bacteroides intestinalis, Candidatus Cryptobacteroides sp., Jilunia laotingensis, Oscillibacter sp. MSJ-31, Bacteroides xylanisolvens, Odoribacter splanchnicus, uncultured Victivallis sp., and Oscillibacter sp. ER4, as shown in Table 1.
[0050] Table 1
[0051] Note: Control refers to normal (control) samples; LDA is the result of linear discriminant analysis.
[0052] 2.3 Variable Selection and Model Building Based on a differentially abundant microbial species table (behavioral samples, listed as microbial species), this study used the presence or absence of hyperuricemia as the response variable and the relative abundance of all microbial species as candidate predictive variables, employing a stepwise regression method based on the AIC criterion for feature selection. This process, through iterative forward selection and backward elimination, aimed to find a concise combination of microbial features with optimal predictive efficacy. Ultimately, a feature set consisting of four key microbial species (Firmicutes.bacterium.CAG.124, Parapravotella.clara, Clostridia.bacterium, and Bacteroides.intestinalis) was obtained, which will be used in subsequent predictive model construction.
[0053] Using the four key microbial species selected through stepwise regression as input features, this study constructs a random forest classifier to predict the risk of hyperuricemia. Random forest is an ensemble learning algorithm that improves the accuracy and robustness of the model by constructing multiple decision trees and aggregating their predictions (majority voting). During the model training phase, queue 1 is divided into a training set for model construction. The final model is trained on the training set.
[0054] Training set (crowd queue 1) performance: The random forest model showed good learning ability on the training set, with an AUC value of 71.68%, indicating that the model can effectively capture patterns in the data.
[0055] Test set (crowd cohort 2) performance: In the test set, the final model's prediction accuracy was 66.12% ( Figure 1 The performance metrics of the test set and the training set are close, indicating that the model does not show significant overfitting and has good generalization performance and potential clinical application value.
[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The application of a reagent for detecting the abundance of microorganisms in feces in the preparation of products for diagnosing hyperuricemia, characterized in that, The microorganisms include Firmicutes.bacterium.CAG.124, Parapravotella.clara, Clostridia.bacterium, and Bacteroides.intestinalis.
2. The application according to claim 1, characterized in that, Compared with healthy subjects, the abundance of Firmicutes bacterium CAG.124 was downregulated in the feces of patients with hyperuricemia.
3. The application according to claim 1, characterized in that, Compared with healthy subjects, the abundance of Pararaprevotella clara was downregulated in the feces of patients with hyperuricemia.
4. The application according to claim 1, characterized in that, Compared with healthy subjects, the abundance of Clostridia bacterium was downregulated in the feces of patients with hyperuricemia.
5. The application according to claim 1, characterized in that, Compared with healthy subjects, the abundance of Bacteroides intestinalis was downregulated in the feces of patients with hyperuricemia.
6. The application according to claim 1, characterized in that, The abundance of microorganisms in feces was detected by microbial sequencing and bioinformatics analysis.
7. A device for diagnosing hyperuricemia, characterized in that, Includes a data acquisition module and a classification module; The data acquisition module is used to acquire microbial abundance data of the feces of the subjects to be tested; The classification module is used to input the microbial abundance data of the test subject's feces into a pre-trained classification model. The classification model classifies the microbial abundance data of the test subject's feces to predict whether the test subject has hyperuricemia or to assess whether the test subject is a susceptible population for hyperuricemia. The classification model was trained using the following method: a. Obtain microbial abundance data in the feces of healthy subjects and patients with hyperuricemia; b. Use the microbial abundance data to train a classification model and obtain a pre-trained classification model; The microorganisms include Firmicutes.bacterium.CAG.124, Parapravotella.clara, Clostridia.bacterium, and Bacteroides.intestinalis.
8. The apparatus according to claim 7, characterized in that, The abundance of microorganisms in feces was detected by microbial sequencing and bioinformatics analysis.
9. The apparatus according to claim 7, characterized in that, The classification model includes the random forest model.