A method and apparatus for determining absolute genome abundance
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-27
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]但是,不同简并引物对不同微生物的扩增效率也存在不一致性,这种扩增效率的差异会导致最终的定量结果出现偏差,无法准确反映样本中各微生物的实际丰度
[0015]本发明提供的一种基因组绝对丰度的确定方法和装置,利用预训练的扩增效率预测模型,能够基于第一基因片段中的GC含量对扩增效率进行预测,以进一步基于不同基因组片段分别对应的扩增效率对各个基因组片段的第一16S拷贝数进行了校正,能够去除各个微生物扩增效率差异在定量结果中所带来的偏差,基于校正后的第二16S拷贝数实现了对基因测序结果中各个微生物菌群绝对丰度的确定。
Smart Images

Figure CN122575485A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, and in particular to a method and apparatus for determining the absolute abundance of a genome. Background Technology
[0002] In the field of microbial research, 16S rRNA gene sequencing is a commonly used analytical method, widely applied to the study of microbial community structure and diversity. In the study of complex microbial community samples, since the microbial community includes a variety of different microorganisms with varying 16S rRNA gene sequences, current techniques utilize degenerate primers to increase the binding chances of these primers to the 16S sequences of different microorganisms, thereby improving the coverage of PCR amplification.
[0003] However, different degenerate primers exhibit inconsistent amplification efficiencies for different microorganisms. This difference in amplification efficiency can lead to deviations in the final quantitative results, failing to accurately reflect the actual abundance of each microorganism in the sample. Therefore, current gene sequencing results based on degenerate primer amplification are inaccurate, and a method to determine the absolute abundance of each microbial community in gene sequencing results is urgently needed. Summary of the Invention
[0004] This invention provides a method and apparatus for determining the absolute abundance of a genome. Utilizing a pre-trained amplification efficiency prediction model, the amplification efficiency can be predicted based on the GC content in a first gene fragment. Furthermore, the first 16S copy number of each genome fragment is corrected based on the amplification efficiency corresponding to different genome fragments. This removes the bias caused by differences in the amplification efficiency of various microorganisms in the quantitative results. Based on the corrected second 16S copy number, the absolute abundance of each microbial community in the gene sequencing results is determined.
[0005] This invention provides a method for determining the absolute abundance of a genome, comprising the following steps: obtaining multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment from the gene sequencing results; inputting the GC content in each of the first genome fragments into a pre-trained amplification efficiency prediction model, and outputting the amplification efficiency corresponding to each first genome fragment; wherein, the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of second genome fragments with known GC content under standard PCR reaction conditions; for each first genome fragment: correcting the first 16S copy number according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genome fragment; determining the absolute abundance of each first genome fragment based on the second 16S copy number and the total genome copy number.
[0006] Optionally, determining the absolute abundance of each first genome fragment based on the second 16S copy number and the total genome copy number includes: obtaining 16S copy number information corresponding to each first genome fragment; and for each first genome fragment: obtaining the genome copy number corresponding to the first genome fragment based on the corresponding 16S copy number information and the second 16S copy number.
[0007] Optionally, obtaining the 16S copy number information corresponding to each of the first genome fragments includes: in response to the existence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, obtaining the 16S copy number information corresponding to the operational taxonomic unit sequence from the ribosomal RNA operon copy number database; and in response to the absence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, determining the 16S copy number information corresponding to the first genome fragment using a genetic distance algorithm.
[0008] Optionally, determining the 16S copy number information corresponding to the first genome fragment using the genetic distance algorithm includes: determining the target genome fragment that is most closely related to the first genome fragment using the genetic distance algorithm; wherein the target genome fragment is a genome fragment with known 16S copy number information; and using the 16S copy number information corresponding to the target genome fragment as the 16S copy number information corresponding to the first genome fragment.
[0009] Optionally, the step of obtaining the genome copy number corresponding to the first genome fragment based on the corresponding 16S copy number information and the second 16S copy number for each first genome fragment includes: for each first genome fragment, using the ratio between the second 16S copy number and the 16S copy number information corresponding to the first genome fragment as the genome copy number corresponding to the first genome fragment.
[0010] Optionally, the step of correcting the first 16S copy number according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genome fragment for each first genome fragment includes: for each first genome fragment, using the ratio between the first 16S copy number and the amplification efficiency as the second 16S copy number corresponding to the first genome fragment.
[0011] The present invention also provides an apparatus for determining the absolute abundance of a genome, comprising the following modules: The acquisition module is used to acquire multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment included in the gene sequencing results; The prediction module is used to input the GC content of each of the first genomic fragments into a pre-trained amplification efficiency prediction model and output the amplification efficiency corresponding to each of the first genomic fragments; wherein, the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of a second genomic fragment with known GC content under standard PCR reaction conditions; A 16S copy number correction module is used to correct the first 16S copy number for each of the first genome fragments according to the amplification efficiency, so as to obtain a second 16S copy number corresponding to the first genome fragment. The determination module is used to determine the absolute abundance of each of the first genome fragments based on the second 16S copy number and the total genome copy number.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for determining absolute abundance of the genome as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for determining absolute abundance of the genome as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method for determining absolute abundance of genomes as described above.
[0015] The present invention provides a method and apparatus for determining the absolute abundance of a genome. It utilizes a pre-trained amplification efficiency prediction model to predict amplification efficiency based on the GC content in a first gene fragment. Furthermore, it corrects the first 16S copy number of each genome fragment based on the amplification efficiency corresponding to different genome fragments. This can remove the bias caused by the difference in amplification efficiency of each microorganism in the quantitative results. Based on the corrected second 16S copy number, the absolute abundance of each microbial community in the gene sequencing results is determined. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the method for determining absolute genome abundance provided by the present invention.
[0018] Figure 2 This is a schematic diagram of the specific process of step 104 provided by the present invention.
[0019] Figure 3 This is a schematic diagram of a specific process for determining the absolute abundance of the genome provided by the present invention.
[0020] Figure 4 This is a schematic diagram of the structure of the device for determining the absolute abundance of the genome provided by the present invention.
[0021] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] Figure 1 This is one of the flowcharts illustrating the method for determining absolute genome abundance provided by this invention, such as... Figure 1 As shown, the method includes the following: Step 101: Obtain the gene sequencing results of the sample to be tested; the gene sequencing results include multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment; Among them, the sample to be tested is a microbial community sample. A microbial community often contains multiple bacterial groups, and different bacterial groups correspond to different gene fragments. Therefore, the gene sequencing results of the sample to be tested will contain multiple first gene fragments.
[0024] For microbial communities, degenerate primers are typically used for PCR amplification to specifically amplify 16S rRNA. Degenerate primers are primer sequences composed of a mixture of different sequences representing all possible bases encoding a single amino acid. In studies of complex microbial samples, due to the differences in 16S rRNA gene sequences among different microorganisms, degenerate primers can increase the binding chances with different microbial 16S sequences, improving amplification coverage and enabling the detection and analysis of a wider variety of microorganisms.
[0025] However, for primers, whether degenerate or single, as a key element of PCR amplification, their design must follow a series of strict principles: (1) Primer length is generally controlled between 15 and 30 bases, with 18 to 24 bases being commonly used. This is because primers that are too short are not specific enough and are easy to bind to non-target sequences, leading to non-specific amplification. Primers that are too long not only increase the cost of synthesis, but may also form complex secondary structures, which also affect the specificity of amplification.
[0026] (2) GC content is also an important factor in primer design. Ideally, the GC content of primers should be between 40% and 60%. If the GC content is too high, the annealing temperature of the primers will increase, which may lead to difficulty in binding the primers to the template; while if the GC content is too low, the stability of the primers will decrease, which is also not conducive to PCR amplification. It is understandable that when the GC content deviates from the appropriate range, it is necessary to adjust the reaction conditions such as the annealing temperature to ensure the amplification effect, but this also increases the complexity and uncertainty of the experiment.
[0027] (3) The Tm value of the primers should generally be controlled between 55-65℃, and the difference between the Tm values of the upstream and downstream primers should not exceed 4℃, so as to ensure that the primers and template DNA can bind stably and synchronously. If the Tm values of the two primers are too different, during the PCR reaction, one primer may bind to the template too early or too late, resulting in reduced amplification efficiency or even failure to obtain the expected amplification product.
[0028] (4) Primer design should also avoid forming secondary structures, such as hairpin structures or dimers. Hairpin structures cause the primer to fold itself, making it unable to bind effectively to the template; primer dimers will consume the primer and reaction substrates such as dNTPs, reducing the amplification efficiency of the target product. The base distribution should also be as random as possible to avoid the appearance of consecutive identical base sequences, especially more than four consecutive purines or pyrimidines, to prevent non-specific binding of the primer to the template.
[0029] GC content has a significant impact on amplification efficiency during PCR, and the inventors have found that there is usually a linear or non-linear correlation between GC content and amplification efficiency. Therefore, in this embodiment of the invention, amplification efficiency was predicted based on the GC content in the first gene fragment, so as to further correct the 16S copy number based on the amplification efficiency.
[0030] Specifically, the double helix structure of DNA consists of two antiparallel polynucleotide chains linked together by hydrogen bonds between bases. Three hydrogen bonds form between G (guanine) and C (cytosine), while two hydrogen bonds form between A (adenine) and T (thymine). This difference in the number of hydrogen bonds makes the DNA double-stranded structure in GC-rich regions more stable. During PCR amplification, GC-rich 16S sequences are more difficult to unwind due to their higher double-strand stability, requiring higher temperatures and energy. When PCR reaction conditions cannot meet these unwinding requirements, the double-stranded DNA cannot be fully unwound, making it difficult for primers to bind to the template strand, thus reducing amplification efficiency. Furthermore, GC content also affects the specificity and stability of primer-template binding. Excessively high GC content may lead to non-specific binding between primers and template, increasing the probability of primer dimer formation, further consuming primers and reaction substrates, and affecting the amplification of the target sequence.
[0031] Step 102: Input the GC content of each first genome fragment into the pre-trained amplification efficiency prediction model, and output the amplification efficiency corresponding to each first genome fragment; wherein, the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of second genome fragments with known GC content under standard PCR reaction conditions.
[0032] The amplification efficiency for each first genomic fragment output ranges from 0% to 100%, such as 0%, 20%, 50%, 70%, 100%, etc.
[0033] Specifically, the amplification efficiency of a second genomic fragment with a known GC content under standard PCR reaction conditions was obtained through the following process: PCR amplification of a second genomic fragment with a known GC content was performed under precisely controlled reaction conditions (including annealing temperature, Mg...). 2+ Based on the concentration of GC, dNTP concentration, primer concentration, and cycle number, the amount of amplified product is accurately determined using qPCR after amplification to obtain the amplification efficiency at that GC content. During the PCR amplification stage, multiple parallel groups can be set up to perform multiple parallel PCR amplifications of the same second genomic fragment with a known GC content, and quantitative detection is performed separately to ensure the reliability of the amplification efficiency.
[0034] It should be noted that, in this embodiment of the invention, by collecting a large number of 16S rRNA gene sequences from different microorganisms (which can cover a wide range of GC contents), and performing PCR amplification on these 16S rRNA gene sequences respectively, and by quantitatively analyzing the amplification results, the amplification efficiency corresponding to the genomic fragments with different GC contents can be obtained. This can be used as training data to train the model, thereby obtaining a pre-trained amplification efficiency prediction model.
[0035] For example, the amplification efficiency prediction model can be a machine learning model such as a neural network model or a vector machine model. The training data includes genomic fragments with different GC contents and the amplification efficiency corresponding to each GC content. By continuously adjusting the model parameters, an amplification efficiency prediction model that can accurately predict amplification efficiency based on GC content can be obtained. In addition to machine learning models, mathematical modeling methods such as linear regression or nonlinear regression can also be used to analyze genomic fragments with different GC contents and amplification efficiencies, and construct a functional relationship model between GC content and amplification efficiency.
[0036] In this embodiment of the invention, the accuracy and effectiveness of the pre-trained amplification efficiency prediction model were also verified. Specifically, a series of unknown samples with unknown GC content and microbial flora were selected. The GC content of the 16S sequence in the unknown samples was determined using sequencing technology, and this result was input into the pre-trained amplification efficiency prediction model to predict the amplification efficiency of the unknown samples under standard PCR conditions. Subsequently, actual PCR amplification experiments were performed on these unknown samples, with reaction conditions strictly controlled to be consistent with those used in model construction. After amplification, the actual amplification efficiency was measured using qPCR and compared with the amplification efficiency predicted by the model to verify the accuracy of the amplification efficiency predicted by the amplification efficiency model trained in this embodiment of the invention.
[0037] Step 103: For each first genome fragment: Correct the first 16S copy number according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genome fragment; For example, suppose the gene sequencing results include a first genome fragment A and a first genome fragment B, where the first 16S copy number of first genome fragment A is 5000, and the first 16S copy number of first genome fragment B is also 5000. However, due to the different amplification efficiencies of first genome fragment A and first genome fragment B during PCR amplification, the first 16S copy number of 5000 for first genome fragment A and the first 16S copy number of 5000 for first genome fragment B are not directly comparable. Therefore, it is necessary to correct the first 16S copy number based on amplification efficiency. For example, if step 102 yields an amplification efficiency of 100% for the first genomic fragment A, then in one PCR cycle, the 16S atoms in fragment A will double, exhibiting a quantity increase of 1 → 2 → 4 → 8 → 16 → ... over multiple cycles. However, if the amplification efficiency for the first genomic fragment B is 90%, then in one PCR amplification cycle, the 16S atoms in fragment B will only increase 1.9 times, exhibiting a quantity increase of 1 → 1.9 → 1.9 over multiple cycles. 2 → 1.9 3 → 1.9 4 →……。 It is evident that as the number of cycles increases, the amplification efficiency directly affects the 16S copy number. Therefore, in this embodiment of the invention, step 103 may specifically include: for each first genomic fragment: the ratio between the first 16S copy number and the amplification efficiency is used as the second 16S copy number corresponding to the first genomic fragment. That is, the second 16S copy number corresponding to the first genomic fragment A is 5000 / 100%=5000, while the second 16S copy number corresponding to the first genomic fragment B is 5000 / 90%=5555.55.
[0038] Step 104: Determine the absolute abundance of each first genome fragment based on the second 16S copy number and the total genome copy number.
[0039] It is understood that the 16S rRNA gene is only a small fixed gene segment in the genome fragment, and different genome fragments may contain different numbers of 16S rRNA genes. Therefore, the 16S copy number is not equivalent to the genome copy number. Therefore, in the embodiments of the present invention, not only is the 16S copy number in each genome fragment corrected by using amplification efficiency, but the genome copy number corresponding to the genome fragment is also calculated by referring to the number of 16S included in each genome fragment, so as to obtain the absolute abundance of each first genome fragment.
[0040] In an optional embodiment, step 104 specifically includes: obtaining the 16S copy number information corresponding to each first genome fragment; for each first genome fragment: obtaining the genome copy number corresponding to the first genome fragment based on the corresponding 16S copy number information and the second 16S copy number. For known microbial communities, the 16S copy number information can be directly obtained, while for unknown microbial communities, it is necessary to determine the corresponding 16S copy number information through estimation. Therefore, in a further optional embodiment, step 104 can be as follows: Figure 2 As shown, it specifically includes: Step 201: In response to the existence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, obtain the 16S copy number information corresponding to the operational taxonomic unit sequence from the ribosomal RNA operon copy number database; The ribosomal RNA operons(rrn) Database (rrnDB) is an important tool in the field of microbial research for the precise analysis of microbial community composition and abundance. It collects 16S copy number information for bacteria and archaea from NCBI whole-genome data. This data comes from a large number of sequenced microbial genomes and has undergone rigorous processing and validation, ensuring high accuracy and reliability. Therefore, for known microbial communities, 16S copy number information can be directly retrieved from the rrnDB database.
[0041] Specifically, due to the sheer volume of genomic fragments and the fact that many similar fragments share the same 16S copy number information, the rrnDB database stores data by Operational Taxonomic Unit (OTU). Each OTU represents multiple genomic fragments that can be clustered. By uploading OTUs to the rrnDB database's online analysis platform (https: / / rrndb.umms.med.umich.edu / ), species annotation results can be directly obtained, assigning corresponding 16S copy number information to each OTU.
[0042] Furthermore, this embodiment of the invention validates the method described above for obtaining 16S copy number information using the rrnDB database, and further performing copy number correction and calculating abundance based on the 16S copy number information. Specifically, it includes: selecting a series of standard strain samples with known copy numbers. These standard strains cover different microbial phyla, including common strains such as *Escherichia coli*, *Bacillus subtilis*, and *Staphylococcus aureus*, and their 16S copy numbers have been accurately determined through previous genome sequencing and validation. Following standard experimental procedures, DNA extraction, 16S gene amplification, and high-throughput sequencing are performed on these standard strain samples to obtain uncorrected species abundance data. Copy number correction is then performed on the sequencing data using the rrnDB database to obtain corrected species abundance data. By comparing the species abundance data before and after correction, and calculating the fold change in abundance of each strain before and after correction, the results are as follows: In one set of experiments, for Escherichia coli with a 16S copy number of 7, its abundance in the sequencing data was 30% before correction, while after copy number correction, its abundance was adjusted to 4.3%; for Bacillus subtilis with a copy number of 10, its abundance was 10% before correction, and after correction, its abundance was adjusted to 1%, which is basically consistent with its actual abundance.
[0043] As can be seen, the present invention, through the method described above of obtaining 16S copy number information using the rrnDB database and further performing copy number correction and calculating abundance based on the 16S copy number information, can effectively eliminate abundance deviation caused by 16S copy number differences and obtain abundance data that reflects the true quantity.
[0044] Step 202: In response to the absence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, the 16S copy number information corresponding to the first genome fragment is determined using the genetic distance algorithm; For unknown microbial communities, this embodiment of the invention utilizes a genetic distance algorithm to estimate the 16S copy number information corresponding to the first genome fragment by comparing it with the phylogenetic relationships of other known microbial communities. Specifically, in an optional embodiment, step 202 includes: using a genetic distance algorithm to determine the target genome fragment that is most closely related to the first genome fragment; wherein, the target genome fragment is a genome fragment with known 16S copy number information; and using the 16S copy number information corresponding to the target genome fragment as the 16S copy number information corresponding to the first genome fragment.
[0045] For example, assuming strain A has no copy number information and strain B has a known copy number, and the genetic distance between strains A and B is d, the copy number of strain B can be weighted and adjusted based on the genetic distance d to estimate the relative copy number of strain A. For instance, if the copy number of strain B is n, and the genetic distance d between strains A and B is small, it indicates that they are closely related, and the relative copy number of strain A can be estimated as n×(1-d), (0≤d≤1); conversely, if the genetic distance d is large, the estimated relative copy number of strain A will decrease accordingly. Furthermore, 16S copy number information can be estimated by constructing a genetic distance matrix containing strains with known copy numbers and strains with no copy number information. By analyzing this matrix, multiple closely related strains with known copy numbers can be found for each strain with no copy number information, and their copy numbers and genetic distances can be comprehensively considered to perform a more accurate relative abundance estimation. This genetic distance-based correction strategy can, to some extent, compensate for the lack of copy number information in the rrnDB database and improve the accuracy of abundance estimation for strains without copy number information.
[0046] In an optional embodiment, the present invention verified step 202 through the following process: First, various representative microbial strains were obtained, including common strains with known copy numbers such as *Escherichia coli* and *Bacillus subtilis*, as well as some simulated strains without copy number information in the rrnDB database. These strains were mixed in different proportions to construct simulated samples with different community structures. For example, a sample containing 30% *Escherichia coli*, 20% *Bacillus subtilis*, and 50% simulated strains without copy number information was constructed to simulate a complex microbial community composition. 16S rRNA gene sequencing and analysis were performed on the simulated samples to obtain uncorrected strain abundance data. The abundance of the simulated samples was corrected using a genetic distance calculation correction method to obtain corrected abundance data. The abundance data before and after correction were compared with the actual proportion of added bacterial species. The results are as follows: In one set of experiments, for a simulated sample, the RMSE of the abundance estimate of the bacterial species without copy number information before correction was 0.2, while after correction by genetic distance calculation, the RMSE decreased to 0.1, indicating that the correction method can effectively improve the accuracy of abundance estimation.
[0047] Step 203: For each first genome fragment: Based on the corresponding 16S copy number information and the second 16S copy number, obtain the genome copy number corresponding to the first genome fragment.
[0048] Specifically, for each first genome segment: the ratio between the second 16S copy number and the 16S copy number information corresponding to the first genome segment is used as the genome copy number corresponding to the first genome segment.
[0049] The following is based on Figure 3 Taking an example, we will specifically describe one implementation of the method for determining the absolute abundance of the genome provided in this embodiment of the invention, such as... Figure 3 As shown, it includes: Step 301: Construct an amplification efficiency prediction model based on the amplification efficiency of second genomic fragments with known GC content under standard PCR reaction conditions; Step 302: Obtain the gene sequencing results of the sample to be tested; the gene sequencing results include multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment; Step 303: Input the GC content of each first genome fragment into the pre-trained amplification efficiency prediction model, and output the amplification efficiency corresponding to each first genome fragment; Step 304: For each first genome fragment: Correct the first 16S copy number according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genome fragment; Step 305: Determine whether the rrnDB database contains the OTU corresponding to the first genome; If so, proceed with steps 306 and 308; Step 306: Obtain the 16S copy number information corresponding to the OTU from the rrnDB database; If not, proceed to steps 307 and 308; Step 307: Use the genetic distance algorithm to determine the 16S copy number information corresponding to the first genome segment; Step 308: For each first genome fragment: Based on the corresponding 16S copy number information and the second 16S copy number, obtain the genome copy number corresponding to the first genome fragment.
[0050] In summary, the method for determining the absolute abundance of the genome provided in this embodiment of the invention utilizes a pre-trained amplification efficiency prediction model to predict the amplification efficiency based on the GC content in the first gene fragment. Furthermore, it corrects the first 16S copy number of each genome fragment based on the amplification efficiency corresponding to different genome fragments, thus removing the bias caused by differences in the amplification efficiency of various microorganisms in the quantitative results. Based on the corrected second 16S copy number, the absolute abundance of each microbial community in the gene sequencing results is determined.
[0051] The apparatus for determining absolute genome abundance provided by the present invention will be described below. The apparatus for determining absolute genome abundance described below can be referred to in correspondence with the method for determining absolute genome abundance described above.
[0052] like Figure 4As shown, the genome absolute abundance determination device 400 provided in this embodiment of the invention includes: The acquisition module 401 is used to acquire the gene sequencing results of the sample to be tested; the gene sequencing results include multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment; The prediction module 402 is used to input the GC content of each of the first genomic fragments into a pre-trained amplification efficiency prediction model, and output the amplification efficiency corresponding to each of the first genomic fragments; wherein, the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of a second genomic fragment with known GC content under standard PCR reaction conditions; The 16S copy number correction module 403 is used to correct the first 16S copy number for each of the first genome fragments according to the amplification efficiency, so as to obtain the second 16S copy number corresponding to the first genome fragment. The determination module 404 is used to determine the absolute abundance of each of the first genome fragments based on the second 16S copy number and the total genome copy number.
[0053] In an optional embodiment of the present invention, the determining module 404 is further configured to: obtain 16S copy number information corresponding to each of the first genome fragments; and for each of the first genome fragments: obtain the genome copy number corresponding to the first genome fragment based on the corresponding 16S copy number information and the second 16S copy number.
[0054] In an optional embodiment of the present invention, the determining module 404 is further configured to: in response to the existence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, obtain 16S copy number information corresponding to the operational taxonomic unit sequence from the ribosomal RNA operon copy number database; and in response to the absence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, determine the 16S copy number information corresponding to the first genome fragment using a genetic distance algorithm.
[0055] In an optional embodiment of the present invention, the determining module 404 is further configured to determine the target genome fragment most closely related to the first genome fragment using a genetic distance algorithm; wherein the target genome fragment is a genome fragment with known 16S copy number information; and the 16S copy number information corresponding to the target genome fragment is used as the 16S copy number information corresponding to the first genome fragment.
[0056] In an optional embodiment of the present invention, the determining module 404 is further configured to, for each first genome fragment, use the ratio between the second 16S copy number and the 16S copy number information corresponding to the first genome fragment as the genome copy number corresponding to the first genome fragment.
[0057] In an optional embodiment of the present invention, the 16S copy number correction module 404 is further configured to, for each first genomic fragment, use the ratio between the first 16S copy number and the amplification efficiency as the second 16S copy number corresponding to the first genomic fragment.
[0058] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 501, a communications interface 502, a memory 503, and a communications bus 504, wherein the processor 501, the communications interface 502, and the memory 503 communicate with each other through the communications bus 504. The processor 501 can call logical instructions in the memory 503 to execute a method for determining the absolute abundance of the genome. This method includes: acquiring gene sequencing results of a sample to be tested; the gene sequencing results include multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment; inputting the GC content in each of the first genome fragments into a pre-trained amplification efficiency prediction model, and outputting the amplification efficiency corresponding to each first genome fragment; wherein the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of second genome fragments with known GC content under standard PCR reaction conditions; for each first genome fragment: correcting the first 16S copy number according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genome fragment; and determining the absolute abundance of each first genome fragment based on the second 16S copy number and the total genome copy number.
[0059] Furthermore, the logical instructions in the aforementioned memory 503 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0060] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the method for determining the absolute abundance of the genome provided by the above methods. The method includes: acquiring the gene sequencing results of the sample to be tested; the gene sequencing results include multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment; inputting the GC content in each of the first genome fragments into a pre-trained amplification efficiency prediction model, and outputting the amplification efficiency corresponding to each first genome fragment; wherein the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of second genome fragments with known GC content under standard PCR reaction conditions; for each first genome fragment: correcting the first 16S copy number according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genome fragment; and determining the absolute abundance of each first genome fragment according to the second 16S copy number and the total genome copy number.
[0061] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for determining the absolute abundance of the genome provided by the methods described above. This method includes: acquiring gene sequencing results of a sample to be tested; the gene sequencing results including a plurality of first genome fragments, a total genome copy number, and a first 16S copy number corresponding to each of the genome fragments; inputting the GC content in each of the first genome fragments into a pre-trained amplification efficiency prediction model, and outputting the amplification efficiency corresponding to each of the first genome fragments; wherein the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of second genome fragments with known GC content under standard PCR reaction conditions; for each first genome fragment: correcting the first 16S copy number according to the amplification efficiency to obtain a second 16S copy number corresponding to the first genome fragment; and determining the absolute abundance of each first genome fragment based on the second 16S copy number and the total genome copy number.
[0062] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0063] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for determining the absolute abundance of a genome, characterized in that, include: Obtain the gene sequencing results of the sample to be tested; the gene sequencing results include multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment; The GC content of each of the first genomic fragments is input into a pre-trained amplification efficiency prediction model, and the amplification efficiency corresponding to each first genomic fragment is output. The amplification efficiency prediction model is obtained based on the amplification efficiency analysis of a second genomic fragment with known GC content under standard PCR reaction conditions. For each of the first genome fragments: the first 16S copy number is corrected according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genome fragment; The absolute abundance of each of the first genome fragments is determined based on the second 16S copy number and the total genome copy number.
2. The method for determining absolute genomic abundance according to claim 1, characterized in that, Determining the absolute abundance of each first genome fragment based on the second 16S copy number and the total genome copy number includes: Obtain the 16S copy number information corresponding to each of the first genome fragments; For each of the first genome fragments: the genome copy number corresponding to the first genome fragment is obtained based on the corresponding 16S copy number information and the second 16S copy number.
3. The method for determining absolute genomic abundance according to claim 2, characterized in that, The step of obtaining the 16S copy number information corresponding to each of the first genome fragments includes: In response to the existence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, the 16S copy number information corresponding to the operational taxonomic unit sequence is obtained from the ribosomal RNA operon copy number database; In response to the absence of an operational taxonomic unit sequence corresponding to the first genome fragment in the ribosomal RNA operon copy number database, the 16S copy number information corresponding to the first genome fragment is determined using a genetic distance algorithm.
4. The method for determining absolute genomic abundance according to claim 3, characterized in that, The step of determining the 16S copy number information corresponding to the first genomic fragment using a genetic distance algorithm includes: A genetic distance algorithm is used to determine the target genome fragment that is most closely related to the first genome fragment; wherein, the target genome fragment is a genome fragment with known 16S copy number information; The 16S copy number information corresponding to the target genome fragment is used as the 16S copy number information corresponding to the first genome fragment.
5. The method for determining absolute genomic abundance according to claim 2, characterized in that, For each of the first genome fragments: the genome copy number corresponding to the first genome fragment is obtained based on the corresponding 16S copy number information and the second 16S copy number, including: For each of the first genome fragments: the ratio between the second 16S copy number and the 16S copy number information corresponding to the first genome fragment is used as the genome copy number corresponding to the first genome fragment.
6. The method for determining absolute genomic abundance according to claim 1, characterized in that, For each of the first genomic fragments: the first 16S copy number is corrected according to the amplification efficiency to obtain the second 16S copy number corresponding to the first genomic fragment, including: For each of the first genomic fragments: the second 16S copy number corresponding to the first genomic fragment is determined by the ratio between the first 16S copy number and the amplification efficiency.
7. A device for determining the absolute abundance of a genome, characterized in that, include: The acquisition module is used to acquire the gene sequencing results of the sample to be tested; the gene sequencing results include multiple first genome fragments, the total genome copy number, and the first 16S copy number corresponding to each genome fragment; The prediction module is used to input the GC content of each of the first genomic fragments into a pre-trained amplification efficiency prediction model and output the amplification efficiency corresponding to each of the first genomic fragments; wherein, the amplification efficiency prediction model is obtained based on the amplification efficiency analysis of a second genomic fragment with known GC content under standard PCR reaction conditions; A 16S copy number correction module is used to correct the first 16S copy number for each of the first genome fragments according to the amplification efficiency, so as to obtain a second 16S copy number corresponding to the first genome fragment. The determination module is used to determine the absolute abundance of each of the first genome fragments based on the second 16S copy number and the total genome copy number.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for determining absolute genomic abundance as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for determining the absolute abundance of the genome as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for determining the absolute abundance of the genome as described in any one of claims 1 to 6.