Microbial pollution source tracing method based on Sanger sequencing overlapping peak analysis
By constructing a standard gene sequence library of pollution sources and analyzing overlapping peaks to generate base combination sequences, the problem of identification failure of multiple base overlapping peaks in Sanger sequencing was solved, enabling precise tracing of pollution sources and improving the accuracy of strain identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-31
AI Technical Summary
Existing Sanger sequencing technology cannot effectively resolve contamination sources when faced with multi-base overlapping peaks, leading to identification failures, increased costs, and project delays, and it cannot accurately trace the source of contamination.
By constructing a standard gene sequence library of pollution sources, analyzing overlapping peaks to generate all possible base combination sequences, comparing them with microbial databases, and calculating the reliability based on peak height ratios, the main strains and strains requiring traceability can be accurately identified, thus tracing the source of pollution.
This system enables systematic analysis of overlapping peaks, improving the accuracy and efficiency of strain identification, accurately locating contamination sources, optimizing experimental procedures, and reducing costs and the risk of delays.
Abstract
Description
Technical Field
[0001] This invention relates to the field of microbial sequencing and analysis technology, specifically to a method for tracing microbial contamination sources based on Sanger sequencing overlap peak analysis. Background Technology
[0002] In the field of microbial identification and classification, Sanger sequencing technology is currently the most widely used and accepted gold standard method for determining the sequence of specific genetic markers—such as the internal transcribed spacer region of fungi or the 16S rRNA gene of prokaryotes. This technology, through PCR amplification and capillary electrophoresis sequencing of the DNA of the test strain, can obtain clear, continuous peak patterns and highly accurate base sequences. Homology comparison with existing authoritative databases then enables precise identification and classification of the strain. Due to its standardized operating procedures, high technological maturity, and relatively low cost per test, this method is widely used in many important scenarios such as clinical diagnosis, food safety monitoring, environmental microbiological investigation, industrial strain identification, and microbial traceability, forming a fundamental and core technical support within microbiology laboratories.
[0003] However, in practice, this technology is often interfered with by various experimental factors, leading to abnormal sequencing results. The most typical and challenging problem is the appearance of "multi-base overlap peaks" in the sequencing peak diagram, where two or more different bases (A, T, C, G) appear simultaneously at the same electrophoretic migration position. Systematic analysis reveals that the root causes of this phenomenon can be mainly categorized into two types: first, biological mixing, where researchers fail to isolate pure cultures during plating and colony picking, resulting in colonies that are themselves a mixture of two or more microorganisms; second, process contamination, where cross-contamination of nucleic acids occurs between different samples or with the environment during a series of experimental steps, such as genomic DNA extraction, PCR reaction system preparation, amplification product purification, or sequencing library construction, resulting in the final sequencing template containing target gene fragments from different sources. When such overlapping signals appear in the sequencing peak diagram, the direct consequence is that a clear and unique DNA sequence cannot be obtained through conventional base reading software. This makes it impossible to carry out subsequent database comparison and strain identification work effectively, which seriously affects the accuracy, reliability and timeliness of the identification report.
[0004] Faced with these challenges, the current industry standard approach is passively discarding and repeating experiments. Once overlapping peaks are detected, the usual practice is to declare the sequencing a failure and start over from the original sample or preserved strain, repeating the entire process of single-colony isolation, DNA extraction, PCR amplification, and even Sanger sequencing. This start-from-beginner approach has significant drawbacks: First, it greatly increases manpower, reagent, and time costs, extending project cycles and potentially causing serious delays in time-sensitive scenarios; second, it essentially discards problematic data, failing to extract any valuable information from the generated sequencing results; most importantly, this method can only determine "the presence of contamination or mixing," but cannot reveal which microorganisms or where the contamination originated. In other words, current technology lacks the ability to analyze and trace contamination events themselves. This makes it impossible for researchers to pinpoint whether the contamination originated from the environment, reagents, cross-samples, or improper initial separation, thus hindering targeted improvements to experimental procedures, optimization of quality control systems, and effective at tracing responsibility and analyzing root causes of quality disputes or infection control issues. This predicament of "knowing there is a problem but not knowing what the problem is" has become a major technical bottleneck in improving the testing quality and traceability capabilities of microbiology laboratories. Summary of the Invention
[0005] The purpose of this invention is to provide a method for tracing microbial contamination sources based on Sanger sequencing overlap peak analysis. This method not only solves the technical bottleneck of overlap peak analysis, but also achieves accurate identification and contamination tracing, significantly improving the efficiency, accuracy and process control capabilities of microbial detection.
[0006] A method for tracing microbial contamination sources based on Sanger sequencing overlap peak analysis includes the following steps:
[0007] S1. Select the characteristic gene regions for identification based on the type of microorganism to be tested;
[0008] S2. Systematically collect samples from various potential pollution sources involved in the experiment, perform Sanger sequencing based on the selected characteristic gene regions, obtain standard gene sequences for each potential pollution source, and construct a standard gene sequence library for pollution sources.
[0009] S3. Obtain sequencing data of the microbial sample to be tested and determine overlapping peaks: Perform Sanger sequencing on the microbial sample to be tested based on the selected characteristic gene region to obtain a sequencing peak diagram; if there are multi-base overlapping peaks in the sequencing peak diagram, proceed to S4; otherwise, follow the conventional single sequence identification.
[0010] S4. Analyze overlapping peaks to identify strain composition: Analyze the multi-base overlapping peaks to generate all possible base combination sequences, compare the base combination sequences with the microbial gene database, and screen out effective base combination sequences with homology reaching the first threshold.
[0011] If all valid base combinations correspond to the same known strain, then the microbial sample to be tested is determined to be a pure culture of that strain.
[0012] If the effective base combination sequence corresponds to multiple known strains, then the main strain and the strain to be traced are distinguished, and S5 is executed;
[0013] S5. Source tracing: The effective base combination sequence of the strain to be traced in S4 is compared with the standard gene sequence library of the source of pollution.
[0014] If the homology between the effective base combination sequence of a strain to be traced and the standard gene sequence of a certain pollution source in the pollution source standard gene sequence library reaches the second threshold, then the strain to be traced is determined to originate from that specific pollution source.
[0015] If the effective base combination sequence of a strain to be traced does not match in the standard gene sequence library of the pollution source, then the strain to be traced is determined to be an inherent microbial component in the microbial sample to be tested or a mixed strain of non-process-related contamination.
[0016] As a preferred embodiment of the present invention, in step S1, the potentially contaminated sample includes an experimental environment sample, an experimental equipment control sample, and other strain samples detected at the same time.
[0017] As a preferred technical solution of the present invention, in step S3, the sequencing peak diagram is first evaluated for quality, and the sequence fragments with sequencing quality value Q≥20 are selected before the multi-base overlapping peak identification is performed.
[0018] As a preferred embodiment of the present invention, in step S3, the criterion for determining the polybasic overlap peak is:
[0019] At the same base position in the sequencing peak diagram, two or more characteristic peaks with a height not less than 50% of the height of the main peak appear, and each characteristic peak corresponds to a different base.
[0020] As a preferred embodiment of the present invention, in step S4, if the effective base combination sequence corresponds to multiple known strains, then step S5 is executed, specifically including:
[0021] When the effective base combination sequence corresponds to multiple known strains, the sequence confidence level of each known strain is calculated, and the known strain with the highest confidence level is determined as the main strain, while the remaining known strains are the strains to be traced.
[0022] Then, step S5 is performed using the strain to be traced as a potential mixed sample.
[0023] As a preferred embodiment of the present invention, the method for calculating the credibility includes:
[0024] For each effective base combination sequence belonging to the same known strain, the product of the peak height ratio of the corresponding base at each multi-base overlap peak position of the effective base combination sequence is calculated as the confidence level of the effective base combination sequence.
[0025] The total sequence confidence of a known strain is obtained by summing the confidence scores of all valid base combination sequences belonging to the same known strain.
[0026] As a preferred technical solution of the present invention, when generating all possible base combination sequences, for any of the multi-base overlapping peak positions, bases with a peak height ratio of less than 10% at that position are excluded, and base combination sequences containing such bases are not generated.
[0027] In a preferred embodiment of the present invention, in step S4, the first threshold is 97%.
[0028] In a preferred embodiment of the present invention, in step S5, the second threshold is 99%.
[0029] As can be seen from the above technical solutions, the technical solution of the present invention provides a method for tracing microbial contamination sources based on Sanger sequencing overlap peak analysis, which has the following beneficial effects:
[0030] First, this invention fundamentally solves the technical problem of identification failure caused by multiple base overlap peaks in sequencing peak diagrams. By systematically analyzing the overlap peaks and enumerating base combinations, it is possible to mine and reconstruct possible strain sequences from abnormal data that are traditionally considered invalid. This enables full utilization of existing sequencing data and avoids the waste of time, manpower and reagent costs caused by re-experimentation.
[0031] Second, this invention significantly improves the accuracy and reliability of strain identification. By performing high-threshold homology comparisons between the generated sequences and authoritative databases, and by introducing a sequence reliability calculation model based on peak height ratio, the main target strain and confounding strains in the sample can be scientifically distinguished, leading to clear and reliable identification conclusions. This effectively overcomes the limitations of traditional methods in dealing with mixed signals.
[0032] Third, this invention enables precise location and systematic tracing of pollution sources. By pre-constructing a gene sequence library of pollution sources covering the environment, equipment, and samples from the same period, identified contaminating strains can be rapidly compared with potential pollution sources, thus clearly determining whether the pollution originates from the experimental process or is inherent in the sample itself. This function provides direct technical support for laboratories to conduct root cause analysis of pollution, optimize operating procedures, and improve quality management systems.
[0033] Furthermore, this method is highly versatile and easy to implement. It can be seamlessly integrated with conventional microbial databases without requiring changes to existing Sanger sequencing procedures or the addition of specialized equipment. It is suitable for identification scenarios of various microorganisms such as fungi and bacteria and has good prospects for widespread application.
[0034] It should be understood that all combinations of the foregoing concepts and the additional concepts described in more detail below can be considered part of the inventive subject matter of this disclosure, provided that such concepts do not contradict each other. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below in conjunction with the embodiments of this invention. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the described embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art to which this invention pertains.
[0036] The terms "first," "second," and similar words used in the specification and claims of this patent application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, unless the context clearly indicates otherwise, the singular forms of "an," "a," or "the," etc., do not indicate a quantity limitation, but rather indicate the presence of at least one. Terms such as "comprising" or "including" mean that the element or object preceding "comprising" encompasses the features, integrals, steps, operations, elements, and / or components listed following "comprising" or "including," and do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0037] This invention addresses the lack of analytical and tracing capabilities in existing technologies for contamination events, which prevents researchers from pinpointing whether contamination originates from the environment, reagents, cross-samples, or improper initial separation. Consequently, it hinders targeted improvements to experimental procedures and optimization of quality control systems, and makes it impossible to address quality disputes or infection control issues. This invention provides a method for tracing microbial contamination sources based on Sanger sequencing overlap peak analysis.
[0038] The method for tracing microbial contamination sources includes the following specific steps:
[0039] S1. Based on the type of strain to be identified, such as fungi or prokaryotes, select the corresponding characteristic gene regions to design specific primers; for example, select the ITS gene region of fungi or the 16S rRNA gene region of prokaryotes to design specific primers.
[0040] S2. Construct a standard gene sequence library for pollution sources.
[0041] Specifically, the process involves systematically collecting samples from all possible potential sources of contamination involved in the experiment. These potential contamination samples include, but are not limited to, samples from the experimental environment, control samples from experimental equipment, and samples from other strains tested concurrently. Based on the characteristic gene regions selected in step S1, target DNA is extracted from the potential contamination source samples. Using the extracted target DNA as a template, PCR amplification is performed using the obtained specific primers to obtain the target gene amplification products. After purification, the target gene amplification products are bidirectionally sequenced using Sanger sequencing technology to obtain the gene sequences of each potential contamination source. These sequences are then archived and stored to create a library of corresponding contamination source standard gene sequences, which serves as a source tracing and comparison database.
[0042] S3. Obtain sequencing data of the microbial sample to be tested and determine overlapping peaks.
[0043] Microbial samples to be tested are picked from the culture medium to obtain samples of strains to be identified; target DNA is extracted from the samples of strains to be identified based on the characteristic gene regions selected in step S1; using the extracted target DNA as a template, PCR amplification is performed using the specific primers obtained in step S1 to obtain the target gene amplification product of the samples of strains to be identified; after purifying the target gene amplification product of the samples of strains to be identified, bidirectional sequencing is performed using Sanger sequencing technology to obtain sequencing peak diagram files and raw sequencing data of the strains to be identified.
[0044] A Python script was used to assess the quality of the sequencing peak diagrams of the obtained strains to be identified, screening out sequence fragments with a sequencing quality value Q ≥ 20. Then, existing sequencing peak diagram analysis software was used to identify whether multi-base overlapping peaks existed in the valid sequence fragments. Specifically, the criteria for determining the presence of multi-base overlapping peaks were: at the same base position in the sequencing peak diagram, two or more characteristic peaks with heights not less than 50% of the height of the main peak appeared, and each of these characteristic peaks corresponded to a different base. The main peak was the peak with the highest signal intensity.
[0045] If the sequencing peaks show multi-base overlapping peaks, proceed to step S4. Otherwise, identify the sequence according to existing conventional single-sequence identification methods or procedures.
[0046] S4. Analyze overlapping peaks to identify strain species.
[0047] Information analysis is performed on the identified multi-base overlapping peaks, recording all possible base types and peak height percentages corresponding to each multi-base overlapping peak position. Based on the positional order of the multi-base overlapping peaks, all possible base combination sequences are generated using methods such as recursive enumeration. Specifically, for any base whose peak height percentage in a multi-base overlapping peak is less than 10%, it can be excluded to reduce the number of invalid base combination sequences; that is, no base combination sequence containing that base is generated.
[0048] For example, if the number of multi-base overlapping peaks is set to n, and the number of possible bases corresponding to each multi-base overlapping peak position is k (i=1,2,3…n), then the total number of base combination sequences generated is k×k×k…×k. Generally, the raw peak diagram or data from sequencing will give the signal value of each base A / T / C / G at each position. In the case of a single peak, only one base has the highest signal value at that position, while the other three signal values are extremely low or zero; in the case of overlapping peaks, there is not only one base with a signal value at that position, and all of them have relatively high values.
[0049] All generated base combination sequences were compared with existing microbial gene databases for homology. Base combination sequences with homology reaching a first threshold with known strain gene sequences in authoritative microbial gene databases were selected and recorded as valid base combination sequences. Base combination sequences with homology below the first threshold were excluded and recorded as invalid base combination sequences. Authoritative microbial gene databases include, but are not limited to, NCBI GenBank and UNITE databases; the first threshold is preferably 97%.
[0050] If the effective base combination sequence corresponds to the same known strain, the sample of the strain to be identified is determined to be a pure culture of that known strain.
[0051] If the effective base combination sequence corresponds to multiple known strains, then the dominant strain and the strain requiring tracing are distinguished. Specifically, the method for distinguishing the dominant strain and the strain requiring tracing includes calculating the sequence confidence level corresponding to each known strain, and determining the known strain with the highest confidence level as the dominant strain. The gene sequence of the dominant strain matches the expected target or phenotypic characteristics of the sample to be identified, and therefore it is identified as the target microorganism to be analyzed in the sample. The remaining known strains are the strains requiring tracing, and the source of these strains is further analyzed as potential mixed samples.
[0052] Optionally, the confidence level is calculated as follows: for each effective base combination sequence belonging to the same known strain, the peak height ratio of the corresponding base at each multi-base overlap peak position of the effective base combination sequence is calculated, and the product result is used as the confidence level of the effective base combination sequence; the confidence levels of all effective base combination sequences belonging to the same known strain are summed to obtain the total sequence confidence level of the known strain.
[0053] S5. Tracing pollution sources.
[0054] The effective base combination sequence of the strain to be traced is compared with the standard gene sequence library of the pollution source for homology comparison.
[0055] If the homology between the effective base sequence of a strain to be traced and the standard gene sequence of a pollution source in the pollution source standard gene sequence library reaches a second threshold, then the strain is determined to originate from the corresponding pollution source. Preferably, the second threshold is 99%.
[0056] If the effective base combination of a certain strain to be traced does not match in the standard gene sequence library of the pollution source, then the strain is determined to be an inherent mixed strain in the microbial sample to be tested.
[0057] Example 1
[0058] Fungal strain identification and contamination source tracing
[0059] This embodiment uses the identification of suspected mold strains isolated from a winery as an example to illustrate the implementation process of the present invention in detail:
[0060] (1) Sequencing sample preparation and Sanger sequencing: The strain to be identified is a fungus. Specific primers were designed for the ITS gene region. The upstream primer sequence is 5'-TCCGTAGGTGAACCTGCGG-3', and the downstream primer sequence is 5'-TCCTCCGCTTATTGATATGC-3'. Genomic DNA of the suspected fungal strain was extracted using the CTAB method. PCR amplification was performed using the DNA as a template. The PCR reaction system (50 μL) consisted of: 2×Taq PCR MasterMix 25 μL, upstream and downstream primers (10 μmol / L) 2 μL each, template DNA 5 μL, and ddH2O 16 μL. PCR reaction conditions: 94℃ pre-denaturation for 5 min; 94℃ denaturation for 30 s, 58℃ annealing for 30 s, 72℃ extension for 45 s, for a total of 35 cycles; and 72℃ final extension for 10 min. After amplification, the amplified products were verified by agarose gel electrophoresis and purified using a PCR product purification kit. The purified amplified products were sent to a sequencing company for Sanger bidirectional sequencing to obtain sequencing peak files and raw sequencing data. All subsequent analysis steps were completed using Python scripts.
[0061] (2) Sequencing peak quality assessment and overlapping peak identification: The sequencing peak file is read by Python script, and a quality filtering module is written to set the quality threshold Q=20 to automatically filter out effective sequence fragments with a length of 550bp; the script has a built-in overlapping peak identification algorithm (secondary peak height ≥ 50% of the main peak). After analysis, the multi-base overlapping peaks at the 120bp, 245bp and 360bp positions of the sequence are accurately located, and the peak height percentage is automatically calculated: at the 120bp position, A (60%), T (40%); at the 245bp position, C (70%), G (30%); at the 360bp position, A (55%), C (45%).
[0062] (3) Overlapping peak information analysis and base combination enumeration: After reading the above overlapping peak positions and base proportion data, the Python script calls the Cartesian product algorithm (simulating recursive enumeration logic) to automatically generate all possible base combination sequences, resulting in 2×2×2=8 base combination sequences: Sequence 1 (120A-245C-360A), Sequence 2 (120A-245C-360C), Sequence 3 (120A-245G-360A), Sequence 4 (120A-245G-360C), Sequence 5 (120T-245C-360A), Sequence 6 (120T-245C-360C), Sequence 7 (120T-245G-360A), and Sequence 8 (120T-245G-360C). At the same time, the script automatically records the peak height proportion product corresponding to each base combination sequence, which is the basis for confidence calculation.
[0063] (4) Homology alignment and effective sequence screening of base combination sequences: The Python script loads localized FASTA format files from the UNITE and NCBIGenBank databases, has a built-in sequence homology calculation module (number of identical bases / shortest sequence length), sets a homology threshold of 97%, and performs batch alignment of 8 base combination sequences. The results show that the script automatically determines that sequence 1 has 99.8% homology with Aspergillus flavus, sequence 5 has 99.5% homology with Penicillium expansum, and the homology of the remaining 6 base combination sequences is below the first threshold, and they are marked as invalid base combination sequences.
[0064] (5) Strain identification: The Python script extracts the peak height ratio data of the effective base combination sequence and automatically calculates the confidence level: the confidence level of sequence 1 = 60% × 70% × 55% = 23.1%, and the confidence level of sequence 5 = 40% × 70% × 55% = 15.4%. After sorting the confidence levels in descending order, the script determines that the main strain of the strain to be identified is Aspergillus flavus and the potential mixed strain is Penicillium expansum, and outputs the identification report.
[0065] (6) Tracing the source of contamination: The Python script loaded the sequencing data (FASTA format) of the experimental bench wiping sample, the pipette tip control sample, and the Penicillium expansum standard strain sample tested at the same time. The script automatically compared sequence 5 with the standard gene sequences of each contamination source. The results showed that sequence 5 had 99.9% homology with the Penicillium expansum standard strain and less than 90% homology with other contamination sources. The script combined the experimental operation timestamp data to automatically associate the sample with the Penicillium expansum standard strain's concurrent treatment records. Finally, it was determined that the source of contamination was the Penicillium expansum standard strain tested at the same time, and the cause of contamination was cross-contamination during pipetting operations.
[0066] Example 2
[0067] Identification of prokaryotic strains and tracing of pollution sources
[0068] This embodiment uses the identification of suspected bacterial strains isolated from a winery as an example to illustrate in detail the entire implementation process of this invention based on Python scripts:
[0069] (1) Sequencing Sample Preparation and Sanger Sequencing: The strain to be identified was a prokaryote. Universal primers were designed based on the 16S rRNA gene region. The upstream primer sequence was 5'-AGAGTTTGATCCTGGCTCAG-3', and the downstream primer sequence was 5'-GGTTACCTTGTTACGACTT-3'. Genomic DNA of the strain to be identified was extracted using a bacterial genomic DNA extraction kit. PCR amplification was performed using the DNA as a template. The PCR reaction system (50 μL) consisted of: 25 μL of 2×Taq PCR MasterMix, 2 μL each of upstream and downstream primers (10 μmol / L), 3 μL of template DNA, and 18 μL of ddH2O. PCR reaction conditions were: 95℃ pre-denaturation for 5 min; 95℃ denaturation for 30 s, 55℃ annealing for 30 s, 72℃ extension for 1 min, for a total of 30 cycles; and a final extension at 72℃ for 10 min. After purification, the amplified products were subjected to Sanger bidirectional sequencing to obtain sequencing peak files and raw sequencing data. All subsequent analysis steps were completed using Python scripts.
[0070] (2) Sequencing peak quality assessment and overlapping peak identification: The Python script reads the sequencing peak file and filters out effective sequence fragments with a Q value ≥ 20 and a length of 1450bp through the quality filtering module; the built-in overlapping peak identification algorithm accurately identifies the multi-base overlapping peaks at positions 520bp and 880bp and automatically calculates the peak height percentages: at position 520bp, G (65%) and A (35%); at position 880bp, T (75%) and C (25%).
[0071] (3) Overlapping peak information analysis and base combination enumeration: Based on the above overlapping peak data, the Python script calls the Cartesian product algorithm to automatically generate all possible base combination sequences, resulting in 2×2=4 base combination sequences: sequence 1 (520G-880T), sequence 2 (520G-880C), sequence 3 (520A-880T), and sequence 4 (520A-880C). At the same time, the peak height ratio data of each base combination sequence is recorded.
[0072] (4) Homology comparison and effective sequence screening of base combination sequences: The Python script loaded localized files from the EzBioCloud database and the NCBI GenBank database, set the homology threshold to 97%, and performed batch homology calculations on four base combination sequences. The results showed that the script automatically determined that sequence 1 had 99.7% homology with Escherichia coli, and sequence 3 had 99.6% homology with Staphylococcus aureus. The remaining two base combination sequences had homology below the threshold and were marked as invalid base combination sequences.
[0073] (5) Strain identification: The Python script extracts the peak height ratio data of the effective base combination sequence and automatically calculates the confidence level: the confidence level of sequence 1 = 65% × 75% = 48.75%, and the confidence level of sequence 3 = 35% × 75% = 26.25%. After sorting by confidence level, the script determines that the main strain is Escherichia coli and the potential mixed strain is Staphylococcus aureus.
[0074] (6) Tracing the source of contamination: The Python script loaded the sequencing data of the blank control of culture medium, the control after sterilization of the inoculation loop, and the air sample of the ward environment. The sequence 3 was automatically compared with the standard sequence of each source of contamination. The results showed that it had 99.8% homology with the control sample after sterilization of the inoculation loop. The script combined the sterilization record timestamp data and finally determined that the source of contamination was cross-contamination caused by incomplete sterilization of the inoculation loop.
[0075] In this embodiment of the invention, the letters A, T, C, and G are abbreviations for the four bases of deoxyribonucleic acid (DNA), where A represents adenine, T represents thymine, C represents cytosine, and G represents guanine. In Sanger sequencing, each base is labeled with a fluorescent signal of a specific wavelength and manifests as an independent signal peak. The peak height or peak area typically reflects the signal intensity of the base at that position.
[0076] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A method for tracing a source of microbial contamination based on Sanger sequencing peak overlap analysis, characterized in that, The method comprises the following steps: S1, selecting a characteristic gene region for identification according to the type of the microorganism to be tested; S2, systematically collecting various potential contamination source samples involved in the experiment, performing Sanger sequencing based on the selected characteristic gene region, obtaining standard gene sequences of each potential contamination source, and constructing a contamination source standard gene sequence library; S3, obtaining sequencing data of the microorganism sample to be tested and judging overlapping peaks: performing Sanger sequencing on the microorganism sample to be tested based on the selected characteristic gene region to obtain a sequencing peak graph; when there are multiple-base overlapping peaks in the sequencing peak graph, perform S4; otherwise, follow the conventional single sequence identification; S4, analyzing overlapping peaks to identify strain composition: analyzing the multiple-base overlapping peaks to generate all possible base combination sequences, comparing the base combination sequences with a microorganism gene database, and screening effective base combination sequences with a homology reaching a first threshold value; If all effective base combination sequences correspond to the same known strain, it is determined that the microorganism sample to be tested is a pure culture of the strain; If the effective base combination sequences correspond to multiple known strains, the main strain and the strain to be traced are distinguished, and S5 is performed; S5, contamination source tracing: comparing the effective base combination sequences of the strain to be traced in S4 with the contamination source standard gene sequence library for homology comparison; If the effective base combination sequence of a strain to be traced has a homology reaching a second threshold value with the standard gene sequence of a contamination source in the contamination source standard gene sequence library, it is determined that the strain to be traced is derived from the specific contamination source; If the effective base combination sequence of a strain to be traced has no match in the contamination source standard gene sequence library, it is determined that the strain to be traced is an inherent microbial component of the microorganism sample to be tested or a mixed strain of non-process contamination.
2. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 1, characterized in that, In S1, the potential contamination samples include experimental environment samples, experimental instrument control samples, and other strain samples detected at the same period.
3. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 1, characterized in that, In S3, the base sequence fragments with a sequencing quality value Q≥20 are first screened out for quality evaluation of the sequencing peak graph, and then the multiple-base overlapping peaks are identified.
4. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 2, characterized in that, In S3, the determination standard of the multiple-base overlapping peaks is: In the same base position of the sequencing peak graph, two or more characteristic peaks with a peak height not lower than 50% of the main peak height appear, and each characteristic peak corresponds to a different base.
5. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 1, characterized in that, In S4, if the effective base combination sequences correspond to multiple known strains, S5 is performed, specifically including: When the effective base combination sequences correspond to multiple known strains, the sequence reliability of each known strain is calculated, and the known strain with the highest reliability is determined as the main strain, and the remaining known strains are used as strains to be traced; And the strain to be traced is used as a potential mixed sample to perform the S5 step.
6. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 5, characterized in that, The calculation method of the reliability includes: For each effective base combination sequence belonging to the same known strain, the product of the peak height proportion of the corresponding base at each multiple-base overlapping peak position of the effective base combination sequence is calculated as the reliability of the effective base combination sequence. Summing up the reliabilities of all valid base combination sequences belonging to the same known strain, the total sequence reliability of the known strain is obtained.
7. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 6, characterized in that, In generating all possible base combination sequences, for any said multi-base overlapping peak position, the base with peak height less than 10% of the position is excluded, and the base combination sequence containing the base is not generated.
8. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 1, characterized in that, In the S4, the first threshold value is 97%.
9. The Sanger sequencing overlap peak analysis-based microbial contamination source tracing method according to claim 1, characterized in that, In the S5, the second threshold value is 99%.