Urban sewage pan-coronavirus monitoring method based on high-throughput sequencing

By combining calcium flocculation-citric acid enrichment with coronavirus RNA amplification and high-throughput sequencing, the problem of simultaneous monitoring of multiple genera/subgenera of coronaviruses in urban sewage has been solved, enabling low-cost and rapid coronavirus diversity analysis, which is suitable for public health monitoring and risk early warning.

CN122038652APending Publication Date: 2026-05-15NANHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANHUA UNIV
Filing Date
2026-02-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies for monitoring coronaviruses in urban sewage are difficult to simultaneously monitor and characterize the diversity of multiple genera/subgenera of coronaviruses, and are costly and time-consuming, making it difficult to meet the needs of large-scale, routine monitoring.

Method used

Viral particles were enriched using the calcium flocculation-citric acid method, and pan-coronavirus targeted amplification was performed by combining conserved regions of the coronavirus RNA-dependent RNA polymerase gene. Rapid classification and tracing of coronavirus sequences were achieved through one-step semi-nested reverse transcription PCR amplification and high-throughput sequencing, combined with bioinformatics analysis.

Benefits of technology

With low sequencing data volume, it enables rapid monitoring and localization of potential sources of multiple coronavirus lineages, reduces detection costs, improves the high throughput and reliability of monitoring, and fills the monitoring blind spots of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122038652A_ABST
    Figure CN122038652A_ABST
Patent Text Reader

Abstract

The invention discloses an urban sewage pan-coronavirus monitoring method based on high-throughput sequencing. The method comprises the following steps: step 1, sample pretreatment and virus enrichment; step 2, carrying out targeted amplification on the coronavirus; step 3, construction of an amplicon library and high-throughput sequencing; and step 4, bioinformatics analysis and diversity output. According to the method, the relative proportion of coronavirus sequences in a library is increased through targeting amplification and amplicon sequencing, and on the premise of meeting genus / subgenus classification and diversity analysis requirements, the sequencing data volume required by each sample can be remarkably reduced, so that the detection cost of a single sample is reduced, and the reliability of a detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pathogen detection, and more particularly to a method for monitoring pan-coronaviruses in urban sewage based on high-throughput sequencing. Background Technology

[0002] Coronaviridae are single-stranded positive-sense RNA viruses with high genetic diversity and mutation potential. They are widely distributed in humans and various animal hosts, posing a risk of cross-host transmission and triggering public health events. Urban wastewater can collect biological information from human and animal sources, making it an important window for conducting population-scale pathogen monitoring and risk early warning.

[0003] Current methods for monitoring coronaviruses in urban wastewater largely rely on targeted detection technologies such as real-time quantitative PCR (qPCR). These methods typically depend on pre-defined targets (e.g., specific sites of a known strain or a few known lineages). When faced with the broad diversity of coronavirus families, unknown mutations, or the introduction of zoonotic coronaviruses, they are prone to insufficient coverage and "monitoring blind spots," making it difficult to achieve simultaneous monitoring and diversity characterization of multiple genera / subgenera of coronaviruses within the same system.

[0004] Virome / metagenomic sequencing can theoretically reveal the diversity of viral lineages in a sample, but it often presents significant cost-effectiveness issues in wastewater samples: wastewater has a complex matrix, a high proportion of non-target nucleic acids (bacterial, bacteriophage, host, and environmental nucleic acids, etc.), while target coronaviruses are often found in low abundance. To obtain effective sequences for coronavirus classification and source tracing analysis, sequencing depth often needs to be significantly increased, resulting in high single-sample testing costs, heavy analytical workload, and long cycles, making it difficult to meet the practical needs of large-scale, routine monitoring for "low cost, speed, and high throughput."

[0005] Therefore, there is an urgent need for a pan-coronavirus monitoring method for complex urban sewage matrices: one that can rapidly obtain the sequence information of coronaviruses in samples and analyze their diversity and potential sources with a low amount of sequencing data (thus reducing costs), so as to be used for public health monitoring and risk warning. Summary of the Invention

[0006] The purpose of this invention is to provide a method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing. This invention can rapidly obtain the sequence information of coronaviruses in samples and analyze their diversity and potential sources with relatively low sequencing data volume, thereby enabling its use for public health monitoring and risk early warning.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows: A method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing includes the following steps: Step 1, Sample Pretreatment and Virus Enrichment: After pretreatment to remove large particulate impurities, the urban sewage samples were dissolved and enriched using calcium flocculation-citric acid CFCD to obtain a virus enrichment solution, and total nucleic acid was extracted. Step 2, Pancoronavirus Targeted Amplification: Using conserved regions of the coronavirus RNA-dependent RNA polymerase gene as amplification targets, amplification is performed using specific degenerate primers covering multiple genera / subgenera of coronaviruses; using the total nucleic acid as a template, one-step semi-nested reverse transcription PCR amplification is performed to obtain the target amplified fragment for sequencing. Step 3, Amplicon Library Construction and High-Throughput Sequencing: Purify, construct, index, and sequence prime target amplicon fragments to obtain amplicon sequencing data; Step 4: Bioinformatics Analysis and Diversity Output: The amplicon sequencing data is subjected to quality control, splicing / assembly, and sequence identification. Combined with alignment and phylogenetic analysis, the coronavirus sequences are classified, clustered, and traced to their origins, forming a coronavirus diversity spectrum in the sample.

[0008] In a further improvement, the total nucleic acid extraction method in step one is as follows: (1.1) Add CaCl2 and Na2HPO4 to the supernatant of the wastewater after removing large particulate impurities, so that CaCl2 and Na2HPO4 form calcium phosphate flocs under mixed conditions to adsorb virus particles; the final concentration of CaCl2 and Na2HPO4 is 1-50 mM; the reaction time is 5-30 min; the reaction pH is 6.5-9.0; (1.2) Centrifuge to collect the flocculated sediment and discard the supernatant; (1.3) Add citrate buffer to dissolve the flocculent precipitate to obtain a solution containing virus particles; the concentration of citrate buffer is 10-200 mM, pH 4.0-7.0; dissolve by blowing / vortexing until the precipitate is completely dissolved; (1.4) Filter the solution through a filter membrane with a pore size of 0.22 to 0.8 μm to reduce the background of large particles and obtain a virus enrichment solution; (5) Extract total nucleic acid from the virus enrichment solution.

[0009] In a further improvement, in step two, the specific degenerate primer combination includes primer pan-CoV_outF, primer pan-CoV_R, and primer Pan-CoV_InF; wherein the nucleotide sequence of primer pan-CoV_outF is 5'-CCAARTTYTAYGGHGGITGG-3'; the nucleotide sequence of primer pan-CoV_R is 5'-TGTTGIGARCARAAYTCATGIGG-3', and the nucleotide sequence of primer Pan-CoV_InF is 5'-GGTTGGGAYTAYCCHAARTGTGA-3'.

[0010] In a further improvement, step two, the one-step semi-nested reverse transcription PCR amplification, includes the following steps: (2.1) First round of one-step RT-PCR: Using the total nucleic acid as a template, reverse transcription and initial amplification were performed using pan-CoV_outF and pan-CoV_R to obtain the first round of products. The reaction system used one-step RT-PCR reagents; the annealing temperature was set to 45-60℃; the number of cycles was 30-45. (2.2) Second round of PCR: Using the first round product as a template, Pan-CoV_InF and pan-CoV_R are used for the second round of amplification to obtain the second round product; the number of cycles can be 20 to 35, and the annealing temperature is set to 45 to 60℃; (2.3) Detection of amplification products: Electrophoresis of the second-round products confirms the presence of the target length band; the second-round products are purified to remove primer dimers and non-specific fragments to obtain the target amplification fragment for sequencing.

[0011] Further improvements are made to step three, which includes the following steps: (3.1) Library construction: The target amplified fragments used for sequencing are repaired at the ends, A-tails are added, sequencing adapters are ligated, and sample indexes are added to form a sample library; (3.2) Multisample pooled sequencing: Multiple sample libraries are mixed in equal molar amounts and then sequenced to obtain amplicon sequencing data.

[0012] Further improvements include implementing low-cost constraints during sequencing: while ensuring classification and diversity output, the effective data volume per sample can be controlled within a preset threshold; the preset threshold is that the effective reads per sample are no higher than 1×10^7, or the effective base count per sample is no higher than 2Gbp.

[0013] A further improvement, step four includes the following steps: (4.1) Data quality control: Remove adapters, low-quality reads, and short sequences from amplicon sequencing data; and perform basic quality assessment; (4.2) Assembly: MEGAHIT or coronaSPAdes were used to assemble the filtered reads in a reference genome-free manner to obtain high-quality contigs / representative sequences for analysis; (4.3) Screening of candidate coronavirus sequences: Representative sequences are compared with coronavirus reference databases to obtain a set of RNA-dependent RNA polymerase RdRp related sequences as screening sequences; (4.4) Phylogenetic verification and classification: Perform multiple sequence alignment between the screened sequences and the reference sequences related to the α / β / γ / δ coronaviruses and construct a phylogenetic tree; set the number of bootstraps to evaluate the node support rate; output the classification results of the genus / subgenus or more subdivided levels; the number of bootstraps ≥ 500; (4.5) Diversity and source tracing output: Statistically count the number of sequences, relative proportions or read support of different taxa on a sample basis; combine phylogenetic localization to infer the potential host origin category, form data results for monitoring and early warning, and form the coronavirus diversity spectrum in the sample.

[0014] Further improvements include a "assembly / splicing correction + phylogenetic topological consistency verification" mechanism for low-abundance or potentially new lineage sequences: sequences must simultaneously meet the alignment screening threshold and the rationality of their phylogenetic position to be considered valid for inclusion in diversity statistics, thereby reducing misjudgments caused by environmental background noise.

[0015] Compared with the prior art, the present invention has at least the following effects (which can be quantified and verified in comparative examples): 1) Low cost: By increasing the relative proportion of coronavirus sequences in the library through "targeted amplification + amplicon sequencing", the amount of sequencing data required per sample can be significantly reduced while meeting the requirements for genus / subgenus classification and diversity analysis, thereby reducing the cost of single sample detection. 2) Rapid and high-throughput: Stable amplicon is obtained by one-step semi-nested RT-PCR, and combined with multi-sample indexing pooling of high-throughput sequencing, which enables rapid parallel monitoring of a large number of sewage samples. 3) Broad spectrum: It can obtain sequence information of multiple coronavirus lineages (including but not limited to α-coronavirus, β-coronavirus, γ-coronavirus, δ-coronavirus and other unassigned lineages) in a single sample, and can be used for phylogenetic localization and inference of potential host origin; 4) Reliability: Through a multi-confirmation mechanism of “assembly / splicing correction - comparison screening - phylogenetic topology verification”, the misjudgment caused by background noise of environmental samples and chimerism / contamination is reduced, and the credibility of classification and source tracing conclusions is improved. Attached Figure Description

[0016] Figure 1 This is a graph showing the electrophoresis results of a one-step semi-nested RT-PCR targeted amplification.

[0017] Figure 2 This section compares three virus enrichment methods. The enrichment results for (a) Influenza A virus (IAV) and (b) Coxsackievirus A6 (CVA6) are shown. The left figure shows the gene copy number (copies / µL) measured by different methods, and the right figure shows the corresponding virus recovery rate (%). Experiments were repeated three times (n = 3). Method abbreviations: Control, virus stock solution (unconcentrated); CFCD, calcium flocculation-citric acid dissolution method; AlCl3, aluminum salt coagulation precipitation method; PEG, polyethylene glycol precipitation method. An asterisk indicates statistically significant intergroup comparisons (* p < 0.05, ** p < 0.01, *** p < 0.001).

[0018] Figure 3 This is a graph for high-throughput monitoring and rapid reporting.

[0019] Figure 4 This is the clustering diagram of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0021] Example 1: Low-cost and rapid diversity monitoring of pancoronaviruses in urban wastewater samples 1) Sample collection and preprocessing (1) Collect sewage samples from Hengyang City, Hunan Province, preferably instantaneous or mixed samples from the same monitoring point; the sample volume can be 50 mL to 2 L (e.g., 200 mL to 1 L).

[0022] (2) The sample should be stored at 4℃ for a short time and processed as soon as possible; before processing, it can be centrifuged at low speed to remove large solid particles (e.g., 3000-8000 g, 5-20 min), and the supernatant can be used for subsequent enrichment.

[0023] 2) Virus enrichment (CFCD) and nucleic acid extraction (1) Add CaCl2 and Na2HPO4 to the supernatant to form calcium phosphate flocs under mixed conditions to adsorb virus particles. The final concentration of CaCl2 and Na2HPO4 can be set to any value or combination in the range of 1 to 50 mM; the reaction time can be 5 to 60 min; the reaction pH can be 6.5 to 9.0 (which can be adjusted according to the physicochemical properties of the wastewater).

[0024] (2) Centrifuge to collect the flocculated sediment (e.g., 3000-12000 g, 5-30 min), and discard the supernatant.

[0025] (3) Add citrate buffer to dissolve the precipitate to obtain a solution containing virus particles. The citrate buffer can be 10-200 mM, pH 4.0-7.0; dissolution can be carried out by blowing / vortexing until the precipitate is completely dissolved.

[0026] (4) Filter the solution through a filter membrane to reduce the background of large particles such as bacteria (e.g., 0.22-0.8 μm, preferably 0.45 μm) to obtain a virus enrichment solution.

[0027] (5) Use the TaKaRa MiniBEST Viral RNA / DNA Extraction Kit to extract total nucleic acid from the virus enrichment solution; optional process control (internal control) can be added to evaluate recovery rate and inhibition.

[0028] 3) Pancoronavirus one-step semi-nested RT-PCR targeted amplification (1) Target region: The conserved region in the coronavirus RdRp gene located between specific primer binding sites is selected as the amplification target to balance coverage and phylogenetic resolution. The length of the amplified fragment is preferably about 559-602 bp.

[0029] (2) Primer system: A specific degenerate primer combination was used, including at least the outer upstream primer Pan-CoV_InF, the inner upstream primer pan-CoV_outF, and the downstream primer pan-CoV_R; wherein Pan-CoV_InF and pan-CoV_outF are located in adjacent regions of the RdRp conserved region, and pan-CoV_R is a universal downstream primer; each primer has degenerate bases set at one or more sites to expand the coverage. The sequence of primer pan-CoV_outF is 5'-CCAARTTYTAYGGHGGITGG-3'; the sequence of primer pan-CoV_R is 5'-TGTTGIGARCARAAYTCATGIGG-3'; and the sequence of primer Pan-CoV_InF is 5'-GGTTGGGAYTAYCCHAARTGTGA-3'.

[0030] (3) First round of one-step RT-PCR: Using the nucleic acid obtained in step 2 as a template, reverse transcription and initial amplification were performed using pan-CoV_outF and pan-CoV_R. The reaction system can use TaKaRa one-step RT-PCR reagent; the annealing temperature can be set to 54.5℃; and the number of cycles can be 30.

[0031] (4) Second round PCR (semi-nested): Using the first round product as a template, the second round of amplification is performed using pan-CoV_outF and pan-CoV_R; the number of cycles can be 35.

[0032] (5) Detection of amplified products: Electrophoresis confirms the presence of a band of the target length (e.g., Figure 1 (As shown in the image) The amplification products were purified to remove primer dimers and non-specific fragments. To reduce the risk of contamination, a negative control and partitioned operation were preferably included, and a dUTP / UNG system was used as a contamination prevention measure.

[0033] 4) Amplicon library construction and high-throughput sequencing (1) Library construction: The purified amplification products are repaired at the ends, A tails are added, sequencing adapters are ligated and sample indexes are added.

[0034] (2) Multi-sample pooled sequencing: Multiple sample libraries are mixed in equal molar amounts and then sequenced; the sequencing platform is Illumina, etc.; the read length can be PE150.

[0035] (3) Low cost constraint: Under the premise of satisfying classification and diversity output, the effective data volume of a single sample can be controlled within a preset threshold, for example, the effective reads per sample are no more than 1×10^7, or the effective base number per sample is no more than 2 Gbp.

[0036] 5) Bioinformatics analysis and diversity output (1) Data quality control: remove connectors, remove low-quality reads, remove short sequences; and perform basic quality assessment.

[0037] (2) Assembly: The filtered reads were assembled without a reference genome using MEGAHIT or coronaSPAdes to obtain high-quality contigs / representative sequences for analysis.

[0038] (3) Screening of candidate coronavirus sequences: Compare and screen representative sequences with coronavirus reference databases (e.g., BLASTn / BLASTx or similar comparison tools) to obtain a set of RdRp related sequences.

[0039] (4) Phylogenetic verification and classification: Perform multiple sequence alignment between the selected sequences and α / β / γ / δ coronaviruses and other relevant reference sequences and construct a phylogenetic tree; set the number of bootstraps (e.g. ≥500 or 1000 times) to evaluate node support rate; output classification results at the genus / subgenus or more subdivided levels.

[0040] (5) Diversity and source tracing output: Statistically count the number of sequences, relative proportions or read support of different taxa on a sample basis; combine phylogenetic localization to infer the potential host origin category (such as human, rodent, bird or other mammal origin, etc.) to form data results that can be used for monitoring and early warning.

[0041] (6) Dual identification mechanism: For low abundance or potential new lineage sequences, the "assembly / splicing correction + phylogenetic topological consistency verification" mechanism is adopted: that is, the sequence must simultaneously meet the alignment screening threshold and the rationality of the phylogenetic position (and the consistency of repeated experiments when necessary) to be effectively detected and included in the diversity statistics, thereby reducing misjudgment caused by environmental background noise.

[0042] Example 2: Process adaptation under different sample volumes / inhibitor backgrounds Based on Example 1, to verify the advantages of the method of the present invention, influenza A virus (IAV, representing an enveloped virus) and Coxsackievirus A6 (CVA6, representing a non-enveloped virus) were selected as model viruses. Equal amounts of virus were added to wastewater samples, and enrichment and extraction were performed using the calcium flocculation-citric acid dissolution method (CFCD) of the present invention, the traditional aluminum chloride coagulation method (AlCl3), and the polyethylene glycol precipitation method (PEG), respectively. Untreated stock virus solution was used as a control. The results are as follows: Figure 2 As shown, the CFCD method significantly improved the recovery rate of IAV compared to the AlCl3 and PEG methods (p<0.05); the recovery rate of CVA6 was also superior to that of traditional methods. This indicates that the CFCD method not only has high recovery efficiency but also good universality for viruses with different structures.

[0043] Example 3: High-throughput monitoring and rapid reporting Samples collected from multiple monitoring sites in the same city within the same time window were processed in parallel using indexed pooled sequencing. Under a predetermined low data volume threshold, the coronavirus diversity profiles of each site were output, and trend charts or early warning thresholds were generated. For example... Figure 3 As shown, this method not only successfully detected multiple coronaviruses covering the four established genera α-, β-, δ-, and γ-CoV, verifying its excellent broad-spectrum coverage capability; more importantly, Figure 4 The highlighted red portion shows the specific sequences (such as those related to the bighead carp coronavirus HnCoV) detected in this study that cluster outside the four established genera mentioned above. These sequences form independent branches on the phylogenetic tree, strongly demonstrating that this invention has the ability to detect unclassified or potential coronaviruses in complex environmental samples, effectively filling the blind spots of existing monitoring technologies.

[0044] Comparative Example 1: Cost and detection capability comparison between shotgun virome / metagenomic sequencing and amplicon sequencing of this invention. 1) Sample: Wastewater sample from the same source and of the same volume as in Example 1, using the same nucleic acid extraction method.

[0045] 2) Method A (Comparative Example): Without RdRp targeted amplification, directly construct the shotgun sequencing library and run it on the machine; set the sequencing data volume as follows: D1: 1×10^7 valid reads / sample; D2: 4×10^7 valid reads / sample; D3: 1×10^8 valid reads / sample.

[0046] 3) Method B (Example): Perform RdRp targeted amplicon library construction and sequencing according to Example 1; set the sequencing data volume to 1.5×10^7 reads per sample.

[0047] 4) Evaluation indicators: (1) Effective detection rate of coronavirus (whether RdRp fragments that can be used for classification can be obtained). (2) The percentage of coronavirus reads or the target enrichment multiple; (3) Number of resolvable lineages (genea / subgenera / branch number) and diversity index (optional); (4) The minimum amount of sequencing data required to obtain "no less than N classifiable RdRp sequences / or cover no less than M taxonomic units"; (5) Cost per sample (which can be estimated based on reagent + sequencing costs) and reporting cycle.

[0048] 5) Comparative conclusion: Under the condition of achieving the same "number of classifiable sequences / number of lineages", the amount of sequencing data required by the present invention is significantly lower than that of the shotgun method, thereby significantly reducing the cost per sample and shortening the reporting cycle, as shown in Table 1.

[0049] Table 1. Comparison of Costs and Detection Capabilities of Different Methods

[0050] Comparative Example 2: Comparison of qPCR single-target detection and the diversity output capability of this invention 1) Sample: Same as in Example 1.

[0051] 2) Method A (Comparative Example): Detection was performed using a qPCR system targeting the N gene of the single coronavirus SARS-CoV-2, outputting only Ct values ​​or positive / negative results.

[0052] 3) Method B (Example): Perform pancoronavirus amplicon sequencing as described in Example 1 and output diversity profiles.

[0053] 4) Evaluation indicators: (1) Whether it can simultaneously identify multiple genera / subgenera of coronaviruses; (2) Whether it can output phylogenetic location and potential source category; (3) Coverage capability for unknown or non-preset target spectrum (whether there is a detection blind zone).

[0054] 5) Comparative conclusions: qPCR has insufficient coverage of lineages for non-preset targets and is difficult to output diversity profiles; the present invention can output multi-lineage information in a single sample and is suitable for broad-spectrum monitoring of unknown risks. The specific results are shown in Table 2.

[0055] Table 2 Comparison of the diverse output capabilities of different methods

[0056] Comparative Example 3: Comparison of sensitivity and stability between single-round RT-PCR and semi-nested RT-PCR in low-abundance wastewater samples 1) Samples: Same as in Example 1, with gradient A (stock solution), gradient B (1:10 dilution) and gradient C (1:100 dilution) selected.

[0057] 2) Method A (Comparative Example): Amplification and library construction were performed using only a single-round one-step RT-PCR (F1+R) method.

[0058] 3) Method B (Example): Semi-nested amplification (first round F1+R, second round F2+R) and library construction and sequencing were used.

[0059] 4) Evaluation indicators: amplification success rate, number of classifiable sequences, repeatability (consistency between different batches / repeated experiments), and false positive control (negative control contamination rate).

[0060] 5) Comparative conclusions: The semi-nested strategy has a higher amplification success rate and more stable diversity output under the background of low load and high inhibition in wastewater, reducing the cost waste caused by repeated detection and invalid sequencing, as shown in Table 3.

[0061] Table 3. Comparison of sensitivity and stability between single-round RT-PCR and semi-nested RT-PCR in low-abundance wastewater samples.

[0062] The embodiments described herein merely illustrate a few possible embodiments of this patent application, and their level of detail should not be construed as limiting the scope of the invention. Those skilled in the art will recognize that various modifications and alterations can be made without departing from the fundamental principles of the invention. Therefore, the scope of protection of this invention should be based on the claims.

Claims

1. A method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing, characterized in that, Includes the following steps: Step 1, Sample Pretreatment and Virus Enrichment: After pretreatment to remove large particulate impurities, the urban sewage samples were dissolved and enriched using calcium flocculation-citric acid CFCD to obtain a virus enrichment solution, and total nucleic acid was extracted. Step 2, Pancoronavirus Targeted Amplification: Using conserved regions of the coronavirus RNA-dependent RNA polymerase gene as amplification targets, amplification is performed using specific degenerate primers covering multiple genera / subgenera of coronaviruses; using the total nucleic acid as a template, one-step semi-nested reverse transcription PCR amplification is performed to obtain the target amplified fragment for sequencing. Step 3, Amplicon Library Construction and High-Throughput Sequencing: Purify, construct, index, and sequence prime target amplicon fragments to obtain amplicon sequencing data; Step 4: Bioinformatics Analysis and Diversity Output: The amplicon sequencing data is subjected to quality control, splicing / assembly, and sequence identification. Combined with alignment and phylogenetic analysis, the coronavirus sequences are classified, clustered, and traced to their origins, forming a coronavirus diversity spectrum in the sample.

2. The method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing as described in claim 1, characterized in that, In step one, the total nucleic acid extraction method is as follows: (1.1) Add CaCl2 and Na2HPO4 to the supernatant of the wastewater after removing large particulate impurities, so that CaCl2 and Na2HPO4 form calcium phosphate flocs under mixed conditions to adsorb virus particles; the final concentration of CaCl2 and Na2HPO4 is 1-50 mM; the reaction time is 5-30 min; the reaction pH is 6.5-9.0; (1.2) Centrifuge to collect the flocculated sediment and discard the supernatant; (1.3) Add citrate buffer to dissolve the flocculent precipitate to obtain a solution containing virus particles; the concentration of citrate buffer is 10-200 mM, pH 4.0-7.0; dissolve by blowing / vortexing until the precipitate is completely dissolved; (1.4) Filter the solution through a filter membrane with a pore size of 0.22 to 0.8 μm to reduce the background of large particles and obtain a virus enrichment solution; (5) Extract total nucleic acid from the virus enrichment solution.

3. The method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing as described in claim 1, characterized in that, In step two, the specific degenerate primer combination includes primer pan-CoV_outF, primer pan-CoV_R, and primer Pan-CoV_InF; wherein the nucleotide sequence of primer pan-CoV_outF is 5'-CCAARTTYTAYGGHGGITGG-3'; the nucleotide sequence of primer pan-CoV_R is 5'-TGTTGIGARCARAAYTCATGIGG-3', and the nucleotide sequence of primer Pan-CoV_InF is 5'-GGTTGGGAYTAYCCHAARTGTGA-3'.

4. The method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing as described in claim 1, characterized in that, In step two, the one-step semi-nested reverse transcription PCR amplification includes the following steps: (2.1) First round of one-step RT-PCR: Using the total nucleic acid as a template, reverse transcription and initial amplification were performed using pan-CoV_outF and pan-CoV_R to obtain the first round of products. The reaction system used one-step RT-PCR reagents; the annealing temperature was set to 45-60℃; the number of cycles was 30-45. (2.2) Second round of PCR: Using the first round product as a template, Pan-CoV_InF and pan-CoV_R are used for the second round of amplification to obtain the second round product; the number of cycles can be 20 to 35, and the annealing temperature is set to 45 to 60℃; (2.3) Detection of amplification products: Electrophoresis of the second-round products confirms the presence of the target length band; the second-round products are purified to remove primer dimers and non-specific fragments to obtain the target amplification fragment for sequencing.

5. The method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing as described in claim 1, characterized in that, The steps in step three are as follows: (3.1) Library construction: The target amplified fragments used for sequencing are repaired at the ends, A-tails are added, sequencing adapters are ligated, and sample indexes are added to form a sample library; (3.2) Multisample pooled sequencing: Multiple sample libraries are mixed in equal molar amounts and then sequenced to obtain amplicon sequencing data.

6. The method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing as described in claim 5, characterized in that, When performing sequencing, a low-cost constraint is imposed: under the premise of satisfying the classification and diversity output, the effective data volume of a single sample can be controlled within a preset threshold; the preset threshold is that the effective reads per sample are no more than 1×10^7, or the effective base number per sample is no more than 2 Gbp.

7. The method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing as described in claim 1, characterized in that, Step four includes the following steps: (4.1) Data quality control: Remove adapters, low-quality reads, and short sequences from amplicon sequencing data; and perform basic quality assessment; (4.2) Assembly: The filtered reads were assembled from scratch using MEGAHIT or coronaSPAdes without a reference genome to obtain high-quality contigs / representative sequences for analysis; (4.3) Screening of candidate coronavirus sequences: Representative sequences are compared with coronavirus reference databases to obtain a set of RNA-dependent RNA polymerase RdRp related sequences as screening sequences; (4.4) Phylogenetic verification and classification: Perform multiple sequence alignment between the screened sequences and the reference sequences related to the α / β / γ / δ coronaviruses and construct a phylogenetic tree; set the number of bootstraps to evaluate the node support rate; output the classification results of the genus / subgenus or more subdivided levels; the number of bootstraps ≥ 500; (4.5) Diversity and source tracing output: Statistically count the number of sequences, relative proportions or read support of different taxa on a sample basis; combine phylogenetic localization to infer the potential host origin category, form data results for monitoring and early warning, and form the coronavirus diversity spectrum in the sample.

8. The method for monitoring pan-coronaviruses in urban wastewater based on high-throughput sequencing as described in claim 7, characterized in that, For low-abundance or potentially new lineage sequences, an "assembly / splicing correction + phylogenetic topological consistency verification" mechanism is adopted: that is, the sequence must simultaneously meet the alignment screening threshold and the rationality of the phylogenetic position in order to be effectively detected and included in the diversity statistics, thereby reducing misjudgments caused by environmental background noise.