Method for determining regional viral evolutionary trends based on sewage sequencing and related devices

By combining wastewater sequencing and linear deconvolution models with infection distribution and neutralizing antibody data, the shortcomings of virus evolution rate and transmission risk assessment in wastewater epidemiology have been addressed, enabling accurate dynamic prediction and monitoring of virus evolution trends.

CN122090928APending Publication Date: 2026-05-26RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI
Filing Date
2025-12-30
Publication Date
2026-05-26

Smart Images

  • Figure CN122090928A_ABST
    Figure CN122090928A_ABST
Patent Text Reader

Abstract

This application provides a method and related equipment for determining regional viral evolution trends based on wastewater sequencing. The method includes determining the viral evolution rate based on multiple first target viral lineages and their abundance, as well as multiple second target viral lineages and their abundance; performing risk analysis on predetermined target viruses present in multiple first water samples based on multiple target viral mutation site information to obtain viral transmission risk; determining the change in the immune level of the population in the predetermined target area against the predetermined target viral lineage based on infection distribution data and neutralizing antibody titer data for each immune source to obtain the prediction result of the dominant lineage succession of the viral lineage over time; and using the viral evolution rate, viral transmission risk, and the prediction result of the dominant lineage succession of each viral lineage over time as the viral evolution trend information corresponding to the predetermined target area, thus solving the technical problem that the determination of regional viral evolution trends in the prior art is inaccurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and related equipment for determining the regional virus evolution trend based on wastewater sequencing. Background Technology

[0002] In recent years, wastewater-based epidemiology (WBE) has become an important approach for regional infectious disease surveillance. By enriching, sequencing, and analyzing viral nucleic acids in wastewater, a prospective and unbiased representative assessment of population infection status can be achieved. Metagenomic sequencing technology can directly resolve viral sequence information in environmental wastewater samples, thereby identifying viral types and their characteristic mutations within the region. This technology has been gradually applied to the surveillance of emerging infectious diseases and the screening of variant strains.

[0003] However, most existing studies remain at the level of static sequence detection and phylogenetic analysis, failing to describe the rate of viral evolution, adaptive enhancement, and lineage succession over time in real-world transmission environments. In actual epidemic transmission, viral lineage replacement is a complex and dynamic process driven by both herd immunity decline and immune escape advantage. Currently, wastewater virome data analysis, phylogenetic studies, and immune dynamics modeling remain fragmented, unable to comprehensively reflect the direction of viral evolution and the decisive role of immune changes in strain competition, thus failing to reflect the true evolutionary process of viruses. Summary of the Invention

[0004] In view of this, the purpose of this application is to propose a method and related equipment for determining the regional viral evolution trend based on sewage sequencing, so as to overcome all or part of the shortcomings of the prior art.

[0005] To achieve the above objectives, this application provides a method for determining regional viral evolution trends based on wastewater sequencing, comprising: acquiring multiple first water body samples corresponding to a predetermined target region within a predetermined time period; and acquiring multiple second water body samples for comparison with the multiple first water body samples; performing mutation detection on each of the multiple first water body samples and the multiple second water body samples using wastewater sequencing to obtain target virus mutation site information corresponding to the water body sample; determining the target virus lineage and lineage abundance corresponding to the water body sample based on the target virus mutation site information using a pre-constructed linear deconvolution model; and determining the predetermined target region based on the target virus lineage and lineage abundance corresponding to the multiple first water body samples, and the target virus lineage and lineage abundance corresponding to the multiple second water body samples. The study analyzes the virus evolution rate corresponding to the domain, and based on the mutation site information of multiple target viruses corresponding to the multiple first water samples, performs risk analysis on the predetermined target viruses present in the multiple first water samples to obtain the virus transmission risk corresponding to the predetermined target area. It also acquires infection distribution data of the predetermined target virus lineages present in the multiple first water samples over time and neutralizing antibody titer data for each immune source. Based on the infection distribution data and neutralizing antibody titer data for each immune source, it determines the change in the immune level of the population in the predetermined target area against the predetermined target virus lineage, and obtains the prediction result of the dominant lineage succession of the virus lineage over time. Finally, it uses the virus evolution rate, the virus transmission risk, and the prediction result of the dominant lineage succession of each virus lineage over time as the virus evolution trend information corresponding to the predetermined target area.

[0006] Optionally, determining the virus evolution rate corresponding to the predetermined target region based on the target virus lineages and lineage abundances corresponding to the plurality of first water body samples, and the target virus lineages and lineage abundances corresponding to the plurality of second water body samples, includes: determining the first average nucleotide genetic distance corresponding to the plurality of first water body samples using predetermined software based on the target virus lineages and lineage abundances corresponding to the plurality of first water body samples; determining the second average nucleotide genetic distance corresponding to the plurality of second water body samples using predetermined software based on the target virus lineages and lineage abundances corresponding to the plurality of second water body samples; and calculating the virus evolution rate corresponding to the predetermined target region based on the first average nucleotide genetic distance and the second average nucleotide genetic distance.

[0007] Optionally, the multiple target virus mutation site information includes multiple target virus mutation sites; the step of performing risk analysis on the predetermined target virus present in the multiple first water body samples based on the multiple target virus mutation site information corresponding to the multiple first water body samples to obtain the virus transmission risk corresponding to the predetermined target area includes: for each of the multiple target virus mutation sites, performing structural stability analysis on the target virus mutation site to obtain a first conclusion related to its structural stability; performing receptor binding energy analysis on the target virus mutation site to obtain a second conclusion related to its receptor binding energy; performing antibody binding energy analysis on the target virus mutation site to obtain a third conclusion related to its antibody binding energy; and determining the virus transmission risk based on all the first conclusions, all the second conclusions, and all the third conclusions.

[0008] Optionally, determining the change in the immune level of the population within the predetermined target area against the predetermined target viral lineage based on the infection distribution data and the neutralizing antibody titer data for each immune source, and obtaining the prediction result of the dominant lineage succession of the viral lineage over time, includes: standardizing the neutralizing antibody titer data for each immune source to obtain the target neutralizing antibody titer data for each immune source; determining the linear decay value of the neutralizing antibody titer over time based on multiple target neutralizing antibody titer data; determining the cross-protective efficacy of each immune source over time based on the linear decay value of the neutralizing antibody titer over time; determining the cross-protective efficacy of each immune source against the target viral lineage over time based on the cross-protective efficacy of each immune source over time and the infection distribution data; and determining the prediction result of the dominant lineage evolution of the target viral lineage over time based on the cross-protective efficacy corresponding to all immune sources over time.

[0009] Optionally, the target virus mutation site information includes multiple target virus mutation sites. The step of determining the target virus lineage and lineage abundance corresponding to the water sample using a pre-constructed linear deconvolution model based on the target virus mutation site information includes: determining a set of target virus variants containing at least one target virus variant corresponding to each target virus mutation site using a pre-constructed database of mutation sites and strain relationships based on the target virus mutation site information; constructing a mutation site matrix based on the set of target virus variants and multiple target virus mutation sites corresponding to each target virus variant; performing sequencing data statistics on multiple target virus mutation sites corresponding to the water sample to obtain the mutation allelic count and sequencing coverage depth corresponding to each target virus mutation site, and based on the mutation allelic count and sequencing coverage depth corresponding to each target virus mutation site... By varying allelic counts and sequencing coverage depth, the substitution frequency corresponding to each target virus mutation site is calculated. The substitution frequencies corresponding to each target virus mutation site are then filtered to obtain sample observation sequences containing multiple target substitution frequencies. Based on the observation sequences, the mutation site matrix, and the target virus variant set, a predetermined algorithm is used to solve the linear deconvolution model to obtain the initial target virus lineage and the initial lineage abundance corresponding to the initial target virus lineage for the water sample. Based on the initial lineage abundance corresponding to the initial target virus lineage, the target substitution frequency, and the mutation site matrix, a goodness of fit is calculated. In response to determining that the goodness of fit is greater than a predetermined goodness of fit, the initial target virus lineage and its corresponding initial lineage abundance are used as the target virus lineage and lineage abundance corresponding to the target virus mutation site information.

[0010] Optionally, the step of using wastewater sequencing to detect mutations in the water sample to obtain target virus mutation site information corresponding to the water sample includes: extracting viral nucleic acid from the water sample; performing multiplex polymerase chain reaction amplification on the extracted viral nucleic acid based on a predetermined target virus set to obtain viral nucleic acid to be detected; sequencing each viral nucleic acid to obtain the viral gene sequence corresponding to each predetermined target virus present in the water sample; performing quality improvement processing on each viral gene sequence; performing primer cutting, gene comparison, and mutation detection on each viral gene sequence after quality improvement processing to obtain the viral gene sequence and viral mutation site corresponding to each predetermined target virus; and using the viral gene sequence and viral mutation site corresponding to each predetermined target virus as the target virus mutation site information corresponding to the water sample.

[0011] Optionally, the quality improvement process for each viral gene sequence includes: for each viral gene sequence, removing a predetermined adapter sequence and a predetermined terminal sequence to obtain a first viral gene sequence; performing read filtering on the first viral gene sequence to obtain a second viral gene sequence; and performing gene knockout on the second viral gene sequence based on a predetermined common viral gene sequence to obtain a viral gene sequence that has undergone the quality improvement process.

[0012] Based on the same inventive concept, this application also provides a device for determining regional virus evolution trends based on wastewater sequencing, comprising: an acquisition module configured to acquire multiple first water body samples corresponding to a predetermined target area within a predetermined time period, and to acquire multiple second water body samples compared with the multiple first water body samples; a detection module configured to perform mutation detection on each of the multiple first water body samples and the multiple second water body samples using wastewater sequencing to obtain target virus mutation site information corresponding to the water body sample; a first determination module configured to determine the target virus lineage and lineage abundance corresponding to the water body sample based on the target virus mutation site information using a pre-constructed linear deconvolution model; and a second determination module configured to determine the target virus lineage and lineage abundance corresponding to the multiple first water body samples, and the target virus lineage and lineage abundance corresponding to the multiple second water body samples. The system comprises: a first target area corresponding to the virus evolution rate; an analysis module configured to perform risk analysis on the predetermined target viruses present in the multiple first water samples based on the mutation site information of multiple target viruses corresponding to the multiple first water samples, thereby obtaining the virus transmission risk corresponding to the predetermined target area; a second determination module configured to acquire infection distribution data of the predetermined target virus lineages present in the multiple first water samples over time and neutralizing antibody titer data of each immune source; based on the infection distribution data and neutralizing antibody titer data of each immune source, determine the change in the immune level of the population in the predetermined target area against the predetermined target virus lineages, thereby obtaining the prediction result of the dominant lineage succession of the virus lineage over time; and a third determination module configured to use the virus evolution rate, the virus transmission risk, and the prediction result of the dominant lineage succession of each virus lineage over time as the virus evolution trend information corresponding to the predetermined target area.

[0013] Based on the same inventive concept, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0014] Based on the same inventive concept, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the method described above.

[0015] As can be seen from the above, the method and related equipment for determining regional virus evolution trends based on wastewater sequencing provided in this application include: acquiring multiple first water body samples corresponding to a predetermined target area within a predetermined time period; acquiring multiple second water body samples for comparison with the multiple first water body samples; performing mutation detection on each of the multiple first water body samples and the multiple second water body samples using wastewater sequencing to obtain target virus mutation site information corresponding to the water body sample; determining the target virus lineage and lineage abundance corresponding to the water body sample using a pre-constructed linear deconvolution model based on the target virus mutation site information; and determining the target virus lineage and lineage abundance corresponding to the multiple first water body samples and the target virus lineage and lineage abundance corresponding to the multiple second water body samples based on the target virus lineage and lineage abundance of the multiple first water body samples. The study analyzes the virus evolution rate corresponding to a predetermined target area; based on the mutation site information of multiple target viruses corresponding to multiple first water samples, it performs risk analysis on the predetermined target viruses present in the multiple first water samples to obtain the virus transmission risk corresponding to the predetermined target area; it acquires infection distribution data of the predetermined target virus lineages present in the multiple first water samples over time and neutralizing antibody titer data for each immune source; based on the infection distribution data and neutralizing antibody titer data for each immune source, it determines the change in the immune level of the population in the predetermined target area against the predetermined target virus lineages to obtain the prediction result of the dominant lineage succession of the virus lineages over time; and it uses the virus evolution rate, the virus transmission risk, and the prediction result of the dominant lineage succession of each virus lineage over time as the virus evolution trend information corresponding to the predetermined target area. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the method for determining regional virus evolution trends based on wastewater sequencing, as described in this application. Figure 2 This is a schematic diagram of the structure of the wastewater sequencing-based regional virus evolution trend determination device according to an embodiment of this application; Figure 3This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0020] As described in the background section, wastewater epidemiology has become an important approach for regional infectious disease surveillance in recent years. By processing viral nucleic acids in wastewater, a prospective and unbiased representative assessment of population infection status can be achieved. Metagenomic sequencing technology can directly analyze viral sequence information in environmental wastewater samples, thereby identifying viral types and their characteristic mutations within a region. This technology has been gradually applied to the surveillance of emerging infectious diseases and the screening of variant strains. However, most existing studies remain at the level of static sequence detection. Static sequence detection mainly involves a one-time or periodic determination of the viral nucleic acid sequence in wastewater samples, followed by comparison with known viral sequences to determine the viral type and mutation status. While this provides some valuable information, it cannot quantify the rate of viral evolution or determine the transmission risk of the virus, making it difficult to reflect the true evolutionary process of the virus.

[0021] In addition to the above, existing research also struggles to describe the adaptive enhancement and lineage succession patterns of viruses over time in real-world transmission environments. In actual epidemic transmission, viral lineage replacement is a complex dynamic process driven by both herd immunity decline and immune escape advantage. Currently, wastewater virome data analysis, phylogenetic studies, and immune dynamics modeling remain fragmented, failing to comprehensively reflect the direction of viral evolution and the decisive role of immune changes in strain competition. Existing technologies obtain viral sequences through wastewater sampling, viral nucleic acid extraction, and high-throughput sequencing. Combining Fastp quality control, Bowtie2 / BWA alignment, and Samtools / iVar mutation detection can identify characteristic mutation sites, determining the existence of specific variant strains and enabling quantitative analysis of static viral lineages and visualization of mutation hotspots in wastewater.

[0022] In summary, existing research has the following shortcomings, making it difficult to reflect the true evolutionary process of viruses: 1) Insufficient detection sensitivity and spectral resolution Low viral concentrations and strong host background interference in wastewater make conventional sequencing methods sensitive to sample quality, leading to the easy omission of low-abundance viruses or emerging strains. At the same time, for viruses with high genetic diversity and complex lineage structures, existing methods are unable to accurately identify and quantify the coexistence of multiple lineages and changes in relative abundance in a single analysis, resulting in insufficient analysis of cross-host transmission and mixed infections.

[0023] 2) Its functionality is limited to static detection and lacks the ability to analyze dynamic evolution. Existing analyses mostly focus on mutation statistics and lineage classification, failing to quantify viral evolution rates, adaptive enhancement trends, or lineage replacement mechanisms, and thus failing to reflect the true evolutionary process of viral populations.

[0024] 3) Lack of a transmission risk assessment framework based on mutation characteristics Quantitative correlation indicators between viral mutation sites and immune escape potential and enhanced transmissibility have not yet been established, making it impossible to predictively identify and warn of potentially high-risk strains.

[0025] 4) Dynamic modeling of immune decline and herd immunity levels has not been achieved. Current immune dynamics models mainly rely on clinical data and cannot utilize wastewater metagenomic signals to invert regional immune changes; they also lack quantitative analysis of the driving role of immune decline in strain succession.

[0026] 5) The system is fragmented and has poor scalability. Virus detection, evolutionary analysis, and immune modeling remain independent processes, lacking a unified interface and automated workflow, making it difficult to support real-time updates, regional assessments, and cross-virus system promotion and application.

[0027] In view of this, embodiments of this application propose a method for determining regional virus evolution trends based on wastewater sequencing, referring to... Figure 1 This includes the following steps: Step 101: Obtain multiple first water body samples corresponding to a predetermined target area within a predetermined time period, and obtain multiple second water body samples for comparison with the multiple first water body samples.

[0028] In this step, existing wastewater epidemiology studies struggle to describe the evolutionary trends of viruses in real-world transmission environments. To accurately study wastewater epidemiology, multiple first water body samples are obtained within a predetermined time period corresponding to a predetermined target area. The predetermined target area is the study area containing wastewater, determined based on actual needs. Multiple second water body samples are obtained for comparison with the first water body samples. These second water body samples can be water samples collected within the predetermined target area in a previous predetermined time period, which is adjacent to the predetermined time period in which the first water body samples were collected. Alternatively, the second water body samples can be water samples collected in other predetermined areas within the predetermined time period, where the predetermined target area and the other predetermined areas are different. For example, the first and second water body samples may represent samples from different ecological niches, or samples from different predetermined administrative regions, such as provinces, cities, or districts (counties). This application collects water samples from real-world transmission environments, laying a data foundation for accurately analyzing the virus evolution trends corresponding to the predetermined target area.

[0029] Step 102: For each of the plurality of first water samples and the plurality of second water samples, perform mutation detection on the water sample using wastewater sequencing to obtain the target virus mutation site information corresponding to the water sample.

[0030] In this step, the water sample refers to a collection of liquids and their constituent substances collected from the aquatic environment. The water sample contains multiple viral / microbial subtypes or contaminants from different sources. To utilize the water sample to study viral evolution trends, it is necessary to determine whether there are mutation sites in the water sample related to a predetermined target virus. The predetermined target virus is determined according to predetermined requirements; for example, it is a virus with a hazard level greater than a predetermined hazard level. First, mutation detection is performed on the water sample to obtain the target virus mutation site information corresponding to the water sample. This target virus mutation site information consists of multiple target virus mutation sites present in the water sample. These target virus mutation sites are the specific locations in the viral genome where base substitutions, insertions, deletions, or structural rearrangements have occurred, achieving the goal of comprehensively determining the target virus mutation site information corresponding to the water sample.

[0031] Step 103: Based on the target virus mutation site information, determine the target virus lineage and lineage abundance corresponding to the water sample using a pre-constructed linear deconvolution model.

[0032] In this step, the target virus mutation site information reflects the specific location of mutation in the viral genome. By analyzing the target virus mutation site information, different viral variants can be distinguished. To accurately determine the target virus lineage and lineage abundance corresponding to the water sample, this application utilizes a pre-constructed linear deconvolution model to analyze the target virus mutation site information. The linear deconvolution model in this application is a linear model. Solving the linear deconvolution model based on the target virus mutation site information can effectively separate overlapping viral lines, accurately quantify the relative proportion of each viral variant, avoid the cumbersome steps of designing primers for each viral variant as required by traditional methods (such as quantitative real-time PCR), and reduce errors caused by primer cross-reaction. Furthermore, based on the accurately quantified relative proportion of each viral variant, the lineage abundance corresponding to the water sample can be accurately determined. The target virus lineage includes multiple viral variants, and the lineage abundance is used to describe the relative content of each viral variant in the mixed viral variants.

[0033] This application achieves the identification and relative abundance quantification of multi-lineage viruses in water samples by using a linear deconvolution model combined with target virus mutation site information. It breaks through the dependence of traditional targeted detection modes on high viral load and single-lineage scenarios, and can accurately reflect low abundance, new occurrence and multi-lineage coexistence states, thereby significantly improving the sensitivity and representativeness of environmental virus monitoring.

[0034] Step 104: Based on the target virus lineages and lineage abundances corresponding to the plurality of first water samples, and the target virus lineages and lineage abundances corresponding to the plurality of second water samples, determine the virus evolution rate corresponding to the predetermined target region.

[0035] In this step, lineage abundance reflects the distribution proportion of each viral variant in the mixed viral variants. Since the viral variants in the viral variant set are constantly evolving, static analysis of only the first water body sample corresponding to the predetermined target area cannot reveal the evolutionary trend of the viral variants contained in the first water body sample. Therefore, this application determines the viral evolution rate corresponding to the predetermined target area by processing the target viral lineages and lineage abundances corresponding to multiple first water body samples and multiple second water body samples. Target viral lineages and lineage abundances can reflect the genetic diversity of the viral population. If the abundance of the same viral variant changes significantly in the two water body samples, it indicates that its gene frequency is changing rapidly in the population. The first and second water body samples have temporal or spatial heterogeneity; a single sample can only provide a "static snapshot," while comparing two samples can capture dynamic processes. By dynamically changing the lineage and abundance of target viruses, the genetic structure of the viral population can be reflected in time or space, thereby inferring its evolution rate and dynamically characterizing the viral evolution trend and lineage succession mechanism in the predetermined target area, so as to accurately determine the viral evolution rate corresponding to the predetermined target area.

[0036] Step 105: Based on the multiple target virus mutation site information corresponding to the multiple first water body samples, perform risk analysis on the predetermined target viruses present in the multiple first water body samples to obtain the virus transmission risk corresponding to the predetermined target area.

[0037] In this step, the viral evolution rate primarily reflects the speed of change in the viral gene sequence over time, demonstrating the alteration of its genetic material. However, using only the viral evolution rate as information on viral evolutionary trends is insufficient to fully represent these trends. For example, some viruses, despite their slow evolution rate, may possess high transmissibility; relying solely on the evolution rate could underestimate their potential threat. Therefore, it is necessary to quantify the transmission risk posed by the target virus. Based on the mutation site information of multiple target viruses corresponding to multiple first-body water samples, a risk analysis is performed on the target viruses present in these samples to obtain the viral transmission risk corresponding to the target area. By quantifying the viral transmission risk, the transmission risk faced by the target area can be accurately identified.

[0038] Step 106: Obtain infection distribution data of the predetermined target virus lineage and neutralizing antibody titer data of each immune source in the plurality of first water samples over time; based on the infection distribution data and neutralizing antibody titer data of each immune source, determine the change in the immune level of the population in the predetermined target area against the predetermined target virus lineage, and obtain the prediction result of the dominant lineage succession of the virus lineage over time.

[0039] In this step, infection distribution data of predetermined target viral lineages and neutralizing antibody titers for each immune source are acquired over time in multiple first-body water samples. The infection distribution data visually reflects the activity level, prevalence, and transmission trends of different target viral lineages within a specific area over different time periods. Analyzing this data reveals which target viral lineages are currently dominant, their transmission speed, and the existence of potential transmission hotspots. The neutralizing antibody titers for each immune source reveal the body's defense capabilities against different target viral lineages from the perspective of population immunity. The titers of neutralizing antibodies produced through different immunization routes, such as natural infection and vaccination, directly relate to population susceptibility to the virus. High titers of neutralizing antibodies generally indicate stronger immune protection, reducing the risk of infection and mitigating disease severity. Immune sources include vaccination programs, previously infected variants, and breakthrough infections.

[0040] Based on infection distribution data and neutralizing antibody titer data for each source of immunity, the changes in the immune level of the population within a predetermined target area against a predetermined target viral lineage are determined, yielding a prediction of the dominant lineage succession over time. Infection distribution data reflects the spread of the virus in the population, while neutralizing antibody titer data reflects the population's immunity to the virus. Combining these two data points quantifies the dynamic relationship between population immunity and virus transmission, providing a prediction of the dominant lineage succession over time. Dynamically determining the predicted dominant lineage succession over time ensures the accuracy of this prediction.

[0041] Among them, the prediction of the dominant lineage succession over time is a forecast of the changing trend of the relative competitive advantage of different target viral lineages within a predetermined target area. Simply put, it predicts which target viral lineages(s) will become the dominant viral lineages in the spread and prevalence of the target area at different points in the future.

[0042] Step 107: The virus evolution rate, the virus transmission risk, and the prediction results of the dominant lineage succession of each virus lineage over time are used as the virus evolution trend information corresponding to the predetermined target area.

[0043] In this step, since the prediction results of virus evolution rate, virus transmission risk and the dominant lineage succession of each virus lineage over time are all accurate, the prediction results of virus evolution rate, virus transmission risk and the dominant lineage succession of each virus lineage over time are used as the virus evolution trend information corresponding to the predetermined target area, so as to achieve the purpose of accurately determining the virus evolution trend information corresponding to the predetermined target area.

[0044] The above scheme obtains multiple first water body samples corresponding to a predetermined target area within a predetermined time period, and multiple second water body samples compared with the multiple first water body samples; for each of the multiple first water body samples and the multiple second water body samples, mutation detection is performed on the water body sample using wastewater sequencing to obtain the target virus mutation site information corresponding to the water body sample; based on the target virus mutation site information, the target virus lineage and lineage abundance corresponding to the water body sample are determined using a pre-constructed linear deconvolution model; based on the target virus lineage and lineage abundance corresponding to the multiple first water body samples, and the target virus lineage and lineage abundance corresponding to the multiple second water body samples, the virus evolution rate corresponding to the predetermined target area is determined; based on the... Information on multiple target virus mutation sites corresponding to multiple first water body samples is used to conduct risk analysis on the predetermined target viruses present in the multiple first water body samples to obtain the virus transmission risk corresponding to the predetermined target area; infection distribution data of the predetermined target virus lineages present in the multiple first water body samples over time and neutralizing antibody titer data for each immune source are obtained; based on the infection distribution data and neutralizing antibody titer data for each immune source, the change in the immune level of the population in the predetermined target area against the predetermined target virus lineage is determined to obtain the prediction result of the dominant lineage succession of the virus lineage over time; the virus evolution rate, the virus transmission risk, and the prediction result of the dominant lineage succession of each virus lineage over time are used as the virus evolution trend information corresponding to the predetermined target area.

[0045] In some embodiments, determining the virus evolution rate corresponding to the predetermined target region based on the target virus lineages and lineage abundances corresponding to the plurality of first water body samples, and the target virus lineages and lineage abundances corresponding to the plurality of second water body samples, includes: determining a first average nucleotide genetic distance corresponding to the plurality of first water body samples using predetermined software based on the target virus lineages and lineage abundances corresponding to the plurality of first water body samples; determining a second average nucleotide genetic distance corresponding to the plurality of second water body samples using predetermined software based on the target virus lineages and lineage abundances corresponding to the plurality of second water body samples; and calculating the virus evolution rate corresponding to the predetermined target region based on the first average nucleotide genetic distance and the second average nucleotide genetic distance. In this embodiment, based on the target virus lineage and lineage abundance corresponding to the first water body sample, a first average nucleotide genetic distance corresponding to the first water body sample is determined using predetermined software. Similarly, based on the target virus lineage and lineage abundance corresponding to multiple second water body samples, a second average nucleotide genetic distance corresponding to multiple second water body samples is determined using the same predetermined software. The predetermined software is MEGA X (Molecular Evolutionary Genetics Analysis X). The time axis is divided into sliding or discrete windows (weekly / monthly / quarterly). Within each window t, based on the target virus lineage and lineage abundance corresponding to the first water body sample, the first average nucleotide genetic distance corresponding to the first water body sample is calculated using MEGA X. Likewise, based on the target virus lineage and lineage abundance corresponding to the second water body samples, the second average nucleotide genetic distance corresponding to the second water body samples is calculated using MEGA X.

[0046] The viral evolution rate corresponding to a predetermined target region is calculated based on the first and second average nucleotide genetic distances. Viral evolution is the result of the accumulation of genetic variation within a population, and the target viral lineage and lineage abundance reflect the competition and transmission dynamics of viral variants. Analyzing only the genetic distance of a single viral variant ignores the evolutionary patterns at the population level. Calculating the average genetic distance using the target viral lineage and lineage abundance allows for a more comprehensive capture of the overall degree of variation in the viral population. Nucleotide genetic distance directly reflects the degree of variation in the viral genome. By comparing the average genetic distances of different samples, the evolutionary rate of the viral population can be inferred from the differences in genetic distance sequences, quantifying the evolutionary speed of the viral population in time or space, thus achieving the goal of accurately determining the overall evolutionary rate of the viral population. The viral evolution rate corresponding to the predetermined target region is determined using the following formula: Formula 1, where, For the rate of viral evolution, The first average nucleotide genetic distance within window t. denoted as the second average nucleotide genetic distance within window t.

[0047] When R(t) > 0, it indicates that the virus in this window is evolving at an accelerated rate relative to the reference window; when R(t) < 0, it indicates that the virus in this window is evolving at a decelerated rate relative to the reference window. The evolution rate can eliminate the influence of sequencing depth, sample size, etc. on absolute distance, making it comparable between different time periods or cities.

[0048] In some embodiments, the multiple target virus mutation site information includes multiple target virus mutation sites; the step of performing risk analysis on predetermined target viruses present in the multiple first water body samples based on the multiple target virus mutation site information corresponding to the multiple first water body samples to obtain the virus transmission risk corresponding to the predetermined target area includes: for each of the multiple target virus mutation sites, performing structural stability analysis on the target virus mutation site to obtain a first conclusion related to its structural stability; performing receptor binding energy analysis on the target virus mutation site to obtain a second conclusion related to its receptor binding energy; performing antibody binding energy analysis on the target virus mutation site to obtain a third conclusion related to its antibody binding energy; and determining the virus transmission risk based on all the first conclusions, all the second conclusions, and all the third conclusions. In this embodiment, this application constructs a quantitative mapping relationship between viral mutation characteristics and the strength of transmission risk, thereby supporting the early identification and warning of potentially high-risk mutations. Based on the representative (or abundance-weighted consensus) target protein structure and the reference contrast protein structure of each target viral mutation site, changes in structural stability, receptor binding energy, and antibody binding energy are quantitatively calculated, thereby obtaining a strain-level transmission risk assessment. The reference contrast protein structure is the protein structure of the initial site (i.e., the pre-mutation site) corresponding to the target viral mutation site.

[0049] Structural stability analysis was performed on the target viral mutation sites to obtain the first conclusions related to their structural stability. Specifically, the target protein structure and the contrast protein structure corresponding to the target viral mutation sites were obtained. Based on these structures, the free energy corresponding to the target viral mutation sites was calculated. The protein structures were obtained from PDB (Protein Structure Database) or AlphaFold2 (used to predict the three-dimensional structure of proteins from amino acid sequences). The free energy corresponding to the target viral mutation sites was determined using the following formula: ΔG = G _Mut -G _WT Formula 2, where G _Mut G represents the free energy corresponding to the target protein structure. _WT To compare the free energies corresponding to the protein structures, FoldX was used to calculate the two types of free energies. When the free energy corresponding to the target viral mutation site is greater than zero, the first conclusion is structural instability; when the free energy corresponding to the target viral mutation site is less than zero, the first conclusion is structural stabilization.

[0050] Receptor binding energy analysis was performed on the target viral mutation sites to obtain a second conclusion related to their receptor binding energies. Specifically, the receptor binding energies of the target protein structure corresponding to the target viral mutation site and the predetermined host receptor were determined, as well as the receptor binding energies of the comparison protein structure and the predetermined host receptor were determined. The target protein structure and the comparison protein structure respectively form complexes with the predetermined host receptor (such as ACE2, sialic acid receptor, etc.), and the above two receptor binding energies are determined using Rosetta or FoldX (software tools for protein structure prediction, design, and energy calculation). The receptor binding energy corresponding to the target viral mutation site is determined using the following formula: E rec =E Mut-rec -E WT-rec Formula 3, where E Mut-rec E represents the receptor binding energy corresponding to the target protein structure. WT-rec To compare the receptor binding energies corresponding to the protein structures, the second conclusion is that when the receptor binding energy corresponding to the target viral mutation site is greater than zero, the receptor affinity decreases, resulting in a lower potential invasion efficiency; conversely, when the receptor binding energy corresponding to the target viral mutation site is less than zero, the second conclusion is that the receptor affinity increases, resulting in a higher potential invasion efficiency.

[0051] Antibody binding energy analysis was performed on the target viral mutation sites to obtain a third conclusion related to their antibody binding energies. Specifically, the antibody binding energies of the target protein structure corresponding to the target viral mutation site and the predetermined set of neutralizing antibodies were determined, as well as the antibody binding energies of the comparison protein structure and the predetermined set of neutralizing antibodies were determined. For each antibody a∈A in the predetermined set of neutralizing antibodies, the above two antibody binding energies were determined using FoldX AnalyseComplex (the core functional module in FoldX software used to calculate the interaction energies between two molecules or molecular populations). The antibody binding energy corresponding to the target viral mutation site was determined using the following formula: Formula 4, where... This represents the antibody binding energy corresponding to the target protein structure. To compare the antibody binding energies corresponding to protein structures, A represents a predetermined set of neutralizing antibodies, A={a1,…,aM}, whose structure is derived from PDB or AlphaFold2. When the antibody binding energy corresponding to the viral variant is greater than zero, the third conclusion is that overall antibody binding capacity decreases, and immune escape is enhanced; when the antibody binding energy corresponding to the viral variant is less than zero, the third conclusion is that antibody binding is enhanced, and immune pressure increases.

[0052] Based on the first, second, and third conclusions corresponding to all target viral mutation sites, the viral transmission risk is quantified to obtain the viral transmission risk corresponding to the predetermined target region. By using indicators such as changes in protein structural stability, receptor affinity, and antibody binding energy, a quantitative correlation between mutation characteristics and transmissibility is established, supporting the early identification and warning of potentially high-risk strains, and achieving the goal of quantitatively assessing the transmission risk and immune breakthrough potential of variant strains.

[0053] In some embodiments, determining the change in the immune level of the population within the predetermined target area against a predetermined target viral lineage based on the infection distribution data and neutralizing antibody titer data for each immune source, and obtaining a prediction result of the dominant lineage succession of the viral lineage over time, includes: standardizing the neutralizing antibody titer data for each immune source to obtain target neutralizing antibody titer data for each immune source; determining the linear decay value of the neutralizing antibody titer over time based on multiple target neutralizing antibody titer data; determining the cross-protective efficacy of each immune source over time based on the linear decay value of the neutralizing antibody titer over time; determining the cross-protective efficacy of each immune source against the target viral lineage over time based on the cross-protective efficacy of each immune source over time and the infection distribution data; and determining the prediction result of the dominant lineage evolution of the target viral lineage over time based on the cross-protective efficacy corresponding to all immune sources over time. In this embodiment, to quantify the synergistic driving of dominant lineage iteration by immune decline and viral mutation, this application, based on the herd immunity dynamics framework of neutralizing antibody titers, assesses the ability of target viral variants to overcome the herd immunity barrier by quantitatively describing changes in immunity acquisition, antibody decline, and cross-protection levels, and further drives a lineage competition model to predict the replacement process of dominant variants. This application proposes coupling the herd immunity decline pattern with lineage competition dynamics to construct a quantitative prediction model for variant replacement driven by changes in immune levels. It conducts predictions of prevalent strains based on wastewater viromes, quantifies the synergistic driving of dominant lineage iteration by immune decline and viral mutation, reveals and predicts that the lineage replacement process driven by immune decline can explain the mechanism of dominant lineage formation and provide predictions of future replacement trends. This analysis elucidates the driving role of herd immunity decline in strain replacement, providing forward-looking support for public health strategy formulation.

[0054] Neutralizing antibody titer data for each immune source were standardized to obtain target neutralizing antibody titer data for each immune source. Due to significant differences in the absolute values ​​of antibody titers across different experimental systems, the average titer from recovered patients was used as a unified reference for logarithmic standardization. An initial neutralization level matrix of viral variant i × immune type k in the target viral lineage was constructed: Formula 5, where, This represents the neutralizing titer value of viral variant i under immune type k. This represents the average titer value among recovered patients.

[0055] Following immunization, neutralizing antibody levels decrease exponentially over time, and this decline can be approximated as linear in the logarithmic titer space. Therefore, based on multiple target neutralizing antibody titer data, the linear decline value of neutralizing antibody titers over time is determined. The decline is calculated using the following formula: Formula 6, where, For viral variant i at time intervals The linear decay value within, Let be the initial neutralization level matrix for virus variant i. For the decay time interval, The characteristic timescale of antibody decline for different immune types (empirical rule: severe cases > mild cases > symptomatic cases > asymptomatic cases; differences also exist among different vaccine platforms).

[0056] Based on the linear decay value of neutralizing antibody titers over time, the cross-protective efficacy of each immune source over time was determined. According to the empirical mapping relationship between neutralizing titer and protective efficacy, the decayed titers were converted into cross-protective efficacy. Formula 7, where, To achieve the titer threshold required for 50% infection protection, The slope parameter (typically ≈3, calibrated according to regional immunization efficacy) enables a biological interpretation from immunization intensity (titer) to infection protection capability.

[0057] Based on the time-varying cross-protective efficacy and infection distribution data of each immune source, the time-varying cross-protective efficacy of each immune source against the target viral lineage was determined. To integrate the immune contributions from different immune sources at different time points, the temporal distribution of immune events was convolved with cross-immune decay to obtain the time function of population k-level immunity against viral variant i, where... The effective immune barrier strength of population k against viral variant i at time t is characterized, i.e., the time-varying cross-protective efficacy of each immune source against the target viral lineage. Daily discrete integration is used in the engineering to generate a full-time trajectory, thereby revealing the mechanism by which immune decline and antigen escape jointly drive lineage replacement. The immune barrier strength is determined by the following formula: Formula 8, where the temporal distribution of immune events is based on... This indicates that inferences can be made based on case surveillance data, vaccination schedules, or questionnaires. If necessary, a kernel density estimation function can be used for smooth fitting, where H represents the Hill function. For viral variant i at time intervals The antibody decay value within the body. Represented as .

[0058] Based on the cross-protective efficacy over time corresponding to all sources of immunity, the predicted evolution of the dominant lineage of the target virus lineage over time was determined. The "breakability" of the immune barrier against viral variant i was defined as: Formula 9, where, The strength of the immune barrier for viral variant i at time point t can be determined using formula eight.

[0059] Setting the baseline replication advantage of viral variants The time-varying fitness is obtained as follows: Formula 10, where β is the immune pressure weight. Let the proportion of mutant strains be... (satisfy Lineage replacement follows reproducer dynamics: Formula 11, where, This represents the rate at which the proportion of viral variant i changes over time. The time fitness of viral variant i at time point t. Let represent the proportion of variant i at time point t, and let j represent the index of all other competing variants in the set of variants. This represents the percentage of the j-th other viral variant in the set of viral variants. Let be the time fitness corresponding to the j-th other viral variant. The predicted evolution of the dominant lineage of the target virus lineage over time is determined using Formula 11. If the proportion of a certain viral variant increases rapidly in a short period, it indicates that its transmission efficiency is significantly higher than other viral variants, possessing strong viral transmissibility. Based on the rate of change of the proportion of viral variants over time, the circulating viral variant is identified among all viral variants, where the transmissibility of the circulating viral variant is greater than the predetermined viral transmissibility.

[0060] It can simulate the timing of the emergence and replacement rate of dominant lineages under different immune decline scenarios, and can fit the number of parameters using historical data. , , , , To obtain prediction curves and goodness of fit, a model linking regional immune decline and relative fitness was constructed to simulate the timing of the emergence and replacement rate of epidemic variants, providing a scientific basis for intervention strategies and the allocation of public health resources.

[0061] In some embodiments, the target virus mutation site information includes multiple target virus mutation sites. The step of determining the target virus lineage and lineage abundance corresponding to the water sample using a pre-constructed linear deconvolution model based on the target virus mutation site information includes: determining a set of target virus variants containing at least one target virus variant corresponding to each target virus mutation site using a pre-constructed database of mutation sites and strain relationships based on the target virus mutation site information; constructing a mutation site matrix based on the set of target virus variants and multiple target virus mutation sites corresponding to each target virus variant; performing sequencing data statistics on multiple target virus mutation sites corresponding to the water sample to obtain the mutation allelic count and sequencing coverage depth corresponding to each target virus mutation site, and based on the mutation allelic count and sequencing coverage depth corresponding to each target virus mutation site... The mutation allelic count and sequencing coverage depth are used to calculate the substitution frequency corresponding to each target virus mutation site. The substitution frequencies corresponding to each target virus mutation site are then filtered to obtain sample observation sequences containing multiple target substitution frequencies. Based on the observation sequences, the mutation site matrix, and the target virus variant set, a predetermined algorithm is used to solve the linear deconvolution model to obtain the initial target virus lineage and the initial lineage abundance corresponding to the initial target virus lineage for the water sample. Based on the initial lineage abundance corresponding to the initial target virus lineage, the target substitution frequency, and the mutation site matrix, a goodness of fit is calculated. In response to determining that the goodness of fit is greater than a predetermined goodness of fit, the initial target virus lineage and its corresponding initial lineage abundance are used as the target virus lineage and lineage abundance corresponding to the target virus mutation site information. In this embodiment, the target virus mutation site information includes multiple target virus mutation sites. The accumulation of one or more mutation sites leads to significant changes in the viral genome, thereby forming a viral variant. For each mutation site, a search is performed in a pre-constructed mutation site and strain relationship database to find the corresponding viral variant, resulting in a viral variant set containing at least one viral variant corresponding to each mutation site. The mutation site and strain relationship database is obtained by associating and storing each historical mutation site corresponding to a predetermined administrative region to which the predetermined target region belongs, along with the corresponding viral variants. It should be noted that, to improve processing efficiency, duplicate viral variants in the viral variant set are deduplicated. The viral variant set can be represented as... By utilizing a database of mutation points and strain relationships corresponding to a predetermined target area, the set of viral variants can be determined, thereby achieving the goal of comprehensively identifying the viral variants corresponding to the predetermined target area.

[0062] A mutation site matrix is ​​constructed based on a set of viral variants and the corresponding mutation points of each variant, representing the set of viral variants and each variant in matrix form. Specifically, an empty matrix is ​​pre-constructed, and the set of viral variants and the corresponding mutation points of each variant are encoded according to a predetermined encoding rule. Each element in the mutation site matrix can be a binary value or a more complex code to represent the specific type or degree of mutation. The encoded set of viral variants and the corresponding mutation points of each variant are then filled into the empty matrix to obtain the mutation site matrix. For example, an empty two-dimensional matrix is ​​constructed, where rows represent different viral variants and columns represent different mutation point positions or types. The encoded set of viral variants and the corresponding mutation points of each variant are then filled into the empty two-dimensional matrix. The mutation points corresponding to the set of viral variants can be represented as follows: ,in, Representing mutation points, the matrix of mutation sites Here, p represents the total number of viral variants. The mutation site matrix, also known as the mutation barcode matrix, clearly presents the distribution and combination characteristics of mutation points in different variants through its row and column structure. It quantifies the set of viral variants and the corresponding mutation points for each variant. The mutation site matrix integrates the mutation point data of all viral variants in matrix form, ensuring no crucial information is omitted. Furthermore, the standardized structure of the mutation site matrix eliminates analytical obstacles caused by format differences.

[0063] Sequencing data from multiple target viral mutation sites corresponding to water samples were statistically analyzed to obtain the mutation allelic count and sequencing coverage depth for each target viral mutation site. The mutation allelic count is the number of sequencing reads of mutant alleles (such as base substitutions, insertions, or deletions) detected at a specific genomic site; the sequencing coverage depth is the number of times each base is sequenced. Based on the mutation allelic count and sequencing coverage depth for each target viral mutation site, the substitution frequency for each target viral mutation site was calculated. A mutation site specifically refers to an actual change in genetic material occurring at a genomic site, and the substitution frequency refers to the proportion of a particular mutation form (such as base substitution or amino acid mutation) occurring in a population or sample at the genomic site corresponding to the mutation site. It reflects the dynamic changes of the mutation site during propagation. The substitution frequency for each target viral mutation site was calculated using the following formula: Formula 12, where, Indicates the target viral mutation site The substitution frequency, Target viral mutation sites mutation allelic counts, Target viral mutation sites sequencing coverage depth.

[0064] Replacement frequency fluctuates less in the short term and reflects long-term trends. By using replacement frequency, the dynamic changes of target virus mutation sites during the transmission process can be accurately quantified, so that replacement frequency can serve as an objective indicator and avoid the interference of subjective judgment on the assessment of target virus lineage and lineage abundance.

[0065] Using predetermined mutation information, the substitution frequency corresponding to each target viral mutation site is screened to obtain observed sequences containing multiple target substitution frequencies. The predetermined mutation information includes sequencing depth, base quality, and mutation frequency. Sequencing depth is the average number of times each base in the genome is sequenced, i.e., the ratio of total sequencing data (number of bases) to genome size. Base quality refers to the probability that each base is correctly identified during sequencing. Mutation frequency refers to the proportion of mutant alleles (such as base substitutions, insertions, or deletions) appearing in the sequencing reads at a specific genomic site. Predetermined mutation information reflects the reliability of sequencing; therefore, using the predetermined mutation information corresponding to each target viral mutation site, the substitution frequency corresponding to each target viral mutation site is screened. Specifically, each sub-information in the predetermined mutation information is compared with its corresponding threshold. In response to determining that each sub-information is greater than or equal to its corresponding threshold, the substitution frequency of the target viral mutation site corresponding to the mutation information is determined as the target substitution frequency. For example, the threshold corresponding to sequencing depth is 10, the threshold corresponding to base quality is 30, and the threshold corresponding to mutation frequency is 1%. By screening all substitution frequencies, the high quality of the target substitution frequency was ensured.

[0066] The observation sequence is composed of all target substitution frequencies, and the observation sequence can be represented as follows: Based on the observed sequences, mutation site matrix, and target virus variant set, a predetermined algorithm is used to solve the linear deconvolution model to obtain the initial target virus lineage and its corresponding initial lineage abundance. For example, the predetermined algorithm can be a non-negative least squares (NNLS) algorithm, a sparse constraint regression algorithm (Lasso (minimum absolute shrinkage and selection operator regression) or Elastic Net (elastic network regression)), variational inference or the EM algorithm (Expectation-Maximization Algorithm), or a Bayesian mixture model algorithm for solving the abundance distribution and confidence interval. Samples are mixed proportionally from different lineages to establish the linear deconvolution model. ,in, For the observation sequence, This is a matrix of mutation sites. This represents the proportion of viral variants among all viral variants. This represents the error term. A predetermined algorithm is used to solve the problem, and the solution is then subjected to minimum normalization to obtain the initial spectral abundance. The initial target virus lineage corresponding to the initial lineage abundance is determined using the following formula: Formula Thirteen.

[0067] The observed sequence, composed of all target substitution frequencies, reflects the genetic characteristics observed at each observation site. The proportion of viral variants among all viral variants is an unknown variable, representing the abundance information of different lineages. The mutation site matrix characterizes the features of different viral variants at each mutation site. By establishing a linear deconvolution model, the observed data are explicitly linked to the unknown lineage abundance and error term. Lineage abundance itself has a non-negative physical meaning. The NNLS algorithm is used to solve the linear deconvolution model. The NNLS algorithm can find the solution that minimizes the error between the observed sequence and the model prediction, under the constraint that the proportion of each variant cannot be negative. The sum of the proportions of all viral variants must be 1. The solution obtained by the NNLS algorithm may not strictly satisfy this constraint, so normalization is required. After normalization, the initial lineage abundance truly reflects the actual proportion of viral variants and accurately represents the relative abundance of different lineages in the viral population. The viral genome contains a large number of mutation sites, and the changes at these sites are interconnected and complex. The linear deconvolution model in this application can link these complex mutation information with the abundance of different variants. Taking into account the intrinsic relationship between the mutation site matrix and viral variants, it can accurately determine the initial lineage abundance and its corresponding initial target viral lineage.

[0068] Signals from low-abundance viral variants are easily masked by those from high-abundance variants, making accurate identification difficult using traditional analytical methods. The linear deconvolution model in this application treats the observed mixed signal as a linear combination of signals from each independent lineage. This model effectively separates overlapping viral lineages. Therefore, the linear deconvolution model in this application enhances the resolution of low-abundance lineages, enabling accurate identification and relative abundance estimation of multiple lineages coexisting in environmental samples, and providing real-time quantitative data on virus transmission levels.

[0069] After determining the initial lineage abundance, the accuracy of the initial lineage abundance needs to be verified. Based on each target substitution frequency, the initial lineage abundance, and the mutation site matrix, the residual of the target viral mutation site corresponding to each target substitution frequency is calculated, and the residual is determined by the following formula: Formula Fourteen, where, Indicates the target viral mutation site corresponding to the target substitution frequency. The residuals. Based on the residuals of the target viral mutation sites corresponding to each target substitution frequency and each target substitution frequency, the goodness of fit is calculated and determined by the following formula: Formula 15, where, The average of all target substitution frequencies. Target viral mutation sites The corresponding target substitution frequency. If the goodness of fit is greater than the predetermined goodness of fit, it indicates that the initial lineage abundance is accurate. The initial lineage abundance and its corresponding initial target virus lineage are determined as the lineage abundance and its corresponding target virus lineage, ensuring that the lineage abundance and target virus lineage are accurate.

[0070] It should be noted that when the goodness of fit is less than or equal to the predetermined goodness of fit, the target viral mutation sites are resampled using bootstrap (a statistical method based on resampling) to repeatedly estimate the initial lineage abundance. We obtained a 95% confidence interval for the initial phylogenetic abundance.

[0071] In some embodiments, the step of using wastewater sequencing to detect mutations in the water sample to obtain target virus mutation site information corresponding to the water sample includes: extracting viral nucleic acid from the water sample; performing multiplex polymerase chain reaction amplification on the extracted viral nucleic acid based on a predetermined target virus set to obtain viral nucleic acid to be detected; sequencing each viral nucleic acid to obtain the viral gene sequence corresponding to each predetermined target virus present in the water sample; performing quality improvement processing on each viral gene sequence; performing primer cutting, gene comparison, and mutation detection on each viral gene sequence after quality improvement processing to obtain the viral gene sequence and viral mutation site corresponding to each predetermined target virus; and using the viral gene sequence and viral mutation site corresponding to each predetermined target virus as the target virus mutation site information corresponding to the water sample. In this embodiment, viral nucleic acid is extracted from water samples. Exemplarily, commercial nucleic acid extraction kits are used to obtain viral RNA or DNA. To improve the accuracy of viral nucleic acid extraction from water samples, the water samples are pre-enriched before extraction. Exemplarily, methods such as polyethylene glycol (PEG) precipitation, aluminum salt coagulation, ultrafiltration, and membrane concentration are used to pre-enrich viral particles in the water samples. For the commonly observed low viral load in water samples, to further improve whole-genome coverage, a multiplex PCR (Polymerase Chain Reaction) amplification system targeting the target virus is constructed on the template after nucleic acid extraction. When the viral load is less than a predetermined load, viral load amplification is performed. Specifically, primer design is based on reference genomes of multiple representative lineages of the target virus. Conserved regions are screened as primer binding sites through multiple sequence alignment, and a tiling amplification strategy is used to cover key regions of the genome. For primer sites with variations, degenerate primers or backup primers can be designed to ensure lineage compatibility of the amplification.

[0072] Based on a predetermined target virus set, the extracted viral nucleic acids are amplified using multiplex polymerase chain reaction (PCR) to obtain the viral nucleic acids to be detected. First, the target virus species to be detected are determined, forming a predetermined target virus set. Then, viral nucleic acids are extracted from the sample, and multiplex polymerase chain reaction (PCR) technology is used. Multiple pairs of primers specific to different target viral nucleic acids are simultaneously added to the same reaction system. These primers guide the nucleic acid polymerase to amplify the extracted viral nucleic acids in large quantities, ultimately obtaining a sufficient amount of viral nucleic acids for subsequent detection and analysis. Each viral nucleic acid to be detected is sequenced to obtain the viral gene sequence corresponding to each predetermined target virus present in the water sample. After purification and library construction, the viral nucleic acids to be detected are sequenced using a high-throughput sequencing platform. For example, a high-throughput sequencing platform (Illumina, MGI, or Nanopore) is used for sequencing to generate viral gene sequences in FASTQ format (a standard file format for storing raw sequencing data). To ensure the effectiveness of subsequent processing, each viral gene sequence undergoes quality improvement processing. For example, quality control is performed using Fastp software (a preprocessing tool designed specifically for sequencing data in FASTQ format).

[0073] For each improved viral gene sequence, primer cutting, gene comparison, and mutation detection are performed to obtain the viral gene sequence and viral mutation sites corresponding to each predetermined target virus present in the water sample. The Qreads (improved viral gene sequences) are aligned to the target virus reference genome (Covid, Influenza, RSV, or other viral genomes) using Minimap2 or Bowtie2 (sequence alignment tools), generating a BAM file (binary alignment map file). Primer sequence cutting is then performed using iVar (a toolkit specifically for virus sequencing data processing) based on the BED (a text format) primer file.

[0074] Using a pre-defined toolkit, the coverage of each virus gene sequence after quality enhancement was statistically analyzed. For example, the pre-defined toolkit was Samtools, a key tool in bioinformatics for analyzing viral gene sequence coverage, used to process SAM and BAM files associated with high-throughput sequencing data. The goal was to achieve a coverage of ≥10× for at least 50% of sites. Using Samtools' mpileup and iVar functions to call mutations, high-confidence SNPs / Indels were screened based on depth >10, quality >30, and allelic frequency ≥1%, to obtain the viral gene sequence and viral mutation sites of each predetermined target virus present in the water sample. This ensured the accuracy of the obtained viral gene sequences and viral mutation sites corresponding to each predetermined target virus present in the water sample.

[0075] It should be noted that multiple viral mutation sites belong to metagenomic data. Metagenomic data refers to the massive genetic information obtained by high-throughput sequencing of all microbial genomic DNA extracted directly from environmental samples (such as soil, water, intestines, air, etc.). Metagenomic data is used to identify viral strains in a predetermined target area and quantify the infection status.

[0076] In some embodiments, the quality improvement process for each viral gene sequence includes: for each viral gene sequence, removing a predetermined adapter sequence and a predetermined terminal sequence to obtain a first viral gene sequence; performing read filtering on the first viral gene sequence to obtain a second viral gene sequence; and performing gene knockout on the second viral gene sequence based on a predetermined common viral gene sequence to obtain a viral gene sequence that has undergone the quality improvement process. In this embodiment, to fix the DNA fragment onto the sequencing chip for sequencing, specific adapter sequences are typically added to both ends of the DNA fragment. These adapter sequences are not part of the viral gene sequence itself. To ensure accurate subsequent processing of the viral gene sequence, predetermined adapter sequences are removed from the viral gene sequence. The predetermined end sequences may contain non-specific sequences or low-quality sequence portions generated during sequencing, which can interfere with the determination of the true viral gene sequence. Therefore, predetermined end sequences are removed from the viral gene sequence. It should be noted that the predetermined adapter sequences and predetermined end sequences in this application are determined based on historical experience. Furthermore, other predetermined data in this application are also determined based on historical experience and will not be elaborated further. For example, the predetermined end sequence is a segment with an average quality value less than a predetermined quality (e.g., Q30).

[0077] To remove short reads from the first viral gene sequence, read filtering was performed to obtain the second viral gene sequence. Reads in the first viral gene sequence with base pair lengths shorter than a predetermined base pair length were removed, for example, 50 bp. The second viral sequence carries the following sequence information: Q20 / Q30 distribution, GC content, and read length statistics. The Q20 / Q30 distribution represents the proportion of reads with different quality levels in the sequence; the GC content refers to the proportion of guanine (G) and cytosine (C) in the sequence; the read length statistics report provides detailed information about read lengths, including the average length, median length, maximum length, minimum length, and the distribution of the number of reads in different length intervals. The second viral sequence is then comprehensively evaluated based on this sequence information.

[0078] During sample collection, processing, and sequencing, common viral gene sequences or other non-target viral gene sequences may be introduced, causing sample contamination. To address this issue, a second viral gene sequence is removed based on a predetermined common viral gene sequence, resulting in a quality-enhanced viral gene sequence. To avoid host contamination interference, Bowtie2 (a tool specifically designed for aligning sequencing reads with long reference genome sequences) is used to align the reads to human / common bacterial reference genomes and remove matching reads, resulting in a set of valid reads (Qreads) for viral analysis—the quality-enhanced viral gene sequence. By performing quality enhancement on each viral gene sequence, the quality of each viral gene sequence is improved.

[0079] In another embodiment provided in this application, the mutation rate ratio corresponding to the viral variant strain can also be determined. Using a predetermined viral gene sequence corresponding to the viral variant strain, mutation statistics are performed on the viral gene sequence corresponding to the viral variant strain to obtain the number of non-synonymous mutations and the number of synonymous mutations corresponding to the viral variant strain. Specifically, using the predetermined viral gene sequence, the viral gene sequence corresponding to the viral variant strain is divided into functional regions to obtain multiple functional regions. For example, multiple functional regions include spike protein S, nucleocapsid protein N, and non-structural proteins nsp1–nsp16, etc. Multiple functional regions can be represented as... ,in, This represents the total number of functional regions. For each functional region, the target mutation point within that region is mapped to its corresponding codon, and the mutation type is distinguished to obtain the number of nonsynonymous mutations and synonymous mutations corresponding to that functional region. The target mutation point is defined as a mutation with a confidence level greater than a predetermined confidence level. The number of nonsynonymous mutations (resulting in amino acid changes) corresponding to the viral variant is counted. And, count the number of synonymous mutations (that do not change amino acids): .

[0080] Based on the number of non-synonymous mutations and the number of synonymous mutations corresponding to the viral variant, the mutation rate ratio corresponding to the viral variant is calculated and determined using the following formula: Formula sixteen, where, The mutation rate of non-synonymous mutations. The mutation rate is the mutation rate of synonymous mutations.

[0081] when >1. The segment is determined to be in a positive selection. <1 indicates a purification selection. ≈1 represents a neutral evolution. This calculation is repeated over a sliding time window to form a temporal dN / dS trajectory, which is used to identify when and where to choose to increase pressure.

[0082] This application uses a sliding time window to calculate genetic distance and dN / dS, continuously tracking changes in viral evolution rate and positive selection hotspots to reveal the time nodes and evolutionary driving forces behind the formation of dominant lineages, and dynamically analyzing viral evolutionary trends and adaptive enhancement mechanisms. Furthermore, this application dynamically calculates changes in genetic distance and dN / dS indices of the viral genome using a sliding time window, revealing changes in viral evolution rate, selective pressure regions, and adaptive enhancement trends during transmission.

[0083] In another example provided in this application, a phylogenetic tree can be constructed using high-quality SNPs to quantitatively characterize the rate of viral evolution at different stages.

[0084] First, using the high-confidence SNP data (depth > 10×, quality > 30, mutation frequency > 1%) obtained from the preceding module, the vcf 2 phylip tool (used to convert VCF (Variant Call Format) genomic variation data to PHYLIP format) was used to convert VCF (the standard file format for storing genomic variation information) to PHYLIP format (an input file format accepted by commonly used phylogenetic analysis software) as input for SNP-based phylogenetic analysis. MAFFT (multiple sequence alignment software) was used to perform multiple sequence alignment (MSA) on the whole genome sequence, employing L-INS-i or FFT-NS-i modes suitable for viral hypervariable regions (these two modes are alignment strategies designed by MAFFT for different sequence characteristics). Then, trimAl (a tool for removing low-confidence fragments) was used to remove low-confidence fragments from the alignment matrix, obtaining a clean alignment matrix with stable quality. Based on this matrix, a maximum likelihood (ML) phylogenetic tree is constructed using IQ-TREE (a phylogenetic tree construction software), and the optimal alternative model is determined through the built-in model selector (ModelFinder) to ensure the reliability of the topology and branch length. The outputs of the above steps include: high-quality SNP alignment (high-quality SNP alignment results), purified MSA (purified multiple sequence alignment results), and ML phylogenetic tree.

[0085] In another embodiment provided in this application, this application uses urban wastewater metagenomic sequencing data as its core and constructs a comprehensive assessment system for regional viral infection and evolution trends through a four-stage technical link: "strain identification and quantification—evolutionary rate characterization—mutation identification and risk assessment—herd immunity dynamics modeling," which is used to implement the method provided in this application. This application belongs to the interdisciplinary technical field of wastewater epidemiology, bioinformatics, and infectious disease dynamics modeling, specifically involving a joint assessment system for viral evolution monitoring and immune dynamics based on urban wastewater metagenomic sequencing. This system integrates functions such as environmental virome data acquisition and analysis, lineage characteristic mutation identification, evolutionary rate calculation, herd immunity decline modeling, and variant strain replacement prediction, and can comprehensively quantify and assess regional viral evolution trends, immune decline effects, and strain lineage replacement processes. Relying on the broad coverage and forward-looking advantages of wastewater epidemiology, this system can achieve early warning and evolutionary monitoring of infectious disease epidemic risks without relying on clinical sampling, providing a scientific basis for public health decision-making. This application establishes an integrated technical system encompassing multi-lineage identification and quantification, evolution rate characterization and risk assessment, and immune-driven lineage substitution simulation, enabling wastewater virus monitoring to leap from "static existence judgment" to "dynamic evolution interpretation and trend prediction." The scope of protection for this application is not limited to specific algorithmic tools or kinetic equations; any joint assessment of virus evolution and immune-driven processes based on wastewater metagenomic data should be considered within the scope of protection of this application.

[0086] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.

[0087] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0088] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides a device for determining the regional virus evolution trend based on wastewater sequencing.

[0089] refer to Figure 2The device for determining regional virus evolution trends based on wastewater sequencing includes: The acquisition module 10 is configured to acquire multiple first water body samples corresponding to a predetermined target area within a predetermined time period, and to acquire multiple second water body samples compared with the multiple first water body samples.

[0090] The detection module 20 is configured to perform mutation detection on each of the plurality of first water samples and the plurality of second water samples using wastewater sequencing to obtain the target virus mutation site information corresponding to the water sample.

[0091] The first determining module 30 is configured to determine the target virus lineage and lineage abundance corresponding to the water sample based on the target virus mutation site information and using a pre-constructed linear deconvolution model.

[0092] The second determining module 40 is configured to determine the virus evolution rate corresponding to the predetermined target region based on the target virus lineage and lineage abundance corresponding to the plurality of first water samples, and the target virus lineage and lineage abundance corresponding to the plurality of second water samples.

[0093] The analysis module 50 is configured to perform risk analysis on the predetermined target virus present in the plurality of first water samples based on the multiple target virus mutation site information corresponding to the plurality of first water samples, and obtain the virus transmission risk corresponding to the predetermined target area.

[0094] The third determining module 60 is configured to acquire infection distribution data of a predetermined target virus lineage present in the plurality of first water samples over time and neutralizing antibody titer data of each immune source; based on the infection distribution data and neutralizing antibody titer data of each immune source, determine the change in the immune level of the population in the predetermined target area against the predetermined target virus lineage, and obtain the prediction result of the dominant lineage succession of the virus lineage over time.

[0095] The fourth determining module 70 is configured to use the virus evolution rate, the virus transmission risk, and the prediction results of the dominant lineage succession of each virus lineage over time as the virus evolution trend information corresponding to the predetermined target area.

[0096] Using the aforementioned apparatus, multiple first water body samples corresponding to a predetermined target area within a predetermined time period are acquired, as well as multiple second water body samples compared with the multiple first water body samples; for each of the multiple first water body samples and the multiple second water body samples, mutation detection is performed on the water body sample using wastewater sequencing to obtain target virus mutation site information corresponding to the water body sample; based on the target virus mutation site information, the target virus lineage and lineage abundance corresponding to the water body sample are determined using a pre-constructed linear deconvolution model; based on the target virus lineage and lineage abundance corresponding to the multiple first water body samples, and the target virus lineage and lineage abundance corresponding to the multiple second water body samples, the virus evolution rate corresponding to the predetermined target area is determined; based on the... Information on multiple target virus mutation sites corresponding to multiple first water body samples is used to conduct risk analysis on the predetermined target viruses present in the multiple first water body samples to obtain the virus transmission risk corresponding to the predetermined target area; infection distribution data of the predetermined target virus lineages present in the multiple first water body samples over time and neutralizing antibody titer data for each immune source are obtained; based on the infection distribution data and neutralizing antibody titer data for each immune source, the change in the immune level of the population in the predetermined target area against the predetermined target virus lineage is determined to obtain the prediction result of the dominant lineage succession of the virus lineage over time; the virus evolution rate, the virus transmission risk, and the prediction result of the dominant lineage succession of each virus lineage over time are used as the virus evolution trend information corresponding to the predetermined target area.

[0097] In some embodiments, the second determining module 40 is further configured to: determine a first average nucleotide genetic distance corresponding to the plurality of first water samples based on the target virus lineage and lineage abundance corresponding to the plurality of first water samples using predetermined software; determine a second average nucleotide genetic distance corresponding to the plurality of second water samples based on the target virus lineage and lineage abundance corresponding to the plurality of second water samples using predetermined software; and calculate the virus evolution rate corresponding to the predetermined target region based on the first average nucleotide genetic distance and the second average nucleotide genetic distance.

[0098] In some embodiments, the plurality of target viral mutation site information includes a plurality of target viral mutation sites; the analysis module 50 is further configured to perform structural stability analysis on each of the plurality of target viral mutation sites to obtain a first conclusion related to the structural stability of the target viral mutation site; perform receptor binding energy analysis on the target viral mutation site to obtain a second conclusion related to the receptor binding energy of the target viral mutation site; perform antibody binding energy analysis on the target viral mutation site to obtain a third conclusion related to the antibody binding energy of the target viral mutation site; and determine the virus transmission risk based on all the first conclusions, all the second conclusions, and all the third conclusions.

[0099] In some embodiments, the third determining module 60 is further configured to: standardize the neutralizing antibody titer data for each immune source to obtain target neutralizing antibody titer data for each immune source; determine the linear decay value of the neutralizing antibody titer over time based on multiple target neutralizing antibody titer data; determine the cross-protective efficacy of each immune source over time based on the linear decay value of the neutralizing antibody titer over time; determine the cross-protective efficacy of each immune source against the target viral lineage over time based on the cross-protective efficacy of each immune source over time and the infection distribution data; and determine the predicted result of the dominant lineage evolution of the target viral lineage over time based on the cross-protective efficacy corresponding to all immune sources over time.

[0100] In some embodiments, the first determining module 30 is further configured to: determine a set of target virus variants containing at least one target virus variant corresponding to each target virus variant based on the target virus variant site information and using a pre-constructed database of variant sites and strain relationships; construct a mutation site matrix based on the set of target virus variants and multiple target virus variant sites corresponding to each target virus variant; perform sequencing data statistics on multiple target virus variant sites corresponding to the water sample to obtain the mutation allelic count and sequencing coverage depth corresponding to each target virus variant site; and calculate the substitution frequency corresponding to each target virus variant site based on the mutation allelic count and sequencing coverage depth corresponding to each target virus variant site. The substitution frequency corresponding to each target virus mutation site is screened to obtain a sample observation sequence containing multiple target substitution frequencies; based on the observation sequence, the mutation site matrix, and the target virus variant set, the linear deconvolution model is solved using a predetermined algorithm to obtain the initial target virus lineage and the initial lineage abundance corresponding to the initial target virus lineage for the water sample; based on the initial lineage abundance corresponding to the initial target virus lineage, the target substitution frequency, and the mutation site matrix, the goodness of fit is calculated; in response to determining that the goodness of fit is greater than a predetermined goodness of fit, the initial target virus lineage and its corresponding initial lineage abundance are used as the target virus lineage and lineage abundance corresponding to the target virus mutation site information.

[0101] In some embodiments, the detection module 20 is further configured to extract viral nucleic acids from the water sample, amplify the extracted viral nucleic acids using multiple polymerase chain reaction based on a predetermined target virus set to obtain viral nucleic acids to be detected; sequence each viral nucleic acid to be detected to obtain the viral gene sequence corresponding to each predetermined target virus present in the water sample; perform quality improvement processing on each viral gene sequence; perform primer cutting, gene comparison, and mutation detection on each viral gene sequence after quality improvement processing to obtain the viral gene sequence and viral mutation site corresponding to each predetermined target virus; and use the viral gene sequence and viral mutation site corresponding to each predetermined target virus as the target virus mutation site information corresponding to the water sample.

[0102] In some embodiments, the detection module 20 is further configured to remove a predetermined adapter sequence and a predetermined terminal sequence from each viral gene sequence to obtain a first viral gene sequence; perform read filtering on the first viral gene sequence to obtain a second viral gene sequence; and perform gene knockout on the second viral gene sequence based on a predetermined common viral gene sequence to obtain a viral gene sequence that has undergone the quality improvement process.

[0103] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.

[0104] The apparatus described above is used to implement the corresponding wastewater sequencing-based regional virus evolution trend determination method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0105] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method for determining the regional virus evolution trend based on wastewater sequencing as described in any of the above embodiments.

[0106] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0107] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0108] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0109] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0110] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0111] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0112] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0113] The electronic devices described above are used to implement the corresponding wastewater sequencing-based regional virus evolution trend determination method in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0114] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the wastewater sequencing-based method for determining regional virus evolution trends as described in any of the above embodiments.

[0115] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0116] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the regional virus evolution trend determination method based on sewage sequencing as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0117] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to execute the wastewater sequencing-based regional virus evolution trend determination method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0118] It should be noted that the embodiments of this application can also be further described in the following ways: It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0119] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0120] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0121] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0122] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0123] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0124] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0125] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A method for determining regional viral evolution trends based on wastewater sequencing, characterized in that, include: Acquire multiple first water body samples corresponding to a predetermined target area within a predetermined time period, and acquire multiple second water body samples compared with the multiple first water body samples; For each of the plurality of first water samples and the plurality of second water samples, sewage sequencing is used to detect mutations in the water sample to obtain the target virus mutation site information corresponding to the water sample. Based on the target virus mutation site information, the target virus lineage and lineage abundance corresponding to the water sample are determined using a pre-constructed linear deconvolution model. Based on the target virus lineage and lineage abundance corresponding to the plurality of first water body samples, and the target virus lineage and lineage abundance corresponding to the plurality of second water body samples, the virus evolution rate corresponding to the predetermined target area is determined. Based on the mutation site information of multiple target viruses corresponding to the multiple first water samples, a risk analysis is performed on the predetermined target viruses present in the multiple first water samples to obtain the virus transmission risk corresponding to the predetermined target area. Acquire infection distribution data and neutralizing antibody titer data for each immune source of the predetermined target virus lineage present in the plurality of first water samples over time; based on the infection distribution data and neutralizing antibody titer data for each immune source, determine the change in the immune level of the population in the predetermined target area against the predetermined target virus lineage, and obtain the prediction result of the dominant lineage succession of the virus lineage over time. The virus evolution rate, the virus transmission risk, and the predicted results of the dominant lineage succession of each virus lineage over time are used as the virus evolution trend information corresponding to the predetermined target area.

2. The method according to claim 1, characterized in that, The step of determining the virus evolution rate corresponding to the predetermined target area based on the target virus lineages and lineage abundances corresponding to the plurality of first water body samples, and the target virus lineages and lineage abundances corresponding to the plurality of second water body samples, includes: Based on the target virus lineage and lineage abundance corresponding to the plurality of first water body samples, the first average nucleotide genetic distance corresponding to the plurality of first water body samples is determined using predetermined software. Based on the target virus lineage and lineage abundance corresponding to the plurality of second water body samples, the second average nucleotide genetic distance corresponding to the plurality of second water body samples is determined using predetermined software. The viral evolution rate corresponding to the predetermined target region is calculated based on the first average nucleotide genetic distance and the second average nucleotide genetic distance.

3. The method according to claim 1, characterized in that, The information on multiple target virus mutation sites includes multiple target virus mutation sites; The step of performing a risk analysis on predetermined target viruses present in the multiple first water body samples based on multiple target virus mutation site information corresponding to the multiple first water body samples, and obtaining the virus transmission risk corresponding to the predetermined target area, includes: For each of the multiple target virus mutation sites, a structural stability analysis is performed on the target virus mutation site to obtain a first conclusion related to its structural stability. By performing receptor binding energy analysis on the target viral mutation sites, a second conclusion related to the receptor binding energy of the target viral mutation sites is obtained. Antibody binding energy analysis was performed on the target viral mutation sites to obtain a third conclusion related to the antibody binding energy of the target viral mutation sites. Based on all the first, second, and third conclusions, the risk of virus transmission is determined.

4. The method according to claim 1, characterized in that, Based on the infection distribution data and neutralizing antibody titer data for each immune source, the change in the immune level of the population within the predetermined target area against the predetermined target viral lineage is determined, and the prediction results of the dominant lineage succession of the viral lineage over time are obtained, including: The neutralizing antibody titer data for each immune source were standardized to obtain the target neutralizing antibody titer data for each immune source; Based on multiple target neutralizing antibody titer data, the linear decay value of neutralizing antibody titer over time was determined; Based on the linear decay value of the neutralizing antibody titer over time, the cross-protective efficacy of each immune source over time is determined. Based on the cross-protective efficacy of each immune source over time and the infection distribution data, the cross-protective efficacy of each immune source against the target viral lineage over time is determined. Based on the cross-protective efficacy of all immune sources over time, the predicted evolution of the dominant lineage of the target virus lineage over time is determined.

5. The method according to claim 1, characterized in that, The target virus mutation site information includes multiple target virus mutation sites. The step of determining the target virus lineage and lineage abundance corresponding to the water sample based on the target virus mutation site information using a pre-constructed linear deconvolution model includes: Based on the target virus mutation site information, a set of target virus variants containing at least one target virus variant corresponding to each target virus mutation site is determined using a pre-constructed database of mutation sites and strain relationships. A mutation site matrix is ​​constructed based on the set of target virus variants and multiple target virus mutation sites corresponding to each target virus variant. Sequencing data statistics were performed on multiple target virus mutation sites corresponding to the water samples to obtain the mutation allelic count and sequencing coverage depth corresponding to each target virus mutation site. Based on the mutation allelic count and sequencing coverage depth corresponding to each target virus mutation site, the substitution frequency corresponding to each target virus mutation site was calculated. The substitution frequency corresponding to each target viral mutation site was screened to obtain a sample observation sequence containing multiple target substitution frequencies; Based on the observed sequence, the mutation site matrix, and the target virus variant set, the linear deconvolution model is solved using a predetermined algorithm to obtain the initial target virus lineage corresponding to the water sample and the initial lineage abundance corresponding to the initial target virus lineage. Based on the initial lineage abundance corresponding to the initial target virus lineage, the target substitution frequency, and the mutation site matrix, the goodness of fit is calculated. In response to determining that the goodness of fit is greater than a predetermined goodness of fit, the initial target virus lineage and its corresponding initial lineage abundance are used as the target virus lineage and lineage abundance corresponding to the target virus mutation site information.

6. The method according to claim 1, characterized in that, The mutation detection of the water sample using wastewater sequencing to obtain the target virus mutation site information corresponding to the water sample includes: Viral nucleic acid was extracted from the water sample, and the extracted viral nucleic acid was amplified by multiple polymerase chain reaction based on a predetermined target virus set to obtain the viral nucleic acid to be detected. Sequencing of each viral nucleic acid to be detected yields the viral gene sequence corresponding to each predetermined target virus present in the water sample. Each viral gene sequence underwent quality improvement processing; Each viral gene sequence that has undergone the aforementioned quality improvement process is subjected to primer splicing, gene comparison, and mutation detection to obtain the viral gene sequence and viral mutation sites corresponding to each predetermined target virus. The viral gene sequence and viral mutation sites corresponding to each predetermined target virus are used as the target virus mutation site information corresponding to the water sample.

7. The method according to claim 6, characterized in that, The quality enhancement process for each viral gene sequence includes: For each viral gene sequence, a predetermined adapter sequence and a predetermined terminal sequence are removed from the viral gene sequence to obtain a first viral gene sequence; The first viral gene sequence was filtered to obtain the second viral gene sequence. Based on a predetermined common viral gene sequence, the second viral gene sequence is subjected to gene knockout to obtain a viral gene sequence that has undergone the quality improvement process.

8. A device for determining regional virus evolution trends based on wastewater sequencing, characterized in that, include: The acquisition module is configured to acquire multiple first water body samples corresponding to a predetermined target area within a predetermined time period, and to acquire multiple second water body samples compared with the multiple first water body samples. The detection module is configured to perform mutation detection on each of the plurality of first water samples and the plurality of second water samples using wastewater sequencing to obtain target virus mutation site information corresponding to the water sample. The first determining module is configured to determine the target virus lineage and lineage abundance corresponding to the water sample based on the target virus mutation site information using a pre-built linear deconvolution model. The second determining module is configured to determine the virus evolution rate corresponding to the predetermined target area based on the target virus lineage and lineage abundance corresponding to the plurality of first water samples, and the target virus lineage and lineage abundance corresponding to the plurality of second water samples. The analysis module is configured to perform risk analysis on predetermined target viruses present in the multiple first water body samples based on multiple target virus mutation site information corresponding to the multiple first water body samples, and obtain the virus transmission risk corresponding to the predetermined target area; The third determining module is configured to acquire infection distribution data of a predetermined target virus lineage present in the plurality of first water samples over time and neutralizing antibody titer data of each immune source; based on the infection distribution data and neutralizing antibody titer data of each immune source, determine the change in the immune level of the population in the predetermined target area against the predetermined target virus lineage, and obtain the prediction result of the dominant lineage succession of the virus lineage over time. The fourth determining module is configured to use the virus evolution rate, the virus transmission risk, and the prediction results of the dominant lineage succession of each virus lineage over time as the virus evolution trend information corresponding to the predetermined target area.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method described in any one of claims 1 to 7.