Techniques for assessing evolutionary stability of genomic fragments
By calculating the genetic differentiation parameters of genomic fragments using the GDT model, the problem of the evolutionary rate differences at different locations of the genome not being reflected in existing technologies has been solved, enabling more accurate estimation of evolutionary rates and improving the design effectiveness of vaccines and treatments.
Patent Information
- Application Number
- CN202480047333.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-19
- Filing Date
- 2024-07-17
- Publication Date
- 2026-02-17
AI Technical Summary
Existing conventional methods cannot accurately reflect the differences in evolutionary rates at different locations in the genome, especially the high evolutionary selection pressure in epitope regions, resulting in insufficient design efficacy for vaccines and treatments.
By employing the GDT model, the evolutionary rate can be directly estimated by calculating the genetic differentiation parameters of genomic fragments, avoiding unrealistic assumptions in conventional techniques and enabling efficient calculation of the evolutionary rate of genomic fragments.
It provides more accurate evolutionary rate estimates, helps select stable candidate epitopes for vaccines and treatments, and improves the effectiveness of vaccines and treatments.
Smart Images

Figure CN121548858A_ABST
Abstract
Description
Technical Background
[0002] This disclosure relates to the analysis of genomic fragments (including protein fragments), and particularly to the assessment of the evolutionary stability or rate of evolution of genomic fragments.
[0003] Many diseases are caused by pathogens, such as viruses or other protein-based antigens. Specific sites on the proteins of such pathogens (called epitopes) can be neutralized by antibodies binding to them. Epitopes are short amino acid chains. It is well known that the immune system can learn to produce an appropriate immune response after exposure to a pathogen. For example, B-cell epitopes are recognized by secretory antibodies or B-cell receptors, while T-cell epitopes are presented by the major histocompatibility complex (MHC) and recognized by T-cell receptors. This principle forms the basis for vaccine development and antibody therapy development targeting pathogenic diseases. However, once a pathogen mutates, certain antibodies may lose their effectiveness, making individuals susceptible to new strains and / or rendering existing treatments ineffective. For example, mutations in the RNA polymerase in influenza viruses can lead to resistance to antiviral drugs such as favipiravir and oseltamivir. Therefore, understanding the evolutionary stability of pathogens helps in designing effective vaccines or treatments.
[0004] In cancer vaccines, selecting appropriate epitopes is crucial for effective treatment. Novel epitopes are more likely to appear in unstable genomic regions, which can be assessed from an evolutionary perspective. Therefore, studying the evolutionary stability of tumor cell fragments (such as oncogenes and tumor suppressor genes (TSGs)) can help identify novel epitopes and develop cancer vaccines.
[0005] In diagnosis and prognosis, circulating tumor DNA (ctDNA) is considered a promising biomarker for cancer diagnosis. Assessing the evolutionary stability of ctDNA has the potential to serve as an alternative approach to tumor phylogenetics to infer tumor aberrations and progression stages.
[0006] Evolutionary stability can be assessed by estimating the evolutionary rate based on observed mutation rates. Conventional methods are primarily based on constructing phylogenetic trees of strains. Examples of these methods include root tip regression (RTT), parameterization of maximum likelihood (ML) trees, and Bayesian variant methods. Phylogenetic trees are typically constructed using large genomic fragments, such as the full sequences of major surface proteins (typically 50 to 200 amino acids). Therefore, the resulting evolutionary rates are appropriate for the overall evolutionary rate of the proteins. Summary of the Invention
[0007] Heterogeneity in evolutionary rates has been observed across different locations in the genome. In pathogens like viruses, epitopes are typically under high evolutionary selection pressure (because evading antibodies provides a survival advantage) and tend to mutate at higher rates than other regions. Evolutionary rates obtained using conventional techniques fail to explain these differences and may not accurately reflect the evolutionary rates of actual or potential epitopes. Furthermore, a better understanding of which parts of a protein are evolutionarily stable could help in selecting more stable candidate epitopes for drug or vaccine development. Estimating the evolutionary stability of the tumor genome through ctDNA analysis can also provide an indicator of tumor growth rate.
[0008] Some embodiments described herein relate to techniques for estimating evolutionary rates, applicable to genomic fragments (or amino acid sequences) of any length, ranging from a single codon (or amino acid position) to an entire genome or protein. In some embodiments, calculating the evolutionary rate of a genomic fragment may include: obtaining a reference sequence of the target genome at an initial time; obtaining strain-specific sequences of each genome of one or more strains of the target genome over a period of time after the initial time; selecting genomic fragments for analysis; calculating genetic differentiation parameters of the fragments based on the genetic distance between the fragments in the reference sequence and their corresponding fragments in each strain-specific sequence; and calculating the evolutionary rate of the fragments based on the genetic differentiation parameters.
[0009] The following detailed description, together with the accompanying drawings, will provide a better understanding of the nature and advantages of the claimed invention.
[0010] Attached Figure Description
[0011] Figure 1 A flowchart illustrating a process for determining the evolutionary rate of genomic fragments according to some implementation schemes is shown.
[0012] Figure 2 A table showing the results of a comparison between the conventional process for determining the rate of evolution and the process according to some implementation schemes.
[0013] Figure 3 A graph showing the fragment-specific evolution rate of the SARS-CoV-2 spike protein as a function of nucleotide position, as determined according to some implementation schemes.
[0014] Figure 4 A table shows the fastest-evolving and slowest-evolving (or most conserved) 90-nucleotide fragments of the spike protein, determined using a procedure according to the implementation scheme.
[0015] Figure 5 Candidate epitopes of the SARS-CoV-2 spike protein are listed, along with a table of evolutionary stability quantified according to the implementation scheme.
[0016] Detailed description
[0017] For purposes of illustration and description, the following description presents exemplary embodiments of the invention. The invention is not intended to be exhaustive or to limit the claimed invention to the precise forms described; rather, it will be understood by those skilled in the art that various modifications and variations can be made. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling those skilled in the art to best utilize the invention in the various embodiments and make various modifications suitable for the intended particular use.
[0018] Some of the embodiments described herein relate to techniques for estimating evolutionary rates, which can be applied to genomic fragments (or amino acid sequences) of any length, including fragments corresponding to B-cell or T-cell epitopes. Some embodiments use models referred to herein as “genetic differentiation over time” or “GDT” models.
[0019] It is evident that GDT models avoid the unrealistic assumptions inherent in conventional techniques. Examples of such assumptions include a constant evolutionary rate across different genomic segments or periods, or the independence of mutations at genomic locations. GDT models can be directly calculated based on genetic distance and probability as a function of time, without needing to fit the entire phylogeny of the viral strain; therefore, GDT models are computationally more efficient than conventional techniques.
[0020] According to some implementation schemes, GDT models can be based on probabilistic estimates of substitution rates within genomic segments. In the example presented in this paper, the number of substitutions is described by a Poisson random variable. Specifically, it is assumed that the gene sequence (e.g., the DNA or RNA sequence of a pathogen) has multiple positions. For genomic fragments containing consecutive nucleotides in a sequence ( k Replacement rate This can be defined as the substitution of each nucleotide site per unit time that is common between sites in the fragment. A fragment can begin at any position in the sequence and can have any length. If the fragment... The length is 1 (one nucleotide site). It is the site-specific substitution rate, and if the sequence Length is J , This refers to the rate of replacement at the genome level. Therefore, It can be understood as a fragment. The evolution rate parameter.
[0021] In some implementations, the GDT model provides a method for evaluation The technology. For example, if It is a fragment The replacement time at any site in the middle can be assumed to be Event rate parameter It exhibits an exponential distribution. Therefore, in time... At The probability of observing substitution at any site can be expressed as:
[0022] As long as the replacement rate Small (consistent with known pathogens), fragment k The expected number of replacements We can then use the Poisson approximation of the binomial distribution:
[0023] in It is a fragment The number of sites in the sequence. Therefore, from the initial time to time... fragments Total number of replacements The expected value is the parameter A Poisson random variable, where the expected value is:
[0024] By obtaining the baseline sample sequence at the start time, and at a later time By obtaining multiple sample sequences and comparing these sequences to determine the number of substitutions in each sample, parameters for a specific pathogen can be determined empirically. The "genetic differentiation" parameter Defined as initial time and time Between fragments The sample mean of the number of replacements in the middle. Can be used as The maximum likelihood estimate. Therefore, according to equations (1) and (2), the replacement rate It can be calculated as:
[0025] Figure 1A flowchart illustrating a process 100 for calculating genomic fragment replacement rates according to some implementation schemes is shown. Parts of process 100 can be implemented using a computer system. At box 102, sequence data of a “reference” genome representing the target (e.g., a virus or other pathogen or tumor cell) is obtained at an initial time. The initial time can be, but is not necessarily, the time when the target was first observed or isolated; furthermore, any point in time where the genomic sequence of the target is available can be chosen as the initial time. The gene sequence data can be obtained using appropriate techniques, including known DNA or RNA sequencing techniques, access to databases of known sequences, or other techniques.
[0026] At box 104, genome sequence data of one or more target strains are obtained at a later time period, which can encompass, for example, weeks, months, or years after the initial time. There may be, but is not required, an interval between the initial time and the start of the later time period. Similar to box 102, any time period in which genome sequences are available can be selected, and any suitable technique can be applied to obtain the gene sequence data.
[0027] At box 106, you can select a fragment within the genome. For example, a fragment encoding an amino acid sequence that is known or a candidate antibody-binding epitope can be selected. The length of such a fragment can be, for example, 12 to 18 nucleotides. More generally, the location of the fragment in the genome and the length of that fragment can be selected as needed. ).
[0028] At box 108, the genetic differentiation parameters of this fragment can be calculated. In various implementation schemes, each strain is targeted at a later time. It can be done by identifying the strain Fragments in the original genome ( Genetic distance between To calculate Then, the sample mean of genetic distance is calculated using the following method:
[0029] in Indicates the time period The number of observed strains. Various distance metrics can be used to define genetic distance. Examples are described below.
[0030] At box 110, evolutionary rates, such as replacement rates, can be calculated using genetic differentiation parameters. For example, equation (4) above can be used to calculate the evolutionary rate. In some implementations, evolutionary stability can be calculated in addition to or instead of the evolutionary rate. Evolutionary stability can be quantified as a parameter inversely proportional to the evolutionary rate (a higher evolutionary rate corresponds to lower evolutionary stability, and vice versa). For example, the evolutionary stability of a fragment can be quantified as the evolutionary rate of that fragment. The negative logarithm. (According to equation (4), it is clear that it is not necessary to calculate the evolution rate first.) Evolutionary stability can then be calculated.
[0031] For a given sequence and time period Boxes 106, 108, and 110 of process 100 can be applied to any number of fragments of arbitrary length, including overlapping fragments. Therefore, the evolutionary stability of a genome can be characterized based on the evolutionary rate of its fragments. Furthermore, process 100 can be repeated on the same fragment for different initial and later times of selection, which can reveal changes in the evolutionary rate as a function of time.
[0032] As mentioned above, at box 108, the later strain fragments and the initial genome ( Genetic distance between Used to calculate genetic differentiation parameters Genetic distance can be defined using various distance metrics. For example, in some implementations, genetic distance is based on substitution mutations that only consider one base being replaced by another. (This approach is referred to below as "GDT-type S" or "GDT-S"). If The gene sequence representing the strain, in which This indicates a location in the genome sequence where there are no gaps, and Indicates the strain Location in the gene sequence The nucleotides at the location, then the strain and Genetic distance between It can be defined as:
[0033] in This is a function that returns 1 if the parameter is true, and 0 otherwise. Equation (6) corresponds to the fragment based on the genome. Calculate the Hamming distance.
[0034] As another example, in some implementations, genetic distance is based on taking into account substitution mutations and vacancies, i.e., insertions or deletions. (This approach is referred to below as "GDT-type SG" or "GDT-SG"). As used herein, "vacancy length" refers to the number of consecutive positions where an insertion or deletion exists. For vacancy lengths greater than 1, counting mutations using Hamming distance, considering each vacancy, can overestimate the impact of the vacancy. Therefore, a penalty term can be used to calculate the vacancy length. For example, genetic distance... It can be defined as:
[0035] in It is the length of consecutive vacancies (i.e., the number of inserted or deleted bases), and It is an extended open shot penalty, which can be selected as needed, for example, .
[0036] Equations (6) and (7) are examples of calculating genetic distance. Other metrics can be substituted.
[0037] To illustrate the performance of the GDT process, the GDT-S and GDT-SG implementations of Process 100 were applied to the whole SARS-CoV-2 genome and compared with results obtained using conventional Bayesian evolutionary analysis sampling tree (BEAST) software based on phylogenetic trees. Both models were applied to a set of 360 SARS-CoV-2 sequences isolated in California during three time periods: Time Period 1 from May to October 2020; Time Period 2 from March to August 2021; and Time Period 3 from November to April 2022. The reference strain (ancestor) for all time periods was the Wuhan-Hu-1 genome. (All sequences were downloaded from the Global Influenza Sharing Initiative (GISAID) database).
[0038] Figure 2 Table 200 presents a comparison of the evolutionary rates obtained for each time period using the standard BEAST procedure and the GDT-S and GDT-SG procedures according to some implementation schemes. As shown in the figure, in time period 1 (the early stages of the SARS-CoV-2 epidemic), all three models obtained similar evolutionary rates. In time periods 2 and 3, as more genetic variations emerged, as a result of different treatments of insertion and deletion events, GDT-SG began to provide slightly higher estimates than BEAST and GDT-S. Table 2 also shows the computation time used by each procedure (executed using the same hardware); notably, for a sample size of 120 genetic sequences, the GDT procedure is more than 30,000 times faster than BEAST. The computational efficiency of the GDT procedure opens the prospect of efficiently and effectively analyzing thousands or even millions of sequences.
[0039] According to some implementation schemes, the evolutionary stability of the entire genome can be described based on the fragment-based evolutionary rate determination provided by process 100. For example, process 100 can be applied to use fragment length measurements... and step length The sliding window defines multiple segments. Boxes 104, 106, and 108 of process 100 can be applied from position 1, position... ,Location equal length at the beginning The fragment. For example, using fragment length. And step length The implementation of the sliding window method of process 100 is applied to the SARS-CoV-2 virus. (In this example, the fragment length is...) Step size The sequences are long, resulting in overlapping fragments. (However, this is not necessary). To reflect mutational activity in subsequent fluctuations of COVID-19 infections caused by the alpha, beta, delta, and omega-3 variants, the examples used sequence samples isolated in California from the period of March 2020 to February 2022, totaling 45,413 sequence samples.
[0040] Figure 3 Graph 300 shows the evolutionary rate of the SARS-CoV-2 spike protein as a function of nucleotide position, obtained using the sliding window implementation of process 100. Graph 300 shows that the evolutionary rate is highest in the SD2 (shown in 302), RBD (shown in 304), and NTD (shown in 306) regions, while the S2 subunit (shown in 308) is relatively conserved. Notably, the high-rate regions contain some nucleotide positions where mutations from the Omeprone BA.5 variant appeared three months after the end of the analysis period. Further quantification of the results... Figure 4 Table 400 lists the 10 fastest-evolving and 10 slowest-evolving (or most conserved) 90-nucleotide segments of the spike protein. It is worth noting that genome-wide methods (such as BEAST) do not support the analysis of local evolutionary rates within the genome.
[0041] According to some implementation schemes, the fragment-based evolutionary rate determination provided by process 100 can be used to evaluate the stability of epitopes of arbitrary length as candidates for developing broad-spectrum antibodies and / or vaccines. The evolutionary stability of a fragment can be quantified as the evolutionary rate of that fragment (e.g., calculated using process 100). The negative logarithm of ). In various implementations, candidate epitopes may include known epitopes recorded in existing databases (e.g., the Immune Epitope Database and Analysis Resource Library (IEDB)) and / or calculated predicted epitopes.
[0042] For example, IEBB lists 1,416 candidate epitopes of SARS-CoV-2 linear peptides that have known T-cell immune responses in human hosts. The evolutionary rates of these peptides are calculated to identify candidate epitopes with high stability (lowest evolutionary rate) and high response frequency (determined based on known T-cell immune responses). A total of 32 candidate epitopes were identified. Figure 5 A partial list of identified epitopes is presented as Table 500, which includes the protein containing the candidate epitope; the amino acid sequence of the epitope; the length of the amino acid sequence; the start and end positions of the epitope within the protein; and stability, which is quantified as an evolutionary rate using an implementation of process 100. The negative logarithm. This type of information can be used to select target epitopes or combinations of target epitopes for the development of vaccines or antibody therapies.
[0043] While the invention has been described with reference to specific embodiments, those skilled in the art will understand that corresponding changes and modifications are also possible. For example, the size (number of nucleotides) and position of each fragment, as well as the number of fragments, can be selected as needed. The analyses described herein can also be applied to amino acid sequences, not just nucleotide sequences. The techniques described herein can be used to determine the evolutionary rate of any genome, protein, or fragment thereof. As mentioned above, the ability to assess the evolutionary rate or evolutionary stability of small fragments of pathogen genomes or proteins can help identify target epitopes for vaccines and / or immunotherapies, and can also help predict trends in future infection and / or morbidity.
[0044] All processes described herein are illustrative and are subject to modification. Operations may be performed in a different order than described, to the extent logically permissible; operations may be omitted or combined; and operations not explicitly described above may be added.
[0045] The data analysis and computation operations described herein can be implemented in a computer system, which can be a conventionally designed computer system, such as a desktop computer, laptop computer, tablet computer, mobile device (e.g., smartphone), etc. Such a system may include one or more processors to execute program code (e.g., a general-purpose microprocessor that can be used as a central processing unit (CPU) and / or a dedicated processor that can provide enhanced parallel processing capabilities, such as a graphics processing unit (GPU)); memory and other storage devices for storing program code and data; user input devices (e.g., keyboard, pointing devices (e.g., mouse or touchpad), microphone); user output devices (e.g., display device, speaker, printer); combined input / output devices (e.g., touchscreen display); signal input / output ports; network communication interfaces (e.g., wired network interfaces (e.g., Ethernet interfaces) and / or wireless network communication interfaces (e.g., Wi-Fi)); etc.
[0046] Computer programs incorporating the features of this invention, which can be implemented using program code, can be encoded and stored on various computer-readable storage media; suitable media include magnetic disks or magnetic tapes, optical storage media (e.g., optical discs (CDs) or DVDs (Digital Universal Disk), flash memory, and other non-transitory media. (It should be understood that "storage" of data is different from transmitting data using transient media such as carrier waves). Computer-readable media encoded with program code can include internal storage media of compatible electronic devices and / or external storage media readable by electronic devices that execute the code. In some cases, the program code may be provided to electronic devices via Internet download or other transmission paths.
[0047] While this document describes various circuits and components with reference to specific boxes, it should be understood that these boxes are defined for ease of description and are not intended to imply a specific physical arrangement of component portions. These boxes do not need to correspond to physically different components, and the same physical components can be used to implement various aspects of multiple such boxes. Components described as dedicated or fixed-function circuits can be configured to perform operations by providing an appropriate arrangement of circuit components (e.g., logic gates, registers, switches, etc.); automated design tools can be used to generate appropriate arrangements of circuit components that implement the operations described herein. Components described as processors or microprocessors can be configured to perform the operations described herein by providing appropriate program code. Depending on how the initial configuration is obtained, individual boxes may be reconfigurable or non-reconfigurable. Embodiments of the invention can be implemented in a variety of devices, including electronic devices implemented using a combination of circuitry and software.
[0048] Therefore, although the invention has been described with respect to specific embodiments, it should be understood that the invention is intended to cover all modifications and equivalents within the scope of the appended claims.
Claims
1. A computer-implemented method comprising: obtaining, at an initial time, a reference sequence of a genome of a target; obtaining, over a time period following the initial time, strain-specific sequences of respective genomes of one or more strains of the target; selecting a segment of the genome for analysis; computing a genetic differentiation parameter for the segment based on a genetic distance between the segment in the reference sequence and a corresponding segment in each of the strain-specific sequences; and computing an evolutionary rate or evolutionary stability parameter for the segment based on the genetic differentiation parameter.
2. The computer-implemented method of claim 1, wherein the target is a pathogen.
3. The computer-implemented method of claim 2, wherein the selected segment corresponds to a candidate epitope of the pathogen.
4. The computer-implemented method of claim 1, further comprising: performing the acts of selecting a segment of the genome, computing the genetic differentiation parameter, and computing a respective evolutionary rate or evolutionary stability parameter for each of a plurality of different segments of the genome.
5. The computer-implemented method of claim 4, wherein the different segments are selected according to a sliding window.
6. The computer-implemented method of claim 4, wherein the target is a pathogen, and the selected different segments include candidate epitopes of the pathogen.
7. The computer-implemented method of claim 6, further comprising: ranking the candidate epitopes based at least in part on a comparison of the respective evolutionary rate or evolutionary stability parameters for the different segments.
8. The computer-implemented method of claim 4, further comprising: selecting a target epitope or a combination of target epitopes for development of a vaccine or antibody therapy based at least in part on the respective evolutionary rate or evolutionary stability parameters for the different segments.
9. The computer-implemented method of claim 1, wherein, Computing the genetic differentiation parameter comprises computing an average of a distance measure between the segment in the reference sequence and the corresponding segment in each of the strain-specific sequences.
10. The computer-implemented method of claim 9, wherein the distance measure is a Hamming distance based on a number of substitutions.
11. The computer-implemented method of claim 9, wherein, The distance measure comprises a Hamming distance based on a number of substitutions modified by a penalty term computed based on a number of consecutive sites of nucleotide deletion or insertion.
12. The computer-implemented method of claim 1, wherein the target is a genomic segment of a tumor cell.
13. The computer-implemented method of claim 12, wherein, The selected segment corresponds to a candidate neoepitope of the tumor cell.
14. The computer-implemented method of claim 1, wherein, The target is ctDNA derived from a tumor cell.
15. The computer-implemented method of claim 14, wherein, The segment is selected to estimate a stage of tumor growth or cancer progression.
16. A system comprising: a memory; and a processor coupled to the memory and configured to perform the method of any one of claims 1-15.
17. A computer-readable storage medium storing program code instructions that, when executed by a processor in a computer system, cause the processor to perform the method of any one of claims 1-15.