Bacterial ancient genome analysis method, device and equipment
By using a phylogenetic tree-based approach to extract and assemble homologous fragments and supplement outgroup gene sequences, the problem of speed and accuracy in bacterial paleogeny analysis in existing technologies has been solved, enabling efficient analysis of large amounts of bacterial data and accurate construction of bacterial evolutionary relationships.
Patent Information
- Application Number
- CN202511536798.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing technologies struggle to quickly and accurately analyze large volumes of bacterial community data to reconstruct their ancient genomes at various developmental stages, hindering the precise construction of bacterial evolutionary relationships, especially when dealing with complex evolutionary events.
By identifying branch nodes based on the phylogenetic tree of bacterial strains, extracting homologous fragments, performing localization analysis, and splicing orthologous fragments, and supplementing the ancient genome skeleton by combining outgroup gene sequences, the MHB algorithm and skeleton completion algorithm are used to maximize the preservation of bacterial ancestral information.
It enables rapid and accurate analysis of large volumes of bacterial community data, maximizes the preservation of bacterial ancestral information, reduces information loss, can handle complex evolutionary events, and is suitable for the analysis of bacterial ancient genomes.
Smart Images

Figure CN121034397A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of genome evolution analysis, in particular to a bacterial paleogenomic analysis method, device and equipment. BACKGROUND
[0002] In the research of bacterial evolution, it is of great significance to clarify the genomic evolution trajectory of most bacteria (such as Salmonella typhi) in the history of species evolution and the process of co-evolution with humans. The ideal research scheme is to analyze the bacterial paleogenome at each major development node, and track the evolution path of gene recombination by comparing the paleogenomes of adjacent nodes. Although the rapid development of whole genome sequencing today provides a high-resolution tool for evolutionary tracing, the existing means of studying bacterial evolution still has great limitations, and the accurate construction of bacterial evolutionary relationships still faces multiple obstacles. SUMMARY
[0003] The present application provides a bacterial paleogenomic analysis method, device and equipment to solve the problem of how to quickly and accurately analyze a large amount of bacterial community data to obtain the paleogenome of the bacterial community at each development node.
[0004] The first aspect of the present application provides a bacterial paleogenomic analysis method, comprising the following steps: determining a branch node of a to-be-synthesized paleogene according to a strain phylogenetic tree in bacterial strain data; selecting branch strains from the branch node of the to-be-synthesized paleogene for comparison and extraction of homologous fragments; performing positioning analysis on the homologous fragments of each strain, obtaining orthologous fragments from the homologous fragments according to the positioning analysis results, and splicing the orthologous fragments to obtain a bacterial paleogenomic skeleton; obtaining outgroup gene sequences of outgroup strains of the branch strains corresponding to the branch node, and supplementing the bacterial paleogenomic skeleton according to the outgroup gene sequences of the outgroup strains of the branch strains corresponding to the branch node.
[0005] Optionally, before determining the branch node of the to-be-synthesized paleogene according to the strain phylogenetic tree in the bacterial strain data, the method further comprises: extracting base sequences and / or amino acid sequences of the strains in the bacterial strain data; calculating the evolutionary relationship between organisms according to the base sequences and / or amino acid sequences, and determining the strain phylogenetic tree in the bacterial strain data according to the evolutionary relationship between organisms.
[0006] Optionally, the positioning analysis on the homologous fragments of each strain comprises: identifying the position of the homologous fragments of each strain in the genome; numbering the homologous fragments according to the position of the homologous fragments in the genome; sorting the numbers of all homologous fragments according to the length of the homologous fragments, and analyzing the homology and positional order relationship between the homologous fragments in sequence according to the sorting results, and deleting the homologous fragments with inconsistent homology in the positioning process until all homologous fragments are analyzed.
[0007] Optionally, the orthologous fragments are obtained from the homologous fragments according to the positioning analysis result, and the paleogenomic skeleton of the bacteria is obtained by splicing the orthologous fragments, including: extracting the reserved homologous fragments and the corresponding positional order relationship from the positioning analysis result; taking the reserved homologous fragments as the orthologous fragments in the homologous fragments; and splicing the orthologous fragments according to the positional order relationship to obtain the paleogenomic skeleton of the bacteria.
[0008] Optionally, the paleogenomic skeleton of the bacteria is supplemented according to the outgroup gene sequence of the outgroup strain of the branch strain corresponding to the branch node, including: obtaining the skeleton sequence of the paleogenomic skeleton and the original gene sequence of the branch strain corresponding to the branch node; comparing the original gene sequence with the skeleton sequence and the outgroup gene sequence respectively; and supplementing the supplementable fragments in the comparison result to the paleogenomic skeleton of the bacteria, wherein the supplementable fragments are fragments that do not exist in the comparison of the original gene sequence and the skeleton sequence and exist in the comparison of the outgroup gene sequence and the original gene sequence.
[0009] Optionally, the comparison of the original gene sequence with the skeleton sequence and the outgroup gene sequence further includes: generating an orthologous block file of the skeleton sequence based on the comparison of the original gene sequence and the skeleton sequence, identifying a gap region in the original gene sequence according to the orthologous block file of the skeleton sequence, and taking the gap region as a skeleton block; generating an orthologous block file of the outgroup gene sequence based on the comparison of the original gene sequence and the outgroup gene sequence, identifying the supplementable fragments according to the orthologous block file of the outgroup gene sequence, and taking the supplementable fragments as supplement blocks; and generating the comparison result according to the skeleton block and the supplement blocks.
[0010] Optionally, before the supplementable fragments in the comparison result are supplemented to the paleogenomic skeleton of the bacteria, it further includes: determining a first gap length in the original gene sequence and a second gap length in the skeleton sequence according to the orthologous block file, calculating a ratio of the first gap length and the second gap length; determining whether a supplement condition of the paleogenomic skeleton is met according to the first gap length, the second gap length and the ratio, and if the supplement condition is met, supplementing the paleogenomic skeleton of the bacteria according to the comparison result.
[0011] Optionally, the complementary fragment in the comparison result is complemented to the ancient genome skeleton of the bacteria, comprising: identifying relative position relationship of the complement block and the skeleton block in the comparison result; determining an overlap mode between the complement block and the skeleton block according to the relative position relationship, the overlap mode comprising a complete containing mode, a left end overlap mode, a right end overlap mode and an internal overlap mode, the complete containing mode being that the skeleton block contains the complement block, the left end overlap mode being that the complement block overlaps the left end of the skeleton block, the right end overlap mode being that the complement block overlaps the right end of the skeleton block, and the internal overlap mode being that the complement block contains the skeleton block; if the overlap mode is the complete containing mode, dividing the skeleton block into a first half and a second half with the start position of the complement block as a boundary, and inserting the complement block between the first half and the second half; if the overlap mode is the left end overlap mode, dividing the skeleton block into a first half and a second half with the end position of the complement block as a boundary, and inserting the complement block in front of the second half; if the overlap mode is the right end overlap mode, dividing the skeleton block into a first half and a second half with the start position of the complement block as a boundary, and inserting the complement block behind the first half; and if the overlap mode is the internal overlap mode, replacing the skeleton block with the complement block.
[0012] The second aspect embodiment of the present application provides a bacterial ancient genome analysis device, comprising: a determination module configured to determine a branch node of an ancient gene to be synthesized according to a strain phylogenetic tree in strain data of bacteria; an extraction module configured to select branch strains from between branch nodes of the ancient gene to be synthesized to perform comparison and extraction of homologous fragments; an analysis module configured to perform positioning analysis on the homologous fragments of each strain, obtain orthologous fragments from the homologous fragments according to a positioning analysis result, and splice the orthologous fragments to obtain an ancient genome skeleton of the bacteria; and a complement module configured to obtain outgroup gene sequences of outgroup strains of the branch strains corresponding to the branch node, and complement the ancient genome skeleton of the bacteria according to the outgroup gene sequences of the outgroup strains of the branch strains corresponding to the branch node.
[0013] The third aspect embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor executes the program to implement the bacterial ancient genome analysis method of the above embodiments.
[0014] Therefore, the present application includes the following beneficial effects: The embodiments of the present application can rely on the current genome of the strain, quickly and accurately analyze a large amount of strain data, and then rely on the phylogenetic tree of the strain to solve the ancient genome of the strain at each development node. The position information between genes is considered during solving, the ancestor information of the bacteria is maximized to effectively reduce the loss of the ancestor information of the bacteria.
[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0016] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a bacterial ancient genome analysis method according to an embodiment of this application; Figure 2 This is an example diagram of a system evolution tree according to an embodiment of this application; Figure 3 This is an example diagram of a homologous segment according to an embodiment of this application; Figure 4 This is an example diagram showing the sorting of homologous segments according to embodiments of this application; Figure 5 Example diagram for determining the position of homologous segments according to embodiments of this application; Figure 6 This is an example diagram of an ancient genome skeleton according to an embodiment of this application; Figure 7 This is an example diagram of a genome according to an embodiment of this application; Figure 8 This is an example diagram showing the location of the segment to be supplemented according to an embodiment of this application; Figure 9 This is an example diagram showing the location of the unsupplementary fragment according to an embodiment of this application; Figure 10 This is a fragmentary explanation diagram of an example diagram according to an embodiment of this application; Figure 11 This is an example diagram of a ternary tree according to an embodiment of this application; Figure 12 This is an example diagram of a binary tree according to an embodiment of this application; Figure 13 This is an example diagram of complete internal insertion according to an embodiment of this application; Figure 14 This is an example diagram of left-end overlap according to an embodiment of this application; Figure 15 This is an example diagram of right-end overlap according to an embodiment of this application; Figure 16 This is an example diagram of left-end cross-boundary insertion according to an embodiment of this application; Figure 17 This is an example diagram of right-end cross-boundary insertion according to an embodiment of this application; Figure 18 This is an example diagram illustrating complete cross-boundary coverage according to an embodiment of this application; Figure 19A diagram illustrating an example of the Salmonella genome according to an embodiment of this application; Figure 20 An example diagram of an ancient genome according to an embodiment of this application; Figure 21 This is a diagram illustrating the evolutionary history of the strain genome according to embodiments of this application; Figure 22 This is a schematic diagram of inversions detected in strains and chromosomes according to embodiments of this application; Figure 23 This is a schematic diagram of the inverted termination end according to an embodiment of this application; Figure 24 This is an example diagram of a bacterial paleogenomics analysis device according to an embodiment of this application; Figure 25 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0017] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0018] Examples of ancestor reconstruction methods in related technologies are as follows: (1) Ancestor reconstruction method based on local optimal solution: BPAnalysis is computationally intensive and time-consuming; for a dataset containing 13 genomes and 105 gene fragments, it could take 200 years to obtain results. GRAPPA offers a new implementation of breakpoint analysis, capable of handling larger datasets that BPAnalysis cannot handle. However, GRAPPA still uses a simplified evolutionary model and can only handle a limited number of types of genome rearrangement events. The single cut or join (SCJ) method is still based on genome rearrangement distances. While it has a very low false positive rate and highly conserved genome reconstruction results, it still cannot accurately reconstruct phylogenetic trees and ancestral genomes, and only 60% to 95% of phylogenetic branches and 50% to 85% of ancestral adjacency relationships can be recovered.
[0019] (2) Ancestor reconstruction method based on global optimal solution: While the GASTS method overcomes the key initialization problems of BPAnalysis and GRAPPA, it can only process simplified genomic datasets with the same genomic content and unique coaxial blocks. It cannot account for gene duplication, gene loss, and gain events, and it cannot output intermediate ancestral genomes in the ancestral reconstruction process.
[0020] (3) Ancestor reconstruction using contiguous ancestral regions (CAR): The main methods are InferCARs and InferCARsPro. InferCARsPro requires a given phylogenetic tree and known branch lengths to reconstruct the ancestral genome, which is difficult to satisfy when applied to real-world genome datasets. Furthermore, algorithms based on local parsimony may overlook many true adjacencies that should exist in the ancestral genome. Both methods can only solve small-scale phylogenetic tree problems because they assume a known phylogenetic tree for the input genome. Neither method can handle complex evolutionary events, including insertions, deletions, and duplications. ProCARs is a parameter-free method that does not require branch lengths in the phylogenetic tree. However, ProCARs can only consider genomes with identical genomic content and non-duplicated blocks, and it does not allow insertion and deletion events. Moreover, in ancestor reconstruction on real-world genome datasets, ProCARs' highest resolution is only 100kb, even lower than InferCARs and InferCARsPro.
[0021] In summary, most computational ancestor reconstruction methods in related technologies can only handle genomic datasets with identical content and unique genomic markers. They mostly only handle genomic rearrangement events, which are only a subset of the complex evolutionary events in reality. MGRA2 is one of the few distance / event-based methods reported to handle a wide range of evolutionary events. For homology / proximity-based methods, only PMAG++, Gapped Adjacency, and ANGES handle complex evolutionary events including insertions, deletions, and duplications; however, these methods are mostly designed for eukaryotic genomes and are not applicable to prokaryotic genomes such as bacteria.
[0022] Therefore, for prokaryotes like bacteria, there is still a lack of effective methods for calculating ancient bacterial genome sequences. Existing algorithms suffer from poor accuracy and small size of results when analyzing small circular DNA fragments in bacterial genes, and excessive loss of ancestral information. Tools for tracing or inferring ancient bacterial genomes are currently unavailable, and there is no systematic method for solving this problem. Most methods rely on manual assembly, which makes rapid analysis of large volumes of bacterial genomic data difficult. This is a crucial problem that needs to be addressed in studying the evolutionary trajectory of bacterial communities and exploring the flow of gene gain and loss in bacterial strains.
[0023] To address one of the aforementioned problems, this application proposes a method for rapidly and accurately solving bacterial ancient genomes. Specifically, it proposes a bacterial ancient genome analysis method, apparatus, and device. The apparatus, apparatus, and device of this application embodiment are described below with reference to the accompanying drawings.
[0024] Specifically, Figure 1This is a flowchart of a bacterial ancient genome analysis method according to an embodiment of this application.
[0025] like Figure 1 As shown, this bacterial paleogeny analysis method includes the following steps: In step S101, the branch node of the ancient gene to be synthesized is determined based on the phylogenetic tree of the bacterial strain data.
[0026] It is understood that, in the embodiments of this application, the branch node closest to the ancient gene to be synthesized can be determined based on the phylogenetic tree. The bacterial strain data may include data from multiple strains, such as data from Salmonella.
[0027] In this embodiment of the application, before determining the branch node of the ancient gene to be synthesized based on the phylogenetic tree of the bacterial strain data, the method further includes: extracting the base sequence and / or amino acid sequence of the bacterial strain from the bacterial strain data; calculating the evolutionary relationship between organisms based on the base sequence and / or amino acid sequence; and determining the phylogenetic tree of the bacterial strain data based on the evolutionary relationship between organisms.
[0028] It is understood that the embodiments of this application can be based on the genome of the bacterial community, and cladistic analysis can be performed. Cladecy analysis is a method for studying the evolution and systematic classification of species or sequences. Generally, the research object is the base sequence or amino acid sequence. The evolutionary relationship between organisms is calculated by mathematical statistical algorithms, and a phylogenetic tree is obtained. Based on this, the kinship between species or genes can be determined, and the branch node closest to the ancient gene to be synthesized can be identified.
[0029] For example, for a bacterial community (where A, B, and D are bacterial evolutionary nodes, and C is the ancient gene to be synthesized), the evolutionary relationship could be represented as follows: Figure 2 The phylogenetic tree shown indicates that the closest evolutionary nodes for the ancient gene C to be synthesized are A and B.
[0030] In step S102, branch strains are selected from the branch nodes of the ancient gene to be synthesized for comparison and extraction of homologous fragments.
[0031] It is understood that, in the embodiments of this application, representative strains selected between two sub-branches of the ancient genome to be synthesized can be compared to extract homologous fragments, which refer to gene sequences derived from a common ancestor.
[0032] In step S103, homologous fragments of each strain are localized and analyzed. Based on the results of the localization analysis, orthologous fragments are obtained from the homologous fragments, and the orthologous fragments are spliced together to obtain the ancient genome skeleton of the bacteria.
[0033] It is understandable that, due to various rearrangement events in bacterial genomes during evolution, seemingly homologous segments between two chromosomes are not necessarily truly homologous. Therefore, this application uses the MHB algorithm to locate homologous segments between two bacterial chromosomes, thereby obtaining an orthologous genome backbone. This allows for the use of the strain's current genome to simultaneously consider the positional information between genes during the solution process, maximizing the preservation of bacterial ancestral information and enabling rapid and accurate analysis of large amounts of bacterial community data. Furthermore, the ancient genome of the bacterial community at each developmental node can be solved using the strain's phylogenetic tree.
[0034] In this embodiment of the application, the homologous fragment localization analysis of each strain includes: identifying the position of the homologous fragment in the genome of each strain; numbering the homologous fragments according to their positions in the genome; sorting all the homologous fragment numbers according to their length; analyzing the homology and positional order relationship between the homologous fragments according to the sorting results; deleting homologous fragments with inconsistent homology during the localization process until the analysis of all homologous fragments is completed.
[0035] Understandably, the MHB algorithm includes: numbering homologous fragments of each strain according to their location within the genome, sorting them according to their length, analyzing the homology and positional order of the homologous fragments based on the sorting results, for example, determining the size relationship of the homologous fragments based on the sorting results, analyzing the homologous fragments in descending order, comparing the homology of the homologous fragment to be analyzed with that of the homologous fragments already analyzed, determining the positional order of the homologous fragments if they match, and deleting homologous fragments with inconsistent homology during the localization process until all homologous fragments have been analyzed.
[0036] In this embodiment of the application, the process of obtaining orthologous fragments from homologous fragments based on the location analysis results and splicing the orthologous fragments to obtain the ancient genome skeleton of bacteria includes: extracting the retained homologous fragments and their corresponding positional order relationships from the location analysis results; using the retained homologous fragments as orthologous fragments in the homologous fragments; and splicing the orthologous fragments according to their positional order relationships to obtain the ancient genome skeleton of bacteria.
[0037] It is understood that, after the homologous fragment localization analysis, the embodiments of this application can extract the retained homologous fragments and their corresponding positional order relationships from the localization analysis results, and splice the retained homologous fragments according to the positional order relationships to obtain the ancient genome skeleton of bacteria.
[0038] For example, to further illustrate the localization analysis process of the MHB algorithm, let's take... Figure 3 Taking the homologous trunk between the two chromosomes strA and strB as an example, in order to analyze...Figure 3 The homologous backbone between the two chromosomes strA and strB is shown. Homologous segments in each strain are numbered according to their location within the genome and then sorted according to their length. An example of the sorting results is shown below. Figure 4 As shown, taking six homologous segments as an example, the homologous blocks in the attached figure are homologous segments.
[0039] like Figure 5 As shown, the localization relationship of the longest (L1) and second longest (L2) segments in two genomes is observed and compared. If the localization between the two genomes is consistent (usually), a third long block pair (L3) is included, and the localization relationship of the third long block pair (L3) is compared with that of L1 / L2. If the relationship between the genomes is consistent, L3 is retained; otherwise, L3 is deleted. The composition and positional order of homologous backbone fragments are then updated. Next, the L4 block is further examined, comparing its homology and positional order with all previously identified orthologous blocks. Similarly, if L4 is consistent with all previously identified orthologous blocks, it is retained; otherwise, it is deleted, and the process moves to the next block. These steps are repeated until the shortest homologous block pair is analyzed. Finally, all the obtained orthologous fragments are joined together to obtain the following... Figure 6 The ancient genome skeleton of the bacteria shown.
[0040] In step S104, the exogroup gene sequence of the exogroup strain of the branch strain corresponding to the branch node is obtained, and the ancient genome skeleton of the bacteria is supplemented according to the exogroup gene sequence of the exogroup strain of the branch strain corresponding to the branch node.
[0041] It is understood that the outgroup strains corresponding to the branch nodes are the genomes of closely related outgroup strains or the genomes of the most recent ancestor. In this embodiment, the genomes of closely related outgroup strains or the genomes of the most recent ancestor can be compared with the representative strains of any branch to extract homologous fragments and further repair the ancient genome skeleton. The same comparison and repair of the ancient genome skeleton is carried out for the representative strains of another branch. Through the above steps, the phylogenetic tree is solved from bottom to top, and the oldest gene of the target strain population is obtained by relying on the outgroup strains of the bacterial community.
[0042] In this embodiment of the application, supplementing the ancient genome skeleton of bacteria based on the outgroup gene sequence of the outgroup strain of the branch strain corresponding to the branch node includes: obtaining the skeleton sequence of the ancient genome skeleton and the original gene sequence of the branch strain corresponding to the branch node; comparing the original gene sequence with the skeleton sequence and the outgroup gene sequence respectively; and supplementing the ancient genome skeleton of bacteria with supplementable fragments from the comparison results, wherein the supplementable fragments are fragments that are not present in the comparison between the skeleton sequence and the original gene sequence, and fragments that are present in the comparison between the outgroup gene sequence and the original gene sequence.
[0043] It is understood that the embodiments of this application utilize a backbone completion algorithm to compare and analyze the genomes of two strains to infer their common ancestral gene sequences. In this process, orthologous fragments originate from their common ancestor, while non-homologous fragments may be variants or exogenous fragments. Furthermore, in the evolutionary process of strains, outgroups are more closely related to ancient genes; therefore, outgroup gene sequences at the same level as their ancient genes can be introduced as a reference. By supplementing the outgroup gene sequences and the parts shared by the two strains, the genomic backbone is further improved.
[0044] Specifically, such as Figure 7 As shown, this embodiment of the application obtains sequences A and B of the branch strain corresponding to the branch node, the ancient genome backbone BG obtained from sequences A and B, and the outpopulation gene sequence out. The ancient genome backbone BG can be first aligned with sequence A, and the outpopulation gene sequence out can be aligned with sequence A to find... Figure 8 The areas to be filled are shown in the blank areas in the figure.
[0045] If, during the alignment with the outgroup sequence, a fragment is found in the corresponding region where the outgroup gene sequence `out` is aligned with sequence A or B, but not in the alignment of the backbone with sequence A or B, then this fragment can be added to the corresponding region. If the fragment is found in the alignment with sequence A or B but not in the alignment with the outgroup gene sequence `out`, then this fragment is not inserted (the outgroup strains are older and closer to ancient gene evolutionary relationships; this is to follow the rule). In cases where no fragment can be added, the following applies: Figure 9 As shown.
[0046] It should be noted that, Figure 7 , Figure 8 and Figure 9 Explanation of the fragment as follows Figure 10 As shown.
[0047] In such Figure 11 In case one, the result is obtained by supplementing the single outgroup once. Figure 12 The second scenario shown yields the result after supplementing the outgroup twice. Based on the above, an additional round of supplementation can be added: BG3 is compared to sequence A, the outgroup gene sequence 'out' is compared to sequence A, and BG4 is compared to sequence A, the outgroup gene sequence 'out' is compared to sequence B.
[0048] In this embodiment of the application, comparing the original gene sequence with the backbone sequence and the outgroup gene sequence respectively, further includes: generating an orthologous block file of the backbone sequence based on the comparison between the original gene sequence and the backbone sequence; identifying gap regions in the original gene sequence based on the orthologous block file of the backbone sequence and using the gap regions as backbone blocks; generating an orthologous block file of the outgroup gene sequence based on the comparison between the original gene sequence and the outgroup gene sequence; identifying supplementary fragments based on the orthologous block file of the outgroup gene sequence and using the supplementary fragments as supplementary blocks; and generating a comparison result based on the backbone blocks and supplementary blocks.
[0049] Understandably, when performing sequence completion, the backbone completion algorithm first scans the orthologous block file to be completed, then compares the gap lengths, and uses the orthologous block file to identify regions in the source species that are not fully covered by the backbone genome. The orthologous block file of the backbone sequence can be understood as follows: Figure 6 The skeleton file shown is from the same source.
[0050] It should be noted that skeleton sequences are different from skeleton files. Skeleton files are a type of file stored in the system with the extension .backbone, while skeleton sequences are a sequence file generated based on skeleton files.
[0051] In this embodiment of the application, before supplementing the ancient genome backbone of bacteria with the supplementable fragments from the comparison results, the method further includes: determining the length of the first gap in the original gene sequence and the length of the second gap in the backbone sequence based on the orthologous block file, and calculating the ratio of the length of the first gap to the length of the second gap; determining whether the conditions for supplementing the ancient genome backbone are met based on the length of the first gap, the length of the second gap, and the ratio; if the conditions for supplementing are met, then supplementing the ancient genome backbone of bacteria based on the comparison results.
[0052] It is understood that this application uses two key criteria to determine whether the conditions for supplementing the ancient genome backbone are met. The supplementation conditions are deemed met if the following criteria are met: When the length of the region corresponding to the backbone is zero, while the length of the original sequence region exceeds the minimum length threshold by a certain number of base pairs, it is determined to be a completely missing region; or when the ratio of the backbone length to the length of the corresponding sequence region is less than a preset value, and the length of the source sequence region exceeds the minimum length threshold by a certain number of base pairs, it is determined to be a significantly imbalanced region requiring completion. The preset value can be set to 0.1 or 0.2, etc., and the minimum length threshold is to ensure that the supplemented sequence fragment has biological significance, for example, it can be set to 1000bp or 1100bp, etc.
[0053] For example, when the length of the region corresponding to the backbone is zero (i.e., the second gap length span2 = 0), while the length of the original sequence region (the first gap length span1) exceeds 1000 base pairs, it is determined to be a completely missing region; or when the ratio of the length of the backbone to the length of the corresponding sequence region (span2 / span1) is less than 0.1, and the length of the source sequence region exceeds 1000 base pairs (i.e., span1 > 1000), it is determined to be a significantly unbalanced region, which needs to be filled in (i.e., the backbone block is identified).
[0054] The following is a concrete example illustrating the process of sequence completion based on the skeleton completion algorithm: Table 1 shows an example of a skeleton file.
[0055] Table 1 The coordinates in the skeleton file are absolute coordinates of the source sequence. The relative coordinates used in the sequence file will be discussed later regarding the relative coordinates of the skeleton sequence; all absolute coordinates are relative to the source sequence, i.e., sequence A.
[0056] Example of the recognition process based on the BP algorithm (see attached) Figure 13 The second step describes the types of ternary trees in the text. The first step is to compare the skeleton sequence with sequence A. During this process, a skeleton file will be generated. The contents of the skeleton file are similar to those shown in Table 2.
[0057] Table 2 Based on Table 2, there is a region of approximately 10kb in sequence A between 20000 and 30000, while the corresponding region is only 2 (from 20000 to 20002). This is a typical "gap," which can be called a "skeleton block." Then, the outgroup is compared with sequence A. In this process, a similar file will be generated, as shown in Table 3.
[0058] Table 3 As shown in Table 3, comparing the outgroup with the A sequence generates the homologous block labeled in this file, which exactly covers the "skeleton block" region found above and can be regarded as a "supplementary block".
[0059] Furthermore, embodiments of this application may employ the following notch identification logic: Read the file generated by comparing the skeleton sequence and sequence A.
[0060] When it reads the second line (30001 40000 ...), it calculates the span between the block and the previous block.
[0061] span1 (the gap length in A) = 30000 - 20000 = 10000 span2 (gap length in the skeleton sequence) = 20001 - 20000 = 1 Check if the conditions span1>1000 (10000>1000, satisfied) and spanRatio<0.1 (0 / 10000<0.1, where 0.1 is an example of a preset value, then padding is required).
[0062] This identifies a "skeleton block" to be added, with coordinates (20000, 30000) on A.
[0063] For ease of labeling, the variables in this embodiment are recorded here and defined as follows: newMap1[1] = "Source Sequence Label\22000\28000\Patch Block\Who is the Same Source\n".
[0064] In this embodiment, supplementing the bacterial ancient genome backbone with supplementable fragments from the comparison results includes: identifying the relative positional relationship between the supplementary block and the backbone block in the comparison results; determining the overlap pattern between the supplementary block and the backbone block based on the relative positional relationship, where the overlap patterns include a complete containment pattern, a left-end overlap pattern, a right-end overlap pattern, and an internal overlap pattern. A complete containment pattern is where the backbone block contains the supplementary block; a left-end overlap pattern is where the left end of the supplementary block overlaps with the left end of the backbone block; a right-end overlap pattern is where the right end of the supplementary block overlaps with the right end of the backbone block; and an internal overlap pattern is where the supplementary block contains the backbone block. If the overlap pattern is complete... In full overlap mode, the skeleton block is divided into a front and back half, with the start position of the supplement block as the dividing line, and the supplement block is inserted between the front and back half. If the overlap mode is left overlap mode, the skeleton block is divided into a front and back half, with the end position of the supplement block as the dividing line, and the supplement block is inserted before the back half. If the overlap mode is right overlap mode, the skeleton block is divided into a front and back half, with the start position of the supplement block as the dividing line, and the supplement block is inserted after the front half. If the overlap mode is internal overlap mode, the supplement block replaces the skeleton block.
[0065] Specifically, the overlapping pattern identification of the positions that need to be supplemented can be broadly divided into four patterns.
[0066] (1) Fully contained mode: The supplementary block is completely inside the skeleton block; (2) Left-end overlap mode: The supplementary block starts from the left end of the skeleton block and extends outside the block; (3) Right-end overlap mode: The supplementary block starts from inside the skeleton block and extends to the right end; (4) Internal overlap mode: The skeleton block is completely contained within the supplement block.
[0067] In cases where multiple supplementary blocks may exist at the same insertion site, embodiments of this application can implement an intelligent merging algorithm. By constructing a site identifier (start position_end position), it can automatically detect duplicate insertion sites and merge the information of multiple supplementary blocks into a single string record.
[0068] After analyzing and observing a large amount of bacterial sequence evolution data, six insertion scenarios were identified in the sequences (within four patterns), and solutions were proposed for each scenario: Case 1: such as Figure 13 The case shown is where the supplementary block is completely inside the current skeleton block. In this case, the supplementary sequence needs to be inserted into the middle of the original skeleton block. Specifically, the first half of the output skeleton block (from the start to the start position of the supplementary block), the complete supplementary block sequence is inserted, and finally the second half of the output skeleton block (from the end position of the supplementary block to the end of the skeleton block) is output.
[0069] Scenario 2: such as Figure 14 As shown, the supplementary block starts from the beginning of the skeleton block and ends inside the skeleton block; in this case, there is no need to segment the front end of the skeleton block. Specifically: directly output the supplementary block sequence, output the remaining skeleton block portion (starting from the end position of the supplementary block), and adjust the subsequent coordinates according to the length of the supplementary block.
[0070] Scenario 3: such as Figure 15 As shown, the supplementary block starts inside the skeleton block and ends at the end of the skeleton block. In this case, the first half of the skeleton block and the supplementary block need to be output. Specifically, the first half of the skeleton block is output, the supplementary block sequence is inserted, and a flag is set to indicate that the current skeleton block processing is complete.
[0071] Case 4: Figure 16 As shown, the left end is inserted across boundaries. The supplementary block starts from the outside of the skeleton block and extends into the inside of the skeleton block. In this case, it is necessary to handle the overlapping area across boundaries. Specifically: output the complete sequence of supplementary blocks, output the remaining part of the skeleton block (starting from the end position of the supplementary block), and handle the cross-boundary calculation of coordinates.
[0072] Case 5: Figure 17 As shown, the right-end cross-boundary insertion occurs, with the supplementary block starting inside the skeleton block and extending to its outer edge. Specifically: the first half of the output skeleton block is used to insert the supplementary block sequence, and a flag is set. The supplementary block will affect the processing of subsequent skeleton blocks.
[0073] Case 6: For example Figure 18 As shown, the skeleton block is completely inside the supplement block. In this case, the current skeleton block is replaced by the supplement block.
[0074] Furthermore, in the following embodiments, newMap2[i] and newMap3[i] are the positions of the inserted sequence in the relative coordinate system; realBGst is the relative coordinate starting point of the current skeleton block, and realBGend is the relative coordinate ending point of the current skeleton block; source_st is the absolute coordinate starting point of the current skeleton block, and source_end is the absolute coordinate ending point of the current skeleton block; the precise coordinate mapping method in this embodiment is as follows: 1. Establish the relative coordinate system for the skeleton sequence; realBGst = 1 indicates the starting relative coordinates of the current backbone block; realBGend = 0 indicates the relative coordinates of the end of the current backbone block; The system updates the coordinate system after processing each skeleton block; realBGend = source_end - source_st + realBGst; In the next round: realBGst = realBGend + 1.
[0075] 2. Coordinate transformation during insertion: Convert the coordinates of the skeleton block in the relative coordinate system to the absolute coordinates of the insertion sequence. a = newMap2[i] - realBGst + source_st, representing the absolute coordinates of the starting position of insertion; b = newMap3[i] - realBGst + source_st, representing the absolute coordinates of the insertion end position.
[0076] For example, Block1: [1000-2000] has relative coordinates [1-1001]; Block2: Relative coordinates of [3000-4000] [1002-2002]; Block3: Relative coordinates of [5000-6000] [2003-3003]; If we now need to insert a sequence of length 200 at the relative coordinate 1500; Block2: source_st=3000, source_end=4000; realBGst = 1002 means that Block1 occupies 1001 relative coordinate positions; realBGend = 4000 - 3000 + 1002 = 2002. At this point, we can determine the relationship: 1002 < 1500 < 1700 < 2002 (which means it can be determined as case 1: complete insertion inside). Insertion sequence: newMap2[i]=1500, newMap3[i]=1700; a = 1500 - 1002 + 3000 = 3498; b = 1700 - 1002 + 3000 = 3698; Convert to absolute coordinates to determine the start and end points of the inserted sequence, and supplement the relevant sequences.
[0077] The bacterial paleogenomics analysis method proposed in this application, based on a phylogenetic tree, enables bottom-up analysis. Relying on the extant bacterial genome, the paleogenomics at each level can be solved progressively. Finally, with the help of outgroup strains within the bacterial community, the paleogenomics of the community is obtained. Therefore, this application fills a gap in related technologies for inferring paleogenomics. This tool helps researchers infer the key states of strains at various evolutionary nodes, study the gains and losses of bacterial gene-related changes, and provides a valuable tool for further analyzing the gains, losses, and origins of bacterial drug resistance genes, thus contributing to the depiction of bacterial evolutionary trajectories.
[0078] The following specific examples illustrate the bacterial ancient genome analysis method: Salmonella is an important pathogenic bacterium with over 2,600 serotypes, among which the main pathogenic to humans includes Salmonella typhi (Salmonella typhi). S. Typhi), Salmonella paratyphi ( S. Paratyphi A / B / C) can cause both typhoid and non-typhoid illnesses in humans. Typhoid fever, also known as enteric fever, is a serious invasive infection that primarily affects the blood and deep reticuloendothelial tissue. It is mainly transmitted through contaminated water and food, and typical symptoms include high fever, headache, vomiting, and diarrhea. Systematic analysis of Salmonella is available. S. enterica subsp. enterica The ancient genome of this subspecies was successfully deduced (4.05 Mb, containing 4.28 million nucleotides, encoding 4,187 genes) (e.g. Figure 19 Outer ring). Referring to GC skew, the origin of replication (oriC) and the terminator (terC) divide the AOC genus almost equally into two halves (e.g., Figure 19 The inner circle suggests the ancient conservation of Salmonella chromosomes and potential chromosome arm balance constraints.
[0079] To further analyze the genomic evolutionary history of Salmonella Typhi, this example collected 148 strains.S . Typhi, 8 strains S Paratyphi A and 5 outgroup strains (including S. Cerro strains CFSAN001588 and 87, S Saintpaul strain CFSAN004175, S Manchester ST278 and S The complete chromosome sequence of *Indica RKS3057* can be obtained quickly from a large amount of strain data. This application's embodiments can rapidly solve for all ancient genomes (...). Figure 20 (The computation time is only 48 minutes, and the computational efficiency increases linearly with the number of strains, without exponential growth.) A phylogenetic tree was constructed, and the ancestral chromosome sizes (AOCs) of each phylogenetic node were reconstructed. The results show that... S. Typhi and S. Paratyphi's most recent common ancestor (OPT) branched further from the common ancestor of cluster 2. S Paratyphi (P lineage) and S. Typhi (T lineage) Figure 21 The AOCs of OPT, PT (ancestor of the P lineage), P, and T were 4.04 Mb, 4.33 Mb, 4.37 Mb, and 4.47 Mb, respectively. Figure 21 Existing S. The AOCs size of the Typhi branch was stable (median 4.72 Mb), almost identical to that of existing strains (4.73 Mb). Figures 20-21 This indicates that the chromosome as a whole is stable, with no large-segment insertions or deletions. The AOCs of OPT, PT, P, and T are all equally divided in half by the origin-terminator axis (inferred from GC skewness), while... S. The AOCs at each node of Typhi also exhibited a similar pattern, suggesting the constraining effect of replication arm balance on genome evolution. Figure 21 Ancestor chromosome size refers to the ancient genome.
[0080] A comparison of AOCs according to their evolutionary paths shows that S. Paratyphi A and S. The Typhi chromosome underwent at least two large-scale inversions: one occurred in their common ancestor (called "PT inversion"), and the other occurred in... S. Typhi ancestors (called "T inversions") Figure 21 The T inversion is located within the PT inversion region, and both are centered at the replication end (ter) of the AOCs. The PT inversion exists in the PT ancestor AOCs and in existing AOCs. S.In the chromosome of Paratyphi A strain, and the existing S. More inversions were detected in the AOCs and chromosomes of the Typhi strain, including PT-inv1, PT-inv2, PT-inv3, Ter-1, Ter-2, Ori-1, and Ori-2. Figure 22 Based on the characteristics of inverted combinations, the existing S. Typhi strains can be divided into 6 genomic types (types 0-5) Figure 20 ,22). All large segment inversions were centered at the ter or ori locus, further confirming the role of chromosome arm balance. S. Typhi and S. Constraints on the evolution of the Paratyphi A genome ( Figure 22 ).
[0081] Further analysis of the inversion breakpoints revealed that, except for the ends of T inversions, Ter-1 inversions, and Ter-2 inversions, all other inversions terminated near the ribosomal RNA (rRNA) gene. Figure 23 PT-inv1 represents PT inversion type 1. Combined with the high conservation of Salmonella chromosome conformation and the fragmentary inversions involving major regions of the complete ter domain in Salmonella, the above observations further suggest that chromosome conformation may have a constraining effect on genomic inversions and may also be highly correlated with phenotypic changes such as drug resistance.
[0082] It should be noted that the method in this application embodiment can also be used for the analysis of other prokaryotes, such as Escherichia coli, Pseudomonas aeruginosa, Klebsiella pneumoniae, Akkermansia myxophilus, and Vibrio cholerae.
[0083] Next, the bacterial paleogenetic analysis apparatus according to the embodiments of this application is described with reference to the accompanying drawings.
[0084] Figure 24 This is a block diagram of a bacterial paleogenetic analysis device according to an embodiment of this application.
[0085] like Figure 24 As shown, the bacterial ancient genome analysis device 10 includes: a determination module 100, an extraction module 200, an analysis module 300, and a supplementary module 400.
[0086] The system comprises the following modules: 100, which determines the branch nodes of the ancient gene to be synthesized based on the phylogenetic tree of the bacterial strains; 200, which selects branch strains from the branch nodes of the ancient gene to be synthesized for comparison and extraction of homologous fragments; 300, which performs localization analysis on the homologous fragments of each strain, obtains orthologous fragments from the homologous fragments based on the localization analysis results, and splices the orthologous fragments to obtain the ancient genome skeleton of the bacteria; and 400, which obtains the exogroup gene sequences of the exogroup strains corresponding to the branch nodes, and supplements the ancient genome skeleton of the bacteria based on the exogroup gene sequences of the exogroup strains corresponding to the branch nodes.
[0087] It should be noted that the foregoing explanation of the bacterial paleogenomics analysis method embodiment also applies to the bacterial paleogenomics analysis device of this embodiment, and will not be repeated here.
[0088] The bacterial paleogenetic analysis device proposed in this application can perform bottom-up analysis based on a phylogenetic tree. Relying on the extant bacterial genome, it can solve the paleogenetic genome at each level, and finally, with the help of outgroup strains, obtain the paleogenetic genome of the bacterial community. Therefore, this application fills the gap in related technologies for inferring paleogenetic genomes. This tool can help researchers infer the key states of strains at various evolutionary nodes, study the gains and losses of bacterial gene-related changes, and provide a good tool for further analyzing the gains, losses, and origins of bacterial drug resistance genes, thus helping to depict the evolutionary trajectory of bacteria.
[0089] Figure 25 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 2201, the processor 2202, and the computer program stored on the memory 2201 and executable on the processor 2202.
[0090] When the processor 2202 executes the program, it implements the bacterial paleogeny analysis method provided in the above embodiments.
[0091] Furthermore, electronic devices also include: Communication interface 2203 is used for communication between memory 2201 and processor 2202.
[0092] The memory 2201 is used to store computer programs that can run on the processor 2202.
[0093] The memory 2201 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0094] If the memory 2201, processor 2202, and communication interface 2203 are implemented independently, then the communication interface 2203, memory 2201, and processor 2202 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 25 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0095] Optionally, in a specific implementation, if the memory 2201, processor 2202 and communication interface 2203 are integrated on a single chip, then the memory 2201, processor 2202 and communication interface 2203 can communicate with each other through an internal interface.
[0096] The processor 2202 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0097] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the bacterial paleogeny analysis method described above.
[0098] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0099] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0100] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0101] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.
[0102] Those skilled in the art will understand that all or part of the steps of the methods implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0103] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for analyzing bacterial ancient genomes, characterized in that, Includes the following steps: The branch nodes of the ancient gene to be synthesized are determined based on the phylogenetic tree of the bacterial strain data. Branch strains were selected from the branch nodes of the ancient gene to be synthesized for comparison and extraction of homologous fragments; The homologous fragments of each strain are localized and analyzed. Based on the localization analysis results, orthologous fragments are obtained from the homologous fragments, and the orthologous fragments are spliced together to obtain the ancient genome skeleton of the bacteria. Obtain the outgroup gene sequence of the outgroup strain of the branch strain corresponding to the branch node, and supplement the ancient genome skeleton of the bacteria based on the outgroup gene sequence of the outgroup strain of the branch strain corresponding to the branch node.
2. The bacterial archegenome analysis method according to claim 1, characterized in that, Before determining the branch nodes of the ancient gene to be synthesized based on the phylogenetic tree of bacterial strain data, the following steps are also included: Extract the base sequence and / or amino acid sequence of the bacterial strain from the bacterial strain data; The evolutionary relationships between organisms are calculated based on the base sequence and / or amino acid sequence, and the phylogenetic tree of the bacterial strains in the strain data is determined based on the evolutionary relationships between organisms.
3. The bacterial archegenome analysis method according to claim 1, characterized in that, The localization analysis of the homologous fragments for each strain includes: Identify the location of the homologous fragment within the genome for each strain; The homologous fragments are numbered according to their location within the genome; The homologous segments are numbered according to their lengths. Based on the sorting results, the homology and positional order of the homologous segments are analyzed sequentially. During the localization process, homologous segments with inconsistent homology are deleted until the analysis of all homologous segments is completed.
4. The bacterial archegenome analysis method according to claim 3, characterized in that, The step of obtaining orthologous fragments from the homologous fragments based on the localization analysis results, and splicing the orthologous fragments to obtain the ancient genome skeleton of the bacteria includes: Extract the retained homologous segments and their corresponding positional order from the location analysis results; The retained homologous segments are regarded as orthologous homologous segments among the homologous segments; The ancient genome skeleton of the bacteria was obtained by assembling the orthologous fragments according to the positional order.
5. The bacterial archegenome analysis method according to claim 1, characterized in that, The step of supplementing the ancient genome skeleton of the bacteria based on the outgroup gene sequence of the outgroup strain of the branch strain corresponding to the branch node includes: Obtain the backbone sequence of the ancient genome and the original gene sequence of the strain corresponding to the branch node; The original gene sequence was compared with the backbone sequence and the outpopulation gene sequence, respectively; The supplementary fragments from the comparison results are added to the ancient genome backbone of the bacteria, wherein the supplementary fragments are fragments that are not present in the comparison between the backbone sequence and the original gene sequence, and fragments that are present in the comparison between the outpopulation gene sequence and the original gene sequence.
6. The bacterial archegenome analysis method according to claim 5, characterized in that, The step of comparing the original gene sequence with the backbone sequence and the outpopulation gene sequence respectively also includes: Based on the comparison between the original gene sequence and the backbone sequence, an orthologous block file of the backbone sequence is generated. The gap regions in the original gene sequence are identified according to the orthologous block file of the backbone sequence, and the gap regions are used as backbone blocks. Based on the comparison between the original gene sequence and the outpopulation gene sequence, an orthologous block file of the outpopulation gene sequence is generated. Supplementary fragments are identified based on the orthologous block file of the outpopulation gene sequence, and the supplementary fragments are used as supplementary blocks. The comparison results are generated based on the skeleton block and the supplementary block.
7. The bacterial archegenome analysis method according to claim 6, characterized in that, Before adding supplementary fragments from the comparison results to the ancient genome backbone of the bacteria, the process also includes: Based on the orthologous block file, determine the first gap length in the original gene sequence and the second gap length in the backbone sequence, and calculate the ratio of the first gap length to the second gap length; The ancient genome skeleton of the bacteria is supplemented based on the first gap length, the second gap length, and the ratio. If the supplementation conditions are met, the ancient genome skeleton of the bacteria is supplemented based on the comparison results.
8. The bacterial archegenome analysis method according to claim 6 or 7, characterized in that, The step of adding supplementable fragments from the comparison results to the ancient genome backbone of the bacteria includes: Identify the relative positional relationship between the supplementary block and the skeleton block in the comparison results; The overlap pattern between the supplementary block and the skeleton block is determined based on the relative positional relationship. The overlap pattern includes a complete containment pattern, a left-end overlap pattern, a right-end overlap pattern, and an internal overlap pattern. The complete containment pattern is that the skeleton block contains the supplementary block. The left-end overlap pattern is that the supplementary block overlaps with the left end of the skeleton block. The right-end overlap pattern is that the supplementary block overlaps with the right end of the skeleton block. The internal overlap pattern is that the supplementary block contains the skeleton block. If the overlap mode is the complete containment mode, then the skeleton block is divided into a front half and a back half with the starting position of the supplementary block as the dividing line, and the supplementary block is inserted between the front half and the back half. If the overlap pattern is the left-end overlap pattern, then the skeleton block is divided into a front half and a back half with the end position of the supplementary block as the dividing line, and the supplementary block is inserted in front of the back half. If the overlap pattern is the right-end overlap pattern, then the skeleton block is divided into a front half and a back half, with the starting position of the supplementary block as the dividing line, and the supplementary block is inserted after the front half. If the overlap pattern is the internal overlap pattern, then the supplementary block replaces the skeleton block.
9. A bacterial ancient genome analysis device, characterized in that, include: The determination module is used to identify the branch nodes of the ancient gene to be synthesized based on the phylogenetic tree of the bacterial strain data. The extraction module is used to select branch strains from the branch nodes of the ancient gene to be synthesized for comparison and extraction of homologous fragments; The analysis module is used to perform localization analysis on the homologous fragments of each strain, obtain orthologous fragments from the homologous fragments based on the localization analysis results, and splice the orthologous fragments to obtain the ancient genome skeleton of the bacteria. The supplementary module is used to obtain the exogroup gene sequence of the exogroup strain of the branch strain corresponding to the branch node, and supplement the ancient genome skeleton of the bacteria based on the exogroup gene sequence of the exogroup strain of the branch strain corresponding to the branch node.
10. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the bacterial paleogenetic analysis method according to any one of claims 1-8.
Citation Information
Patent Citations
Population genetic evolution map and construction method thereof
CN110910959A
Generic genome construction method and device based on phylogenetic tree
CN111477281A
Method for analyzing generic genome of bacillus lamellii
CN117174181A