Methods, apparatus and equipment for bacterial paleogenomics analysis

By combining bacterial strain phylogenetic tree and homologous fragment analysis with outgroup gene sequence supplementation, the speed and accuracy problems of bacterial genome data analysis in existing technologies have been solved, achieving rapid and accurate ancient genome reconstruction, preserving bacterial ancestral information, and inferring bacterial evolutionary relationships.

CN121034397BActive Publication Date: 2026-01-30SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511536798.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-01-30
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately analyze large volumes of bacterial genome data, particularly the precise construction of bacterial evolutionary relationships, resulting in excessive loss of ancestral information. Existing algorithms are also unable to effectively handle complex evolutionary events such as insertions, deletions, and duplications.

Method used

By determining the phylogenetic tree of bacterial strains, homologous fragments are extracted, orthologous fragments are spliced ​​together to form an ancient genome backbone, and outgroup gene sequences are used to supplement the backbone. By combining MHB and backbone supplementation algorithms, the preservation of bacterial ancestral information is maximized.

Benefits of technology

It enables rapid and accurate analysis of large volumes of bacterial genome data, maximizes the preservation of bacterial ancestral information, fills the gap in existing technologies for inferring ancient genomes, and helps study the evolutionary trajectory of bacteria.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034397B_ABST
    Figure CN121034397B_ABST
Patent Text Reader

Abstract

This application relates to the field of genome evolution analysis technology, and particularly to a method, apparatus, and device for bacterial ancient genome analysis. The method includes: determining the branch nodes of the ancient gene to be synthesized based on the phylogenetic tree of bacterial strain data; selecting branch strains from the branch nodes of the ancient gene to be synthesized for alignment and extraction of homologous fragments; performing localization analysis on the homologous fragments of each strain, obtaining orthologous fragments from the homologous fragments based on the localization analysis results, and splicing the orthologous fragments to obtain the bacterial ancient genome backbone; obtaining the outgroup gene sequences of the outgroup strains of the corresponding branch strains, and supplementing the bacterial ancient genome backbone based on the outgroup gene sequences of the corresponding branch strains. This solves the problem of how to quickly and accurately analyze large amounts of bacterial community data to obtain the ancient genomes of the bacterial community at various developmental nodes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of genome evolution analysis, in particular to a bacterial paleogenomic analysis method, device and equipment. BACKGROUND

[0002] In the research of bacterial evolution, it is of great significance to clarify the genomic evolution trajectory of most bacteria (such as Salmonella typhi) in the history of species evolution and the process of co-evolution with humans. The ideal research scheme is to analyze the bacterial paleogenome at each major development node, and track the evolution path of gene recombination by comparing the paleogenomes of adjacent nodes. Although the rapid development of whole genome sequencing today provides a high-resolution tool for evolutionary tracing, the existing means of studying bacterial evolution still has great limitations, and the accurate construction of bacterial evolutionary relationships still faces multiple obstacles. SUMMARY

[0003] The present application provides a bacterial paleogenomic analysis method, device and equipment to solve the problem of how to quickly and accurately analyze a large amount of bacterial community data to obtain the paleogenome of the bacterial community at each development node.

[0004] The first aspect of the present application provides a bacterial paleogenomic analysis method, comprising the following steps: determining a branch node of a to-be-synthesized paleogene according to a strain phylogenetic tree in bacterial strain data; selecting branch strains from the branch node of the to-be-synthesized paleogene for comparison and extraction of homologous fragments; performing positioning analysis on the homologous fragments of each strain, obtaining orthologous fragments from the homologous fragments according to the positioning analysis results, and splicing the orthologous fragments to obtain a bacterial paleogenomic skeleton; obtaining outgroup gene sequences of outgroup strains of the branch strains corresponding to the branch node, and supplementing the bacterial paleogenomic skeleton according to the outgroup gene sequences of the outgroup strains of the branch strains corresponding to the branch node.

[0005] Optionally, before determining the branch node of the to-be-synthesized paleogene according to the strain phylogenetic tree in the bacterial strain data, the method further comprises: extracting base sequences and / or amino acid sequences of the strains in the bacterial strain data; calculating the evolutionary relationship between organisms according to the base sequences and / or amino acid sequences, and determining the strain phylogenetic tree in the bacterial strain data according to the evolutionary relationship between organisms.

[0006] Optionally, the positioning analysis on the homologous fragments of each strain comprises: identifying the position of the homologous fragments of each strain in the genome; numbering the homologous fragments according to the position of the homologous fragments in the genome; sorting the numbers of all homologous fragments according to the length of the homologous fragments, and analyzing the homology and positional order relationship between the homologous fragments in sequence according to the sorting result, and deleting the homologous fragments with inconsistent homology in the positioning process until all homologous fragments are analyzed.

[0007] Optionally, the orthologous fragments are obtained from the homologous fragments according to the positioning analysis result, and the paleogenomic skeleton of the bacteria is obtained by splicing the orthologous fragments, including: extracting the reserved homologous fragments and the corresponding positional order relationship from the positioning analysis result; taking the reserved homologous fragments as the orthologous fragments in the homologous fragments; and splicing the orthologous fragments according to the positional order relationship to obtain the paleogenomic skeleton of the bacteria.

[0008] Optionally, the paleogenomic skeleton of the bacteria is supplemented according to the outgroup gene sequence of the outgroup strain of the branch strain corresponding to the branch node, including: obtaining the skeleton sequence of the paleogenomic skeleton and the original gene sequence of the branch strain corresponding to the branch node; comparing the original gene sequence with the skeleton sequence and the outgroup gene sequence respectively; and supplementing the supplementable fragments in the comparison result to the paleogenomic skeleton of the bacteria, wherein the supplementable fragments are fragments that do not exist in the comparison of the original gene sequence and the skeleton sequence and exist in the comparison of the outgroup gene sequence and the original gene sequence.

[0009] Optionally, the comparison of the original gene sequence with the skeleton sequence and the outgroup gene sequence further includes: generating an orthologous block file of the skeleton sequence based on the comparison of the original gene sequence and the skeleton sequence, identifying a gap region in the original gene sequence according to the orthologous block file of the skeleton sequence, and taking the gap region as a skeleton block; generating an orthologous block file of the outgroup gene sequence based on the comparison of the original gene sequence and the outgroup gene sequence, identifying the supplementable fragments according to the orthologous block file of the outgroup gene sequence, and taking the supplementable fragments as supplement blocks; and generating the comparison result according to the skeleton block and the supplement blocks.

[0010] Optionally, before the supplementable fragments in the comparison result are supplemented to the paleogenomic skeleton of the bacteria, it further includes: determining a first gap length in the original gene sequence and a second gap length in the skeleton sequence according to the orthologous block file, calculating a ratio of the first gap length and the second gap length; determining whether a supplement condition of the paleogenomic skeleton is met according to the first gap length, the second gap length and the ratio, and if the supplement condition is met, supplementing the paleogenomic skeleton of the bacteria according to the comparison result.

[0011] Optionally, the complementary fragment in the comparison result is complemented to the ancient genome skeleton of the bacteria, comprising: identifying relative position relationship of the complement block and the skeleton block in the comparison result; determining an overlap mode between the complement block and the skeleton block according to the relative position relationship, the overlap mode comprising a complete containing mode, a left end overlap mode, a right end overlap mode and an internal overlap mode, the complete containing mode being that the skeleton block contains the complement block, the left end overlap mode being that the complement block overlaps the left end of the skeleton block, the right end overlap mode being that the complement block overlaps the right end of the skeleton block, and the internal overlap mode being that the complement block contains the skeleton block; if the overlap mode is the complete containing mode, dividing the skeleton block into a first half and a second half with the start position of the complement block as a boundary, and inserting the complement block between the first half and the second half; if the overlap mode is the left end overlap mode, dividing the skeleton block into a first half and a second half with the end position of the complement block as a boundary, and inserting the complement block in front of the second half; if the overlap mode is the right end overlap mode, dividing the skeleton block into a first half and a second half with the start position of the complement block as a boundary, and inserting the complement block behind the first half; and if the overlap mode is the internal overlap mode, replacing the skeleton block with the complement block.

[0012] The second aspect embodiment of the present application provides a bacterial ancient genome analysis device, comprising: a determination module configured to determine a branch node of an ancient gene to be synthesized according to a strain phylogenetic tree in strain data of bacteria; an extraction module configured to select branch strains from between branch nodes of the ancient gene to be synthesized to perform comparison and extraction of homologous fragments; an analysis module configured to perform positioning analysis on the homologous fragments of each strain, obtain orthologous fragments from the homologous fragments according to a positioning analysis result, and splice the orthologous fragments to obtain an ancient genome skeleton of the bacteria; and a complement module configured to obtain outgroup gene sequences of outgroup strains of the branch strains corresponding to the branch node, and complement the ancient genome skeleton of the bacteria according to the outgroup gene sequences of the outgroup strains of the branch strains corresponding to the branch node.

[0013] The third aspect embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor executes the program to implement the bacterial ancient genome analysis method of the above embodiments.

[0014] Therefore, the present application includes the following beneficial effects:

[0015] The embodiments of the present application can rely on the current genome of the strain, quickly and accurately analyze a large amount of strain data, and then rely on the phylogenetic tree of the strain to solve the ancient genome of the strain at each development node. The position information between genes is considered at the same time to maximize the retention of the ancestral information of the bacteria, so as to effectively reduce the loss of the ancestral information of the bacteria.

[0016] Additional aspects and advantages of the application will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following and / or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0017] The foregoing and / or additional aspects and advantages of the present application are achieved by providing a method for analyzing a bacterial ancient genome, comprising the steps of: obtaining a plurality of ancient genome fragments from a sample; determining a position of each of the plurality of ancient genome fragments; and constructing a phylogenetic tree based on the plurality of ancient genome fragments.

[0018] Figure 1 A flow chart of a method for analyzing a bacterial ancient genome according to an embodiment of the present application;

[0019] Figure 2 An example of a phylogenetic tree according to an embodiment of the present application;

[0020] Figure 3 An example of a homologous fragment according to an embodiment of the present application;

[0021] Figure 4 An example of ordering homologous fragments according to an embodiment of the present application;

[0022] Figure 5 An example of determining a position of a homologous fragment according to an embodiment of the present application;

[0023] Figure 6 An example of an ancient genome skeleton according to an embodiment of the present application;

[0024] Figure 7 An example of a genome according to an embodiment of the present application;

[0025] Figure 8 An example of a position of a fragment to be supplemented according to an embodiment of the present application;

[0026] Figure 9 An example of a position of a fragment without a supplement according to an embodiment of the present application;

[0027] Figure 10 An example of a fragment explanation diagram of an example according to an embodiment of the present application;

[0028] Figure 11 An example of a ternary tree according to an embodiment of the present application;

[0029] Figure 12 An example of a binary tree according to an embodiment of the present application;

[0030] Figure 13 An example of an internal complete insertion according to an embodiment of the present application;

[0031] Figure 14 An example of a left end overlap according to an embodiment of the present application;

[0032] Figure 15 FIG. 8 is a diagram illustrating a right end overlap example according to an embodiment of the present application;

[0033] Figure 16 FIG. 9 is a diagram illustrating a left end cross-border insertion example according to an embodiment of the present application;

[0034] Figure 17 FIG. 10 is a diagram illustrating a right end cross-border insertion example according to an embodiment of the present application;

[0035] Figure 18 FIG. 11 is a diagram illustrating a full cross-border coverage example according to an embodiment of the present application;

[0036] Figure 19 FIG. 12 is a diagram illustrating a Salmonella genome example according to an embodiment of the present application;

[0037] Figure 20 FIG. 13 is a diagram illustrating an ancient genome example according to an embodiment of the present application;

[0038] Figure 21 FIG. 14 is a diagram illustrating a strain genome evolution history analysis example according to an embodiment of the present application;

[0039] Figure 22 FIG. 15 is a diagram illustrating an inversion detected in a strain and a chromosome according to an embodiment of the present application;

[0040] Figure 23 FIG. 16 is a diagram illustrating an inversion end according to an embodiment of the present application;

[0041] Figure 24 FIG. 17 is a diagram illustrating a bacterial ancient genome analysis device example according to an embodiment of the present application;

[0042] Figure 25 FIG. 18 is a diagram illustrating an electronic device structure example according to an embodiment of the present application. DETAILED DESCRIPTION

[0043] Embodiments of the present application are described in detail below with reference to the attached drawings, which are presented as examples in which the same or similar components have the same or similar designations. The embodiments described below are examples intended to provide an explanation of the present application and are not intended to restrict the present application.

[0044] Examples of a related art ancestral reconstruction method are as follows:

[0045] (1) An ancestral reconstruction method based on a local optimal solution:

[0046] BPAnalysis is computationally intensive and time-consuming, and it can take 200 years to get results for a dataset containing 13 genomes and 105 gene fragments. GRAPPA proposes a new implementation of breakpoint analysis that can handle larger data that BPAnalysis cannot solve, but GRAPPA still uses a simplified evolutionary model and can only handle limited types of genome rearrangement events. The single cut or join (SCJ) method is still based on genome rearrangement distance, although it has very low false positive rate and very conservative genome reconstruction results, it still cannot accurately restore the phylogenetic tree and ancestral genome, and only 60% to 95% of the phylogenetic branches and 50% to 85% of the ancestral adjacency relationships can be restored.

[0047] (2) Ancestral reconstruction method based on global optimal solution:

[0048] Although the GASTS method overcomes the key initialization problem of BPAnalysis and GRAPPA, it can only handle simplified genome datasets with the same genome content and unique coaxial blocks, it cannot consider gene duplication, gene loss and gain events, and it cannot output intermediate ancestral genomes during ancestral reconstruction.

[0049] (3) Ancestral reconstruction using continuous ancestral regions (CAR):

[0050] The main representatives are InferCARs and InferCARsPro methods. InferCARsPro requires a given phylogenetic tree and known branch length to reconstruct the ancestral genome, which is difficult to meet when applied to real genome datasets. And the local parsimony-based algorithm can also ignore many real adjacencies that should exist in the ancestral genome. Both methods can only solve small-scale phylogenetic tree problems because they need to assume that the input genome has a known phylogenetic tree. Both methods still cannot handle complex evolutionary events, including insertions, deletions, and duplications. ProCARs is a parameter-free method that does not require branch length of the phylogenetic tree. However, ProCARs can only consider genomes with the same genome content and non-repetitive blocks, and does not allow insertions and deletions. And in the ancestral reconstruction of real genome datasets, the highest resolution of ProCARs is only 100kb, even lower than InferCARs and InferCARsPro.

[0051] In summary, most of the ancestral reconstruction methods in the related art can only handle genome datasets with identical content and unique genomic signatures. They can mostly only handle genome rearrangement events, which are only a part of the complex evolutionary events in real evolution. MGRA2 is one of the few distance / event-based methods that can handle various evolutionary events. For homology / adjacency-based methods, only PMAG++, Gapped Adjacency, and ANGES handle complex evolutionary events including insertions, deletions, and duplications, but the above methods are mostly for eukaryotic genomes and are not suitable for prokaryotic genomes such as bacteria.

[0052] Therefore, there is still a lack of effective method and system for calculating bacterial paleogenomic sequences for bacteria and other prokaryotic organisms. The existing algorithms have poor results and accuracy in analyzing small circular DNA of bacterial genes, and the ancestral information is lost too much. There is currently no bacterial paleogenomic reconstruction or inference tool, and there is no systematic method to solve it. Most of them are manually assembled, and it is difficult to quickly analyze a large batch of bacterial genome data, which is an important problem to be solved in studying the evolutionary trajectory of bacterial communities and exploring the gain and loss of strain genes.

[0053] To solve one of the above problems, the embodiments of the present application propose a method for quickly and accurately solving bacterial paleogenomes, and specifically propose a bacterial paleogenomic analysis method, device and equipment. The device, device and equipment of the embodiments of the present application are described below with reference to the drawings.

[0054] Specifically, Figure 1 A flowchart of a bacterial paleogenomic analysis method according to an embodiment of the present application is shown in FIG. 1.

[0055] As Figure 1 shown, the bacterial paleogenomic analysis method includes the following steps:

[0056] In step S101, a branch node of a paleogene to be synthesized is determined according to a strain phylogenetic tree in bacterial strain data.

[0057] It can be understood that the embodiments of the present application can determine the branch node closest to the paleogene to be synthesized according to the phylogenetic tree, and the bacterial strain data can include data of multiple strains, such as data of Salmonella and the like.

[0058] In the embodiments of the present application, before determining the branch node of the paleogene to be synthesized according to the strain phylogenetic tree in the bacterial strain data, it further includes: extracting the base sequence and / or amino acid sequence of the strain in the bacterial strain data; calculating the evolutionary relationship between organisms according to the base sequence and / or amino acid sequence, and determining the strain phylogenetic tree in the bacterial strain data according to the evolutionary relationship between organisms.

[0059] It can be understood that the embodiments of the present application can be based on the genome of the flora, and then the branch phylogenetic analysis can be performed. The branch phylogenetic analysis is a method for studying the evolution and system classification of species or sequences. Generally, the research object is a base sequence or an amino acid sequence. The evolutionary relationship between organisms is calculated by mathematical algorithm, and the system evolution tree can be used to determine the genetic relationship between each species or gene, and determine the branch node closest to the synthetic ancient gene.

[0060] For example, for the flora (A, B, and D are bacterial evolution nodes, and C is an ancient gene to be synthesized), the evolutionary relationship can be represented as a system evolution tree as shown in Figure 2 According to the system evolution tree relationship, it can be confirmed that the closest evolution nodes of the synthetic ancient gene C are A and B.

[0061] In step S102, branch strains are selected from the branch nodes of the synthetic ancient gene for comparison and extraction of homologous fragments.

[0062] It can be understood that the embodiments of the present application can select representative strains between two sub-branches of the synthetic ancient gene group for comparison and extraction of homologous fragments. The homologous fragments refer to gene sequences derived from a common ancestor.

[0063] In step S103, the homologous fragments of each strain are subjected to positioning analysis, and the orthologous fragments are obtained from the homologous fragments according to the positioning analysis results. The orthologous fragments are spliced to obtain the ancient genome skeleton of the bacteria.

[0064] It can be understood that, due to various rearrangement events in the evolution of the bacterial genome, the homologous fragments between the two chromosomes are not necessarily truly homologous. Therefore, the embodiments of the present application position the homologous fragments between the two bacterial chromosomes based on the MHB algorithm to obtain the orthologous genome skeleton, so that the current genome of the strain can be relied on to consider the position information between genes when solving, maximize the retention of the ancestral information of the bacteria, quickly and accurately analyze a large amount of flora data, and further rely on the system evolution tree of the strain to solve the ancient genome of the flora at each development node.

[0065] In the embodiments of the present application, the positioning analysis of the homologous fragments of each strain includes: identifying the position of the homologous fragments of each strain in the genome; numbering the homologous fragments according to the position of the homologous fragments in the genome; sorting the numbers of all homologous fragments according to the length of the homologous fragments; and analyzing the homology and position sequence relationship between the homologous fragments according to the sorting results. In the positioning process, the homologous fragments with inconsistent homology are deleted until all homologous fragments are analyzed.

[0066] It can be understood that the content of the MHB algorithm includes: numbering the homologous fragments of each strain according to the position of the homologous fragments in the genome, and sorting according to the length of each fragment, and analyzing the homology and position sequence relationship between the homologous fragments in turn according to the sorting result, for example, the sorting result determines the size relationship of the homologous fragments, and the homologous fragments are analyzed in turn in the order from large to small, the homologous fragments to be analyzed are compared with the homologous fragments that have been analyzed for homology, if consistent, the position sequence relationship of the homologous fragments is determined, and the homologous fragments with inconsistent homology are deleted in the positioning process, until all the homologous fragments are analyzed.

[0067] In the embodiment of the present application, the orthologous fragments are obtained from the homologous fragments according to the positioning analysis result, and the ancient genomic skeleton of the bacteria is obtained by splicing the orthologous fragments, including: extracting the retained homologous fragments and the corresponding position sequence relationship from the positioning analysis result; taking the retained homologous fragments as the orthologous fragments in the homologous fragments; and splicing the orthologous fragments according to the position sequence relationship to obtain the ancient genomic skeleton of the bacteria.

[0068] It can be understood that after the positioning analysis of the homologous fragments, the retained homologous fragments and the corresponding position sequence relationship can be extracted from the positioning analysis result, and the retained homologous fragments can be spliced according to the position sequence relationship, so as to obtain the ancient genomic skeleton of the bacteria.

[0069] For example, in order to further illustrate the positioning analysis process of the MHB algorithm, taking the homologous backbone between the two chromosomes strA and strB shown in Figure 3 as an example, in order to analyze the homologous backbone between the two chromosomes strA and strB shown in Figure 3 , the homologous fragments of each strain are numbered according to the position of the homologous fragments in the genome, and sorted according to the length of each fragment, and the sorting result is shown in Figure 4 , taking six homologous fragments as an example, the homologous blocks in the figure are the homologous fragments.

[0070] As shown in Figure 5As shown, the positioning relationship of the longest (LI) and the second longest (L2) fragments in the two genomes is observed and compared. If the positioning is consistent between the two genomes (as is usually the case), the third longest pair (L3) is included, and the positioning relationship of the third longest pair (L3) and LI / L2 and is compared. If the relationship is consistent between the genomes, L3 is retained; if not, L3 is deleted. The composition and positional order of the homologous backbone fragments are then updated. Next, further checking is performed on the L4 block, and the homology and positional order relationship of L4 and all of the orthologous blocks that have been determined previously are compared between the genomes, and similarly, if L4 is consistent with all of the orthologous blocks that have been determined previously, it is retained, and if there is inconsistency, it is deleted and the next one is jumped to. These steps are repeated until the shortest homologous block pair is analyzed. Finally, all of the orthologous fragments obtained are connected, and the bacterial paleogenomic backbone as shown in Figure 6

[0071] In step S104, the outgroup gene sequence of the outgroup strain corresponding to the branch node is obtained, and the paleogenomic backbone of the bacterium is supplemented according to the outgroup gene sequence of the outgroup strain corresponding to the branch node.

[0072] It can be understood that the outgroup strain corresponding to the branch node is the genome of the close outgroup strain or the genome of the most recent ancestor. The embodiments of the present application can compare the genome of the close outgroup strain or the genome of the most recent ancestor with the representative strain of any branch respectively, extract homologous fragments, further repair the paleogenomic backbone, and compare and repair the paleogenomic backbone of the representative strain of the other branch in the same way; the above steps are solved from the bottom of the phylogenetic tree to the top, and the oldest gene of the target strain group is obtained by relying on the outgroup strain of the bacterial community.

[0073] In the embodiments of the present application, the paleogenomic backbone of the bacterium is supplemented according to the outgroup gene sequence of the outgroup strain corresponding to the branch node, including: obtaining the backbone sequence of the paleogenomic backbone and the original gene sequence of the branch node corresponding to the branch strain; comparing the original gene sequence with the backbone sequence and the outgroup gene sequence respectively; and supplementing the supplementable fragments in the comparison results to the paleogenomic backbone of the bacterium, wherein the supplementable fragments are fragments that do not exist in the comparison of the original gene sequence and the backbone sequence, and exist in the comparison of the outgroup gene sequence and the original gene sequence.

[0074] ​It can be understood that the embodiments of the present application utilize the backbone supplement algorithm to compare and analyze the genomes of two strains, and to infer the common ancestral gene sequence. In this process, the orthologous fragments come from the common ancestor, and the non-homologous fragments can be variant or exogenous fragments. In the evolutionary process of the strains, the outgroup is more closely related to the ancestral gene, and therefore, the same level of outgroup gene sequence corresponding to the ancestral gene can be introduced as a reference, and by supplementing the outgroup gene sequence and the part common to the two strains, the genome backbone can be further improved.

[0075] Specifically, as shown in Figure 7 , the embodiments of the present application obtain sequence A and sequence B of the branch strain corresponding to the branch node, the ancestral genome backbone BG obtained by sequence A and sequence B, and the outgroup gene sequence out. The ancestral genome backbone BG can be aligned with sequence A, and the outgroup gene sequence out can be aligned with sequence A, to find the position to be supplemented, such as Figure 8 , as shown in the blank area in the figure.

[0076] If it is found in the alignment that there is a fragment in the alignment of the outgroup sequence out with sequence A or sequence B in the corresponding region, and it is determined that there is no fragment in the alignment of the backbone with sequence A or sequence B, then this fragment can be supplemented to the corresponding region. If it is found in the alignment with sequence A or sequence B that there is this fragment, but not in the alignment with the outgroup gene sequence out, then this fragment is not inserted (the outgroup strain is older, and is more closely related to the ancestral gene in evolutionary relationship, which follows the rules), and there is no supplementable fragment, and the situation is as shown in Figure 9 .

[0077] It should be noted that Figure 7 , Figure 8 and Figure 9 fragments are explained as shown in Figure 10 .

[0078] In case one as shown in Figure 11 , the result is obtained by supplementing the outgroup once, and in case two as shown in Figure 12 , the result can be obtained by supplementing the outgroup twice. On the basis of the above, one round of supplementation can be added, that is, BG3 is aligned with sequence A, the outgroup gene sequence out is aligned with sequence A, BG4 is aligned with sequence A, and the outgroup gene sequence out is aligned with sequence B.

[0079] In the embodiment of the present application, the comparison of the original gene sequence with the backbone sequence and the outgroup gene sequence also includes: generating an orthologous block file of the backbone sequence based on the comparison of the original gene sequence with the backbone sequence, identifying a gap region in the original gene sequence according to the orthologous block file of the backbone sequence, and taking the gap region as a backbone block; generating an orthologous block file of the outgroup gene sequence based on the comparison of the original gene sequence with the outgroup gene sequence, identifying a complementable fragment according to the orthologous block file of the outgroup gene sequence, and taking the complementable fragment as a complement block; and generating a comparison result according to the backbone block and the complement block.

[0080] It can be understood that, when the backbone complement algorithm performs sequence complement, the orthologous block file of the original gene sequence is first scanned, and then a gap length comparison is performed to identify a region in the source species that is not fully covered by the backbone genome using the orthologous block file. The orthologous block file of the backbone sequence can be understood as a homologous backbone file as shown in Figure 6 .

[0081] It should be noted that the backbone sequence and the backbone file are different. The backbone file is a type of file maintained in the system with the extension.backbone, and the backbone sequence is a sequence file generated according to the backbone file.

[0082] In the embodiment of the present application, before the complementable fragment in the comparison result is complemented to the ancient genome backbone of the bacteria, it further includes: determining a first gap length in the original gene sequence and a second gap length in the backbone sequence according to the orthologous block file, calculating a ratio of the first gap length and the second gap length; determining whether a complement condition of the ancient genome backbone is met according to the first gap length, the second gap length, and the ratio, and if the complement condition is met, complementing the ancient genome backbone of the bacteria according to the comparison result.

[0083] It can be understood that, in the embodiment of the present application, two key judgment standards are used to determine whether the complement condition of the ancient genome backbone is met. If the following judgment standards are met, it is determined that the complement condition is met. The two key judgment standards are as follows: when the length of the corresponding region of the backbone is zero and the length of the original sequence region exceeds a minimum length threshold of a number of base pairs, it is determined that the region is a completely missing region; or when the length ratio of the corresponding region of the backbone and the sequence is less than a preset value and the length of the source sequence region exceeds the minimum length threshold, it is determined that the region is a significantly unbalanced region and needs to be complemented. The preset value can be set to 0.1 or 0.2, and the minimum length threshold is to ensure that the complemented sequence fragment has biological significance, for example, it can be set to 1000 bp or 1100 bp.

[0084] For example, when the length of the skeleton corresponding region is zero (i.e. the second gap length span2 = 0) and the length of the original sequence region (the first gap length span1) is more than 1000 base pairs, it is determined as a completely missing region; or when the length ratio of the skeleton and sequence corresponding region (span2 / span1) is less than 0.1 and the length of the source sequence region is more than 1000 base pairs (i.e. span1 > 1000), it is determined as a significantly unbalanced region and needs to be completed (i.e. the skeleton block is identified).

[0085] The process of sequence completion based on the skeleton completion algorithm will be described below through a specific example, as follows:

[0086] The skeleton file example is shown in Table 1.

[0087] Table 1

[0088]

[0089] In the coordinates of the skeleton file, it is the absolute coordinates of the source sequence. The relative coordinates used in the sequence file will be discussed later. The relative coordinates of the skeleton sequence are all relative to the source sequence, i.e. sequence A.

[0090] The identification process is described in accordance with the type of ternary tree in the second step of the BP algorithm Figure 13 First, the comparison between the skeleton sequence and sequence A is performed. In this process, a skeleton file will be generated, and the content of the skeleton file is similar to Table 2.

[0091] Table 2

[0092]

[0093] Based on Table 2, it can be known that there is a region of about 10 kb on sequence A between 20000 and 30000, and the corresponding region is only 2 (from 20000 to 20002), which is a typical "gap" and can be called a "skeleton block". Then the comparison between the outgroup and sequence A is performed, and in this process, a similar file will be generated, and an example is shown in Table 3.

[0094] Table 3

[0095]

[0096] Based on Table 3, it can be known that the homologous block labeled in this file generated by comparing the outgroup and sequence A just covers the "skeleton block" region found above, which can be regarded as a "completion block".

[0097] Further, the application embodiment can adopt the following gap identification logic:

[0098] Read the file generated by comparison of the skeleton sequence and sequence A.

[0099] When reading the second line (30001 40000...), it calculates the gap between the previous block.

[0100] span1 (the length of the gap in A) = 30000 - 20000 = 10000

[0101] span2 (the length of the gap in the skeleton sequence) = 20001 - 20000 = 1

[0102] Check conditions span1>1000 (10000>1000, satisfied) and spanRatio<0.1 (0 / 10000<0.1, 0.1 is an example of a preset value, so it is determined that the patch needs to be filled in).

[0103] Thus, a "skeleton block" to be supplemented is identified, with coordinates on A being (20000, 30000).

[0104] For convenience of marking, the application records the following variables:

[0105] newMap1[1] = "Source sequence annotation \22000\28000\patch block\and who is homologous with \n".

[0106] In the embodiment of the application, the supplementable segment in the comparison result is supplemented to the bacterial ancient genome skeleton, including: identifying the relative positional relationship of the supplement block and the skeleton block in the comparison result; determining the overlap mode between the supplement block and the skeleton block according to the relative positional relationship, the overlap mode including a complete containing mode, a left-end overlap mode, a right-end overlap mode and an internal overlap mode, the complete containing mode being that the skeleton block contains the supplement block, the left-end overlap mode being that the supplement block overlaps the left end of the skeleton block, the right-end overlap mode being that the supplement block overlaps the right end of the skeleton block, and the internal overlap mode being that the supplement block contains the skeleton block; if the overlap mode is the complete containing mode, then taking the start position of the supplement block as a demarcation line, the skeleton block is divided into a first half and a second half, and the supplement block is inserted between the first half and the second half; if the overlap mode is the left-end overlap mode, then taking the end position of the supplement block as a demarcation line, the skeleton block is divided into a first half and a second half, and the supplement block is inserted in front of the second half; if the overlap mode is the right-end overlap mode, then taking the start position of the supplement block as a demarcation line, the skeleton block is divided into a first half and a second half, and the supplement block is inserted behind the first half; and if the overlap mode is the internal overlap mode, then the supplement block is used to replace the skeleton block.

[0107] Specifically, the overlapping pattern identification of the positions that need to be supplemented can be broadly divided into four patterns.

[0108] (1) Fully contained mode: The supplementary block is completely inside the skeleton block;

[0109] (2) Left-end overlap mode: The supplementary block starts from the left end of the skeleton block and extends outside the block;

[0110] (3) Right-end overlap mode: The supplementary block starts from inside the skeleton block and extends to the right end;

[0111] (4) Internal overlap mode: The skeleton block is completely contained within the supplement block.

[0112] In cases where multiple supplementary blocks may exist at the same insertion site, embodiments of this application can implement an intelligent merging algorithm. By constructing a site identifier (start position_end position), it can automatically detect duplicate insertion sites and merge the information of multiple supplementary blocks into a single string record.

[0113] After analyzing and observing a large amount of bacterial sequence evolution data, six insertion scenarios were identified in the sequences (within four patterns), and solutions were proposed for each scenario:

[0114] Case 1: such as Figure 13 The case shown is where the supplementary block is completely inside the current skeleton block. In this case, the supplementary sequence needs to be inserted into the middle of the original skeleton block. Specifically, the first half of the output skeleton block (from the start to the start position of the supplementary block), the complete supplementary block sequence is inserted, and finally the second half of the output skeleton block (from the end position of the supplementary block to the end of the skeleton block) is output.

[0115] Scenario 2: such as Figure 14 As shown, the supplementary block starts from the beginning of the skeleton block and ends inside the skeleton block; in this case, there is no need to segment the front end of the skeleton block. Specifically: directly output the supplementary block sequence, output the remaining skeleton block portion (starting from the end position of the supplementary block), and adjust the subsequent coordinates according to the length of the supplementary block.

[0116] Scenario 3: such as Figure 15 As shown, the supplementary block starts inside the skeleton block and ends at the end of the skeleton block. In this case, the first half of the skeleton block and the supplementary block need to be output. Specifically, the first half of the skeleton block is output, the supplementary block sequence is inserted, and a flag is set to indicate that the current skeleton block processing is complete.

[0117] Case 4: Figure 16As shown, left end cross-border insertion. The supplementary block starts from the outside of the skeleton block and extends to the inside of the skeleton block, which needs to handle the cross-border overlapping area. Specifically: output the complete supplementary block sequence, output the remaining part of the skeleton block (starting from the end position of the supplementary block) and handle the cross-border calculation of coordinates.

[0118] Case 5: As shown, right end cross-border insertion, the supplementary block starts from the inside of the skeleton block and extends to the outside of the skeleton block. Specifically: output the first half of the skeleton block, insert the supplementary block sequence, set the flag bit, and the supplementary block will affect the processing of the subsequent skeleton block. Figure 17

[0119] Case 6: As shown, the skeleton block is completely inside the supplementary block, in which case the current skeleton block is replaced by the supplementary block. Figure 18

[0120] Further, in the following embodiments, newMap2[i], newMap3[i] are the positions of the insertion sequence in the relative coordinate system; realBGst is the relative coordinate starting point of the current skeleton block (Block block), realBGend is the relative coordinate termination point of the current skeleton block; source_st is the absolute coordinate starting point of the current skeleton block, source_end is the absolute coordinate termination point of the current skeleton block; The embodiment of the application accurately maps the coordinate as follows:

[0121] 1. Establish a relative coordinate system for the skeleton sequence;

[0122] realBGst = 1, indicating the starting relative coordinate of the current backbone block;

[0123] realBGend = 0, indicating the end relative coordinate of the current backbone block;

[0124] The system updates the coordinate system every time a skeleton block is processed;

[0125] realBGend = source_end - source_st + realBGst;

[0126] In the next round: realBGst = realBGend + 1.

[0127] 2. Coordinate conversion during insertion, convert the coordinates in the relative coordinate system of the skeleton block to the absolute coordinates of the insertion sequence:

[0128] a = newMap2[i] - realBGst + source_st, indicating the absolute coordinate of the insertion starting position;

[0129] ​​b = newMap3[i] - realBGst + source_st, which represents the absolute coordinate of the insertion end position.

[0130] For example, the relative coordinate [1-1001] of Block1: [1000-2000];

[0131] The relative coordinate [1002-2002] of Block2: [3000-4000];

[0132] The relative coordinate [2003-3003] of Block3: [5000-6000];

[0133] If now a sequence of length 200 needs to be inserted at the relative coordinate 1500 position;

[0134] Block2: source_st=3000, source_end=4000;

[0135] realBGst = 1002, which represents that the previous Block1 occupies 1001 relative coordinate positions;

[0136] realBGend = 4000 - 3000 + 1002 = 2002, at this time the relationship can be judged: 1002<1500<1700<2002 (then it can be judged as case 1: internal complete insertion);

[0137] Insertion sequence: newMap2[i]=1500, newMap3[i]=1700;

[0138] a = 1500 - 1002 + 3000 = 3498;

[0139] b = 1700 - 1002 + 3000 = 3698;

[0140] Convert to absolute coordinates to determine the start and end of the inserted sequence and supplement the relevant sequence.

[0141] The bacterial paleogenomics analysis method proposed in this application, based on a phylogenetic tree, enables bottom-up analysis. Relying on the extant bacterial genome, the paleogenomics at each level can be solved progressively. Finally, with the help of outgroup strains within the bacterial community, the paleogenomics of the community is obtained. Therefore, this application fills a gap in related technologies for inferring paleogenomics. This tool helps researchers infer the key states of strains at various evolutionary nodes, study the gains and losses of bacterial gene-related changes, and provides a valuable tool for further analyzing the gains, losses, and origins of bacterial drug resistance genes, thus contributing to the depiction of bacterial evolutionary trajectories.

[0142] The following specific examples illustrate the bacterial ancient genome analysis method:

[0143] Salmonella is an important pathogenic bacterium with over 2,600 serotypes, among which the main pathogenic to humans includes Salmonella typhi (Salmonella typhi). S. Typhi), Salmonella paratyphi ( S. Paratyphi A / B / C) can cause both typhoid and non-typhoid illnesses in humans. Typhoid fever, also known as enteric fever, is a serious invasive infection that primarily affects the blood and deep reticuloendothelial tissue. It is mainly transmitted through contaminated water and food, and typical symptoms include high fever, headache, vomiting, and diarrhea. Systematic analysis of Salmonella is available. S. enterica subsp. enterica The ancient genome of this subspecies was successfully deduced (4.05 Mb, containing 4.28 million nucleotides, encoding 4,187 genes) (e.g. Figure 19 Outer ring). Referring to GC skew, the origin of replication (oriC) and the terminator (terC) divide the AOC genus almost equally into two halves (e.g., Figure 19 The inner circle suggests the ancient conservation of Salmonella chromosomes and potential chromosome arm balance constraints.

[0144] To further analyze the genomic evolutionary history of Salmonella Typhi, this example collected 148 strains. S . Typhi, 8 strains S Paratyphi A and 5 outgroup strains (including S. Cerro strains CFSAN001588 and 87, S Saintpaul strain CFSAN004175, S Manchester ST278 and S The complete chromosome sequence of *Indica RKS3057* can be obtained quickly from a large amount of strain data. This application's embodiments can rapidly solve for all ancient genomes (...). Figure 20(The computation time is only 48 minutes, and the computational efficiency increases linearly with the number of strains, without exponential growth.) A phylogenetic tree was constructed, and the ancestral chromosome sizes (AOCs) of each phylogenetic node were reconstructed. The results show that... S. Typhi and S. Paratyphi's most recent common ancestor (OPT) branched further from the common ancestor of cluster 2. S Paratyphi (P lineage) and S. Typhi (T lineage) Figure 21 The AOCs of OPT, PT (ancestor of the P lineage), P, and T were 4.04 Mb, 4.33 Mb, 4.37 Mb, and 4.47 Mb, respectively. Figure 21 Existing S. The AOCs size of the Typhi branch was stable (median 4.72 Mb), almost identical to that of existing strains (4.73 Mb). Figures 20-21 This indicates that the chromosome as a whole is stable, with no large-segment insertions or deletions. The AOCs of OPT, PT, P, and T are all equally divided in half by the origin-terminator axis (inferred from GC skewness), while... S. The AOCs at each node of Typhi also exhibited a similar pattern, suggesting the constraining effect of replication arm balance on genome evolution. Figure 21 Ancestor chromosome size refers to the ancient genome.

[0145] A comparison of AOCs according to their evolutionary paths shows that S. Paratyphi A and S. The Typhi chromosome underwent at least two large-scale inversions: one occurred in their common ancestor (called "PT inversion"), and the other occurred in... S. Typhi ancestors (called "T inversions") Figure 21 The T inversion is located within the PT inversion region, and both are centered at the replication end (ter) of the AOCs. The PT inversion exists in the PT ancestor AOCs and in existing AOCs. S. In the chromosome of Paratyphi A strain, and the existing S. More inversions were detected in the AOCs and chromosomes of the Typhi strain, including PT-inv1, PT-inv2, PT-inv3, Ter-1, Ter-2, Ori-1, and Ori-2. Figure 22 Based on the characteristics of inverted combinations, the existing S. Typhi strains can be divided into 6 genomic types (types 0-5) Figure 20 ,22). All large segment inversions were centered at the ter or ori locus, further confirming the role of chromosome arm balance. S. Typhi andS. Constraints on Paratyphi A genome evolution Figure 22 ).

[0146] Further analysis of the inversion breakpoints revealed that, in addition to the ends of T inversions, Ter-1 and Ter-2 inversions, all other inversions terminated near the ribosomal RNA (rRNA) genes Figure 23 PT-inv1 represents PT inversion type 1. In combination with the high conservation of Salmonella chromosome conformations and the phenomenon of fragmentary inversions involving the entire ter major interval, the above observations further suggest that chromosome conformations can have constraints on genome inversions and can be highly relevant to changes in traits such as drug resistance.

[0147] It should be noted that the method of the embodiments of the present application can also be used for analysis of bacteria that are also prokaryotes, such as: Escherichia coli, Pseudomonas aeruginosa, Klebsiella pneumoniae, muciniphila Akkermansia, Vibrio cholerae, and the like.

[0148] Second, the bacterial ancient genome analysis device according to the embodiments of the present application is described with reference to the accompanying drawings.

[0149] Figure 24 is a block schematic diagram of the bacterial ancient genome analysis device according to the embodiments of the present application.

[0150] As shown in Figure 24 , the bacterial ancient genome analysis device 10 includes a determination module 100, an extraction module 200, an analysis module 300, and a supplement module 400.

[0151] The determination module 100 is configured to determine a branch node of an ancient gene to be synthesized according to a phylogenetic tree of strains in strain data of bacteria; the extraction module 200 is configured to select branch strains between the branch node of the ancient gene to be synthesized and perform alignment and extraction of homologous fragments; the analysis module 300 is configured to perform positioning analysis on the homologous fragments of each strain, obtain orthologous fragments from the homologous fragments according to the positioning analysis results, and splice the orthologous fragments to obtain an ancient genome skeleton of the bacteria; and the supplement module 400 is configured to obtain outgroup gene sequences of outgroup strains of the branch strains corresponding to the branch node, and supplement the ancient genome skeleton of the bacteria according to the outgroup gene sequences of the outgroup strains of the branch strains corresponding to the branch node.

[0152] It should be noted that the foregoing explanation and description of the bacterial ancient genome analysis method embodiments also apply to the bacterial ancient genome analysis device of this embodiment, which will not be described here again.

[0153] The bacterial paleogenomic analysis device provided in the embodiments of the present application can realize bottom-up analysis based on a system evolution tree, and can solve each level of paleogenome step by step by relying on the present genome of the bacteria, and finally solve the paleogenome of the bacterial community with the help of the outgroup strain of the bacterial community. Thus, the embodiments of the present application fill the gap of inferring paleogenome in the related art, and the tool can help researchers infer the key state of the strain at each evolution node, study the gain and loss of gene correlation of the bacteria, and provide a good tool for further analyzing the gain and loss and origin of the bacterial drug resistance gene, and help to depict the evolution track of the bacteria.

[0154] Figure 25 The structural schematic diagram of the electronic device provided in the embodiments of the present application is shown. The electronic device can include:

[0155] The memory 2201, the processor 2202, and the computer program stored in the memory 2201 and executable on the processor 2202.

[0156] The processor 2202 implements the bacterial paleogenomic analysis method provided in the above embodiments when executing the program.

[0157] Further, the electronic device further includes:

[0158] The communication interface 2203 is used for communication between the memory 2201 and the processor 2202.

[0159] The memory 2201 is used to store the computer program executable on the processor 2202.

[0160] The memory 2201 can include a high-speed RAM (Random Access Memory, random access memory) memory, and can also include a non-volatile memory, such as at least one disk memory.

[0161] If the memory 2201, the processor 2202 and the communication interface 2203 are independently implemented, the communication interface 2203, the memory 2201 and the processor 2202 can be connected to each other through a bus and complete the communication between each other. The bus can be an ISA (Industry Standard Architecture, industry standard architecture) bus, a PCI (Peripheral Component, peripheral component interconnect) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 25 Only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0162] Optionally, in a specific implementation, if the memory 2201, the processor 2202 and the communication interface 2203 are integrated on a chip, the memory 2201, the processor 2202 and the communication interface 2203 can complete the communication among each other through an internal interface.

[0163] The processor 2202 can be a CPU (Central Processing Unit, central processor), or an ASIC (Application Specific Integrated Circuit, specific integrated circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0164] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the above-mentioned bacterial ancient genome analysis method.

[0165] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the specification, the illustrative description of the above terms is not necessarily for the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples without contradiction.

[0166] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically limited.

[0167] Any process or method descriptions in flow charts or otherwise described herein can be understood as representing code modules, segments, or portions of code that include one or more executable instructions for implementing the specified logic functions (or steps) and / or can be implemented by one or more hardware or software components, either in a computer or in the processor. The preferred embodiments of the present application include additional implementations that can not be described in detail, where the order of steps can be changed, where additional or fewer steps can be performed, where steps can be performed in parallel, where steps can be performed at different times, where the steps can be performed at different locations, where the steps can be performed by different components, etc.

[0168] It should be understood that portions of the application can be implemented in hardware, software, firmware, or combinations thereof. In the above embodiments, steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. As such, if implemented in hardware, as in another embodiment, any of the following technology, known in the art, can be employed for implementing: a hybrid of the techniques described herein; discrete logic circuitry having logic gates for implementing logic functions upon data signals; application specific integrated circuits having appropriate combinational logic gates; programmable gate arrays; field programmable gate arrays, or the like.

[0169] Those skilled in the art of the present technology can understand that all or part of the steps carried out by the above-mentioned embodiments can be completed by programs instructing relevant hardware, and the above-mentioned programs can be stored in a computer readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.

[0170] Although the embodiments of the present application have been shown and described above, it should be understood that the above-mentioned embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-mentioned embodiments within the scope of the present application.

Claims

1. A method for analyzing a bacterial ancient genome, characterized by, The method comprises the following steps: determining a branch node of a to-be-synthesized ancient gene according to a strain phylogenetic tree of strain data of bacteria; selecting branch strains from between the branch node of the to-be-synthesized ancient gene to extract homologous fragments for comparison; performing positioning analysis on the homologous fragments of each strain, obtaining orthologous fragments from the homologous fragments according to the positioning analysis result, and splicing the orthologous fragments to obtain an ancient genome skeleton of the bacteria; obtaining outgroup gene sequences of outgroup strains of the branch strains corresponding to the branch node, and supplementing the ancient genome skeleton of the bacteria according to the outgroup gene sequences of the outgroup strains of the branch strains corresponding to the branch node; the supplementing of the ancient genome skeleton of the bacteria according to the outgroup gene sequences of the outgroup strains of the branch strains corresponding to the branch node comprises: obtaining a skeleton sequence of the ancient genome skeleton and a proto-gene sequence of the branch strains corresponding to the branch node; comparing the proto-gene sequence with the skeleton sequence and the outgroup gene sequence respectively; supplementing a supplementable fragment to the ancient genome skeleton of the bacteria in the comparison result, wherein the supplementable fragment is a fragment that does not exist in the comparison between the skeleton sequence and the proto-gene sequence and exists in the comparison between the outgroup gene sequence and the proto-gene sequence; the comparison of the proto-gene sequence with the skeleton sequence and the outgroup gene sequence respectively further comprises: generating an orthologous block file of the skeleton sequence based on the comparison between the proto-gene sequence and the skeleton sequence, identifying a gap region in the proto-gene sequence according to the orthologous block file of the skeleton sequence, and taking the gap region as a skeleton block; generating an orthologous block file of the outgroup gene sequence based on the comparison between the proto-gene sequence and the outgroup gene sequence, identifying a supplementable fragment according to the orthologous block file of the outgroup gene sequence, and taking the supplementable fragment as a supplement block; generating a comparison result according to the skeleton block and the supplement block.

2. The bacterial ancient genome analysis method according to claim 1, characterized by, Before determining the branch node of the to-be-synthesized ancient gene according to the strain phylogenetic tree of strain data of bacteria, the method further comprises: extracting base sequences and / or amino acid sequences of strains in the strain data of the bacteria; calculating evolutionary relationships between organisms according to the base sequences and / or amino acid sequences, and determining the strain phylogenetic tree of the strain data of the bacteria according to the evolutionary relationships between organisms.

3. The bacterial ancient genome analysis method according to claim 1, characterized by, The positioning analysis on the homologous fragments of each strain comprises: identifying the position of the homologous fragments of each strain in the genome; numbering the homologous fragments according to their positions in the genome; sorting the numbers of all homologous fragments according to their lengths, and analyzing the homology and positional order relationship between homologous fragments in sequence according to the sorting result, and deleting the homologous fragments with inconsistent homology in the positioning process until all homologous fragments are analyzed.

4. The bacterial ancient genome analysis method according to claim 3, characterized by, The obtaining of orthologous fragments from the homologous fragments according to the positioning analysis result and the splicing of the orthologous fragments to obtain the ancient genome skeleton of the bacteria comprise: extracting the retained homologous fragments and the corresponding positional order relationship from the positioning analysis result; The reserved homologous fragments are straight homologous fragments in the homologous fragments; The straight homologous fragments are spliced according to the positional sequence relationship to obtain the ancient genome framework of the bacteria.

5. The bacterial ancient genome analysis method according to claim 1, characterized by, Before the complementable fragments in the comparison result are supplemented to the ancient genome framework of the bacteria, the method further comprises: According to the first gap length in the original gene sequence and the second gap length in the framework sequence, a ratio of the first gap length and the second gap length is calculated; According to the first gap length, the second gap length and the ratio, it is determined whether the supplement condition of the ancient genome framework is met, and if the supplement condition is met, the ancient genome framework of the bacteria is supplemented according to the comparison result.

6. The bacterial ancient genome analysis method according to claim 1 or 5, characterized by, The method for supplementing the complementable fragments in the comparison result to the ancient genome framework of the bacteria comprises: The relative positional relationship between the supplement block and the framework block in the comparison result is identified; According to the relative positional relationship, an overlap mode between the supplement block and the framework block is determined, the overlap mode comprising a complete containing mode, a left end overlap mode, a right end overlap mode and an internal overlap mode, the complete containing mode being that the framework block contains the supplement block, the left end overlap mode being that the supplement block overlaps the left end of the framework block, the right end overlap mode being that the supplement block overlaps the right end of the framework block, and the internal overlap mode being that the supplement block contains the framework block; If the overlap mode is the complete containing mode, the framework block is divided into a first half and a second half with the start position of the supplement block as a boundary, and the supplement block is inserted between the first half and the second half; If the overlap mode is the left end overlap mode, the framework block is divided into a first half and a second half with the end position of the supplement block as a boundary, and the supplement block is inserted in front of the second half; If the overlap mode is the right end overlap mode, the framework block is divided into a first half and a second half with the start position of the supplement block as a boundary, and the supplement block is inserted behind the first half; If the overlap mode is the internal overlap mode, the supplement block is used to replace the framework block.

7. A bacterial ancient genome analysis device characterized by comprising: The method comprises: A determination module is configured to determine a branch node of an ancient gene to be synthesized according to a phylogenetic tree of strains in strain data of bacteria; An extraction module is configured to select a branch strain from between branch nodes of the ancient gene to be synthesized for comparison and extraction of homologous fragments; An analysis module is configured to perform positioning analysis on the homologous fragments of each strain, obtain straight homologous fragments from the homologous fragments according to a positioning analysis result, and splice the straight homologous fragments to obtain an ancient genome framework of the bacteria. The supplementing module is configured to obtain outgroup gene sequences of outgroup strains of the branch strain corresponding to the branch node, and supplement the ancient genome framework of the bacterium according to the outgroup gene sequences of the outgroup strains of the branch strain corresponding to the branch node, wherein the supplementing the ancient genome framework of the bacterium according to the outgroup gene sequences of the outgroup strains of the branch strain corresponding to the branch node comprises: obtaining a framework sequence of the ancient genome framework and a protogene sequence of the branch strain corresponding to the branch node; comparing the protogene sequence with the framework sequence and the outgroup gene sequence respectively; and supplementing a supplementable fragment in a comparison result to the ancient genome framework of the bacterium, wherein the supplementable fragment is a fragment that is not present in the comparison between the framework sequence and the protogene sequence and is present in the comparison between the outgroup gene sequence and the protogene sequence. The comparison of the protogene sequence with the framework sequence and the outgroup gene sequence respectively further comprises: generating an orthologous block file of the framework sequence based on the comparison between the protogene sequence and the framework sequence, identifying a gap region in the protogene sequence according to the orthologous block file of the framework sequence, and taking the gap region as a framework block; generating an orthologous block file of the outgroup gene sequence based on the comparison between the protogene sequence and the outgroup gene sequence, identifying a supplementable fragment according to the orthologous block file of the outgroup gene sequence, and taking the supplementable fragment as a supplement block; and generating a comparison result according to the framework block and the supplement block.

8. An electronic device, comprising: The computer program is stored in the memory and executable on the processor, and the processor executes the program to implement the bacterial ancient genome analysis method according to any one of claims 1-6. The computer program is stored in the memory and executable on the processor, and the processor executes the program to implement the bacterial ancient genome analysis method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Population genetic evolution map and construction method thereof

    CN110910959A

  • Generic genome construction method and device based on phylogenetic tree

    CN111477281A