Integration joint element analysis method, device, electronic equipment and storage medium
By using second-generation sequencing sequence annotation and marker gene identification, the high cost and low efficiency of identifying integrative conjugation elements in bacterial genomes in existing technologies have been solved, enabling rapid and accurate screening and automated processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ICDC CHINA CDC
- Filing Date
- 2026-06-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies are costly and inefficient in identifying integrative binding elements in bacterial genomes, making it difficult to process large numbers of microbial genome sequences efficiently and quickly.
By obtaining second-generation sequencing sequences, annotating contig fragments, and using start and end marker genes for joint judgment, combined with the distribution of core genes, the sequence range of integrative conjugating elements can be determined, enabling rapid and accurate screening.
It reduces data acquisition costs, enables rapid and accurate screening of integrative conjugation elements in a large number of strains' second-generation sequencing data, and supports automated pipeline processing of multiple strains.
Smart Images

Figure CN122493962A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microbial genome information analysis, and more particularly to an integrated conjugation element analysis method, apparatus, electronic device, and storage medium. Background Technology
[0002] An important class of mobile genetic elements—integrative conjugation elements—are widely distributed in bacterial genomes. These elements can integrate into bacterial chromosomes and possess a complete conjugation transfer mechanism, enabling self-transfer between bacterial cells. By mediating the horizontal transfer of various cargo genes, integrative conjugation elements confer a series of beneficial traits to host bacteria, such as antibiotic resistance, pathogenicity, metal resistance, and the ability to degrade compounds, thus playing a crucial role in bacterial evolution.
[0003] Among them, SXT / R391 family integrative conjugating elements are widely distributed in strains such as Vibrio cholerae. The variable region of this element often carries multiple drug resistance genes and can move between the genomes of different bacteria, thereby mediating the widespread dissemination of drug resistance genes. Therefore, identifying SXT / R391 family integrative conjugating elements from bacterial genome sequences is of great significance for addressing the spread of bacterial drug resistance genes.
[0004] Currently, a common method for identifying integrative conjugation elements in bacterial genomes is the online tool ICEfinder. Users upload the genome sequence to be detected, and the tool can identify integrative conjugation elements online. However, this tool is only applicable to completed genome sequences, is costly, and can only accept one genome sequence uploaded manually by the user at a time. When faced with a large number of microbial genome sequences, it is difficult to efficiently and quickly complete the task of identifying integrative conjugation elements.
[0005] Therefore, how to develop a low-cost, accurate, and efficient method for analyzing integrated bonding components is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0006] This invention provides a method, apparatus, electronic device, and storage medium for analyzing integrative conjugation elements. It can directly process second-generation sequencing sequence data, reducing data acquisition costs. Simultaneously, it determines the presence of target integrative conjugation elements by jointly judging start and end marker genes. Furthermore, based on the distribution of core genes in the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment, it determines the sequence range of the target integrative conjugation element. This enables rapid and accurate screening of target integrative conjugation elements and their sequence ranges in a large number of strains' second-generation sequencing sequence data.
[0007] This invention provides a method for analyzing integrated bonding elements, comprising: Obtain the next-generation sequencing sequence of the bacterial strain to be analyzed; Annotate the contig fragments of the second-generation sequencing sequence to obtain the predicted gene; The predicted gene is compared with the core gene of the target integrative binding element for sequence similarity, and candidate genes that match the core gene are selected based on the sequence similarity comparison results. Based on whether the candidate genes simultaneously contain the initiation marker gene and the terminal marker gene of the core gene, it is determined whether the bacterial strain to be analyzed contains the target integrative conjugation element. If the bacterial strain to be analyzed contains a target integrative conjugating element, the sequence range of the target integrative conjugating element is determined based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment.
[0008] According to the method for analyzing integrative conjugating elements provided by the present invention, before determining the sequence range of the target integrative conjugating element based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment, the method further includes: The name of the contig fragment containing each candidate gene, its start and end positions within the contig fragment, and its gene orientation are obtained and annotated in the sequence similarity comparison results to obtain the annotated sequence similarity comparison results; The determination of the sequence range of the target integrative conjugating element based on the distribution of the core gene in the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment includes: When the core genes are all located in the same contig fragment of the next-generation sequencing sequence, the sequence range of the target integrative conjugating element is determined based on the relative positions of the start marker gene and the end marker gene in the annotated sequence similarity comparison results. When the core gene is distributed across multiple contig fragments in the next-generation sequencing sequence, the sequence range of the target integrative conjugating element is determined as follows: For a first contig fragment containing the start marker gene, a first sorting direction is determined based on the relative position of the start marker gene and its adjacent core genes in the labeled sequence similarity comparison results. Based on the first sorting direction and the start marker gene, a first target fragment is extracted from the first contig fragment as the first part of the target integrative conjugation element. For a second contig fragment containing the terminal marker gene, a second sorting direction is determined based on the relative position of the terminal marker gene and its adjacent core gene in the labeled sequence similarity comparison results. Based on the second sorting direction and the terminal marker gene, a second target fragment is extracted from the second contig fragment as the second part of the target integrative conjugation element. For a third contig fragment that does not contain the start marker gene and the end marker gene, but contains at least one other core gene, the third contig fragment is used as the third part of the target integrative conjugation element; Based on the first part, the second part, and the third part, the sequence range of the target integrated contig fragments is determined.
[0009] According to the present invention, an integrated bonding element analysis method further includes: Based on the sequence similarity comparison results after annotation, each gene within the sequence range is labeled to obtain the labeled genes; wherein, genes that match the core gene are labeled as core gene names, and genes that do not match are labeled as putative proteins. Obtain the target information of the labeled genes and organize the target information into structured data; wherein, the target information includes gene name, start and end positions in the contig fragment, annotation color, gene direction and name of the contig to which the gene is located; Based on the structured data, the overall arrangement direction of the core genes in each relevant contig fragment is compared with the arrangement direction of the core genes in the representative element of the target integrative conjugation element. When the arrangement directions are consistent, the gene structure diagram of the target integrative conjugating element is drawn according to the actual arrangement direction of the core genes in the relevant contig fragment; When the arrangement directions are inconsistent, the relevant contig fragments are reversed to draw the gene structure diagram of the target integrative conjugating element.
[0010] According to a method for analyzing integrative conjugating elements provided by the present invention, when the bacterial strain to be analyzed includes multiple strains, the step of comparing the predicted gene with the core gene of the target integrative conjugating element for sequence similarity, and screening candidate genes matching the core gene based on the sequence similarity comparison results, includes: The predicted genes of multiple bacterial strains to be analyzed are merged to construct a merged predicted gene set; The merged predicted gene set is compared with the core gene of the target integrative binding element for sequence similarity, and candidate genes that match the core gene are selected based on the sequence similarity comparison results.
[0011] According to a method for analyzing integrative conjugating elements provided by the present invention, the step of performing sequence similarity comparison between the predicted gene and the core gene of the target integrative conjugating element, and screening candidate genes matching the core gene based on the sequence similarity comparison results, includes: Using the nucleic acid sequence of the core gene of the target integrative conjugating element as the query sequence and the nucleic acid sequence database constructed using the nucleic acid sequence of the predicted gene as the target database, BLASTn sequence alignment is performed to obtain the sequence similarity alignment results; Based on the sequence similarity comparison results, candidate genes that meet the preset similarity conditions are selected; The preset similarity conditions include: sequence consistency of not less than 85%, sequence coverage of not less than 85%, and expected value of less than 0.001.
[0012] According to the present invention, an integrative conjugation element analysis method, wherein the annotation of contig fragments in the second-generation sequencing sequence to obtain predicted genes includes: The Prokka program was used to annotate the contig fragments of the second-generation sequencing sequence to obtain the nucleic acid sequence of the predicted gene, its position information in the contig fragment, and the gene orientation.
[0013] According to the present invention, an integrative conjugation element analysis method is provided, wherein the target integrative conjugation element is an SXT / R391 family integrative conjugation element, and the core gene includes the following 25 genes: xis, int, mobI, traI, traD, traJ, traL, traE, traK, traB, traV, traA, s054, traC, trhF, traW, traU, traN, s063, traF, traH, traG, setC, setD, and setR, wherein the initiation marker gene is xis and the terminal marker gene is setR.
[0014] The present invention also provides an integrated bonding element analysis device, comprising: The sequence acquisition module is used to acquire the next-generation sequencing sequence of the bacterial strain to be analyzed. The contig identification module is used to annotate the contig fragments of the second-generation sequencing sequence to obtain the predicted gene; The gene screening module is used to perform sequence similarity comparison between the predicted gene and the core gene of the target integrative binding element, and to screen out candidate genes that match the core gene based on the sequence similarity comparison results. The element determination module is used to determine whether the bacterial strain to be analyzed contains the target integrative conjugating element based on whether the candidate gene simultaneously contains the initiation marker gene and the terminal marker gene of the core gene. The range determination module is used to determine the sequence range of the target integrative conjugating element if the bacterial strain to be analyzed contains the target integrative conjugating element, based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the integrated bonding element analysis method as described in any of the preceding claims.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the integrated bonding element analysis method as described in any of the preceding claims.
[0017] The present invention provides a method, apparatus, electronic device, and storage medium for analyzing integrative conjugating elements. This method acquires the next-generation sequencing (NGS) sequence of the bacterial strain to be analyzed, annotates the contig fragments of the NGS sequence to obtain predicted genes, then compares the sequence similarity of the predicted genes with the core gene of the target integrative conjugating element and screens candidate genes. The presence of both start and end marker genes in the candidate genes determines whether the target integrative conjugating element is present. If present, the sequence range of the target integrative conjugating element is further determined based on whether the core gene is distributed across the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position of the core gene within the corresponding contig fragment. The integrative conjugating element analysis method of the present invention can directly process NGS sequence data without requiring a complete genome map beforehand, reducing data acquisition costs and making large-scale strain screening economically feasible. Meanwhile, this invention determines the presence of a target integrative conjugating element by jointly identifying start and end marker genes. Furthermore, based on whether the core gene is distributed across the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position of the core gene within the corresponding contig fragment, the sequence range of the target integrative conjugating element is determined. This enables rapid and accurate screening of target integrative conjugating elements and their sequence ranges in a large number of strains' next-generation sequencing data. In addition, the technical solution of this invention can be automated and streamlined for processing multiple strains through script programming. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is one of the flowcharts illustrating the integrated bonding element analysis method provided by the present invention.
[0020] Figure 2 This is the second flowchart of the integrated bonding element analysis method provided by the present invention.
[0021] Figure 3 This is the third flowchart of the integrated bonding element analysis method provided by the present invention.
[0022] Figure 4 This is the fourth flowchart of the integrated bonding element analysis method provided by the present invention.
[0023] Figure 5 This is one of the schematic diagrams of the target integrated bonding element provided by the present invention.
[0024] Figure 6 This is the second schematic diagram of the target integrated bonding element provided by the present invention.
[0025] Figure 7 This is the third of the visualization diagrams of the target integrated bonding element provided by the present invention.
[0026] Figure 8 This is a schematic diagram of the integrated bonding element analysis device provided by the present invention.
[0027] Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0029] This invention proposes an integrated bonding element analysis method, apparatus, electronic device, and storage medium, which are described below in conjunction with... Figures 1-9 Describe it.
[0030] Figure 1 This is one of the flowcharts illustrating the integrated bonding element analysis method provided by the present invention, such as... Figure 1 As shown, the integrated bonding element analysis method includes steps S110, S120, S130, S140 and S150.
[0031] Step S110: Obtain the next-generation sequencing sequence of the bacterial strain to be analyzed.
[0032] Second-generation sequencing sequences refer to contig (contiguous sequences) fragment sequences obtained by assembling the genome of a bacterial strain after sequencing it using a high-throughput sequencing platform.
[0033] Obtain the next-generation sequencing sequence of the bacterial strain to be analyzed. The bacterial strain to be analyzed may include one or more strains.
[0034] In one implementation, the next-generation sequencing data of the bacterial strain to be analyzed can be downloaded from an existing public database (such as EnteroBase). The nucleic acid sequence is in FASTA (a nucleic acid or protein sequence storage format) format.
[0035] In another embodiment, the raw data of the bacterial strain to be analyzed can be obtained by sequencing itself, and after assembly, a sequence file containing one or more contig fragments can be obtained and stored in FASTA format.
[0036] Step S120: Annotate the contig fragment of the second-generation sequencing sequence to obtain the predicted gene.
[0037] Annotate contig fragments of second-generation sequencing sequences, including open reading frame identification, to obtain predicted genes.
[0038] Open reading frame identification (ORB) refers to the process of using bioinformatics software to predict the continuous coding region from a start codon to a stop codon in a DNA sequence that may encode a protein.
[0039] In one implementation, the Prokka (rapid prokaryotic genome annotation) program is used to annotate contig fragments of the next-generation sequencing sequence of the bacterial strain to be analyzed, including open reading frame identification.
[0040] Step S130: Perform sequence similarity comparison between the predicted gene and the core gene of the target integrative binding element, and screen out candidate genes that match the core gene based on the sequence similarity comparison results.
[0041] In one embodiment, the target integrative conjugating element is an SXT / R391 family integrative conjugating element, and its core genes include: integration / splitting module genes xis and int, DNA processing module genes mobI, traI, traD, and traJ, conjugation transfer module genes traL, traE, traK, traB, traV, traA, s054, traC, trhF, traW, traU, traN, traF, traH, and traG, regulatory module genes setC, setD, and setR, and other module genes s063. This embodiment of the invention uses an SXT / R391 family integrative conjugating element as an example for illustration.
[0042] Previous studies have shown that the R391 integrative conjugating element with accession number AY090559.1 in the NCBI (National Center for Biotechnology Information) database is generally considered a representative integrative conjugating element of the SXT / R391 family. In this representative integrative conjugating element, the core genes typically exhibit a linear arrangement from the xis gene to the setR gene: xis, int, mobI, traI, traD, traJ, traL, traE, traK, traB, traV, traA, s054, traC, trhF, traW, traU, traN, s063, traF, traH, traG, setC, setD, setR. The initiation marker gene is xis, and the terminal marker gene is setR.
[0043] In one embodiment, when the bacterial strain to be analyzed includes a single strain, after obtaining the predicted gene of the bacterial strain to be analyzed through annotation (including open reading frame identification, etc.), the predicted gene of the bacterial strain to be analyzed is directly compared with the core gene of the target integrative conjugating element for sequence similarity. Based on the sequence similarity comparison results, candidate genes that match the core gene are selected. Specifically, the nucleic acid sequence of the core gene of the target integrative conjugating element is used as the query sequence, and a nucleic acid sequence database constructed using the nucleic acid sequence of the predicted gene is used as the target database. BLASTn sequence alignment is performed to obtain the sequence similarity comparison results. Based on the sequence similarity comparison results, candidate genes that meet preset similarity conditions are selected. The preset similarity conditions include: sequence identity not less than 85%, sequence coverage not less than 85%, and an expected value less than 0.001.
[0044] In another embodiment, when the bacterial strains to be analyzed include multiple strains, after batch automatic annotation (including open reading frame recognition, etc.), predicted genes for multiple bacterial strains to be analyzed will be obtained. At this time, the predicted genes of multiple bacterial strains to be analyzed are first merged to construct a merged predicted gene set. Then, the merged predicted gene set is compared with the core gene of the target integrative conjugating element for sequence similarity. Based on the sequence similarity comparison results, candidate genes matching the core gene are selected. Specifically, the nucleic acid sequence of the core gene of the target integrative conjugating element is used as the query sequence, and the nucleic acid sequence database constructed from the nucleic acid sequences of the merged predicted gene set is used as the target database. BLASTn sequence alignment is performed to obtain the sequence similarity comparison results. Based on the sequence similarity comparison results, candidate genes that meet preset similarity conditions are selected. The preset similarity conditions include: sequence identity not less than 85%, sequence coverage not less than 85%, and an expected value less than 0.001.
[0045] Step S140: Based on whether the candidate genes simultaneously contain the start marker gene and the end marker gene of the core gene, determine whether the bacterial strain to be analyzed contains the target integrative conjugation element.
[0046] Initiation marker genes and terminal marker genes are two highly conserved and representative genes located at the beginning and end of the target integrative conjugation element, respectively, within the core gene set. Their co-existence can serve as an important basis for determining the presence of the entire target integrative conjugation element.
[0047] The selected candidate genes are analyzed to determine whether they simultaneously contain the initiation marker gene and the terminal marker gene of the target integrative conjugation element. For example, for the SXT / R391 family integrative conjugation element, the initiation marker gene is xis, and the terminal marker gene is setR. If both xis and setR genes are detected in the candidate genes, the genome of the bacterial strain being analyzed is determined to contain the target integrative conjugation element.
[0048] Step S150: If the bacterial strain to be analyzed contains a target integrative conjugating element, the sequence range of the target integrative conjugating element is determined based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment.
[0049] If the bacterial strain to be analyzed contains the target integrative conjugating element, the sequence range of the target integrative conjugating element is further determined based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment.
[0050] Specifically, the sequence range of the target integrative conjugating element is determined based on whether the core gene is located in a single or multiple contig segments, whether the contig segment containing the core gene contains a start marker gene or a terminal marker gene, and the positional relationship between the start marker gene and the terminal marker gene in the contig segment containing the core gene and its adjacent genes.
[0051] The integrative conjugation element analysis method provided in this invention obtains the next-generation sequencing (NGS) sequence of the bacterial strain to be analyzed, annotates the contig fragments of the NGS sequence to obtain predicted genes, then compares the sequence similarity of the predicted genes with the core gene of the target integrative conjugation element and screens candidate genes. The presence of both start and end marker genes in the candidate genes determines whether the target integrative conjugation element is present. If present, the sequence range of the target integrative conjugation element is further determined based on whether the core gene is distributed across the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position of the core gene within the corresponding contig fragment. This integrative conjugation element analysis method can directly process NGS sequence data without requiring a complete genome map beforehand, reducing data acquisition costs and making large-scale strain screening economically feasible. Meanwhile, this embodiment of the invention determines the presence of a target integrative conjugating element by jointly judging the start and end marker genes. Then, based on whether the core gene is distributed in the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment, the sequence range of the target integrative conjugating element is determined. This enables rapid and accurate screening of target integrative conjugating elements and their specific ranges in a large number of strains' next-generation sequencing sequence data. Furthermore, the technical solution of this embodiment of the invention can achieve automated, streamlined processing of multiple strains through script programming.
[0052] Based on any of the above embodiments, before step S150, the integrated bonding element analysis method further includes step S160.
[0053] Step S160: Obtain the name of the contig fragment containing each candidate gene, the start and end positions of the contig fragment, and the gene orientation, and annotate them in the sequence similarity comparison results to obtain the annotated sequence similarity comparison results.
[0054] In one implementation, when annotating contig fragments of a second-generation sequencing sequence (including open reading frame identification), in addition to obtaining the predicted gene, the location information and gene orientation of the predicted gene are also generated to form a complete sequence annotation result.
[0055] Based on the sequence annotation results, the name of the contig fragment containing each candidate gene, the start and end positions of the contig fragment, and the gene orientation are obtained and annotated in the sequence similarity comparison results.
[0056] In one specific implementation, the Prokka program can be used for annotation. The ffn (predicted gene nucleotide sequence file stored in FASTA format) file generated by the program contains the nucleic acid sequences of all predicted genes, and the gff (general feature format) file contains the location information and gene orientation of each predicted gene.
[0057] Because the Prokka program assigns a number to the identified coding sequence, the gene name in the sequence similarity comparison results is also a number provided by Prokka, which does not reflect the contig number to which the gene belongs.
[0058] Therefore, a Python program was used to obtain the name of the contig fragment corresponding to each candidate gene number, the start and end positions of the candidate gene in the contig fragment, and the gene orientation from the gff file in the Prokka results, and then annotated them in the obtained sequence similarity comparison results.
[0059] The integrated conjugation element analysis method provided in this invention extracts the location information of candidate genes and annotates it in the sequence similarity comparison results, enabling the subsequent sequence range determination step to obtain the precise spatial location of each candidate gene. This provides the necessary data foundation for determining the arrangement order and relative distance of core genes on the contig fragment, thereby improving the accuracy of sequence range determination.
[0060] Furthermore, the step "determining the sequence range of the target integrative conjugating element based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment" includes: steps S151 and S152.
[0061] Step S151: When all the core genes are located in the same contig fragment of the second-generation sequencing sequence, the sequence range of the target integrative conjugating element is determined based on the relative positions of the start marker gene and the end marker gene in the sequence similarity comparison results after annotation.
[0062] In this embodiment of the invention, strains identified as containing target family integrative conjugating elements through step S140 can be further applied using the method of this embodiment of the invention. Based on the sequence similarity comparison results after annotation, the distribution of 25 core genes in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment can be analyzed to determine the sequence range of the element in its corresponding contig fragment.
[0063] like Figure 2 As shown, if all core genes are located in the same contig fragment of the next-generation sequencing sequence, then the contig fragment is considered to contain a complete SXT / R391 family integrative conjugating element. In this case, the sequence range is determined directly based on the relative positions of the start marker gene xis and the end marker gene setR in the annotated sequence similarity comparison results.
[0064] Specifically, if the xis gene is to the left of the setR gene, it indicates that the SXT / R391 family of integrative elements is aligned in the forward direction compared to the representative integrative element R391, and the range of integrative elements in the SXT / R391 family extends from the start of the xis gene to the end of the setR gene. If the xis gene is to the right of the setR gene, it indicates that the SXT / R391 family of integrative elements is aligned in the reverse direction compared to the representative integrative element R391, and the range of integrative elements in the SXT / R391 family extends from the start of the setR gene to the end of the xis gene.
[0065] Step S152: When the core gene is distributed across multiple contig fragments in the next-generation sequencing sequence, the sequence range of the target integrative conjugating element is determined as follows: For a first contig fragment containing the start marker gene, a first sorting direction is determined based on the relative position of the start marker gene and its adjacent core genes in the sequence similarity comparison results after annotation. Based on the first sorting direction and the start marker gene, a first target fragment is extracted from the first contig fragment as the first part of the target integrative conjugation element. For a second contig fragment containing the terminal marker gene, a second sorting direction is determined based on the relative position of the terminal marker gene and its adjacent core gene in the annotated sequence similarity comparison results. Based on the second sorting direction and the terminal marker gene, a second target fragment is extracted from the second contig fragment as the second part of the target integrative conjugation element. For a third contig fragment that does not contain the start marker gene and the end marker gene, but contains at least one other core gene, the third contig fragment is used as the third part of the target integrative conjugation element; Based on the first part, the second part, and the third part, the sequence range of the target integrated contig fragments is determined.
[0066] When the core gene is distributed across multiple contig fragments in a next-generation sequencing sequence, the sequence range of the target integrative conjugation element is determined as follows: For a contig fragment containing the start marker gene xis, denoted as the first contig fragment, the sorting direction is determined based on the relative position of the start marker gene xis and its adjacent core gene int in the annotated sequence similarity comparison results, and denoted as the first sorting direction. Then, based on the first sorting direction and the start marker gene xis, the corresponding target fragment (denoted as the first target fragment) is extracted from the first contig fragment as part of the target integrative conjugation element, denoted as the first part. Specifically, it is determined whether the adjacent core gene int is to the left or right of the start marker gene xis. If it is to the left, it means that the SXT / R391 family integrative conjugation element fragment is reversed relative to the corresponding region of the R391 element in the first contig fragment, and the fragment from the left start point of the first contig fragment to the end of the xis gene is taken as part of the SXT / R391 family integrative conjugation element. If it is to the right, it means that the SXT / R391 family integrative conjugation element fragment is forward-oriented relative to the corresponding region of the R391 element in the first contig fragment, and the fragment from the start point of the xis gene to the right end of the first contig fragment is taken as part of the SXT / R391 family integrative conjugation element.
[0067] For a contig fragment containing the terminal marker gene setR, denoted as the second contig fragment, the sorting direction is determined based on the relative position of the terminal marker gene setR and its adjacent core gene setD in the annotated sequence similarity comparison results, denoted as the second sorting direction. Then, based on the second sorting direction and the terminal marker gene setR, the corresponding target fragment (denoted as the second target fragment) is extracted from the second contig fragment as part of the target integrative conjugation element, denoted as the second part. Specifically, it is determined whether the adjacent core gene setD is to the left or right of the terminal marker gene setR. If it is to the left, it indicates that the SXT / R391 family integrative conjugation element fragment is aligned in the forward direction relative to the corresponding region of the R391 element in the second contig fragment, and the fragment from the start of the second contig fragment to the end of the setR gene is taken as part of the SXT / R391 family integrative conjugation element. If it is to the right, it indicates that the SXT / R391 family integrative conjugation element fragment is aligned in the reverse direction relative to the corresponding region of the R391 element in the second contig fragment, and the fragment from the start of the setR gene to the right end of the contig fragment is taken as part of the SXT / R391 family integrative conjugation element.
[0068] For contig fragments that do not contain start and end marker genes but contain at least one other core gene, denoted as the third contig fragment, this third contig fragment is considered the middle part of the SXT / R391 family integrative conjugating element. In this case, the entire third contig fragment is considered part of the target integrative conjugating element and is denoted as the third part. It should be understood that the third part may contain only one contig fragment or may include multiple contig fragments.
[0069] Next, based on the first, second, and third parts obtained above, the sequence range of the target integrated contig element on multiple contig fragments is determined.
[0070] The integrative conjugation element analysis method provided in this invention offers different sequence range determination methods for two scenarios: core genes distributed across the same contig fragment and multiple contig fragments. For the case where the core gene is distributed across multiple contig fragments, the sequence range of the target integrative conjugation element is identified by separately processing contig fragments containing start marker genes, end marker genes, and contig fragments containing only the middle core gene. This overcomes the problem of difficulty in identifying the sequence range of the target integrative conjugation element due to incomplete second-generation sequencing data, achieving accurate determination of the sequence range of the integrative conjugation element and laying the foundation for subsequent characteristic analysis of the target integrative conjugation element, such as carrying drug resistance genes.
[0071] Based on any of the above embodiments Figure 3 This is the third flowchart illustrating the integrated bonding element analysis method provided by the present invention, as shown below. Figure 3 As shown, the integrated bonding element analysis method further includes steps S171, S172, S173, S1741, and S1742.
[0072] It should be noted that the execution order of steps S1741 and S1742 is not important and they can be executed in parallel.
[0073] Step S171: Based on the sequence similarity comparison results after annotation, each gene in the sequence range is annotated to obtain annotated genes; wherein, genes that match the core gene are annotated as core gene names, and genes that do not match are annotated as putative proteins.
[0074] Based on the determination of the sequence range of the target family integrative conjugating element in the contig fragment, each gene within the sequence range is labeled according to the sequence similarity comparison results after labeling, and the labeled genes are obtained.
[0075] Specifically, the annotation results of all genes within the target family's integrative binding element range are extracted from the gff file of the Prokka annotation results. Based on the sequence similarity comparison results after annotation, all genes within the sequence range of the target integrative binding element are labeled. If an open reading frame corresponds to any core gene, it is labeled with the name of the specific core gene; if no core gene is matched (i.e., a non-core gene), it is labeled as a hypothetical protein.
[0076] Step S172: Obtain the target information of the labeled gene and organize the target information into structured data; wherein, the target information includes the gene name, the start and end positions in the contig fragment, the labeling color, the gene direction, and the name of the contig in which the gene is located.
[0077] For each labeled gene, the following target information is extracted: gene name (labeled name), start and end positions in the contig fragment, label color (e.g., light blue for core genes and gray for other genes), and the name of the contig to which the gene belongs. Then, this target information is organized into structured data.
[0078] In one embodiment, the structured data is in the form of an Excel spreadsheet, with each column containing: gene name, start position in the contig fragment, end position in the contig fragment, color, and name of the contig containing the gene.
[0079] Step S173: Based on the structured data, compare the overall arrangement direction of the core genes in each relevant contig fragment with the arrangement direction of the core genes in the representative element of the target integrative conjugation element.
[0080] Based on structured data, the overall arrangement orientation of core genes in each relevant contig fragment is compared with the arrangement orientation of core genes in representative elements of the target integrative conjugation element.
[0081] It should be noted that in actual biological processes, SXT / R391 family integrative conjugative elements integrate into the host chromosome through site-specific recombination, and their overall alignment in the genome may be either forward or reversed. Furthermore, since the contig fragments obtained during genome assembly themselves do not have a fixed orientation and their orientation is reversible, in sequence analysis and visualization results, these integrative conjugative elements may also exhibit two opposite alignments: from the xis gene pointing to the setR gene, or from the setR gene pointing to the xis gene. Therefore, to facilitate visualization and comparative genomics analysis of SXT / R391 family integrative conjugative elements, it is necessary to first compare whether the overall alignment of the core genes in each relevant contig fragment is consistent with the alignment of the corresponding core genes in the representative R391 integrative conjugative element.
[0082] Step S1741: When the arrangement directions are consistent, draw the gene structure diagram of the target integrative conjugating element according to the actual arrangement direction of the core genes in the relevant contig fragment.
[0083] When the arrangement directions are consistent, no direction reversal is performed; instead, the gene structure diagram of the target integrative conjugating element is drawn directly according to the actual arrangement direction of the core genes in the relevant contig fragment.
[0084] In one implementation, the genoPlotR (gene / genome mapping) package in the R language is used to perform plotting, directly drawing the gene structure map of the target integrative conjugating element according to the actual coordinates and gene orientation read from structured data (such as an Excel spreadsheet).
[0085] Specifically, the following script 1 can be used for drawing.
[0086] Script 1: library(genoPlotR) library(dplyr) library(purrr) library(openxlsx) input_file<- "input_excel.xlsx" output_file<- "output_fig.pdf" do_reverse<- FALSE df<- read.xlsx(input_file, colNames = TRUE) required_columns<- c( "Contig", "Gene", "Start", "End", "Direction", "Color" ) if (!all(required_columns %in% colnames(df))) { stop( "The input file must contain columns: ", paste(required_columns, collapse = ", ") ) } df<- df %>% mutate( Contig = as.character(Contig), Gene = as.character(Gene), Start = as.numeric(Start), End = as.numeric(End), Direction = as.numeric(Direction), Color = as.character(Color) ) %>% filter( !is.na(Contig), !is.na(Gene), !is.na(Start), !is.na(End), !is.na(Direction), !is.na(Color), Color != "" ) if (nrow(df) == 0) { stop("No valid gene annotation data was found.") } invalid_directions<- setdiff(unique(df Direction), c(-1, 1)) if (length(invalid_directions)>0) { stop("Direction must be 1 or -1.") } is_valid_color<- function(color_value) { tryCatch( { col2rgb(color_value) TRUE }, error = function(e) FALSE ) } invalid_colors<- unique(df Color[!sapply(df Color, is_valid_color)]) if (length(invalid_colors)>0) { stop( "Invalid color values found: ", paste(invalid_colors, collapse = ", ") ) } genomes_id<- unique(df Contig) if (length(genomes_id) == 0) { stop("No valid Contig was found.") } dna_segs<- map(genomes_id, function(contig_id) { genome_df<- df %>% filter(.data Contig == .env contig_id) dna_seg_df<- data.frame( name = genome_df Gene, start = genome_df Start, end = genome_df End, strand = genome_df Direction, fill = genome_df Color, stringsAsFactors = FALSE ) as.dna_seg( dna_seg_df, fill = dna_seg_df fill, col = dna_seg_df fill, lwd = 0 ) }) %>% set_names(genomes_id) if (do_reverse) { dna_segs_for_plot<- lapply(dna_segs, reverse) names(dna_segs_for_plot)<- names(dna_segs) } else { dna_segs_for_plot<- dna_segs } annotations<- lapply(dna_segs_for_plot, function(dna_seg) { annotation( x1 = middle(dna_seg), text = dna_seg name, rot = 75 ) }) add_blank_line_if_single<- function(dna_segs_for_plot, annotations) { if (length(dna_segs_for_plot) != 1) { return(list( dna_segs_final = dna_segs_for_plot, annotations_final = annotations, pdf_height = 6 )) } dna_seg<- dna_segs_for_plot[[1]] coordinate_range<- range( c(dna_seg start, dna_seg end), na.rm = TRUE ) if (!all(is.finite(coordinate_range)) || diff(coordinate_range) == 0) { coordinate_range<- c(0, 1) } blank_dna_seg_df<- data.frame( name = "", start = coordinate_range[1], end = coordinate_range[1] + max(1, diff(coordinate_range) 0.01), strand = 1, stringsAsFactors = FALSE ) blank_dna_seg<- as.dna_seg( blank_dna_seg_df, fill = "white", col = "white", lwd = 0 ) blank_annotation<- annotation( x1 = coordinate_range[1], text = "", rot = 75 ) first_dna_seg_name<- names(dna_segs_for_plot)[1] if (is.null(first_dna_seg_name) || is.na(first_dna_seg_name)) { first_dna_seg_name<- "" } dna_segs_final<- list( dna_segs_for_plot[[1]], blank_dna_seg ) names(dna_segs_final)<- c(first_dna_seg_name, "") annotations_final<- list( annotations[[1]], blank_annotation ) list( dna_segs_final = dna_segs_final, annotations_final = annotations_final, pdf_height = 5 ) } plot_data<- add_blank_line_if_single( dna_segs_for_plot, annotations ) output_dir<- dirname(output_file) if (!dir.exists(output_dir)) { dir.create(output_dir, recursive = TRUE) } pdf( output_file, width = 15, height = plot_data pdf_height ) tryCatch( { plot_gene_map( dna_segs = plot_data dna_segs_final, annotations = plot_data annotations_final, annotation_cex = 0.39, annotation_height = 5, dna_seg_label_cex = 0.5 ) }, finally = { dev.off() } )。
[0087] In step S1742, when the arrangement directions are inconsistent, the relevant contig fragments are reversed to draw the gene structure diagram of the target integrative conjugating element.
[0088] When the arrangement directions are inconsistent, the sequence of the target integrative conjugating element contained in the relevant contig fragment is reversed so that the locus coordinates change from ascending to descending. Then, the gene structure diagram of the target integrative conjugating element is drawn based on the result of the reversal.
[0089] It should be noted that after the orientation is reversed, the order and orientation of all genes will be reversed accordingly, aligning them with the arrangement orientation of the core genes in the representative element of the target integrative conjugating element. Then, a gene structure map of the target integrative conjugating element is drawn based on the orientation-reversed data.
[0090] In one embodiment, script 2 can be used to reverse the orientation and plot the sequence of the target integrative conjugating element contained in the relevant contig fragment. The difference between script 2 and script 1 is that the sequence of the target integrative conjugating element is reversed, and gene annotation is performed based on the reversed sequence orientation.
[0091] Script 2: library(genoPlotR) library(dplyr) library(purrr) library(openxlsx) input_file<- "input_excel.xlsx" output_file<- "output_fig.pdf" do_reverse<- TRUE df<- read.xlsx(input_file, colNames = TRUE) required_columns<- c( "Contig", "Gene", "Start", "End", "Direction", "Color" ) if (!all(required_columns %in% colnames(df))) { stop( "The input file must contain columns: ", paste(required_columns, collapse = ", ") ) } df<- df %>% mutate( Contig = as.character(Contig), Gene = as.character(Gene), Start = as.numeric(Start), End = as.numeric(End), Direction = as.numeric(Direction), Color = as.character(Color) ) %>% filter( !is.na(Contig), !is.na(Gene), !is.na(Start), !is.na(End), !is.na(Direction), !is.na(Color), Color != "" ) if (nrow(df) == 0) { stop("No valid gene annotation data was found.") } invalid_directions<- setdiff(unique(df Direction), c(-1, 1)) if (length(invalid_directions)>0) { stop("Direction must be 1 or -1.") } is_valid_color<- function(color_value) { tryCatch( { col2rgb(color_value) TRUE }, error = function(e) FALSE ) } invalid_colors<- unique(df Color[!sapply(df Color, is_valid_color)]) if (length(invalid_colors)>0) { stop( "Invalid color values found: ", paste(invalid_colors, collapse = ", ") ) } genomes_id<- unique(df Contig) if (length(genomes_id) == 0) { stop("No valid Contig was found.") } dna_segs<- map(genomes_id, function(contig_id) { genome_df<- df %>% filter(.data Contig == .env contig_id) dna_seg_df<- data.frame( name = genome_df Gene, start = genome_df Start, end = genome_df End, strand = genome_df Direction, fill = genome_df Color, stringsAsFactors = FALSE ) as.dna_seg( dna_seg_df, fill = dna_seg_df fill, col = dna_seg_df fill, lwd = 0 ) }) %>% set_names(genomes_id) if (do_reverse) { dna_segs_for_plot<- lapply(dna_segs, reverse) names(dna_segs_for_plot)<- names(dna_segs) } else { dna_segs_for_plot<- dna_segs } annotations<- lapply(dna_segs_for_plot, function(dna_seg) { annotation( x1 = middle(dna_seg), text = dna_seg name, rot = 75 ) }) add_blank_line_if_single<- function(dna_segs_for_plot, annotations) { if (length(dna_segs_for_plot) != 1) { return(list( dna_segs_final = dna_segs_for_plot, annotations_final = annotations, pdf_height = 6 )) } dna_seg<- dna_segs_for_plot[[1]] coordinate_range<- range( c(dna_seg start, dna_seg end), na.rm = TRUE ) if (!all(is.finite(coordinate_range)) || diff(coordinate_range) == 0) { coordinate_range<- c(0, 1) } blank_dna_seg_df<- data.frame( name = "", start = coordinate_range[1], end = coordinate_range[1] + max(1, diff(coordinate_range) 0.01), strand = 1, stringsAsFactors = FALSE ) blank_dna_seg<- as.dna_seg( blank_dna_seg_df, fill = "white", col = "white", lwd = 0 ) blank_annotation<- annotation( x1 = coordinate_range[1], text = "", rot = 75 ) first_dna_seg_name<- names(dna_segs_for_plot)[1] if (is.null(first_dna_seg_name) || is.na(first_dna_seg_name)) { first_dna_seg_name<- "" } dna_segs_final<- list( dna_segs_for_plot[[1]], blank_dna_seg ) names(dna_segs_final)<- c(first_dna_seg_name, "") annotations_final<- list( annotations[[1]], blank_annotation ) list( dna_segs_final = dna_segs_final, annotations_final = annotations_final, pdf_height = 5 ) } plot_data<- add_blank_line_if_single( dna_segs_for_plot, annotations ) output_dir<- dirname(output_file) if (!dir.exists(output_dir)) { dir.create(output_dir, recursive = TRUE) } pdf( output_file, width = 15, height = plot_data pdf_height ) tryCatch( { plot_gene_map( dna_segs = plot_data dna_segs_final, annotations = plot_data annotations_final, annotation_cex = 0.39, annotation_height = 5, dna_seg_label_cex = 0.5 ) }, finally = { dev.off() } ).
[0092] It should be noted that when the core genes are all located in the same contig fragment of the next-generation sequencing sequence, only the overall orientation of the core genes in that contig fragment is compared to the orientation of the core genes in the representative element R391 integrative conjugation element. When they are consistent, script 1 is used to draw the gene structure of the target integrative conjugation element; when they are inconsistent, script 2 is used to reverse the orientation of the sequence of the target integrative conjugation element contained in the relevant contig fragment and draw it, and gene annotation is performed based on the reversal result. When the core genes are distributed in multiple contig fragments of the next-generation sequencing sequence, the overall orientation of the core genes in each relevant contig fragment is compared to the orientation of the core genes in the corresponding region of the representative element R391 integrative conjugation element; when they are consistent, script 1 is used to draw the gene structure of the target integrative conjugation element; when they are inconsistent, script 2 is used to reverse the orientation of the sequence of the target integrative conjugation element contained in the relevant contig fragment and draw it, and gene annotation is performed based on the reversal result.
[0093] It should also be noted that, given the reversibility of the direction of contig fragments, the above processing method will not change the nucleotide composition and gene content of the sequence, nor will it affect the determination of the existence of related genes, thus helping to improve the consistency and comparability of comparative genomics analysis between different samples.
[0094] The integrative conjugation element analysis method provided in this invention can accurately match predicted genes within a sequence range with core genes by using labeled sequence similarity alignment results, thus achieving the distinction between core genes and putative proteins. Furthermore, by organizing the labeled gene information into structured data, a standardized input format is provided for automated plotting. The arrangement direction of core genes in each relevant contig fragment is then compared with representative elements, and a decision is made based on the comparison results regarding whether to reverse the direction. This achieves a unified display of element structure diagrams across different samples, eliminating visual differences caused by different insertion directions or the reversibility of contig directions, and facilitating cross-sample comparative genomics analysis.
[0095] Based on any of the above embodiments, when the bacterial strain to be analyzed includes multiple strains, step S130 includes: step S131 and step S132.
[0096] Step S131: Merge the predicted genes of multiple bacterial strains to be analyzed to construct a merged predicted gene set.
[0097] In practical applications, researchers typically need to screen a large number of bacterial strains to understand the distribution of SXT / R391 family integrative bonding elements in the population.
[0098] When multiple bacterial strains are being analyzed, batch automatic annotation (including open reading frame identification) is performed on the contig fragments of the next-generation sequencing sequences to obtain predicted genes for each strain. Then, the predicted genes from these multiple strains are merged to construct a merged predicted gene set, which is a single FASTA format file.
[0099] Step S132: Perform sequence similarity comparison between the merged predicted gene set and the core gene of the target integrative binding element, and screen out candidate genes that match the core gene based on the sequence similarity comparison results.
[0100] The sequence similarity of the merged predicted gene set with the core gene of the target integrative binding element is compared, and candidate genes that match the core gene are screened based on the sequence similarity comparison results.
[0101] Specifically, the nucleic acid sequence files of the aforementioned core genes are used as the query sequence, and the local nucleic acid database constructed by merging the predicted gene set is used as the target database. BLASTn (Basic Local Alignment Search Tool for Nucleotide) is used for alignment to obtain the sequence similarity comparison results.
[0102] The integrative conjugation element analysis method provided in this invention merges the predicted genes of multiple strains into a merged predicted gene set, and compares the sequence similarity of the merged predicted gene set with the core gene of the target integrative conjugation element. This allows for the screening of candidate genes for all strains with only one comparison, greatly improving analysis efficiency and making it suitable for the rapid batch identification of large numbers of integrative conjugation elements.
[0103] Based on any of the above embodiments Figure 4 This is the fourth flowchart illustrating the integrated bonding element analysis method provided by the present invention, as shown below. Figure 4 As shown, step S130 includes steps S133 and S134.
[0104] Step S133: Using the nucleic acid sequence of the core gene of the target integrative conjugating element as the query sequence and the nucleic acid sequence database constructed from the nucleic acid sequence of the predicted gene as the target database, perform BLASTn sequence alignment to obtain the sequence similarity alignment result.
[0105] The nucleic acid sequence of the core gene of the target integrative binding element is used as the query sequence.
[0106] In one embodiment, the nucleic acid sequence of the core gene can be obtained as follows: First, obtain the nucleic acid sequences of representative integrative elements and / or typical integrative elements of the SXT / R391 family from public databases; wherein, the representative integrative element of the SXT / R391 family is the R391 integrative element with accession number AY090559.1, and the typical integrative elements of the SXT / R391 family include, but are not limited to: SXT / R391 family integrative elements with accession numbers AY090559.1, AY055428.1, and GQ463144.1; then, extract the nucleic acid sequence corresponding to the core gene using a Python (a programming language) program.
[0107] In one specific implementation, the nucleic acid sequences of three typical SXT / R391 family integrative conjugating elements (accession numbers AY090559.1, AY055428.1, and GQ463144.1, respectively) are downloaded from the NCBI database. The nucleic acid sequences of the above 25 core genes are extracted from them using a Python program, and the nucleic acid sequences of these core genes are merged and saved in a fasta file named core_genes.fasta as the query sequence for subsequent alignment.
[0108] Meanwhile, a nucleic acid sequence database constructed using the nucleic acid sequences of predicted genes serves as the target database.
[0109] In one implementation, when the bacterial strain to be analyzed includes a single strain, the corresponding FASTA file is generated directly from the ffn file obtained by annotating the strain using the Prokka program, and the local BLAST database is constructed using the makeblastdb command with the FASTA file as the target sequence.
[0110] In another implementation, when multiple bacterial strains are to be analyzed, the ffn files obtained after batch automatic annotation of multiple strains using the Prokka program are merged to generate a combined_ffn.fasta file, which contains the nucleic acid sequences of all predicted genes of multiple strains. A local BLAST database, named combined_ffn_database, is constructed using the makeblastdb command with combined_ffn.fasta as the target sequence. The database construction command is: "makeblastdb -in combined_ffn.fasta -dbtype nucl -out combined_ffn_database".
[0111] Here, `makeblastdb` means to create a BLAST (Basic Local Alignment Search Tool) database; `-in combined_ffn.fasta` means to input a FASTA format file containing the nucleic acid sequences of the predicted genes of the bacterial strain to be analyzed; `-dbtype nucl` specifies that the database type is a nucleic acid database; and `-out combined_ffn_database` indicates the name of the output database.
[0112] Then, using the nucleic acid sequence files of the above 25 core genes as query sequences and the local nucleic acid database constructed above as the target database, BLASTn alignment is performed to obtain sequence similarity alignment results.
[0113] In one implementation, the results are output in the outfmt 6 (Output Format 6) format, which facilitates subsequent automated parsing and processing and is beneficial for building a batch analysis workflow.
[0114] The BLASTn search command is: "blastn -query xxx.fasta -db xxx_database -outfmt "6qseqid sseqid pident length qcovs qcovhsp qcovus mismatch gapopen qstart qendsstart send evalue bitscore" -out xxx_blastn.txt".
[0115] Here, `blastn` indicates that a BLASTn search will be performed; `-query xxx.fasta` indicates that the nucleic acid sequence file xxx.fasta of the above 25 core genes is used as the query sequence file; `-db xxx_database` indicates that the local nucleic acid database xxx_database constructed above is used as the target database; `-outfmt “6 qseqid sseqid pident lengthqcovs qcovhsp qcovus mismatch gapopen qstart qend sstart send evaluebitscore”` indicates that the results are output in tabular format 6, and specifies the output format parameters including: query sequence ID (Identifier), target sequence ID, consistency percentage, alignment length, coverage of the query sequence to the target sequence, and a single HSP (High-scoring Segment). The output includes query coverage of the pair (high-scoring fragment pair), query coverage of each unique target sequence, number of mismatches, number of empty pairs, start coordinates on the query sequence, end coordinates on the query sequence, start coordinates on the target sequence, end coordinates on the target sequence, expected value, and standardized alignment score. -outxxx_blastn.txt indicates that the output file is xxx_blastn.txt.
[0116] Step S134: Based on the sequence similarity comparison results, candidate genes that meet preset similarity conditions are selected. These preset similarity conditions include: sequence identity of at least 85%, sequence coverage of at least 85%, and an expected value of less than 0.001.
[0117] After obtaining the sequence similarity alignment results, i.e., the BLASTn output file, the output file is parsed and filtered using preset screening criteria. Predicted genes that meet the preset screening criteria are marked as candidate genes for core genes.
[0118] In one implementation, the preset screening criteria are: identity (referring to the pident of the aforementioned output) ≥ 85%, coverage (referring to the qcovs of the aforementioned output) ≥ 85%, and e-value (expected value) < 0.001. By using these screening criteria, random and low-quality matches are effectively filtered out while ensuring detection sensitivity (allowing a certain degree of natural variation), reducing the false positive rate and improving the accuracy of candidate gene screening.
[0119] The integrated conjugation element analysis method provided in this invention achieves high-throughput batch searching of a large number of predicted genes by constructing a local nucleic acid sequence database and performing BLASTn alignment. Compared with one-by-one alignment or online submission, this invention significantly improves processing efficiency.
[0120] Based on any of the above embodiments, step S120 includes: step S121.
[0121] Step S121: The Prokka program is used to annotate the contig fragment of the second-generation sequencing sequence to obtain the nucleic acid sequence of the predicted gene, its position information in the contig fragment, and the gene orientation.
[0122] The Prokka program is used to annotate contig fragments of the next-generation sequencing sequences of the bacterial genome to be analyzed. The Prokka execution command is: "prokka xxx.fasta --outdir xxx_output --prefix xxx --locustag xxx --kingdom Bacteria".
[0123] Here, xxx.fasta represents the FASTA format file of the input second-generation sequencing sequence, --outdir xxx_output specifies the output folder as xxx_output, --prefix xxx specifies the prefix name of the output file, --locustagxxx specifies the prefix number of each gene, and --kingdom Bacteria specifies the species type as bacteria.
[0124] After the command is executed, multiple files will be generated in the xxx_output directory, including ffn (a file containing the predicted gene nucleotide sequence stored in FASTA format) and gff (general feature format) files. The ffn file contains the nucleic acid sequences of all predicted genes, and the gff file contains the location information and gene orientation of each predicted gene.
[0125] When there are multiple bacterial strains to be analyzed, Prokka annotation commands corresponding to each strain are generated and written line by line into the command file prokka_lines.sh. The command file is then converted and executed under the Linux system to achieve batch automatic annotation of the genome sequences of multiple bacterial strains. The batch automatic execution command is: "sed -i 's / \r / / ' prokka_lines.sh&&bash prokka_lines.sh".
[0126] The method for analyzing integrated conjugating elements provided in this invention uses the Prokka program for annotation. This tool is specifically optimized for bacterial genomes and features fast running speed and high annotation accuracy.
[0127] Furthermore, the solutions of the embodiments of the present invention will be described further through the following examples.
[0128] (1) Download the SXT / R391 family integrative conjugation element FASTA format nucleic acid sequences numbered AY055428.1, AY090559.1, and GQ463144.1 from the NCBI database. Use Python to extract the corresponding nucleic acid sequences of the core genes, including the integration / splitting module genes xis and int, the DNA processing module genes mobI, traI, traD, and traJ, the conjugation transfer module genes traL, traE, traK, traB, traV, traA, s054, traC, trhF, traW, traU, traN, traF, traH, and traG, the regulatory module genes setC, setD, and setR, and other module genes s063. Place the nucleic acid sequences of these 25 genes into a single FASTA file named core_genes.fasta.
[0129] The gene codes for these core genes are shown in Table 1 below.
[0130] Table 1
[0131] The sequences numbered VIB_AA5102AA_AS, VIB_AA5525AA_AS, and VIB_AA5265AA_AS in the EnteroBase database are bacterial next-generation sequencing data. Download these next-generation sequencing data in FASTA format.
[0132] (2) Use the Prokka program to automatically annotate the nucleic acid sequences of strains VIB_AA5102AA_AS, VIB_AA5525AA_AS and VIB_AA5265AA_AS in batches. The Prokka execution command is: "prokka xxx.fasta --outdir xxx_output --prefix xxx --locustag xxx --kingdom Bacteria".
[0133] Write the Prokka comment commands corresponding to each bacterial strain to be analyzed line by line into the command file prokka_lines.sh. Then, on a Linux system, convert and execute the command file using the command: "sed -i 's / \r / / ' prokka_lines.sh&&bash prokka_lines.sh".
[0134] (3) Merge the ffn files from the Prokka results of the VIB_AA5102AA_AS, VIB_AA5525AA_AS, and VIB_AA5265AA_AS strains, and name them combined_ffn.fasta. Run the local BLASTn tool, build a library using the combined_ffn.fasta sequence, and name it combined_ffn_database. Use the fasta sequence file core_genes.fasta of 25 core genes as the query to perform a BLASTn search, and name the output results SXT_R391_blastn.txt.
[0135] The command to create the database is: "makeblastdb -in combined_ffn.fasta -dbtype nucl -out combined_ffn_database".
[0136] The BLASTn search command is: "blastn -query core_genes.fasta -db combined_ffn_database -outfmt "6 qseqid sseqid pident length qcovs qcovhsp qcovus mismatchgapopen qstart qend sstart send evalue bitscore" -out SXT_R391_blastn.txt".
[0137] Here, `blastn` indicates that a BLASTn search will be performed; `-query core_genes.fasta` indicates that the nucleic acid sequence file `core_genes.fasta` of the above 25 core genes is used as the query sequence file; `-db combined_ffn_database` indicates that the local nucleic acid database `combined_ffn_database` constructed above is used as the target database; `-outfmt “6qseqid sseqid pident length qcovs qcovhsp qcovus mismatch gapopen qstart qendsstart send evalue bitscore”` indicates that the results will be output in tabular format 6, and specifies the output format parameters including: query sequence ID (Identifier), target sequence ID, consistency percentage, alignment length, coverage of the query sequence to the target sequence, and a single HSP (High-scoring Segment). The query coverage of the pair (high-scoring fragment pair), the query coverage of each unique target sequence, the number of mismatches, the number of empty pairs, the start coordinates on the query sequence, the end coordinates on the query sequence, the start coordinates on the target sequence, the end coordinates on the target sequence, the expected value, and the standardized alignment score. -out SXT_R391_blastn.txt indicates that the output file is SXT_R391_blastn.txt.
[0138] The BLASTn search results were filtered based on the criteria of identity ≥ 85%, coverage ≥ 85%, and e-value < 0.001. A Python program was used to extract the BLASTn identification results for each strain.
[0139] Based on the above procedure, the identification results of the core gene of the SXT / R391 family integrative conjugation element in strain VIB_AA5102AA_AS are shown in Table 2 below.
[0140] Table 2
[0141] Based on the above procedure, the identification results of the core gene of the SXT / R391 family integrative conjugation element in strain VIB_AA5525AA_AS are shown in Table 3 below.
[0142] Table 3
[0143] Based on the above procedure, the identification results of the core gene of the SXT / R391 family integrative conjugation element in strain VIB_AA5265AA_AS are shown in Table 4 below.
[0144] Table 4
[0145] (4) Since the Prokka program numbers the identified coding sequence, the gene name in the obtained BLASTn sequence similarity comparison results is also the number provided by Prokka. The Python program is used to obtain the contig number corresponding to the gene number of the strain, as well as the start and end positions of the gene in the contig fragment of the strain and the direction of the gene from the gff file in the Prokka results, and to mark it in the BLASTn sequence similarity comparison results.
[0146] In strain VIB_AA5102AA_AS, a total of 3 contig sequences contain the core gene of the SXT / R391 family integrative conjugation element. The identification and labeling results are shown in Table 5 below.
[0147] Table 5
[0148] In strain VIB_AA5525AA_AS, a total of 2 contig sequences contain the core gene of the SXT / R391 family integrative conjugation element. The identification and labeling results are shown in Table 6 below.
[0149] Table 6
[0150] Based on the above results, a total of 1 contig sequence in strain VIB_AA5265AA_AS contains the core gene of the SXT / R391 family integrative conjugation element. The identification and labeling results are shown in Table 7 below.
[0151] Table 7
[0152] (5) In strain VIB_AA5102AA_AS, based on the identification results, xis and int are located on the same contig fragment NODE_7_length_178429_cov_22.6336, and the int gene is located to the left of the xis gene. It is determined that the SXT / R391 family integrative conjugation element fragment in this region is reversed compared to the corresponding region of the R391 element. Therefore, the area from the left end of the NODE_7_length_178429_cov_22.6336 fragment to the end of the xis gene is identified as part of the SXT / R391 family integrative conjugation element (the NODE_7_length_178429_cov_22.6336 fragment from 1 bp to the end of the xis gene is 8041 bp, which is part of the SXT / R391 family integrative conjugation element). (Part of a family integrative conjugating element); setD and setR are on the same contig fragment NODE_18_length_75091_cov_22.7604, and setD is located to the right of setR. It is determined that the SXT / R391 family integrative conjugating element fragment in this region is reversed compared to the corresponding region of the R391 element. Therefore, the area from the start of the setR gene to the right end of the NODE_18_length_75091_cov_22.7604 fragment is considered part of the SXT / R391 family integrative conjugating element (the NODE_18_length_75091_cov_22.7604 fragment from 59042bp to the right end of 75091bp is part of the SXT / R391 family integrative conjugating element). The contig fragment numbered NODE_23_length_55820_cov_24.4889 contains the core genes of other SXT / R391 family integrative binding elements that are not xis and setR, and is considered to be an intermediate fragment of SXT / R391 family integrative binding elements.
[0153] In strain VIB_AA5525AA_AS, based on the identification results, xis and int are located on the same contig fragment NODE_8_length_183464_cov_17.5114, and the int gene is located to the right of the xis gene. This indicates that the SXT / R391 family integrative conjugation element fragment in this region is positively aligned compared to the corresponding region of the R391 element. Therefore, the area from the xis gene in the NODE_8_length_183464_cov_17.5114 fragment to the right end of the contig fragment is considered part of the SXT / R391 family integrative conjugation element (the NODE_8_length_183464_cov_17.5114 fragment starts at 175457 bp and ends at 183465 bp on the right end of the contig fragment). (This refers to a part of the SXT / R391 family integrative conjugating element). setD and setR are on the same contig fragment NODE_10_length_134312_cov_17.554, and setD is located to the left of setR. It is determined that the SXT / R391 family integrative conjugating element fragment in this region is positive relative to the corresponding region of the R391 element. Therefore, the area from the left start of the NODE_10_length_134312_cov_17.554 fragment to the end of the setR gene is considered to be part of the SXT / R391 family integrative conjugating element (the NODE_10_length_134312_cov_17.554 fragment from 1 bp to the end of the setR gene 75263 bp is part of the SXT / R391 family integrative conjugating element).
[0154] In strain VIB_AA5265AA_AS, based on the identification results, all core genes are located in the same contig fragment NODE_4_length_446642_cov_19.4926, and the setR gene is located to the left of the xis gene. It is determined that the SXT / R391 family integrative conjugation element fragment in this region is reversed compared to the corresponding region of the R391 element. Therefore, it is determined that the region from the start of the setR gene to the end of the xis gene is the part of the SXT / R391 family integrative conjugation element (from 59202bp to 149037bp of NODE_4_length_446642_cov_19.4926).
[0155] (6) Extract all coding sequence information within the SXT / R391 family integrative binding element region from the Prokka result gff files of VIB_AA5102AA_AS, VIB_AA5525AA_AS and VIB_AA5265AA_AS. Based on the core gene identification results in BLASTn, if a coding sequence belongs to a core gene, then label it with the core gene name; otherwise, label it as "hypothetical protein".
[0156] (7) Construct an Excel table. The first column of the table is the name of the gene in the SXT / R391 family integrative conjugation element, the second column is the start point of the gene in the contig fragment (determined according to the BLASTn result), the third column is the end point of the gene in the contig fragment (determined according to the BLASTn result), the fourth column is the direction of the gene, the fifth column is the color of the gene (the core gene is marked with sky blue, and other genes are marked with gray), and the sixth column is the name of the contig in which the gene is located.
[0157] Based on the above requirements, the gene distribution of SXT / R391 family integrative conjugation elements in strain VIB_AA5102AA_AS is shown in Table 8 below.
[0158] Table 8
[0159] Based on the above requirements, the gene distribution of SXT / R391 family integrative conjugating elements in strain VIB_AA5525AA_AS is shown in Table 9 below.
[0160] Table 9
[0161] Based on the above requirements, the gene distribution of SXT / R391 family integrative conjugation elements in strain VIB_AA5265AA_AS is shown in Table 10 below.
[0162] Table 10
[0163] In strain VIB_AA5102AA_AS, the SXT / R391 family integrative conjugating elements are located on three contig fragments, thus displayed in three rows. To facilitate visualization and comparative genomics analysis of the SXT / R391 family integrative conjugating elements, the overall orientation of the core genes in each contig fragment was compared with the orientation of the corresponding core genes in the R391 integrative conjugating element. It was found that the orientations of the core genes in the three contig fragments were opposite to those in the corresponding R391 integrative conjugating element. Therefore, script 2 was used to reverse the orientation of the relevant contig fragments before plotting, and gene labeling was performed based on the reversal results. Given the reversibility of contig fragment orientation, this processing method does not change the nucleotide composition or gene content of the sequence, nor does it affect the determination of the presence of related genes, thus improving the consistency and comparability of comparative genomics analysis between different samples. The image is shown below. Figure 5 As shown, where, Figure 5 Genes marked in gray represent "hypothetical proteins".
[0164] In strain VIB_AA5525AA_AS, the SXT / R391 family integrative conjugating elements are located on two contig fragments, and therefore displayed in two rows. To facilitate visualization and comparative genomics analysis of the SXT / R391 family integrative conjugating elements, the overall arrangement orientation of the core genes in each contig fragment was compared with the arrangement orientation of the corresponding core genes in the R391 integrative conjugating element. It was found that the arrangement orientation of the core genes in both contig fragments was consistent with that in the corresponding region of the R391 integrative conjugating element; therefore, script 1 was used for plotting and labeling. The image is shown below. Figure 6 As shown, where, Figure 6 Genes marked in gray represent "hypothetical proteins".
[0165] In strain VIB_AA5265AA_AS, the SXT / R391 family integrative conjugating element is located on a single contig fragment, therefore displayed in one row. To facilitate visualization and comparative genomics analysis of the SXT / R391 family integrative conjugating element, the overall orientation of the core genes in this contig fragment was compared with the orientation of the corresponding core genes in the R391 integrative conjugating element. It was found that the orientation of the core genes in this contig fragment was opposite to that in the corresponding region of the R391 integrative conjugating element. Therefore, script 2 was used to reverse the orientation of the relevant contig fragment before plotting, and gene labeling was performed based on the reversal result. Given the reversibility of the contig fragment orientation, this processing method does not change the nucleotide composition or gene content of the sequence, nor does it affect the determination of the presence of related genes, thus improving the consistency and comparability of comparative genomics analysis between different samples. The image is shown below. Figure 7 As shown, where, Figure 7 Genes marked in gray represent "hypothetical proteins".
[0166] The integrated bonding element analysis apparatus provided by the present invention will be described below. The integrated bonding element analysis apparatus described below and the integrated bonding element analysis method described above can be referred to in correspondence.
[0167] Figure 8 This is a schematic diagram of the integrated bonding element analysis device provided by the present invention, as shown below. Figure 8 As shown, the device includes a sequence acquisition module 810, a contig identification module 820, a gene screening module 830, an element determination module 840, and a range determination module 850; wherein: The sequence acquisition module 810 is used to acquire the second-generation sequencing sequence of the bacterial strain to be analyzed; The contig identification module 820 is used to annotate the contig fragments of the second-generation sequencing sequence to obtain the predicted gene; The gene screening module 830 is used to perform sequence similarity comparison between the predicted gene and the core gene of the target integrative binding element, and to screen out candidate genes that match the core gene based on the sequence similarity comparison results. The element determination module 840 is used to determine whether the bacterial strain to be analyzed contains a target integrative conjugating element based on whether the candidate gene contains both the initiation marker gene and the terminal marker gene of the core gene. The range determination module 850 is used to determine the sequence range of the target integrative conjugating element if the bacterial strain to be analyzed contains the target integrative conjugating element, based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment.
[0168] The integrative conjugation element analysis device provided in this invention obtains the next-generation sequencing (NGS) sequence of the bacterial strain to be analyzed, annotates the contig fragments of the NGS sequence to obtain predicted genes, then compares the sequence similarity of the predicted genes with the core gene of the target integrative conjugation element and screens candidate genes. The presence of both start and end marker genes in the candidate genes determines whether the target integrative conjugation element is present. If present, the sequence range of the target integrative conjugation element is further determined based on whether the core gene is distributed across the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position of the core gene within the corresponding contig fragment. This integrative conjugation element analysis device can directly process NGS sequence data without requiring a complete genome map beforehand, reducing data acquisition costs and making large-scale strain screening economically feasible. Meanwhile, this embodiment of the invention determines the presence of a target integrative conjugating element by jointly judging the start and end marker genes. Then, based on whether the core gene is distributed in the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment, the sequence range of the target integrative conjugating element is determined. This enables rapid and accurate screening of target integrative conjugating elements and their specific ranges in a large number of strains' next-generation sequencing sequence data. Furthermore, the technical solution of this embodiment of the invention can achieve automated, streamlined processing of multiple strains through script programming.
[0169] According to an integrated bonding element analysis apparatus provided by the present invention, the integrated bonding element analysis apparatus further includes: The first annotation module is used to obtain the name of the contig fragment containing each candidate gene, the start and end positions of the contig fragment, and the gene orientation, and to annotate the sequence similarity comparison results to obtain the annotated sequence similarity comparison results.
[0170] The range determination module 850 includes: The first determining subunit is used to determine the sequence range of the target integrative conjugating element based on the relative positions of the start marker gene and the end marker gene in the annotated sequence similarity comparison results when all the core genes are located in the same contig fragment of the second-generation sequencing sequence. The second determining subunit is used to determine the sequence range of the target integrative conjugating element when the core gene is distributed across multiple contig fragments of the next-generation sequencing sequence, by means of: For a first contig fragment containing the start marker gene, a first sorting direction is determined based on the relative position of the start marker gene and its adjacent core genes in the labeled sequence similarity comparison results. Based on the first sorting direction and the start marker gene, a first target fragment is extracted from the first contig fragment as the first part of the target integrative conjugation element. For a second contig fragment containing the terminal marker gene, a second sorting direction is determined based on the relative position of the terminal marker gene and its adjacent core gene in the labeled sequence similarity comparison results. Based on the second sorting direction and the terminal marker gene, a second target fragment is extracted from the second contig fragment as the second part of the target integrative conjugation element. For a third contig fragment that does not contain the start marker gene and the end marker gene, but contains at least one other core gene, the third contig fragment is used as the third part of the target integrative conjugation element; Based on the first part, the second part, and the third part, the sequence range of the target integrated contig fragments is determined.
[0171] According to an integrated bonding element analysis apparatus provided by the present invention, the integrated bonding element analysis apparatus further includes: The second annotation module is used to annotate each gene within the sequence range based on the annotated sequence similarity comparison results to obtain annotated genes; wherein, genes that match the core gene are annotated as core gene names, and genes that do not match are annotated as putative proteins; The information processing module is used to obtain the target information of the labeled genes and process the target information into structured data; wherein, the target information includes the gene name, the start and end positions in the contig fragment, the annotation color, the gene direction, and the name of the contig in which the gene is located; The orientation comparison module is used to compare the overall arrangement orientation of the core genes in each relevant contig fragment with the arrangement orientation of the core genes in a representative element of the target integrative conjugating element based on the structured data. The first drawing module is used to draw the gene structure diagram of the target integrative conjugating element according to the actual arrangement direction of the core genes in the relevant contig fragment when the arrangement direction is consistent. The second drawing module is used to reverse the orientation of the relevant contig fragments when the arrangement orientations are inconsistent, and to draw the gene structure map of the target integrative conjugating element.
[0172] According to the present invention, an integrated conjugation element analysis device is provided, wherein when the bacterial strain to be analyzed comprises multiple strains, the gene screening module 830 is specifically used for: The predicted genes of multiple bacterial strains to be analyzed are merged to construct a merged predicted gene set; The merged predicted gene set is compared with the core gene of the target integrative binding element for sequence similarity, and candidate genes that match the core gene are selected based on the sequence similarity comparison results.
[0173] According to an integrated conjugation element analysis device provided by the present invention, the gene screening module 830 is further specifically used for: Using the nucleic acid sequence of the core gene of the target integrative conjugating element as the query sequence and the nucleic acid sequence database constructed using the nucleic acid sequence of the predicted gene as the target database, BLASTn sequence alignment is performed to obtain the sequence similarity alignment results; Based on the sequence similarity comparison results, candidate genes that meet the preset similarity conditions are selected; The preset similarity conditions include: sequence consistency of not less than 85%, sequence coverage of not less than 85%, and expected value of less than 0.001.
[0174] According to an integrated bonding element analysis device provided by the present invention, the contig identification module 820 is specifically used for: The Prokka program was used to annotate the contig fragments of the second-generation sequencing sequence to obtain the nucleic acid sequence of the predicted gene, its position information in the contig fragment, and the gene orientation.
[0175] According to the present invention, an integrative conjugation element analysis device is provided, wherein the target integrative conjugation element is an SXT / R391 family integrative conjugation element, and the core gene includes the following 25 genes: xis, int, mobI, traI, traD, traJ, traL, traE, traK, traB, traV, traA, s054, traC, trhF, traW, traU, traN, s063, traF, traH, traG, setC, setD, and setR, wherein the initiation marker gene is xis and the terminal marker gene is setR.
[0176] It should be noted that the integrated bonding element analysis device provided in this embodiment of the invention can implement all the method steps implemented in the integrated bonding element analysis method embodiment and achieve the same technical effect. Therefore, the parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0177] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9 As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communications bus 940, wherein the processor 910, the communications interface 920, and the memory 930 communicate with each other through the communications bus 940. The processor 910 can call logical instructions in the memory 930 to execute an integrative conjugation element analysis method, which includes: acquiring the next-generation sequencing sequence of the bacterial strain to be analyzed; annotating the contig fragments of the next-generation sequencing sequence to obtain a predicted gene; performing sequence similarity comparison between the predicted gene and the core gene of the target integrative conjugation element, and screening candidate genes that match the core gene based on the sequence similarity comparison results; determining whether the bacterial strain to be analyzed contains the target integrative conjugation element based on whether the candidate gene simultaneously contains the start marker gene and the end marker gene of the core gene; if the bacterial strain to be analyzed contains the target integrative conjugation element, determining the sequence range of the target integrative conjugation element based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains the start or end marker gene, and the position information of the core gene in the corresponding contig fragment.
[0178] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0179] On the other hand, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for analyzing integrative conjugating elements provided by the above methods. This method includes: obtaining a second-generation sequencing sequence of a bacterial strain to be analyzed; annotating contig fragments of the second-generation sequencing sequence to obtain a predicted gene; performing sequence similarity comparison between the predicted gene and a core gene of a target integrative conjugating element, and screening candidate genes matching the core gene based on the sequence similarity comparison results; determining whether the bacterial strain to be analyzed contains a target integrative conjugating element based on whether the candidate genes simultaneously contain a start marker gene and a terminal marker gene from the core gene; if the bacterial strain to be analyzed contains a target integrative conjugating element, determining the sequence range of the target integrative conjugating element based on the distribution of the core gene in the same or multiple contig fragments, whether the contig fragment containing the core gene contains a start or terminal marker gene, and the position information of the core gene in the corresponding contig fragment.
[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An integrated junction element analysis method characterized by, include: Obtain the next-generation sequencing sequence of the bacterial strain to be analyzed; Annotate the contig fragments of the second-generation sequencing sequence to obtain the predicted gene; The predicted gene is compared with the core gene of the target integrative binding element for sequence similarity, and candidate genes that match the core gene are selected based on the sequence similarity comparison results. Based on whether the candidate genes simultaneously contain the initiation marker gene and the terminal marker gene of the core gene, it is determined whether the bacterial strain to be analyzed contains the target integrative conjugation element. If the bacterial strain to be analyzed contains a target integrative conjugating element, the sequence range of the target integrative conjugating element is determined based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment.
2. The method for analyzing integrated bonding elements according to claim 1, characterized in that, Before determining the sequence range of the target integrative conjugating element based on the distribution of the core gene in the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment, the method further includes: The name of the contig fragment containing each candidate gene, its start and end positions within the contig fragment, and its gene orientation are obtained and annotated in the sequence similarity comparison results to obtain the annotated sequence similarity comparison results; The determination of the sequence range of the target integrative conjugating element based on the distribution of the core gene in the same or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment includes: When the core genes are all located in the same contig fragment of the next-generation sequencing sequence, the sequence range of the target integrative conjugating element is determined based on the relative positions of the start marker gene and the end marker gene in the annotated sequence similarity comparison results. When the core gene is distributed across multiple contig fragments in the next-generation sequencing sequence, the sequence range of the target integrative conjugating element is determined as follows: For a first contig fragment containing the start marker gene, a first sorting direction is determined based on the relative position of the start marker gene and its adjacent core genes in the labeled sequence similarity comparison results. Based on the first sorting direction and the start marker gene, a first target fragment is extracted from the first contig fragment as the first part of the target integrative conjugation element. For a second contig fragment containing the terminal marker gene, a second sorting direction is determined based on the relative position of the terminal marker gene and its adjacent core gene in the labeled sequence similarity comparison results. Based on the second sorting direction and the terminal marker gene, a second target fragment is extracted from the second contig fragment as the second part of the target integrative conjugation element. For a third contig fragment that does not contain the start marker gene and the end marker gene, but contains at least one other core gene, the third contig fragment is used as the third part of the target integrative conjugation element; Based on the first part, the second part, and the third part, the sequence range of the target integrated contig fragments is determined.
3. The method for analyzing integrated bonding elements according to claim 2, characterized in that, The integrated bonding element analysis method further includes: Based on the sequence similarity comparison results after annotation, each gene within the sequence range is labeled to obtain the labeled genes; wherein, genes that match the core gene are labeled as core gene names, and genes that do not match are labeled as putative proteins. Obtain the target information of the labeled genes and organize the target information into structured data; wherein, the target information includes gene name, start and end positions in the contig fragment, annotation color, gene direction and name of the contig to which the gene is located; Based on the structured data, the overall arrangement direction of the core genes in each relevant contig fragment is compared with the arrangement direction of the core genes in the representative element of the target integrative conjugation element. When the arrangement directions are consistent, the gene structure diagram of the target integrative conjugating element is drawn according to the actual arrangement direction of the core genes in the relevant contig fragment; When the arrangement directions are inconsistent, the relevant contig fragments are reversed to draw the gene structure diagram of the target integrative conjugating element.
4. The method for analyzing integrated bonding elements according to claim 1, characterized in that, When the bacterial strain to be analyzed includes multiple strains, the step of performing sequence similarity comparison between the predicted gene and the core gene of the target integrative binding element, and screening candidate genes that match the core gene based on the sequence similarity comparison results, includes: The predicted genes of multiple bacterial strains to be analyzed are merged to construct a merged predicted gene set; The merged predicted gene set is compared with the core gene of the target integrative binding element for sequence similarity, and candidate genes that match the core gene are selected based on the sequence similarity comparison results.
5. The method for analyzing integrated bonding elements according to claim 1, characterized in that, The step of performing sequence similarity comparison between the predicted gene and the core gene of the target integrative binding element, and screening candidate genes that match the core gene based on the sequence similarity comparison results, includes: Using the nucleic acid sequence of the core gene of the target integrative conjugating element as the query sequence and the nucleic acid sequence database constructed using the nucleic acid sequence of the predicted gene as the target database, BLASTn sequence alignment is performed to obtain the sequence similarity alignment results; Based on the sequence similarity comparison results, candidate genes that meet the preset similarity conditions are selected; The preset similarity conditions include: sequence consistency of not less than 85%, sequence coverage of not less than 85%, and expected value of less than 0.
001.
6. The method for analyzing integrated bonding elements according to any one of claims 1 to 5, characterized in that, The annotation of the contig fragments of the second-generation sequencing sequence to obtain the predicted gene includes: The Prokka program was used to annotate the contig fragments of the second-generation sequencing sequence to obtain the nucleic acid sequence of the predicted gene, its position information in the contig fragment, and the gene orientation.
7. The method for analyzing integrated bonding elements according to any one of claims 1 to 5, characterized in that, The target integrative conjugating element is an SXT / R391 family integrative conjugating element, and the core genes include the following 25 genes: xis, int, mobI, traI, traD, traJ, traL, traE, traK, traB, traV, traA, s054, traC, trhF, traW, traU, traN, s063, traF, traH, traG, setC, setD, and setR. The initiation marker gene is xis, and the terminal marker gene is setR.
8. An integrated bonding element analysis device, characterized in that, include: The sequence acquisition module is used to acquire the next-generation sequencing sequence of the bacterial strain to be analyzed. The contig identification module is used to annotate the contig fragments of the second-generation sequencing sequence to obtain the predicted gene; The gene screening module is used to perform sequence similarity comparison between the predicted gene and the core gene of the target integrative binding element, and to screen out candidate genes that match the core gene based on the sequence similarity comparison results. The element determination module is used to determine whether the bacterial strain to be analyzed contains the target integrative conjugating element based on whether the candidate gene simultaneously contains the initiation marker gene and the terminal marker gene of the core gene. The range determination module is used to determine the sequence range of the target integrative conjugating element if the bacterial strain to be analyzed contains the target integrative conjugating element, based on the distribution of the core gene in the same contig fragment or multiple contig fragments, whether the contig fragment containing the core gene contains start or end marker genes, and the position information of the core gene in the corresponding contig fragment.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the integrated bonding element analysis method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the integrated bonding element analysis method as described in any one of claims 1 to 7.