A method for efficiently identifying mobile genetic elements in bacterial plasmids
By establishing a database and using the Blastn tool for searching and screening, combined with homologous sequence alignment and recognition standards, the problems of slow identification speed and high false positive rate in the prior art are solved, and efficient and accurate identification of mobile genetic elements in the cells of bacterial plasmids are achieved.
Patent Information
- Application Number
- CN202411640493.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-18
AI Technical Summary
The existing methods of identifying mobile genetic elements in bacterial plasmid genomes have problems such as slow recognition speed, limited scope of application and high false positive rate, making it difficult to achieve efficient and accurate identification in batch processing.
By establishing a database and using local Blastn tools, the gene sequences of transposons, integrators and insertion sequences are used as query sequences, the genomic sequences of the to be analyzed are searched and screened, and the genomic sequence alignment standards and identification standards are combined to achieve efficient identification of intracellular mobile genetic elements.
It realizes automated batch processing, efficient, accurate and convenient identification of mobile genetic elements in the cell of bacterial plasmids, and improves the accuracy and convenience of identification.
Smart Images

Figure CN119626328B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of microbial bioinformatics analysis, and in particular to a method for efficiently identifying mobile genetic elements in bacterial plasmids. Background Art
[0002] A variety of intracellular mobile genetic elements (IMGEs) are widely present in bacterial plasmid genomes. These elements are DNA fragments that can move within the genome and typically include transposons, integrons, and insertion sequences. These IMGEs have attracted significant attention because they often contain multiple drug resistance genes and virulence factors, and their ability to move within the genome could mediate the widespread dissemination of drug resistance genes. Therefore, identifying the distribution and types of IMGEs within bacterial plasmid genome sequences is crucial for plasmid genome analysis.
[0003] Current methods for identifying mobile genetic elements in bacterial plasmid genomes typically use online tools or localized tools. Using online tools involves submitting the bacterial plasmid genome sequence to a website and waiting to receive the results.
[0004] This solution has the following disadvantages:
[0005] 1. This method is only suitable for searching for mobile genetic elements in a small number of plasmid genome sequences. Since only one genome sequence can be uploaded at a time, it is difficult to efficiently and quickly complete the task of identifying mobile genetic elements when faced with a large number of microbial genome sequences.
[0006] 2. Mobile genetic element recognition is slow. Web-based mobile genetic element recognition is limited by the number of real-time visitors to the website and the network speed, so the recognition speed is usually slow.
[0007] The method of using localized tools for identification is to perform local homology searches and comparisons based on existing integron, transposon, and insertion sequence data. This method also has certain disadvantages:
[0008] This method can only search for integrons, transposons, and insertion sequences separately, resulting in a high false-positive rate for identified mobile genetic elements. In bacterial plasmid genomes, mobile genetic elements are often nested; for example, a transposon contains multiple insertion sequences and integrons. When searching for intracellular mobile genetic elements separately, duplicate identification of the same gene region is common, making it difficult to accurately determine the specific type and distribution of intracellular mobile genetic elements.
[0009] Therefore, there is a need for a method that can realize automated batch processing, is efficient, accurate, and convenient, and can effectively identify mobile genetic elements in bacterial plasmids. Summary of the Invention
[0010] The content of this application is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this application is not intended to identify key features or essential features of the technical solution for which protection is sought, nor is it intended to limit the scope of the technical solution for which protection is sought.
[0011] In order to solve the technical problems mentioned in the above background technology section, some embodiments of the present application provide a method for efficiently identifying mobile genetic elements in bacterial plasmids, comprising the following steps:
[0012] Step 1: Create a database using the genome sequence of the plasmid to be analyzed, and use the gene sequence of the mobile genetic element in the bacterial plasmid as the query sequence to search for the genome sequence of the plasmid to be analyzed in the database;
[0013] Step 2: Screening the search results; identifying the screening results according to the identification criteria to obtain a separate identification result of the type of intracellular mobile genetic element contained in the plasmid to be analyzed;
[0014] Step 3: Merge the identification results of the mobile genetic elements in the plasmids to be analyzed, and obtain the final identification result of the mobile genetic elements in the plasmids to be analyzed according to the identification criteria.
[0015] Furthermore, the intracellular mobile genetic elements include transposons, integrons and insertion sequences.
[0016] Furthermore, in step 1, the gene sequence of the bacterial plasmid intracellular mobile genetic element is collected and obtained through public literature and public databases.
[0017] Furthermore, in step 1, the search method is: running a local Blastn tool in a computer system, using the gene sequences of the transposon, integron and insertion sequence as query sequences, and performing a Blastn search on the gene sequence of the plasmid to be analyzed.
[0018] Furthermore, in step 2, the screening method is: screening the search results according to conventional standards for homologous sequence alignment.
[0019] Furthermore, the conventional standards for homologous sequence alignment are: e-value less than 0.001, identity greater than or equal to 50% and coverage greater than or equal to 50%.
[0020] Furthermore, the identification criteria are:
[0021] In a single plasmid, when there is a complete containment relationship between two intracellular mobile genetic elements, the recognition result with a wider interval is retained;
[0022] When there is an intersection region between two intracellular mobile genetic elements, and the intersection region is less than or equal to 50 bp, it is judged as two different intracellular mobile genetic element identification results;
[0023] When there is an intersection region between two intracellular mobile genetic elements, and the intersection region is greater than 50 bp, the identification result with the longer length is retained;
[0024] When there is an intersection region between two intracellular mobile genetic elements and the non-intersection region is less than 50 bp, they are judged as the same recognition result, and the recognition result with the longer length is retained.
[0025] The beneficial effect of the present application is that it provides a method for efficiently identifying mobile genetic elements in bacterial plasmids that can realize automated batch processing and is efficient, accurate and convenient. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings constituting a part of this application are used to provide a further understanding of this application and make other features, purposes and advantages of this application more apparent. The drawings and descriptions of the exemplary embodiments of this application are used to explain this application and do not constitute an improper limitation on this application.
[0027] In addition, throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the elements and components are not necessarily drawn to scale.
[0028] In the attached figure:
[0029] Figure 1 It is a schematic diagram of the overall process according to an embodiment of the present application;
[0030] Figure 2 is a possible pattern of recognition of mobile genetic elements in bacterial plasmid genomes. DETAILED DESCRIPTION
[0031] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0032] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0034] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0035] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0036] Example, a method for efficiently identifying mobile genetic elements in bacterial plasmids, referring to Figure 1 , including the following steps:
[0037] Step 1: Establish a database using the genome sequence of the plasmid to be analyzed, and use the gene sequence of the mobile genetic element in the bacterial plasmid as a query sequence to search the genome sequence of the plasmid to be analyzed in the database; the gene sequence of the mobile genetic element in the bacterial plasmid is collected from public literature and public databases; the search method is: run a local Blastn tool in a computer system, use the gene sequence of the transposon, integron and insertion sequence as query sequences, and perform a Blastn search on the gene sequence of the plasmid to be analyzed; the specific process is as follows:
[0038] Create a working folder, Plasmid_analysis, and a subfolder, Plasmids. Download the plasmid nucleic acid genomic sequences with accession numbers NZ_CP017387 and NZ_AP022395 from public databases. Merge the two genomic sequences into a single file, named Plasmid_sequences.fasta, and save it in the Plasmids folder.
[0039] Run the local Blastn tool and create a database named Plasmid_sequences_database with the Plasmid_sequences.fasta sequence. Perform Blastn searches with the Insertion / Integron / Transposon sequences (Insertion.fasta / Integron.fasta / Transposon.fasta) as queries, and name the output results Insertion_result.txt, Integron_result.txt, and Transposon_result.txt, respectively.
[0040] Among them, the instruction to establish the database is:
[0041] "makeblastdb-in Plasmid_sequences.fasta-dbtype nucl-parse_seqids-hash_index-out Database\Plasmid_sequences_database";
[0042] The Blastn search command is:
[0043] "blastn-query Insertion.fasta\Integron.fasta\Transposon.fasta-dbDatabase\Plasmid_sequences_database-outfmt" 6qseqid sseqid pident length qcovsqcovhsp qcovus mismatch gapopen qstart qend sstart send evalue bitscore"-outInsertion_result.txt\Integron_result.txt\Transposon_result.txt".
[0044] Step 2: Screen the search results using conventional standards for homologous sequence alignment; identify the screened results based on the identification criteria to obtain a separate identification result for the type of intracellular mobile genetic element contained in the plasmid to be analyzed;
[0045] The conventional standards for homologous sequence alignment are: e-value less than 0.001, identity greater than or equal to 50%, and coverage greater than or equal to 50%. Based on these standards, the search results of transposons, integrons, and insertion sequences of the plasmid to be analyzed are screened in turn. Based on the collected classification information of bacterial plasmid transposons, integrons, and insertion sequences, the results obtained by the Blastn search are annotated with types. During the screening process, if a sequence is aligned to different intracellular mobile genetic elements multiple times, only the result with the largest identity is retained; if the identities are the same, the result with the larger coverage is retained; if both the identity and coverage are the same, only the result with the smaller e-value is retained.
[0046] The identification criteria are:
[0047] In a single plasmid, when there is a complete containment relationship between two intracellular mobile genetic elements, the recognition result with a wider interval is retained;
[0048] When there is no complete inclusion relationship between two intracellular mobile genetic elements, but there is an intersection region, and the intersection region is less than or equal to 50 bp, it is judged as two different intracellular mobile genetic element identification results;
[0049] When there is no complete inclusion relationship between two intracellular mobile genetic elements, but there is an intersection region, and the intersection region is greater than 50 bp, the identification result with the longer length is retained;
[0050] When there is a large crossover region between two intracellular mobile genetic elements and the non-crossover region is less than 50 bp, they are considered to be the same recognition result, and the recognition result with the longer length is retained;
[0051] The recognition results obtained based on the above recognition criteria are:
[0052] Insertion sequence of NZ_AP022395:
[0053]
[0054] Integron of NZ_AP022395:
[0055]
[0056] Transposon of NZ_AP022395:
[0057]
[0058]
[0059] Insertion sequence of NZ_CP030132:
[0060]
[0061]
[0062] Integron of NZ_CP030132:
[0063]
[0064]
[0065] Transposon of NZ_CP030132:
[0066]
[0067]
[0068] Step 3: Merge the insertion sequence, integron, and transposon identification results of each plasmid to be analyzed, and obtain the final identification result of the intracellular mobile genetic element of the plasmid to be analyzed according to the identification criteria;
[0069] The insert sequence, integron, and transposon recognition results of the NZ_AP022395 and NZ_CP030132 plasmids were combined, and the final recognition results were obtained according to the recognition criteria:
[0070] The intracellular mobile genetic elements contained in NZ_AP022395 are:
[0071]
[0072]
[0073] The intracellular mobile genetic elements contained in NZ_CP030132 are:
[0074]
[0075]
[0076] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
[0077] In summary, the method of the present application can improve the accuracy and convenience of identifying mobile genetic elements in bacterial plasmids. Based on data downloaded from public literature and public databases, it can realize automated and batch processing, and achieve efficient identification of mobile genetic elements in bacterial plasmids.
Claims
1. A method for efficiently identifying mobile genetic elements in bacterial plasmids, characterized in that: The method comprises the following steps: step 1: establishing a database with the genome sequence of the plasmid to be analyzed, using the gene sequence of the intracellular mobile genetic element of the bacterial plasmid as the query sequence, and searching the genome sequence of the plasmid to be analyzed in the database; step 2: screening the search results; identifying the screening results according to the identification criteria, and obtaining a separate identification result of the type of the intracellular mobile genetic element contained in the plasmid to be analyzed; step 3: merging the identification results of the intracellular mobile genetic elements of each plasmid to be analyzed, and obtaining the final identification result of the intracellular mobile genetic element of the plasmid to be analyzed according to the identification criteria; The recognition standard is: in a single plasmid, when there is a complete inclusion relationship between two intracellular mobile genetic elements, the recognition result with a larger interval range is retained; when there is an intersection region between two intracellular mobile genetic elements, and the intersection region is less than or equal to 50bp, it is judged as two different identification results of intracellular mobile genetic elements; when there is an intersection region between two intracellular mobile genetic elements, and the intersection region is greater than 50bp, the recognition result with a longer length is retained; when there is an intersection region between two intracellular mobile genetic elements, and the non-intersection region is less than 50bp, it is judged as the same recognition result, and the recognition result with a longer length is retained.
2. The method for efficiently identifying mobile genetic elements in bacterial plasmids according to claim 1, characterized in that: The intracellular mobile genetic elements include transposons, integrons and insertion sequences.
3. The method for efficiently identifying mobile genetic elements in bacterial plasmids according to claim 2, characterized in that: In step 1, the gene sequence of the bacterial plasmid intracellular mobile genetic element is collected and obtained through public literature and public databases.
4. The method for efficiently identifying mobile genetic elements in bacterial plasmids according to claim 3, characterized in that: In step 1, the search method is: running a local Blastn tool in a computer system, using the gene sequences of the transposon, integron and insertion sequence as query sequences, and performing a Blastn search on the gene sequence of the plasmid to be analyzed.
5. The method for efficiently identifying mobile genetic elements in bacterial plasmids according to claim 4, characterized in that: In step 2, the screening method is: screening the search results according to the conventional standards for homologous sequence comparison.
6. The method for efficiently identifying mobile genetic elements in bacterial plasmids according to claim 5, characterized in that: The conventional standards for homologous sequence alignment are: e-value less than 0.001, identity greater than or equal to 50% and coverage greater than or equal to 50%.
Citation Information
Patent Citations
Method for analyzing macro genome integron and mobile element
CN109903810A
Method for detecting movable genetic element based on whole genome data
CN114595234A