A method for detecting mobile genetic elements based on whole-genome data

Through the method based on whole genome data detection, the feature sequence of movable genetic element is extracted and the blast tool is used for comparison, which solves the problem of inconsistent update and naming of movable genetic element database, improves the detection rate and avoids misunderstanding.

CN114595234BActive Publication Date: 2025-06-27HANGZHOU DIANZI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111663118.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-06-27
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The movable genetic component database in the prior art lacks an effective and sustainable management process, which leads to the inability to update and manage new movable components in a timely manner, and the naming and classification are inconsistent, resulting in the same movable genetic component that may be mistaken for different components.

Method used

Using a method based on whole genome data detection, the feature sequence of movable genetic element is obtained through the MGE feature sequence extraction algorithm, and the blast tool is used to compare to obtain the prediction set and result set of movable genetic element, and a unified format is provided for database naming.

Benefits of technology

It effectively solves the problem that the database of movable genetic element cannot be updated in time, improves the detection rate of movable genetic elements, and prevents the same movable genetic element from being mistakenly recognized as different elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595234B_ABST
    Figure CN114595234B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting mobile genetic elements based on whole-genome data. The method uses an MGE feature sequence extraction algorithm to obtain the feature sequences of mobile genetic elements, and obtains a predicted set and a result set of mobile genetic elements through the blast tool. The whole-genome sequencing technology is used to perform individual analysis and determination of the complete gene sequence information of bacteria with unknown genomic sequences, effectively solving the problem that the mobile genetic element database cannot be updated in time, resulting in the inability to detect all mobile genetic elements. Furthermore, the detection rate of mobile elements is improved, and a unified format is provided for the naming of the mobile genetic element database to prevent the same mobile genetic element from being misidentified as different mobile genetic elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological genetic engineering, and particularly to a method for detecting mobile genetic elements based on whole-genome data. Background Art

[0002] Mobile genetic elements are a class of DNA fragments that can move within or between bacterial cells and can encode one or more factors affecting virulence or resistance transmission, including enzymes that mediate their own transfer and integration. Mobile genetic elements include plasmids, insertion sequences, prophages, integrons, and transposons, etc., and there are often various relationships such as nesting, insertion, inversion, truncation, etc. among them to form more complex structures.

[0003] Chinese Patent Publication No. CN107922936A discloses a general method for identifying endogenous physiologically relevant genetic elements that affect an intracellular phenotype of interest. This method sorts and analyzes non-living cells that have been mutagenized based on the phenotype to identify genetic elements, and can identify elements previously unknown to be involved in the phenotype. In the above technical solution, newly emerged mobile elements cannot be updated and managed in the database in a timely manner. Summary of the Invention

[0004] The present invention mainly solves the problem that in the prior art, the mobile genetic element database lacks an effective and sustainable management process, resulting in the inability to update and manage newly emerged mobile elements in the database in a timely manner, and the naming and classification of each mobile element database are not unified. The present invention provides a method for detecting mobile genetic elements based on whole-genome data, uses the MGE feature sequence extraction algorithm to obtain the mobile genetic element feature sequences, and obtains the mobile genetic element prediction set and result set through the blast tool, and provides a unified format for the naming of the mobile genetic element database to prevent the same mobile genetic element from being misidentified as different mobile genetic elements.

[0005] The above technical problem of the present invention is mainly solved by the following technical solution. The present invention includes a method for detecting mobile genetic elements based on whole-genome data, comprising the following steps:

[0006] S1 Obtain the whole-genome sequence of a bacterial strain and the mobile genetic element database;

[0007] S2 Obtain the mobile genetic element feature sequence set through the MGE feature sequence extraction algorithm;

[0008] S3 Compare the mobile genetic element feature sequence set with the whole-genome sequence through the BM string matching algorithm and the blast tool to obtain the mobile genetic element sequence set in this sequence;

[0009] S4 Align the set of mobile genetic element sequences with the mobile genetic element database through the blast tool to obtain the predicted set and the result set of mobile genetic elements;

[0010] S5 Output the result set of mobile genetic elements, and the predicted set is used for experimental verification.

[0011] Use the MGE feature sequence extraction algorithm to obtain the set of mobile genetic element feature sequences, then obtain the set of mobile genetic element sequences through the BM string matching algorithm and the blast tool, and obtain the predicted set and the result set of mobile genetic elements through the blast tool. The result set is used for direct annotation of the whole-genome sequencing data, and the predicted set is used for subsequent experimental verification. If it is verified as a mobile genetic element, it is supplemented to the database, effectively solving the problem that the mobile genetic element database cannot be updated in time, resulting in the inability to detect all mobile genetic elements, and further improving the detection rate of mobile genetic elements.

[0012] Preferably, the whole-genome sequence specifically obtains data from the NCBI database and stores it in the fasta format.

[0013] The database has high authority and high accuracy of the obtained data.

[0014] Preferably, the mobile genetic element database includes an insertion sequence database, an integron database, and a transposon database. The mobile genetic element database obtains the insertion sequence database, the integron database, and the transposon database from the ISFinder database, the INTEGRALL database, and the TheTranspson Registry database respectively, stores them in the fasta format, and formats them into a local library through the blast tool. Using specialized authoritative databases for different mobile genetic elements results in high accuracy of the obtained data. The blast tool alignment is an approximate algorithm for finding fragments with local similarity to the query sequence in a large number of sequences, with high accuracy and wide coverage.

[0015] Preferably, the set of mobile genetic element feature sequences includes the set of inverted terminal repeat sequences of insertion sequences, the set of gene sequences of integrases, the set of integrase attachment site sequences, and the set of direct repeat sequences of transposons. The set of inverted terminal repeat sequences of insertion sequences is obtained from the insertion sequence database through the MGE feature sequence extraction algorithm, the set of gene sequences of integrases and the set of integrase attachment site sequences are obtained from the integron database through the MGE feature sequence extraction algorithm, and the set of direct repeat sequences of transposons is obtained from the transposon database through the MGE feature sequence extraction algorithm.

[0016] The MGE feature sequence extraction algorithm includes feature sequence extraction, feature sequence redundancy removal, and feature sequence set localization. Input the insertion sequence database. According to the gene-level structure, extract the terminal repetitive sequences of the insertion sequences, with a length of 10 to 40 bp, and remove the redundant parts of the extracted sequences. Save the insertion sequence terminal inverted repeat sequence set in fasta format and format it into a local library. Collect the integrase gene sequence set and the integrase attachment site sequence set through the integron database, save them in fasta format, and format them into local libraries respectively. Input the transposon database. According to the gene-level structure, extract the forward sequences of the transposons, with a length of 5 to 9 bp, and remove the redundant parts of the extracted sequences. Save the transposon direct repeat sequence set in fasta format and format it into a local library.

[0017] Preferably, the format of the insertion sequence terminal inverted repeat sequence set is specifically the insertion sequence name, the left side of the terminal repeat sequence, and the right side of the terminal repeat sequence; the format of the integrase gene sequence set and the integrase attachment site sequence set is specifically the NCBI sequence accession number, gene, sequence, and integrase attachment site sequence; the format of the transposon direct repeat sequence set is specifically the transposon name and the direct repeat sequence. It provides a unified format for the naming of the mobile element database to prevent the same mobile genetic element from being misidentified as different mobile genetic elements.

[0018] Preferably, the mobile genetic element sequence set includes an insertion sequence set, a transposon set, an integrase gene and its attachment site set. The insertion sequence set is specifically obtained by aligning the insertion sequence terminal inverted repeat sequence set with the whole genome sequence through a string matching algorithm. The transposon set is specifically obtained by aligning the transposon direct repeat sequence set with the whole genome sequence through a string matching algorithm. The integrase gene and its attachment site set are specifically obtained by aligning the integrase gene sequence set and the integrase attachment site sequence set with the whole genome sequence through the blast tool.

[0019] The string matching algorithm can reduce the time consumed by matching and improve the overall performance. Through the string matching algorithm, after finding the paired insertion sequence terminal inverted repeat sequences in the whole genome sequence, intercept the sequence between the start point and the end point and save it in fasta format to obtain the insertion sequence set; through the string matching algorithm, after finding the paired transposon two-end direct repeat sequences in the whole genome sequence, intercept the sequence between the start point and the end point and save it in fasta format to obtain the transposon set; through the blast tool, after finding the integrase gene sequence and the integrase binding site sequence in the whole genome sequence, intercept the sequence between the start point and the end point and save it in fasta format to obtain the integrase gene and its attachment site set.

[0020] Preferably, the format of the insertion sequence set is specifically the insertion sequence number, starting point, ending point, starting and ending points of the terminal inverted repeat sequence, and the intercepted sequence information; the format of the transposon subset is specifically the transposon number, starting point, ending point, starting and ending points of the direct repeat sequence, and the intercepted sequence information; the format of the integrase gene and its attachment site set is specifically the integrase gene sequence, starting point, ending point, starting and ending points of the integrase attachment site, and sequence information. Providing a unified format can prevent the same mobile genetic element from being misidentified as different mobile genetic elements.

[0021] Preferably, the prediction set includes an insertion sequence prediction set and a transposon prediction set, the result set includes an insertion sequence result set, a transposon result set, and an integrase gene and its attachment site result set. The insertion sequence result set and the insertion sequence prediction set are specifically obtained by comparing the insertion sequence set with the insertion sequence database through the blast tool. The transposon result set and the transposon prediction set are specifically obtained by comparing the transposon subset with the transposon database through the blast tool. The integrase gene and its attachment site result set is specifically obtained by comparing the integrase gene and its attachment site set with the integron database through the blast tool. The result set is used for direct annotation of the whole-genome sequencing data, and the prediction set is used for subsequent experimental verification. If it is verified as a mobile genetic element, it will be supplemented to the database.

[0022] The beneficial effects of the present invention are as follows: The present invention provides a method for detecting mobile genetic elements based on whole-genome data, uses the MGE characteristic sequence extraction algorithm to obtain the characteristic sequences of mobile genetic elements, and obtains the prediction set and result set of mobile genetic elements through the blast tool. It uses the whole-genome sequencing technology to conduct individual analysis and determination of the complete gene sequence information of bacteria with unknown genomic sequences, effectively solving the problem that the mobile genetic element database cannot be updated in time, resulting in the inability to detect all mobile genetic elements, further improving the detection rate of mobile genetic elements, and providing a unified format for the naming of the mobile genetic element database to prevent the same mobile genetic element from being misidentified as different mobile genetic elements. Brief Description of the Drawings

[0023] Figure 1 is a schematic diagram of the logic flow of the present invention.

[0024] Figure 2 is a schematic diagram of a working method of the present invention.

[0025] Figure 3 is a flowchart of the MGE characteristic sequence extraction algorithm of the present invention. Detailed Embodiments

[0026] The technical solution of the present invention will be further specifically described below through embodiments in conjunction with the accompanying drawings.

[0027] Embodiment: A method for detecting mobile genetic elements based on whole-genome data. Taking the whole-genome sequence of Acinetobacter baumannii as an example, the method includes the following steps:

[0028] S1 Obtain the fully assembled whole-genome sequence of Acinetobacter baumannii from the NCBI database, obtain the insertion sequence database from the ISFinder database, obtain the integron database from the INTEGRALL database, and obtain the transposon database from the TheTranspson Registry database;

[0029] S2 Obtain the mobile genetic element characteristic sequence set through the MGE characteristic sequence extraction algorithm; Input the insertion sequence database, including 5550 insertion sequences, with attributes including the insertion sequence name and the source genus. Intercept according to the repeated sequence length at the end of the insertion sequence through the MGE characteristic sequence algorithm, with a length of 10 - 40 bp, retain the attributes, remove the redundant part, and obtain the repeated sequence set at the end of the insertion sequence. Prepare a local library through the blast tool; Input the integron database, including 11957 integron sequences, with attributes including the integron gene source genus, the NCBI accession number of the integron, and the gene names contained in the gene cassette. Obtain the integrase gene sequence set and its attachment site sequence set from it, retain the attributes, and also format it into a local library through blast after removing the redundant part; Input the transposon database, containing 1679 transposons, with attributes including the name, NCBI accession number, and source genus. Obtain the set of direct repeat sequences at both ends of the transposon, with a length of 5 - 9 bp, retain the attributes, and also need to format it into a local library through blast after removing the redundant part;

[0030] S3 uses the BM string matching algorithm and the blast tool to align the mobile genetic element feature sequence set with the whole genome sequence to obtain the mobile genetic element sequence set in this sequence. The mobile genetic element sequence set includes the insertion sequence set, the transposon subset, the integrase gene and its attachment site set. The sequences in the repeated sequence set at the ends of the insertion sequences are relatively short in length, and the sequences are stored in the form of strings. The short sequences in the above set can be matched with the input Acinetobacter baumannii whole genome sequence through the string matching algorithm to obtain the possible insertion sequence set. The direct repeat sequence set at both ends of the transposon is relatively short in length, and the sequences are stored in the form of strings. The short sequences in the above set can be matched with the input Acinetobacter baumannii whole genome sequence through the string matching algorithm to obtain the possible transposon subset. The BM string matching algorithm is used, and the output results of the matching include the start position, length, and the attributes retained in step S2. The integrase gene sequence is relatively long, and it can be aligned with the Acinetobacter baumannii whole genome sequence through the blast tool. The minimum coverage rate and identity are generally 60% and 90%, and the output results include the gene sequence, start position, length, and the attributes retained in step S2.

[0031] S4 aligns the mobile genetic element sequence set obtained in step S3 with the database in step S1 through blast respectively to obtain the mobile genetic element prediction set and the result set.

[0032] S5 outputs the mobile genetic element prediction set and the result set. The result set is used for direct annotation of the whole genome sequencing data, and the prediction set is used for subsequent experimental verification. If it is verified as a mobile genetic element, it will be supplemented to the database.

[0033] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

[0034] Although terms such as mobile genetic element and whole genome sequence are used more frequently in this article, the possibility of using other terms is not excluded. These terms are used only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.

Claims

1. A method for detecting mobile genetic elements based on whole-genome data, characterized in that, It includes the following steps: S1 Obtain the whole genome sequence of the bacterial strain and the mobile genetic element database; S2 Obtain the mobile genetic element characteristic sequence set through the MGE characteristic sequence extraction algorithm. The mobile genetic element characteristic sequence set includes the insertion sequence terminal inverted repeat sequence set, the integrase gene sequence set, the integrase attachment site sequence set, and the transposon direct repeat sequence set; S3 Align the mobile genetic element characteristic sequence set with the whole genome sequence through the BM string matching algorithm and the blast tool to obtain the mobile genetic element sequence set in this sequence; S4 Align the mobile genetic element sequence set with the mobile genetic element database through the blast tool to obtain the prediction set and result set of the mobile genetic element; S5 Output the prediction set and result set of the mobile genetic element.

2. A method for detecting mobile genetic elements based on whole genome data according to claim 1, characterized in that The whole genome sequence is specifically obtained from the NCBI database and stored in fasta format.

3. The method for detecting mobile genetic elements based on whole-genome data according to claim 1, wherein, The mobile genetic element database includes an insertion sequence database, an integron database, and a transposon database. The mobile genetic element database obtains the insertion sequence database, the integron database, and the transposon database from the ISFinder database, the INTEGRALL database, and The TranspsonRegistry database respectively, stores them in fasta format, and formats them into a local database through the blast tool.

4. The method for detecting mobile genetic elements based on whole-genome data according to claim 1 or 3, characterized in that The insertion sequence terminal inverted repeat sequence set is obtained from the insertion sequence database through the MGE characteristic sequence extraction algorithm. The integrase gene sequence set and the integrase attachment site sequence set are obtained from the integron database through the MGE characteristic sequence extraction algorithm. The transposon direct repeat sequence set is obtained from the transposon database through the MGE characteristic sequence extraction algorithm.

5. The method for detecting mobile genetic elements based on whole-genome data according to claim 4, wherein The format of the insertion sequence terminal inverted repeat sequence set is specifically the insertion sequence name, the left side of the terminal repeat sequence, and the right side of the terminal repeat sequence; the format of the integrase gene sequence set and the integrase attachment site sequence set is specifically the NCBI sequence accession number, the gene, the sequence, and the integrase attachment site sequence; the format of the transposon direct repeat sequence set is specifically the transposon name and the direct repeat sequence.

6. The method for detecting mobile genetic elements based on whole-genome data according to claim 4, wherein The mobile genetic element sequence set includes an insertion sequence set, a transposon set, an integrase gene and its attachment site set. The insertion sequence set is specifically obtained by aligning the insertion sequence terminal inverted repeat sequence set with the whole genome sequence through the string matching algorithm. The transposon set is specifically obtained by aligning the transposon direct repeat sequence set with the whole genome sequence through the string matching algorithm. The integrase gene and its attachment site set are specifically obtained by aligning the integrase gene sequence set and the integrase attachment site sequence set with the whole genome sequence through the blast tool.

7. The method for detecting mobile genetic elements based on whole-genome data according to claim 6, wherein The format of the insertion sequence set is specifically the insertion sequence number, starting point, ending point, starting and ending points of the terminal inverted repeat sequence, and the intercepted sequence information; the format of the transposon subset is specifically the transposon number, starting point, ending point, starting and ending points of the direct repeat sequence, and the intercepted sequence information; the format of the integrase gene and its attachment site set is specifically the integrase gene sequence, starting point, ending point, starting and ending points of the integrase attachment site, and the sequence information.

8. A method for detecting mobile genetic elements based on whole-genome data according to claim 6, characterized in that The prediction set includes an insertion sequence prediction set and a transposon prediction set, and the result set includes an insertion sequence result set, a transposon result set, and an integrase gene and its attachment site result set. The insertion sequence result set and the insertion sequence prediction set are specifically obtained by comparing the insertion sequence set with the insertion sequence database through the blast tool. The transposon result set and the transposon prediction set are specifically obtained by comparing the transposon subset with the transposon database through the blast tool. The integrase gene and its attachment site result set is specifically obtained by comparing the integrase gene and its attachment site set with the integron database through the blast tool.

Citation Information

Patent Citations

  • Analysis of identifying genetic elements affecting phenotype

    CN107922936A

  • Individual accurate health preserving method based on gene sequencing technology

    CN110111890A

  • Method and system for analyzing and monitoring viral genome variation

    CN113593639A