Methods and apparatus for mining strain-specific fragments
By comparing the gene similarity between the target strain and other strains of the same species, collinear alignment and fragment merging are performed, solving the accuracy problem of distinguishing strains in traditional methods. This achieves efficient and accurate microbial detection, suitable for low-load samples.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INNER MONGOLIA MENGNIU DAIRY IND (GROUP) CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-05-26
AI Technical Summary
In traditional microbial breeding and plant growth-promoting bacteria research, it is difficult to quickly and accurately distinguish different strains within the same genus. Existing gene detection technologies have limitations and cannot meet the needs of modern microbial research, especially when detecting low-load samples, where sensitivity is insufficient.
By comparing the gene similarity between the target strain and other strains of the same species, a reference strain is determined, and collinear alignment is performed to identify and extract differentially expressed genomic fragments. These fragments are then segmented into candidate fragments, and after meeting preset conditions, they are merged to generate specific fragments. Public nucleic acid databases are used for alignment and filtering to ensure fragment specificity.
It enables rapid screening of the most closely related reference strains from a massive background of strains, accurately locates differentially expressed genomic fragments, improves the accuracy and efficiency of detecting specific target microorganisms, is suitable for low-load samples, enhances computational efficiency and information accuracy, and meets the needs for rapid and accurate detection.
Smart Images

Figure CN122090958A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, and in particular to a method and apparatus for mining strain-specific fragments. Background Technology
[0002] Traditional methods in microbial breeding and plant growth-promoting bacteria research, which rely on morphological observation and physiological and biochemical index analysis, have many shortcomings. These methods are not only time-consuming, labor-intensive, and cumbersome, but they also often fall short in distinguishing closely related microorganisms (such as different strains within the same genus). Their resolution and accuracy are no longer sufficient to meet the needs of modern precision medicine and microbial research.
[0003] To overcome the shortcomings of traditional methods, researchers have developed various molecular detection techniques based on specific gene sequences. For example, plasmids or ribosomal ribonucleic acid (rRNA) genes are used as templates for designing polymerase chain reaction (PCR) primers. Using plasmids for primer design presents several challenges: not all microorganisms contain species-specific plasmids, and some microorganisms lack plasmids altogether. Using rRNA gene regions as templates for PCR detection also presents some issues: although rRNA genes are present in the genomes of all microbial species, species specificity and compatibility with multiplex primers must be considered when designing primers.
[0004] Therefore, how to quickly and accurately mine specific gene fragments and improve the accuracy, reliability and efficiency of detection of specific target microorganisms has become a technical problem that the industry urgently needs to solve. Summary of the Invention
[0005] This application provides a method and apparatus for mining strain-specific fragments, which addresses the technical problem of how to rapidly and accurately mine specific gene fragments to improve the accuracy, reliability, and efficiency of detecting specific target microorganisms.
[0006] This application provides a method for mining strain-specific fragments, including: Compare the gene similarity between the target strain and multiple strains of the same species as the target strain, and determine the strain of the same species corresponding to the highest gene similarity as the reference strain; The genomes of the target strain and the reference strain are collinearly aligned to identify and extract differentially expressed genomic fragments present in the target strain but not in the reference strain. The differentially expressed genomic fragments are segmented into multiple candidate fragments; Candidate fragments that meet preset conditions are identified as specific standard fragments; the preset conditions include that the nucleotide sequence consistency between the candidate fragment and any nucleotide sequence in the public nucleic acid database other than the nucleotide sequence of the target strain itself is less than a preset threshold. Based on the starting coordinates of the specific standard fragment on the genome of the target strain, multiple specific standard fragments that are adjacent or overlapping are merged to generate the specific fragment of the target strain.
[0007] In some embodiments, comparing the gene similarity of the target strain with multiple homologous strains of the target strain, and determining the homologous strain corresponding to the highest gene similarity as the reference strain, includes: Genomic data of the target strain and the plurality of identical strains are obtained from at least one public nucleic acid database; From the genomic data of the target strain and each strain of the same species, a base sequence of a predetermined length is extracted as a genomic fingerprint. The genomic fingerprint of the target strain is compared with the genomic fingerprints of various strains of the same species, and the average nucleotide similarity is used as the gene similarity. The strain of the same species corresponding to the highest average nucleotide identity was determined as the reference strain.
[0008] In some embodiments, the step of performing collinearity alignment between the genomes of the target strain and the reference strain, and identifying and extracting differentially expressed genomic fragments present in the target strain but not in the reference strain, includes: Multiple sequence fragments are located as anchor points between the genomes of the target strain and the reference strain; the sequence fragments satisfy the following conditions: their length is greater than a preset length threshold and their sequences are identical. Using the anchor points as a framework, local sequence alignment is performed on the gap regions between the anchor points to determine the collinearity map between the genomes of the target strain and the reference strain. By comparing the collinearity maps, genomic differential fragments present in the target strain but not in the reference strain are identified and extracted.
[0009] In some embodiments, segmenting the differentially expressed genomic fragments into multiple candidate fragments includes: Based on a preset window length and a preset sliding step size, the differentially expressed genomic fragments are slid-segmented to obtain the multiple candidate fragments; If the length of any differentially expressed genomic fragment is less than the preset window length, the conserved sequence adjacent to the differentially expressed genomic fragment is padded to make up the length.
[0010] In some embodiments, determining candidate fragments that meet preset conditions as specific standard fragments includes: Each candidate fragment is compared with any nucleotide sequence in a public nucleic acid database except for the target strain's own nucleotide sequence, and the candidate fragments that meet the preset conditions are determined as the specific standard fragments; The preset conditions also include coverage and statistical significance meeting preset quantitative standards.
[0011] In some embodiments, merging multiple adjacent or overlapping specific standard fragments based on the starting coordinates of the specific standard fragment on the genome of the target strain includes: Based on the starting coordinates of each specific standard fragment on the genome of the target strain, the specific standard fragments are sorted. If two adjacent specific standard fragments overlap in position or the distance between them is less than a preset interval threshold, the two adjacent specific standard fragments are merged.
[0012] In some embodiments, the method further includes: If the interval between two adjacent specific standard fragments is less than the preset interval threshold, the interval sequences between the two adjacent specific standard fragments are merged together.
[0013] This application provides a strain-specific fragment mining device, comprising: The preliminary screening module is used to compare the gene similarity between the target strain and multiple strains of the same species as the target strain, and to determine the strain of the same species corresponding to the highest gene similarity as the reference strain. The gene comparison module is used to perform collinearity comparison between the genome of the target strain and the genome of the reference strain, and to identify and extract genomic differential fragments that exist in the target strain but not in the reference strain. The fragment segmentation module is used to segment the differentially expressed genomic fragments into multiple candidate fragments; A global filtering module is used to identify candidate fragments that meet preset conditions as specific standard fragments; the preset conditions include that the nucleotide sequence consistency between the candidate fragment and any nucleotide sequence in the public nucleic acid database other than the nucleotide sequence of the target strain itself is less than a preset threshold. The fragment merging module is used to merge multiple specific standard fragments that are adjacent or overlapping in position, based on the starting coordinates of the specific standard fragment on the genome of the target strain, to generate a specific fragment of the target strain.
[0014] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the strain-specific fragment mining method.
[0015] This application provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the strain-specific fragment mining method described above.
[0016] The strain-specific fragment mining method and apparatus provided in this application have the following beneficial effects: (1) Related technologies usually compare the gene similarity of the target strain with all background strains one by one to determine the reference strain. This method is extremely inefficient and requires a huge amount of work. This application directly locates the same strain that is most closely related to the target strain. By comparing the gene similarity between the target strain and the same strain, the massive comparison is transformed into a precise comparison, which brings an exponential improvement in computational efficiency. It enables the preliminary and rapid screening of the reference strain that is most closely related to the target strain and has the most similar genome from a massive number of background strains. (2) By comparing the target strain and the reference strain with high precision, all genomic differential fragments that exist in the target strain but not in the reference strain are accurately located. Since this application performs high-resolution collinearity comparison between the target strain and its single closest relative strain, it can not only identify genomic differential fragments, but also clearly reveal the origin and nature of these genomic differential fragments, thereby providing contextual information with evolutionary background (related technologies cannot reveal the differences of genomic differential fragments in related strains, and therefore cannot provide the evolutionary background of genomic differential fragments), which is crucial for understanding specific biological functions. (3) By segmenting differentially expressed genomic fragments, candidate fragments of uniform length and standardized can be obtained, which facilitates efficient database comparison; by comparing and filtering candidate fragments in public nucleic acid databases, it is ensured that the mined fragments have true specificity; (4) Related technologies usually screen specific fragments, while this application generates specific fragments by merging candidate fragments, thereby realizing the reintegration of fragmented information. It not only splices adjacent or overlapping fragments, but more importantly, it also extracts and incorporates the spacer sequences between these fragments from the original genome, restoring them into biologically more meaningful and continuous specific regions. The resulting specific regions have more complete biological significance, higher downstream application value, and stronger direct usability. (5) Overall, through "preliminary screening - fine comparison - window segmentation - global filtering - intelligent merging", the method ultimately achieves efficient and accurate identification of nucleotide fragments that can uniquely identify specific target strains from massive genomic data, improving the accuracy, reliability, and efficiency of specific target microorganism detection, and bringing about a comprehensive improvement in computational efficiency, information accuracy, and product quality. In addition, this method is suitable for processing low-load samples, improving the sensitivity of microbial detection, meeting the need for rapid and accurate detection of different species and subspecies, and effectively improving the accuracy of microbial detection and the efficiency of disease diagnosis. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the strain-specific fragment mining method provided in this application.
[0020] Figure 2 This is a schematic diagram of the strain-specific fragment mining device provided in this application.
[0021] Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0023] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps, units, or modules is not necessarily limited to those explicitly listed, but may include other steps, units, or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0024] The related technologies have the following shortcomings: (1) In traditional microbial breeding and plant growth-promoting bacteria research, identification methods based on morphological and physiological biochemical characteristics are difficult to accurately distinguish different strains within the same genus, and cannot meet the needs of rapid and accurate identification of specific target microorganisms; (2) There are limitations to using plasmids as primers. Not all microorganisms contain species-specific plasmids, and some microorganisms do not have plasmids. There are also problems with using rRNA gene regions as templates for PCR detection. Although rRNA genes exist in the genomes of all microbial species, species specificity and compatibility of multiple primers need to be considered when designing primers; (3) Existing strain-specific fragment detection methods are powerless in identifying strains that do not have large fragment specific genes; (4) Existing specific common sequence identification methods still need to be improved in terms of identification accuracy and coverage. It is difficult to improve the identification accuracy and coverage of specific common sequences through more optimized clustering algorithms; (5) Existing microbial detection methods are not sensitive enough when processing low-load samples, and cannot meet the needs of rapid and accurate detection of different species and subspecies. This limits the accuracy of microbial detection and the efficiency of disease diagnosis.
[0025] In order to address the shortcomings of related technologies, Figure 1 This is a flowchart illustrating the strain-specific fragment mining method provided in this application, as shown below. Figure 1 As shown, the method includes steps 110, 120, 130, 140 and 150.
[0026] Step 110: Compare the gene similarity between the target strain and multiple strains of the same species as the target strain, and determine the strain of the same species corresponding to the highest gene similarity as the reference strain.
[0027] Specifically, the entity executing the strain-specific fragment mining method provided in this application is a strain-specific fragment mining device or system. This device can be implemented in software, such as a strain-specific fragment mining program; or it can be a device that executes the strain-specific fragment mining method, such as a terminal, computer, or server.
[0028] The target strain is a specific microbial strain from which specific fragments need to be extracted. The target strain can be any microorganism with a complete or nearly complete genome sequence, such as bacteria or fungi. For example, in the field of pathogen detection, the target strain could be a specific serotype or drug resistance spectrum of Klebsiella pneumoniae; in the field of industrial fermentation or probiotics, the target strain could be Bifidobacterium longum with specific desirable traits.
[0029] The same strain refers to other different strains belonging to the same species as the target strain. The genomic data of these strains can come from public nucleic acid databases, such as NCBI (National Center for Biotechnology Information), or from user-owned sequencing databases. The number of strains can be dozens, hundreds, or even thousands. The more strains there are, the more comprehensive the background information, and the more representative the reference strains found will be.
[0030] Gene similarity is a quantitative assessment of the overall genetic similarity between two genomes. There are various methods for calculating gene similarity. For example, it can be calculated based on whole-genome alignment, but this method is computationally intensive and time-consuming. In a specific embodiment, it can be measured by calculating the average nucleotide identity (ANI) between the two genomes. The higher the ANI value, the more similar the genomes of the two strains are.
[0031] When determining the reference strain, the genome of the target strain can be compared with the genomes of all strains of the same species in the database one by one to calculate the gene similarity. Then, the strain with the highest gene similarity can be selected as the reference strain for subsequent analysis.
[0032] Step 120: Perform collinearity comparison between the genomes of the target strain and the reference strain, and identify and extract genomic differential fragments that exist in the target strain but not in the reference strain.
[0033] Specifically, collinearity alignment is a computational process of finding and aligning conserved regions (i.e., regions with identical or highly similar sequences) and structurally variable regions (such as insertions, deletions, inversions, and translocations) between two genomes. Through collinearity alignment, a mapping of the overall correspondence between the two genomes can be constructed. This function can be achieved using various genome alignment and sequence analysis tools, such as MUMmer and ProgressiveMauve. These tools can reveal all sequence differences between two genomes at single-base resolution.
[0034] By analyzing the results of collinearity alignment, sequence fragments present in the target strain genome but absent at corresponding positions in the reference strain genome can be identified. These sequence fragments are known as genomically differentially expressed fragments. Biologically, these fragments may correspond to gene insertions, products of horizontal gene transfer events, or highly variable regions. The length of these fragments can vary, ranging from tens to tens of thousands of bases. The complete sequences of these fragments can be extracted and used as input for subsequent analyses.
[0035] Step 130: Divide the differentially expressed genomic fragments into multiple candidate fragments.
[0036] Specifically, genomic fragments of varying lengths can be processed into short, standardized fragments of uniform length to facilitate subsequent database comparisons and experimental verification.
[0037] In one specific embodiment, a sliding window algorithm is used for segmentation. Specifically, a preset window length (e.g., 100 base pairs (bp)) and a preset sliding step size (e.g., 1 bp, 10 bp, or 50 bp) can be set. Starting from one end of each differentially expressed genomic fragment, the sequence of the first window length is extracted as a candidate fragment; then the window is moved to the other end according to the step size, and a new sequence is extracted as the next candidate fragment, and so on, until the window slides to the end of the differentially expressed fragment.
[0038] The series of short fragments of uniform length generated in the above manner are candidate fragments. The choice of window length needs to balance the specificity of informatics analysis and the feasibility of experimental operation. For example, a length of 100 bp is sufficient to contain enough sequence information to ensure the specificity of alignment, while also providing ample selection space for subsequent PCR primer design (usually 18-25 bp), and its amplification products are also easy to detect.
[0039] Step 140: Identify candidate fragments that meet preset conditions as specific standard fragments; preset conditions include that the nucleotide sequence consistency between the candidate fragment and any nucleotide sequence in the public nucleic acid database other than the target strain's own nucleotide sequence is less than a preset threshold.
[0040] Specifically, in this embodiment, the preset condition is global exclusiveness verification. Specifically, each candidate fragment is used as a query sequence, and a homology search is performed in a large public nucleic acid database (such as NCBI's non-redundant nucleotide database). The search tool can be the blastn program, specifically designed for nucleic acid sequence alignment within the Basic Local Alignment Search Tool (BLAST) bioinformatics tool.
[0041] The preset conditions can specifically include: when a candidate fragment is compared with any nucleotide sequence in the database other than the target strain's own nucleotide sequence, its nucleotide sequence consistency must be less than a preset threshold.
[0042] The preset threshold can be set according to actual needs, for example, it can be 95%. This means that if a candidate fragment finds any non-matching sequence in the database and its consistency is higher than or equal to 95%, the candidate fragment is considered to lack sufficient specificity and should be eliminated. Only those candidate fragments that cannot find any highly similar sequences in the database can pass the screening and be identified as specific standard fragments. This rigorous screening can effectively exclude fragments that belong to the species' core genes, conserved housekeeping genes, or common mobile genetic elements, ensuring that the remaining fragments are truly representative and specific fragments.
[0043] Step 150: Based on the starting coordinates of the specific standard fragments on the genome of the target strain, merge multiple specific standard fragments that are adjacent or overlapping to generate a specific fragment of the target strain.
[0044] Specifically, after segmentation and filtering based on preset conditions, a series of discrete, short, specific standard fragments are obtained. These fragmented information can be reintegrated to restore biologically more meaningful, continuous, specific regions.
[0045] The start coordinates refer to the positional information of each specific standard fragment on the genome of the original target strain, namely the coordinates of the start and stop bases.
[0046] When merging specific standard fragments, the following steps can be followed: First, all specific standard fragments are sorted according to their starting coordinates on the target strain genome. Then, adjacent fragments after sorting are examined in turn: if the starting coordinates of the latter fragment are within the range of the former fragment, or if the coordinates of the two fragments partially overlap, they are merged; if the two fragments do not overlap on the genome but are adjacent in position (i.e., the distance between them is very small), they can also be merged.
[0047] This process is iterative until all adjacent or overlapping fragments that meet the criteria have been integrated. The resulting one or more continuous sequences of varying lengths represent the final target strain-specific fragments. These merged long fragments, compared to individual short fragments, offer greater flexibility and higher analytical value for downstream applications (such as primer and probe design, gene function annotation, etc.), and may represent a complete, specific gene.
[0048] The strain-specific fragment mining method provided in this application has the following beneficial effects: (1) Related technologies usually compare the gene similarity of the target strain with all background strains one by one to determine the reference strain. This method is extremely inefficient and requires a huge amount of work. This application directly locates the same strain that is most closely related to the target strain. By comparing the gene similarity between the target strain and the same strain, the massive comparison is transformed into a precise comparison, which brings an exponential improvement in computational efficiency. It enables the preliminary and rapid screening of the reference strain that is most closely related to the target strain and has the most similar genome from a massive number of background strains. (2) By comparing the target strain and the reference strain with high precision, all genomic differential fragments that exist in the target strain but not in the reference strain are accurately located. Since this application performs high-resolution collinearity comparison between the target strain and its single closest relative strain, it can not only identify genomic differential fragments, but also clearly reveal the origin and nature of these genomic differential fragments, thereby providing contextual information with evolutionary background (related technologies cannot reveal the differences of genomic differential fragments in related strains, and therefore cannot provide the evolutionary background of genomic differential fragments), which is crucial for understanding specific biological functions. (3) By segmenting differentially expressed genomic fragments, candidate fragments of uniform length and standardized can be obtained, which facilitates efficient database comparison; by comparing and filtering candidate fragments in public nucleic acid databases, it is ensured that the mined fragments have true specificity; (4) Related technologies usually screen specific fragments, while this application generates specific fragments by merging candidate fragments, thereby realizing the reintegration of fragmented information. It not only splices adjacent or overlapping fragments, but more importantly, it also extracts and incorporates the spacer sequences between these fragments from the original genome, restoring them into biologically more meaningful and continuous specific regions. The resulting specific regions have more complete biological significance, higher downstream application value, and stronger direct usability. (5) Overall, through “preliminary screening - fine comparison - window segmentation - global filtering - intelligent merging”, the nucleotide fragments that can uniquely identify specific target strains can be efficiently and accurately identified from massive genomic data, which improves the accuracy, reliability and efficiency of specific target microorganism detection and brings about a comprehensive improvement in computational efficiency, information accuracy and product quality.
[0049] Furthermore, this method is suitable for processing low-load samples, improves the sensitivity of microbial detection, meets the need for rapid and accurate detection of different species and subspecies, and effectively improves the accuracy of microbial detection and the efficiency of disease diagnosis.
[0050] It should be noted that each implementation method of this application can be freely combined, rearranged, or executed individually, and does not need to rely on or depend on a fixed execution order.
[0051] In some embodiments, comparing the gene similarity between the target strain and multiple homologous strains of the target strain, and determining the homologous strain corresponding to the highest gene similarity as the reference strain, includes: Obtain genomic data of the target strain and multiple strains of the same species from at least one public nucleic acid database; Base sequences of a predetermined length are extracted from the genomic data of the target strain and various strains of the same species to serve as genomic fingerprints. The genomic fingerprint of the target strain was compared with the genomic fingerprints of various strains of the same species, and the average nucleotide similarity was used as the gene similarity. The strain of the same species corresponding to the highest average nucleotide identity was determined as the reference strain.
[0052] Specifically, accurately identifying specific fragments in the genome of a target strain by directly performing a genome-wide fine alignment or pan-genome analysis presents significant computational challenges, as it requires comparing a massive number of bases one by one, which is time-consuming and labor-intensive. To address this, this application proposes an efficient two-stage analysis strategy, which can be broken down into a first stage of rapid global search and a second stage of precise local alignment. The rapid global search significantly narrows the scope of subsequent fine-grained comparisons through rapid initial screening, thereby fundamentally reducing the overall computational complexity. Specifically, it utilizes genomic information for rapid alignment, thereby efficiently locating the reference strain most closely related to the target strain within a vast strain database.
[0053] Public nucleic acid databases refer to large public repositories that store and provide nucleic acid sequence data and are known to those skilled in the art. Examples include NCBI's GenBank database. Genomic data typically refers to whole genome sequences stored in a specific format (e.g., FASTA format).
[0054] First, download the genome sequence file (genomic data) of the target strain from a public nucleic acid database. Simultaneously, retrieve and download the genome sequence files (genomic data) of all other sequenced strains of the same species that are labeled as the target strain from the same database.
[0055] A genomic fingerprint can be understood as a nucleotide sequence extracted from a complete genome that represents its sequence characteristics. It is not the entire genome, but a highly condensed representation.
[0056] A pre-defined base sequence is often called a k-mer sequence, where k represents the pre-defined length. The choice of k value is a technical trade-off and can be set according to the specific application scenario and computational resources. For example, k values can be set between 15 and 31, with 21 being a commonly used value. Smaller k values have slightly lower specificity but higher tolerance for sequence variations; larger k values have higher specificity but are more sensitive to small variations in the sequence.
[0057] A systematic sketching technique can be used to extract a pre-defined length of base sequence from genomic data as a genomic fingerprint. In one specific embodiment, a hash function (e.g., the MinHash algorithm) is used to calculate all possible k-mer sequences in the genome, and then only those k-mer sequences with the minimum hash value are retained. In this way, a fixed-size and representative genomic fingerprint can be generated for each genome (regardless of its original size). The advantage of this method is that the overlap between the fingerprint sets of two genomes can very accurately reflect their whole-genome similarity.
[0058] Average Nucleotide Identity (ANI) is the gold standard for measuring the overall similarity between two bacterial genomes. In this embodiment, the similarity index obtained by comparing genomic fingerprints can be efficiently and accurately converted into ANI values.
[0059] Traditional ANI value calculation relies on whole-genome alignment, which is computationally intensive. In contrast, efficient ANI value estimation methods utilize sparsification techniques, using only a small subset of representative k-mer sequences as genomic fingerprints. Combined with precise local sequence alignment algorithms, this quantifies the overall similarity between genomes in a very short time, yielding results that highly agree with whole-genome alignment ANI values. This method allows for the screening of hundreds or even thousands of candidate genomes within minutes to hours to identify strains of the same species with a high average nucleotide identity (e.g., greater than a preset value) to the target strain.
[0060] The ANI values of these strains of the same species are sorted, and the strain with the highest ANI value is identified as the reference strain that is most closely related to the target strain.
[0061] The strain-specific fragment mining method provided in this application compares the genomic fingerprints of the target strain and various strains of the same species, using the average nucleotide similarity as the gene similarity. Compared with the traditional whole genome alignment method, it fundamentally reduces the overall computational complexity, ensuring that reliable results are obtained while computational efficiency is improved by orders of magnitude, and greatly improving the efficiency of screening reference strains.
[0062] In some embodiments, collinearity alignment is performed on the genomes of the target strain and the reference strain to identify and extract differentially expressed genomic fragments present in the target strain but not in the reference strain, including: Multiple sequence fragments are located as anchor points between the genomes of the target strain and the reference strain; the sequence fragments must be longer than a preset length threshold and have identical sequences. Using anchor points as a framework, local sequence alignment is performed on the gap regions between anchor points to determine the collinearity map between the genomes of the target strain and the reference strain. By comparing collinearity maps, we can identify and extract genomically differentially expressed fragments that are present in the target strain but not in the reference strain.
[0063] Specifically, after identifying the reference strain in the first stage, the computational workload of the precise local alignment in the second stage is drastically reduced. Because the target strain and the reference strain are highly homologous, and the vast majority of their genome sequences are highly conserved, the specific sequences of the target strain are essentially hidden only in the limited regions of difference between them. Therefore, the task is simplified from cross-alignment with the entire database to a precise comparison of only this pair of closely related genomes—the target strain and the reference strain—allowing for accurate localization of insertion, deletion, mutation, and other variant regions.
[0064] Anchor points can be understood as long deoxyribonucleic acid (DNA) fragments that exist in both the target strain and the reference strain genomes and whose sequences are completely or highly identical. These anchor points form the cornerstone of the collinear relationship between the two genomes.
[0065] To efficiently find these anchor points, an algorithm based on Maximum Unique Match (MUM) can be used. This algorithm utilizes efficient data structures such as suffix trees to quickly find all identical substrings in two long sequences that appear only once in both sequences and cannot be extended further to either side.
[0066] To filter out short, meaningless matches caused by randomness, a preset length threshold needs to be set. This threshold could be set to 20 bp, 50 bp, or dynamically adjusted based on genome complexity. Only sequence fragments longer than the preset length threshold and with consistent sequences will be selected as anchors.
[0067] In one specific embodiment, using the genomic sequences of the target strain and the reference strain as input, the method in this embodiment is executed using the nucmer program in software packages such as MUMmer. This method can efficiently calculate all anchor points that meet the conditions and output their respective position coordinates in the two genomes.
[0068] The backbone refers to the basic correspondence formed by a series of ordered anchor points. The order of these anchor points is usually consistent in both genomes, reflecting the conservation of macroscopic genome structure.
[0069] Interstitial regions are the sequence regions between two adjacent anchor points. These regions may contain small sequence variations, repetitive sequences, or large structural variations.
[0070] For these gap regions, more sensitive but computationally expensive local alignment algorithms can be used for fine-grained comparisons. Because the gap regions are usually short, such high-precision comparisons are computationally feasible.
[0071] By integrating the local alignment results of the anchor point-based backbone and interstitial regions, a complete collinearity map can be generated. This map, in graphical or tabular form, shows in detail how the genome sequence of the target strain corresponds one-to-one with that of the reference strain. The map clearly marks perfectly matched regions, regions with mismatches or gaps, and fragments present only in one of the genomes at single-base resolution.
[0072] By automatically analyzing collinearity map data, we focus on regions that have coordinate records in the genome of the target strain but no corresponding coordinates in the genome of the reference strain. These regions appear as "breakpoints" or "gaps" on the map.
[0073] When the genome map shows a continuous sequence in the target strain, while the corresponding position in the reference strain is a vacancy, this sequence is identified as a genomically differentially expressed fragment present in the target strain but not in the reference strain. The start and end coordinates of these fragments within the target strain's genome are recorded, and their complete DNA sequences are extracted. These extracted sequences of varying lengths will serve as input for subsequent analyses.
[0074] The strain-specific fragment mining method provided in this application locates multiple sequence fragments as anchor points between the genomes of the target strain and the reference strain, avoiding time-consuming alignment across the entire genome. Then, local sequence alignment is performed on the gap regions between the anchor points, thereby systematically constructing a collinearity map of the two genomes as a whole. This ensures that all structural variations, especially insertion fragments unique to the target strain, can be accurately identified at single-base resolution, while significantly reducing the computational resource requirements.
[0075] In some embodiments, differentially expressed genomic fragments are segmented into multiple candidate fragments, including: Based on a preset window length and a preset sliding step size, differentially expressed genomic fragments are slid-segmented to obtain multiple candidate fragments; If the length of any differentially expressed genomic fragment is less than the preset window length, the conserved sequence adjacent to the differentially expressed genomic fragment is padded to the correct length.
[0076] Specifically, the preset window length refers to the fixed length used for truncating sequences. The selection of this length needs to take into account both bioinformatics requirements and experimental validation needs.
[0077] In one specific embodiment, the preset window length can be set to 100 bp. At the informatics level, this length ensures that the extracted sequence contains a sufficient number of specific variant sites (such as SNPs and Indels) to improve the robustness of identification, while also facilitating the simultaneous calculation of the fragment's GC content (the ratio of guanine (G) to cytosine (C) bases), sequence complexity (to avoid simple repetitions), and specificity score (to verify its uniqueness in the background strain library through local alignment). At the experimental level, 100 bp provides ideal space for PCR primer design, allowing for flexible selection of two 18-25 bp primer sequences with high specificity and high melting temperature, while ensuring that the amplification product length falls within the optimal range for conventional electrophoresis detection or probe hybridization, thus achieving the best balance between specificity, amplification efficiency, and experimental compatibility.
[0078] The preset sliding step size refers to the distance the window moves each time. The choice of step size affects the redundancy and coverage density of generated candidate fragments.
[0079] In one specific embodiment, the preset sliding step size can be set to 1 bp. This setting can generate the most comprehensive set of candidate fragments, ensuring that no possible specific combinations are missed, but it will generate a large amount of redundant data. In another specific embodiment, to balance computational efficiency and coverage, the preset sliding step size can be set to 10 bp, 30 bp, or 50 bp. A larger step size can significantly reduce the number of candidate fragments and speed up subsequent analysis.
[0080] Each extracted differentially expressed genomic fragment is segmented independently. Starting from one end of the fragment, a sequence of a predetermined window length is extracted as the first candidate fragment. Then, the window moves to the other end by a predetermined sliding step, extracting a new sequence as the second candidate fragment. This process is repeated until the end of the window reaches or crosses the boundary of the fragment. In this way, a long differentially expressed genomic fragment is segmented into a series of overlapping, uniformly long candidate fragments.
[0081] During collinearity alignment, very short but uniquely specific insertion fragments, such as a genomically differentially expressed fragment of only 40 bp, may be found. If the preset window length is 100 bp, this 40 bp fragment alone cannot generate a candidate fragment of standard length. Directly discarding these short fragments would result in the loss of potentially specific information.
[0082] To address this issue, this application provides a supplementation strategy. When encountering a differentially expressed genomic fragment shorter than a preset window length, it automatically extracts adjacent conserved sequences from one or both sides of the fragment and splices them onto the shorter fragment until the total length of the spliced fragment reaches the preset window length. For example, for a 40bp differentially expressed fragment, a 60bp conserved sequence can be extracted from its upstream side, and the two can be spliced together to form a 100bp candidate fragment.
[0083] A conserved sequence in a neighboring region refers to a sequence region in the target strain genome that is immediately upstream or downstream of the short genomically differentially expressed fragment and that was identified as shared with the reference strain in the aforementioned collinearity alignment.
[0084] To identify such hybrid fragments in subsequent analysis, a special tag or annotation can be added to the fragment when it is generated to indicate that it contains partially conserved sequences.
[0085] The strain-specific fragment mining method provided in this application systematically transforms diverse and varying original differential fragments into a standardized candidate fragment set with uniform length, quantified features, and direct applicability to subsequent computational analysis and experimental design. By length padding, the method solves the problem of processing short specific fragments, avoids information loss, greatly enhances the sensitivity and applicability of the method, and ensures that even strain specificity caused by minor insertion events can be successfully captured and analyzed.
[0086] In some embodiments, determining candidate fragments that meet preset conditions as specific standard fragments includes: Each candidate fragment is compared with any nucleotide sequence in the public nucleic acid database except for the target strain's own nucleotide sequence, and the candidate fragments that meet the preset conditions are identified as specific standard fragments; The preset conditions also include coverage and statistical significance meeting preset quantitative standards.
[0087] Specifically, after obtaining candidate fragments, in order to thoroughly verify their strain-level specificity and exclude homologous sequences that may exist in a broader genetic background, this application introduces a rigorous filtering process based on global homology alignment. The core of this process is to place all candidate fragments in a public nucleic acid database (such as NCBI) containing a massive number of known sequences for systematic searching and comparison.
[0088] This can be done using the blastn program in the BLAST sequence alignment tool. For public nucleic acid databases, the NCBI Non-Redundant Nucleotide Database (nr / nt) can be selected. This database integrates sequences from multiple sources such as GenBank, RefSeq, and PDB, and is one of the most comprehensive collections of nucleic acid sequences currently available.
[0089] All generated candidate fragments were used as query sequences, and BLASTN alignment was performed using the NR / NT database as the target database.
[0090] Since the candidate fragments themselves originate from the target strain, the best match for each candidate fragment during alignment must be its original sequence in the target strain's genome (i.e., a 100% homology, 100% coverage "self-hit"). When analyzing the alignment results, this self-hit result is automatically ignored or skipped; only the alignment results of the candidate fragment with any other nucleotide sequence in the database (i.e., any nucleotide sequence other than the target strain's own nucleotide sequence) are evaluated.
[0091] The preset conditions can be determined based on the following indicators: (1) Nucleotide sequence identity: This metric measures the percentage of identical bases in two aligned sequence fragments. For a fragment to be considered nonspecific, its identity must be higher than or equal to a preset threshold. This threshold can be set to 95%. This means that if a candidate fragment finds a non-self sequence in the database and its alignment identity reaches 95% or higher, it is at risk of being eliminated.
[0092] (2) Coverage: This metric measures the proportion of the matched region to the total length of the query sequence (i.e., candidate fragment), and is usually called query coverage. This metric is introduced to exclude highly consistent matches that are only due to a small local similarity. The preset quantification standard can require a coverage of greater than or equal to 80%. This means that only when a highly consistent match covers most of the candidate fragment is the match considered to be meaningful evidence sufficient to refute its specificity.
[0093] (3) Statistical Significance: This indicator is used to assess whether an alignment result is due to genuine biological homology or simply random occurrence. It is usually measured by the expected value (E-value). The smaller the E-value, the lower the probability that the alignment result is due to chance, and the higher its statistical significance. The preset quantitative standard can require the E-value to be less than 1e-5 (i.e., 0.00001). This means that only statistically significant matching results are adopted as the basis for judgment, thereby effectively filtering out false matches caused by low-complexity sequences or chance.
[0094] In a specific embodiment, the preset condition can be: a candidate fragment can be ultimately identified as a specific standard fragment if and only if, in the blastn search results (excluding self-matches), no non-self-matching result simultaneously meets the following three conditions: nucleotide sequence similarity ≥ 95%, query coverage ≥ 80%, and E-value < 1e-5. Any candidate fragment with at least one such match will be considered a non-specific fragment and discarded.
[0095] In another specific embodiment, a preset condition can be determined solely based on nucleotide sequence consistency. The preset condition could be: any candidate fragment that can be found in a public nucleic acid database with a significant nucleotide sequence consistency higher than 95%, excluding its own or very closely related sequences, will be considered a non-specific fragment and discarded.
[0096] Screening candidate fragments using pre-defined conditions has several important implications. First, it can effectively exclude conserved housekeeping genes or core genome fragments of a species. Although these sequences may show differences in direct close comparisons, they are actually shared in a broader range of populations within the same species or genus. Second, it can identify and remove conserved modules of common mobile genetic elements or horizontally transferred genes. These sequences may be obtained independently in different strains and therefore do not have stable identification value. Finally, it elevates fragment specificity testing from "one-to-one close comparisons" to "one-to-one broad screening of the entire library," greatly improving the reliability and specificity of subsequent experimental markers from a computational perspective, ensuring that the fragments ultimately retained are truly unique "molecular fingerprints."
[0097] The strain-specific fragment mining method provided in this application can effectively distinguish between genuine homologous matches and false matches caused by short fragment random similarity or low-complexity sequences. This ensures that the uniqueness of each specific standard fragment that passes the final screening has undergone the most rigorous test in computational biology, thereby improving the accuracy, reliability and efficiency of detecting specific target microorganisms.
[0098] In some embodiments, based on the starting coordinates of the specific standard fragments on the genome of the target strain, multiple specific standard fragments that are adjacent or overlapping are merged, including: Based on the starting coordinates of each specific standard fragment on the genome of the target strain, the specific standard fragments are sorted. If two adjacent specific standard fragments overlap in position or the distance between them is less than a preset interval threshold, the two adjacent specific standard fragments will be merged.
[0099] Specifically, after rigorous screening based on global homology alignment, several independent fragments that fully meet the specificity criteria can be obtained. These discrete fragments are "high-purity" candidates for the target strain's specific sequences, but in order to more completely and efficiently characterize their genomic features and serve subsequent applications, final information integration is required.
[0100] In the previous screening process, each candidate fragment (including the specific standard fragment that finally passed the screening) retained its starting coordinates in the original target strain genome sequence, that is, the starting base number.
[0101] All specific standard fragments that pass the global filter are collected and arranged in ascending order according to their starting coordinates. The resulting list of specific standard fragments is arranged sequentially from one end of the genome to the other. Adjacent specific standard fragments are two fragments that are immediately next to each other in the sorted list.
[0102] The first merging scenario involves overlapping positions. This occurs when the sliding window is segmented with a step size smaller than the window length. For example, a 100bp window sliding with a step size of 10bp will generate a large number of overlapping candidate segments. If these overlapping segments all pass the global filter, they will be adjacent in the sorted list. Check if the starting coordinates of the next segment (the (i+1)th segment) are less than or equal to the ending coordinates of the previous segment (the i-th segment). If so, the two segments overlap.
[0103] The second merging scenario involves intervals smaller than a preset threshold. This scenario aims to bridge biologically contiguous but computationally broken regions for various reasons. For example, a small, non-specific sequence (such as a short repetitive sequence) might be inserted into a large, specific region, or the algorithm's conservatism might cause it to be segmented. The interval is calculated as the distance between the termination coordinates of the previous segment (the i-th segment) and the starting coordinates of the next segment (the (i+1)-th segment).
[0104] The preset interval threshold is typically smaller than the length of the specific standard fragment itself. For example, for a 100 bp fragment, the preset interval threshold can be set to 50 bp. This means that if the genomic distance between two non-overlapping specific standard fragments is less than 50 bp, they are also considered to originate from the same contiguous specific region and should be merged. This threshold can also be set to other reasonable values.
[0105] When any of the above conditions are met, the two adjacent specific standard fragments are merged. A simple merging method is to create a new, longer sequence with the start coordinates of the previous fragment and the end coordinates of the next fragment. This new sequence then replaces the original two short fragments.
[0106] This process is iterative. The newly merged long segment continues as an element in the list, compared with its next adjacent segment to see if the merging condition is met. This loop continues until the entire list has been traversed and no adjacent segments can be merged.
[0107] Ultimately, through this series of sorting, judgment, and iterative merging, the original series of discrete, specific standard fragments were integrated into several (or one) continuous, target strain-specific fragments of varying lengths.
[0108] The significance of merging adjacent specific standard fragments is that: (1) Restoring biological authenticity: restoring the artificial fragments generated by the algorithm (affected by sliding window cutting) into complete regions that are more likely to reflect real biological events (such as horizontal gene transfer, genome rearrangement or local high variation regions); (2) Optimize subsequent experimental design: A continuous long specific region (e.g., 500bp) provides greater flexibility for experimental verification than several scattered 100bp fragments; researchers can freely choose the optimal PCR primer pair, design hybridization probes or plan sequencing verification intervals within this region, avoiding the limitations of designing on small fragments and improving the robustness of detection. (3) Enhance the value of analysis and annotation: The longer continuous sequence after merging is more conducive to gene prediction, functional domain search or regulatory element analysis, thereby providing clues for understanding the possible biological functions of this specific region and upgrading a simple identifier into a feature region with potential functional significance.
[0109] The strain-specific fragment mining method provided in this application reassembles the specific information that has been artificially fragmented due to the sliding window algorithm, thereby restoring a longer and more biologically complete specific genomic region. This avoids the limitations of designing on a single short fragment and improves the accuracy and reliability of detecting specific target microorganisms.
[0110] In some embodiments, the method further includes: If the interval between two adjacent specific standard fragments is less than a preset interval threshold, the interval sequences between the two adjacent specific standard fragments are merged together.
[0111] Specifically, a spacer sequence refers to a nucleotide sequence located between two adjacent specific standard fragments in the original genome of the target strain. For example, it is a nucleotide sequence located after the termination coordinate of the i-th fragment and before the start coordinate of the (i+1)-th fragment.
[0112] When merging two non-overlapping but adjacent specific standard fragments, a simple approach is to retain only their individual sequences and discard the spacer sequences between them. However, while these spacer sequences may have failed to independently become specific standard fragments in the previous global filtering step due to their shortness or the presence of small non-specific components, they connect the two specific regions and are themselves an integral part of the entire continuous specific region. Directly discarding them would result in information loss and generate an artificially spliced sequence that is biologically inauthentic.
[0113] When it is determined that the i-th segment and the (i+1)-th segment should be merged, the merging operation is no longer a simple splicing of the sequences of the two segments. Instead, the sequence of the i-th segment, the intermediate spacer sequence, and the sequence of the (i+1)-th segment are seamlessly spliced together according to their original order in the genome.
[0114] The newly generated merged fragment has a sequence content that is completely consistent with the complete region from the start of the i-th fragment to the end of the (i+1)-th fragment on the original genome of the target strain, without any loss of information.
[0115] The strain-specific fragment mining method provided in this application ensures that the final generated specific fragment is a biologically real and continuous sequence by retaining and incorporating the spacer sequence during merging.
[0116] The method provided in this application will be described below using the extraction of specific fragments from Klebsiella pneumoniae as an example.
[0117] Step 1: By comparing the genetic similarity of target strain A with all strains within the same species, identify the reference strain B within the same species that is most similar to the target strain. Step 1 specifically includes steps 101 to 104.
[0118] Step 101: Download the genome sequence data of target strain A. Target strain A is selected from Klebsiella pneumoniae, and its genome sequence data was obtained from the NCBI database.
[0119] Step 102: Extract the coding sequence of genome A. Use a coding sequence extraction tool (coding sequence extractor) to extract the coding sequence of genome A, and remove non-coding sequences and spacer sequences.
[0120] Step 103: Use the genome-wide average nucleotide similarity calculation tool (fastANI v1.33) to perform gene similarity analysis between strain A and all strains within the same species. Compare the genome of strain A with other strains in the Klebsiella pneumoniae database and calculate the gene similarity index (ANI) value.
[0121] Step 104: Based on the similarity analysis results, select the strain with the highest genome similarity to A as reference strain B. In this embodiment, the strain with a genome similarity index of 99.5% to A is selected as reference strain B.
[0122] Step 2: Use a DNA and protein sequence comparison tool (MUMmer v4.0.0) to compare the genomic fragments of strains A and B and obtain the similarity ratio. Step 2 specifically includes steps 201 to 203.
[0123] Step 201: Set the MUMmer program parameters. Set the maximum ratio to 0.8, the minimum ratio to 0.5, the maximum beginning insertion to 100, and the minimum ending insertion to 50.
[0124] Step 202: Perform alignment using the MUMmer program. Use the MUMmer program to perform alignment analysis of genomes A and B.
[0125] Step 203: Analyze the alignment results and retain the specific fragments in A compared to B. Fragments with a similarity less than 0.8 in the alignment results are retained as specific fragments.
[0126] Step 3: Divide the sequence in A into segments using a sliding window of 100bp. Step 3 specifically includes steps 301 to 303.
[0127] Step 301: Determine the sliding window size as 100bp and set the sliding step size as 50bp.
[0128] Step 302: Perform sliding slices on genome A to obtain multiple 100bp fragments.
[0129] Step 303: Save the sliced sequence as a FASTA format file.
[0130] Step 4: Align the file with the NR library using the BLASTN program, and filter out fragments with a nucleotide sequence identity greater than 95%. Step 4 specifically includes steps 401 to 403.
[0131] Step 401: Establish alignment parameters for the blastn program. Set the alignment algorithm to blastn, the word length to 11, the maximum number of alignment results to 10, and the minimum confidence value to 80.
[0132] Step 402: Perform alignment using the blastn program. Use the blastn program to align the sliced sequences with the NR database.
[0133] Step 403: Analyze the alignment results and screen out fragments with nucleotide sequence similarity of 95% or less. Retain fragments with nucleotide sequence similarity of less than 95% from the alignment results.
[0134] Step 5: Merge the 100bp fragments that meet the standards to form a specific fragment for strain A. Step 5 specifically includes steps 501 to 504.
[0135] Step 501: Perform sequence verification on the 100bp fragment that meets the standard. Use a gene sequence verification tool (SeqCheck software) to verify the nucleic acid sequence integrity of the fragment.
[0136] Step 502: Sort the segments that passed the verification. Sort the segments using a sequence assembly and alignment tool (SeqMan software).
[0137] Step 503: Merge adjacent segments to form continuous, specific segments. Use SeqMan software to merge adjacent segments.
[0138] Step 504: Use the merged fragment sequence as the final specific fragment result. Save the merged specific fragment sequence as a FASTA format file.
[0139] The method provided in this application will be described below using the example of mining specific fragments of Escherichia coli.
[0140] Step 1: By comparing the genetic similarity of target strain A with all strains within the same species, identify the reference strain B that is most similar to the target strain within the species. Step 1 specifically includes steps 101 to 104.
[0141] Step 101: Download the genome sequence data of target strain A. Target strain A is selected from Escherichia coli, and its genome sequence data was obtained from the GenBank database.
[0142] Step 102: Extract the coding sequence of genome A. Use the getorf program from the open-source bioinformatics software toolkit (EMBOSS) to extract the coding sequence of genome A.
[0143] Step 103: Use the genome-wide average nucleotide similarity calculation tool (fastANI v1.33) to perform gene similarity analysis between strain A and all strains within the same species. Compare the genome of strain A with other strains in the Escherichia coli database and calculate the gene similarity index (ANI) value.
[0144] Step 104: Based on the similarity analysis results, select the strain with the highest genome similarity to A as reference strain B. In this embodiment, the strain with a genome similarity index of 99.6% to A is selected as reference strain B.
[0145] Step 2: Use a DNA and protein sequence comparison tool (MUMmer v4.0.0) to compare the genomic fragments of strains A and B and obtain the similarity ratio. Step 2 specifically includes steps 201 to 203.
[0146] Step 201: Set the MUMmer program parameters. Set the maximum ratio to 0.85, the minimum ratio to 0.6, the maximum beginning insertion to 150, and the minimum ending insertion to 80.
[0147] Step 202: Perform alignment using the MUMmer program. Use the MUMmer program to perform alignment analysis of genomes A and B.
[0148] Step 203: Analyze the alignment results and retain the specific fragments in A compared to B. Fragments with a similarity of less than 0.85 in the alignment results are retained as specific fragments.
[0149] Step 3: Divide the sequence in A into segments using a sliding window of 100bp. Step 3 specifically includes steps 301 to 303.
[0150] Step 301: Determine the sliding window size as 100bp and set the sliding step size as 30bp.
[0151] Step 302: Perform sliding slices on genome A to obtain multiple 100bp fragments.
[0152] Step 303: Save the sliced sequence as a FASTA format file.
[0153] Step 4: Align the file with the NR library using the BLASTN program, and filter out fragments with a nucleotide sequence identity greater than 95%. Step 4 specifically includes steps 401 to 403.
[0154] Step 401: Establish alignment parameters for the blastn program. Set the alignment algorithm to blastn, the word length to 12, the maximum number of alignment results to 20, and the minimum confidence value to 85.
[0155] Step 402: Perform alignment using the blastn program. Use the blastn program to align the sliced sequences with the NR database.
[0156] Step 403: Analyze the alignment results and screen out fragments with nucleotide sequence similarity of 95% or less. Retain fragments with nucleotide sequence similarity of less than 95% from the alignment results.
[0157] Step 5: Merge the 100bp fragments that meet the standards to form a specific fragment for strain A. Step 5 specifically includes steps 501 to 504.
[0158] Step 501: Perform sequence verification on the 100bp fragment that meets the standard. Use the Seq verification tool of the bioinformatics analysis platform (DNAStar) to verify the nucleic acid sequence integrity of the fragment.
[0159] Step 502: Sort the validated fragments. Use DNAStar's Seq assembly tool to sort the fragments.
[0160] Step 503: Merge adjacent fragments to form continuous, specific fragments. Use DNAStar's Seq assembly tool to splice and merge adjacent fragments.
[0161] Step 504: Use the merged fragment sequence as the final specific fragment result. Save the merged specific fragment sequence as a FASTA format file.
[0162] The method provided in this application will be described below using the extraction of specific fragments from the Bifidobacterium longum strain BBMN68 as an example.
[0163] Step 1: By comparing the genetic similarity of target strain A with all strains within the same species, identify the reference strain B within the same species that is most similar to the target strain. Step 1 specifically includes steps 101 to 104.
[0164] Step 101: Download the genome sequence data of target strain A. Target strain A is *Bifidobacterium longum* BBMN68, and its genome sequence data was obtained from the NCBI database.
[0165] Step 102: Extract the coding sequence of genome A. Use the prokaryotic genome annotation tool (Prokka v1.14.6) to annotate genome A and extract all predicted coding sequences.
[0166] Step 103: Use the genome-wide average nucleotide similarity calculation tool (fastANI v1.33) to perform gene similarity analysis between strain A and all strains within the same species. Compare the genome of strain A with other strains in the Bifidobacterium longum database and calculate the gene similarity index (ANI) value.
[0167] Step 104: Based on the similarity analysis results, select the strain with the highest genome similarity to A as reference strain B. In this embodiment, the strain with a genome similarity index of 99.4% to A is selected as reference strain B.
[0168] Step 2: Use a DNA and protein sequence comparison tool (MUMmer v4.0.0) to compare the genomic fragments of strains A and B and obtain the similarity ratio. Step 2 specifically includes steps 201 to 203.
[0169] Step 201: Set the MUMmer program parameters. Set the maximum ratio to 0.82, the minimum ratio to 0.55, the maximum beginning insertion to 120, and the minimum ending insertion to 60.
[0170] Step 202: Perform alignment using the MUMmer program. Use the MUMmer program to perform alignment analysis of genomes A and B.
[0171] Step 203: Analyze the alignment results and retain the specific fragments in A compared to B. Fragments with a similarity less than 0.82 in the alignment results are retained as specific fragments.
[0172] Step 3: Divide the sequence in A into segments using a sliding window of 100bp. Step 3 specifically includes steps 301 to 303.
[0173] Step 301: Determine the sliding window size as 100bp and set the sliding step size as 40bp.
[0174] Step 302: Perform sliding slices on genome A to obtain multiple 100bp fragments.
[0175] Step 303: Save the sliced sequence as a FASTA format file.
[0176] Step 4: Align the file with the NR library using the BLASTN program, and filter out fragments with a nucleotide sequence identity greater than 95%. Step 4 specifically includes steps 401 to 403.
[0177] Step 401: Establish alignment parameters for the blastn program. Set the alignment algorithm to blastn, the word length to 12, the maximum number of alignment results to 15, and the minimum confidence value to 90.
[0178] Step 402: Perform alignment using the blastn program. Use the blastn program to align the sliced sequences with the NR database.
[0179] Step 403: Analyze the alignment results and screen out fragments with nucleotide sequence similarity of 95% or less. Retain fragments with nucleotide sequence similarity of less than 95% from the alignment results.
[0180] Step 5: Merge the 100bp fragments that meet the standards to form a specific fragment for strain A. Step 5 specifically includes steps 501 to 504.
[0181] Step 501: Perform sequence verification on the 100bp fragment that meets the standard. Use the sequence verification function of bioinformatics software (GeneiousPrime 2023.2) to verify the nucleic acid sequence integrity of the fragment.
[0182] Step 502: Sort the verified fragments. Sort the fragments using Geneious Prime's sequence assembly tool.
[0183] Step 503: Merge adjacent segments to form continuous, specific segments. Use Geneious Prime's sequence assembly tool to merge adjacent segments.
[0184] Step 504: Use the merged fragment sequence as the final specific fragment result. Save the merged specific fragment sequence as a FASTA format file.
[0185] The strain-specific fragment mining method provided in this application has the following beneficial effects: (1) By analyzing the strain genome, the limitations of traditional methods that rely solely on morphological and physiological biochemical characteristics for identification are overcome, and the accuracy and reliability of strain identification are significantly improved. (2) By using a 100bp fragment sliding window to cut the fragment and combining it with the blastn program for comparison, the specific fragments of the strain can be effectively identified, avoiding the limitations of using plasmids as primers. It also provides the possibility of primer design for strains that do not have large fragment specific genes, solving the problem of species specificity and multiple primer compatibility when designing primers for rRNA gene regions, and improving the specificity and sensitivity of the detection method. (3) The method is suitable for processing low-load samples, which improves the sensitivity of microbial detection, meets the need for rapid and accurate detection of different species and subspecies, and effectively improves the accuracy of microbial detection and the efficiency of disease diagnosis. (4) By combining gene similarity analysis and specific fragment mining, the shortcomings of the same species-specific common sequence identification methods in related technologies in terms of accuracy and coverage are overcome, and a more comprehensive and accurate strain-specific analysis solution is provided.
[0186] The apparatus provided in the embodiments of this application is described below. The apparatus described below can be referred to in correspondence with the method described above.
[0187] Figure 2 This is a schematic diagram of the strain-specific fragment mining device provided in this application, as shown below. Figure 2 As shown, the device includes: The preliminary screening module 210 is used to compare the gene similarity between the target strain and multiple strains of the same species as the target strain, and to determine the strain of the same species corresponding to the highest gene similarity as the reference strain. The gene comparison module 220 is used to perform collinearity comparison between the genome of the target strain and the genome of the reference strain, and to identify and extract genomic differential fragments that exist in the target strain but not in the reference strain. Fragment segmentation module 230 is used to segment differentially expressed genomic fragments into multiple candidate fragments; The global filtering module 240 is used to identify candidate fragments that meet preset conditions as specific standard fragments. The preset conditions include that the nucleotide sequence consistency between the candidate fragment and any nucleotide sequence in the public nucleic acid database other than the nucleotide sequence of the target strain itself is less than a preset threshold. The fragment merging module 250 is used to merge multiple specific standard fragments that are adjacent or overlapping based on the starting coordinates of the specific standard fragments on the genome of the target strain, thereby generating a specific fragment of the target strain.
[0188] The strain-specific fragment mining device provided in this application has the following beneficial effects: (1) Related technologies usually compare the gene similarity of the target strain with all background strains one by one to determine the reference strain. This method is extremely inefficient and requires a huge amount of work. This application directly locates the same strain that is most closely related to the target strain. By comparing the gene similarity between the target strain and the same strain, the massive comparison is transformed into a precise comparison, which brings an exponential improvement in computational efficiency. It enables the preliminary and rapid screening of the reference strain that is most closely related to the target strain and has the most similar genome from a massive number of background strains. (2) By comparing the target strain and the reference strain with high precision, all genomic differential fragments that exist in the target strain but not in the reference strain are accurately located. Since this application performs high-resolution collinearity comparison between the target strain and its single closest relative strain, it can not only identify genomic differential fragments, but also clearly reveal the origin and nature of these genomic differential fragments, thereby providing contextual information with evolutionary background (related technologies cannot reveal the differences of genomic differential fragments in related strains, and therefore cannot provide the evolutionary background of genomic differential fragments), which is crucial for understanding specific biological functions. (3) By segmenting differentially expressed genomic fragments, candidate fragments of uniform length and standardized can be obtained, which facilitates efficient database comparison; by comparing and filtering candidate fragments in public nucleic acid databases, it is ensured that the mined fragments have true specificity; (4) Related technologies usually screen specific fragments, while this application generates specific fragments by merging candidate fragments, thereby realizing the reintegration of fragmented information. It not only splices adjacent or overlapping fragments, but more importantly, it also extracts and incorporates the spacer sequences between these fragments from the original genome, restoring them into biologically more meaningful and continuous specific regions. The resulting specific regions have more complete biological significance, higher downstream application value, and stronger direct usability. (5) Overall, through “preliminary screening - fine comparison - window segmentation - global filtering - intelligent merging”, the nucleotide fragments that can uniquely identify specific target strains can be efficiently and accurately identified from massive genomic data, which improves the accuracy, reliability and efficiency of specific target microorganism detection and brings about a comprehensive improvement in computational efficiency, information accuracy and product quality.
[0189] Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application, such as... Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor, communications interface, and memory communicate with each other via the communications bus. The processor can invoke logical commands stored in the memory to execute the methods described in the above embodiments, for example: The gene similarity of the target strain with multiple strains of the same species is compared, and the strain with the highest gene similarity is identified as the reference strain. The genomes of the target strain and the reference strain are collinearly aligned to identify and extract genomically differentially expressed fragments present in the target strain but not in the reference strain. These differentially expressed fragments are segmented into multiple candidate fragments. Candidate fragments that meet preset conditions are identified as specific standard fragments. These preset conditions include that the nucleotide sequence consistency between the candidate fragment and any nucleotide sequence in a public nucleic acid database, excluding the target strain's own nucleotide sequence, is less than a preset threshold. Based on the starting coordinates of the specific standard fragments on the target strain's genome, multiple specific standard fragments that are adjacent or overlapping are merged to generate the target strain's specific fragment.
[0190] Furthermore, the logical commands in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0191] The processor in the electronic device provided in this application embodiment can call logical instructions in the memory to implement the above method. Its specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effect, which will not be repeated here.
[0192] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments.
[0193] The specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effects, so it will not be repeated here.
[0194] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0195] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0196] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for mining strain-specific fragments, characterized in that, include: Compare the gene similarity between the target strain and multiple strains of the same species as the target strain, and determine the strain of the same species corresponding to the highest gene similarity as the reference strain; The genomes of the target strain and the reference strain are collinearly aligned to identify and extract differentially expressed genomic fragments present in the target strain but not in the reference strain. The differentially expressed genomic fragments are segmented into multiple candidate fragments; Candidate fragments that meet preset conditions are identified as specific standard fragments; the preset conditions include that the nucleotide sequence consistency between the candidate fragment and any nucleotide sequence in the public nucleic acid database other than the nucleotide sequence of the target strain itself is less than a preset threshold. Based on the starting coordinates of the specific standard fragment on the genome of the target strain, multiple specific standard fragments that are adjacent or overlapping are merged to generate the specific fragment of the target strain.
2. The method for mining strain-specific fragments according to claim 1, characterized in that, The step of comparing the gene similarity of the target strain with multiple strains of the same species, and determining the strain of the same species corresponding to the highest gene similarity as the reference strain, includes: Genomic data of the target strain and the plurality of identical strains are obtained from at least one public nucleic acid database; From the genomic data of the target strain and each strain of the same species, a base sequence of a predetermined length is extracted as a genomic fingerprint. The genomic fingerprint of the target strain is compared with the genomic fingerprints of various strains of the same species, and the average nucleotide similarity is used as the gene similarity. The strain of the same species corresponding to the highest average nucleotide identity was determined as the reference strain.
3. The method for mining strain-specific fragments according to claim 1, characterized in that, The step of performing collinear alignment between the genomes of the target strain and the reference strain, and identifying and extracting differentially expressed genomic fragments present in the target strain but not in the reference strain, includes: Multiple sequence fragments are located as anchor points between the genomes of the target strain and the reference strain; the sequence fragments satisfy the following conditions: their length is greater than a preset length threshold and their sequences are identical. Using the anchor points as a framework, local sequence alignment is performed on the gap regions between the anchor points to determine the collinearity map between the genomes of the target strain and the reference strain. By comparing the collinearity maps, genomic differential fragments present in the target strain but not in the reference strain are identified and extracted.
4. The method for mining strain-specific fragments according to claim 1, characterized in that, The step of segmenting the differentially expressed genomic fragments into multiple candidate fragments includes: Based on a preset window length and a preset sliding step size, the differentially expressed genomic fragments are slid-segmented to obtain the multiple candidate fragments; If the length of any differentially expressed genomic fragment is less than the preset window length, the conserved sequence adjacent to the differentially expressed genomic fragment is padded to make up the length.
5. The method for mining strain-specific fragments according to claim 1, characterized in that, The step of identifying candidate fragments that meet preset conditions as specific standard fragments includes: Each candidate fragment is compared with any nucleotide sequence in a public nucleic acid database except for the target strain's own nucleotide sequence, and the candidate fragments that meet the preset conditions are determined as the specific standard fragments; The preset conditions also include coverage and statistical significance meeting preset quantitative standards.
6. The method for mining strain-specific fragments according to claim 1, characterized in that, The process of merging multiple adjacent or overlapping specific standard fragments based on their starting coordinates on the genome of the target strain includes: Based on the starting coordinates of each specific standard fragment on the genome of the target strain, the specific standard fragments are sorted. If two adjacent specific standard fragments overlap in position or the distance between them is less than a preset interval threshold, the two adjacent specific standard fragments are merged.
7. The method for mining strain-specific fragments according to claim 6, characterized in that, The method further includes: If the interval between two adjacent specific standard fragments is less than the preset interval threshold, the interval sequences between the two adjacent specific standard fragments are merged together.
8. A strain-specific fragment mining device, characterized in that, include: The preliminary screening module is used to compare the gene similarity between the target strain and multiple strains of the same species as the target strain, and to determine the strain of the same species corresponding to the highest gene similarity as the reference strain. The gene comparison module is used to perform collinearity comparison between the genome of the target strain and the genome of the reference strain, and to identify and extract genomic differential fragments that exist in the target strain but not in the reference strain. The fragment segmentation module is used to segment the differentially expressed genomic fragments into multiple candidate fragments; A global filtering module is used to identify candidate fragments that meet preset conditions as specific standard fragments; the preset conditions include that the nucleotide sequence consistency between the candidate fragment and any nucleotide sequence in the public nucleic acid database other than the nucleotide sequence of the target strain itself is less than a preset threshold. The fragment merging module is used to merge multiple specific standard fragments that are adjacent or overlapping in position, based on the starting coordinates of the specific standard fragment on the genome of the target strain, to generate a specific fragment of the target strain.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the strain-specific fragment mining method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the strain-specific fragment mining method according to any one of claims 1 to 7.