Method for screening specific sequence of genome and related device
By comparing resequencing data with genome sequences and combining third-generation and second-generation sequencing technologies, genome-specific sequences are screened out, solving the accuracy problem of screening differential sequences among multiple species strains, improving detection efficiency and reducing costs, and making it suitable for identification and breeding of multiple species strains.
Patent Information
- Application Number
- CN202511581039.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies struggle to accurately screen for genomically differential sequences among multiple species, leading to a large workload and high cost for subsequent experimental verification.
By comparing the resequencing data and genome sequences of the target strain of the first species with those of other strains, specific sequences of the target strain were identified. By combining high-precision resequencing data with long-fragment genome data, and utilizing the advantages of third-generation sequencing and second-generation sequencing, genome-specific sequences were screened out.
It significantly improves the accuracy of structural variation detection among closely related strains, reduces false positive results, lowers the workload and cost of verification, is applicable to multi-species strain identification, and supports precision breeding and biodiversity conservation.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and more specifically to a method and related apparatus for screening genome-specific sequences. Background Technology
[0002] With the rapid development of high-throughput sequencing technology, genome screening has become an important tool for germplasm resource analysis and genetic research. This technology can obtain genetic variation information across the entire genome of an organism in a single operation, providing a rich data foundation for population genetics research and breeding practices.
[0003] Genome screening technology has multiple applications in agriculture, forestry, and biodiversity conservation, specifically in the following aspects: enabling high-throughput identification of varieties / strains; assisting in species classification and identification; completing precise tracing of parental origins; analyzing the genetic background and pedigree of individuals; achieving genetic tracing of specific genes or chromosomes; assessing the genetic contribution of individuals or populations; constructing large-scale mutant libraries for functional genomics research; revealing the genetic diversity characteristics of biological populations; and providing key information for marker-assisted breeding. Summary of the Invention
[0004] In view of this, the present invention aims to provide a method and related apparatus for screening genome-specific sequences, to solve the technical problem in the prior art of accurately screening for truly differentially expressed genome sequences among multiple species strains. The method of the present invention has high accuracy and can significantly reduce the number of candidate sequences required for subsequent experimental verification, thereby reducing development costs and improving the efficiency and reliability of strain identification.
[0005] The first aspect of this application provides a method for screening genome-specific sequences, including:
[0006] The resequencing data of the target strain of the first species is compared with the genome sequence of the target strain of the first species to obtain the first comparison result;
[0007] The resequencing data of at least one other strain of the first species are compared with the genome sequence of the target strain of the first species to obtain a second comparison result;
[0008] The first comparison result of the target strain is compared with the second comparison result of each other strain to obtain the result corresponding to each other strain.
[0009] Based on the results, a genomic region that exists in the genome sequence of the target strain and does not exist in at least one other strain is identified, and the genomic region is identified as a specific sequence of the target strain relative to the at least one other strain.
[0010] Comparing the resequencing data of the target strain of the first species with its genome sequence can effectively verify the integrity and accuracy of genome assembly. This method can identify potential regional sequence deletions and exclude assembly errors caused by exogenous DNA contamination, thereby ensuring the reliability of the reference genome and providing an accurate basis for subsequent analysis.
[0011] In some embodiments, the method further includes, before comparing the resequencing data of the first species with the genome sequence of the first species:
[0012] The sequencing quality of the resequencing data of the target strain and the sequencing quality of the resequencing data of at least one other strain were determined using sequencing quality assessment tools.
[0013] Remove sequencing reads with sequencing quality below a preset value from the resequencing data of the target strain and the resequencing data of at least one other strain;
[0014] Remove sequence adapters from the resequencing data of the target strain and the resequencing data of at least one other strain.
[0015] In some specific embodiments, the preset value of the sequencing quality is Q15, and sequencing reads with a sequencing quality lower than Q15 in the resequencing data are removed.
[0016] In some embodiments, the genome sequence of the target strain is obtained from third-generation sequencing, while the resequencing data of the other strains is obtained from second-generation sequencing. The genome sequence of the target strain and the resequencing data of other strains can be obtained from public databases (such as NCBI, ENA, CNGB, etc.) or obtained through sequencing using conventional experimental methods in the art; this invention does not limit the scope of these methods.
[0017] In some specific embodiments, the genome sequence of the target strain is selected from genome sequences assembled at the chromosome level using ONT or HIFI data as a backbone. This genome sequence exhibits better continuity compared to genome sequences obtained through other methods.
[0018] In some embodiments, comparing the first comparison result of the target strain with the second comparison result of each of the other strains to obtain the result corresponding to each of the other strains includes:
[0019] The overlap between the resequencing data of the target strain and the genome sequence of the target strain is calculated based on the first comparison result to obtain a first calculation result; at least one first region with an overlap of 0 is extracted from the genome sequence of the target strain based on the first calculation result.
[0020] Based on the second comparison result, the overlap between the resequencing data of at least one other strain and the genome sequence of the target strain is calculated to obtain a second calculation result; based on the second calculation result, at least one second region with an overlap of 0 is extracted from the genome sequence of the target strain.
[0021] Filter regions in each of the first and second regions whose sequence length is not less than a preset threshold;
[0022] The first selected region is compared with the second selected region to obtain the results corresponding to each of the other strains.
[0023] In some embodiments, extracting regions with zero overlap includes:
[0024] Merge contiguous areas with the same coverage depth;
[0025] The area includes uncovered areas;
[0026] Obtain long, uncovered segments.
[0027] In some embodiments, the preset threshold is 80bp, and sequences with a length greater than or equal to 80bp are obtained through screening.
[0028] Preferably, the preset threshold is 200 bp to screen for differentially expressed fragments of sufficient length, facilitating the design of specific PCR primers to achieve the detection effect of "the target strain having an amplified band while other strains do not." If a sufficient number of fragments longer than 200 bp cannot be obtained, the threshold is lowered. For example, the threshold can be lowered to 180 bp, 160 bp, 140 bp, 120 bp, 100 bp, 80 bp, or any value between the two. As a feasible example, in the identification of genome-specific sequences, primers are designed on both sides of the differentially expressed fragments, and PCR amplification produces bands of different lengths, thereby achieving the molecular marker effect of "the target strain having amplified bands of different sizes compared to other strains."
[0029] In some embodiments, the method for screening genome-specific sequences further includes:
[0030] Remove regions in the specific sequence where the proportion of repetitive sequences is greater than a preset ratio or the length of the repetitive sequences is greater than a preset length.
[0031] In some embodiments, the preset ratio is 10% to 30%, and the preset length is 40 to 90 bp. For example, the preset ratio is 10%, 20%, or 30%; and the preset length is 40, 50, 60, 70, 80, or 90 bp.
[0032] In some specific embodiments, the preset ratio is 20% or 30%, and the preset length is 50bp or 80bp. The lower threshold for the proportion of repetitive sequences ensures that the selected specific fragments have sufficient sequence uniqueness, which is convenient for designing highly specific PCR primers; while the appropriate length threshold ensures that these fragments have reliable amplification efficiency and detection sensitivity in experimental operations.
[0033] Preferably, the preset ratio is 20% and the preset length is 50 bp, so as to screen out specific fragments that do not contain or contain only a small number of repetitive sequences, which facilitates the subsequent design of specific PCR primers and accurately screens out the real genomic differential sequences between multiple species and strains.
[0034] A second aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the method for screening genome-specific sequences described in the first aspect or any implementation thereof.
[0035] A third aspect of this application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0036] The memory is used to store computer programs;
[0037] The processor is used to execute the computer program to enable the electronic device to implement the method for screening genome-specific sequences as described in the first aspect or any implementation thereof.
[0038] A fourth aspect of this application provides a computer storage medium carrying one or more computer programs that, when executed by an electronic device, enable the electronic device to perform a method for screening genome-specific sequences as described in the first aspect or any implementation thereof.
[0039] This invention relates to a method and related apparatus for screening genome-specific sequences. This method significantly improves the accuracy of detecting structural variations among closely related strains by integrating high-precision resequencing data with long-fragment genome data. It effectively reduces false positives, significantly decreases validation workload and costs, is applicable to multi-species strain identification, and provides reliable support for the rapid development of specific molecular markers. It has broad application prospects in precision breeding and biodiversity conservation. Attached Figure Description
[0040] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0041] Figure 1 This is a schematic diagram of the principle of this technical solution (taking strain A and strain B as examples);
[0042] Figure 2 This is a flowchart illustrating a method for screening genome-specific sequences according to an embodiment of this application;
[0043] Figure 3 This is a data preprocessing flowchart of another embodiment of this application;
[0044] Figure 4 This is a flowchart of step S300 in another embodiment of this application;
[0045] Figure 5 A schematic diagram of the composition structure of an electronic device provided in an embodiment of this application;
[0046] Figure 6 This is a flowchart illustrating the implementation process of the proposed solution.
[0047] Figure 7 Electrophoresis diagram of the specific primer verification experiment for Jianli 2 / Furui 2. Detailed Implementation
[0048] This invention provides a method and related apparatus for screening genome-specific sequences. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired results. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can clearly modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.
[0049] The terminology used in the embodiments section of this application is for explaining specific embodiments of this application only and is not intended to limit this application. Unless otherwise defined in this invention, scientific and technical terms related to this invention should have the meanings understood by those skilled in the art.
[0050] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application.
[0051] The terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not expressly listed or inherent to those processes, methods, products, or apparatus.
[0052] When used herein, the term “and / or” includes the meaning of “and,” “or,” and “all or any other combination of elements linked by the term.”
[0053] The term "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items.
[0054] It should be understood that in the various embodiments of this application, the order of the above processes does not imply the order of execution. Some or all steps may be executed in parallel or sequentially. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0055] The numerical ranges and parameters involved in this invention have been presented as precisely as possible in the specific embodiments. However, any numerical value inevitably contains standard deviations due to individual test methods. Therefore, unless otherwise expressly stated, it should be understood that all numerical ranges or specific data used in this disclosure may have a reasonable deviation within a certain range, such as ±10%, ±5%, ±1%, or ±0.5%.
[0056] Taking strain A and strain B as examples, the basic principle of this invention is as follows: Figure 1 As shown:
[0057] The resequencing data of strains A and B were aligned to the genome of strain A. Regions in the genome of strain A that were aligned with resequencing reads of strain A and whose length met a certain threshold condition, but were not aligned with resequencing reads of strain B, were extracted as specific sequences of strain A relative to strain B.
[0058] Resequencing data is typically obtained using second-generation sequencing technologies such as Illumina, characterized by high accuracy but short sequences. Genome sequences are generally assembled from third-generation sequencing data, characterized by long and continuous sequences, but prone to assembly errors. If genome A contains regions aligned with resequencing reads A, the probability of assembly errors is low. If no resequencing reads B are found in this region, it can be inferred that this region is missing from the genome of strain B.
[0059] This patent combines the advantages of genome resequencing data (short but highly accurate fragments) and genome data (long fragments, but prone to assembly errors). Through the screening method provided in this patent, it rapidly and accurately identifies genome-specific sequences among different strains. The screened specific sequences can be used to design specific primer sets, which, after PCR validation, can be developed into molecular markers for distinguishing different strains. This method significantly reduces the complexity and workload of subsequent experimental validation, eliminating the need for primer design and PCR experiments on a large number of candidate sequences. Only a small number of high-quality sequences obtained through screening need to be validated, thus ensuring accuracy while significantly reducing experimental costs.
[0060] The test materials used in this invention are all common commercial products and can be purchased on the market.
[0061] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0062] like Figure 2 As shown, the method for screening genome-specific sequences provided in this application may include:
[0063] S100. Compare the resequencing data of the target strain of the first species with the genome sequence of the target strain of the first species to obtain the first comparison result.
[0064] A strain is a group of individuals within a species that share specific genetic similarities; it is a taxonomic unit within a species. For example, in the species *Cyprinus carpio*, *Furui Carp No. 2* and *Jian Carp No. 2* are different strains; in the mouse species *Musmusculus*, C57BL / 6J and BALB / c are different strains.
[0065] Resequencing data is a collection of sequence reads generated by sequencing the genome of an individual strain. It is typically used to compare with a reference genome to detect variations such as single nucleotide polymorphisms, insertions, and deletions.
[0066] The genome sequence is the reference genome of the target strain. It is usually assembled using technologies such as third-generation sequencing and represents the complete or near-complete nucleotide sequence of the strain's genome.
[0067] The genome sequence of the target strain is obtained from third-generation sequencing, while the resequencing data of the other strains is obtained from second-generation sequencing.
[0068] Third-generation sequencing refers to sequencing technologies capable of generating long reads, such as PacBio single-molecule real-time sequencing from Pacific Biosciences or nanopore sequencing from Oxford Nanopore. Their read lengths can reach several kb or even Mb levels, effectively traversing repetitive regions in the genome to obtain more complete and continuous genome assembly results.
[0069] Next-generation sequencing (NGS) refers to high-throughput, short-read sequencing technology based on the principle of sequencing-by-synthesis. Its read length is typically 50-300 bp, and it is characterized by high throughput, relatively low cost, and extremely high base accuracy, making it ideal for resequencing for variant detection and coverage analysis.
[0070] Compared to genome sequences, resequencing data consists of short sequence fragments from a specific individual that cover a portion of the genome, while genome sequences are continuous long sequences used as alignment benchmarks.
[0071] Specifically, sequence alignment algorithms can be used for comparison. The alignment process includes: aligning the header of the resequencing data with the header of the genome sequence; determining whether the relative regions in the resequencing data and the genome sequence are consistent; if they are inconsistent, shifting the header of the resequencing data by one position and returning to the step of determining whether the relative regions in the resequencing data and the genome sequence are consistent, until the entire resequencing data is traversed or a matching position is found. This is essentially an alignment strategy based on seed expansion.
[0072] In practice, the resequencing data and genome sequence of the target strain are imported into alignment software, such as the BWA-MEM algorithm in BWA (Burrows-Weler Aligner), the alignment function provided by Hisat2, or the local alignment function provided by Bowtie2, to achieve an efficient and accurate sequence alignment process.
[0073] The first alignment result is a file in SAM or BAM format. This file records in detail the alignment position, alignment quality, and CIGAR string of each resequencing read on the target strain's genome sequence. This information is used to accurately describe the match between the read and the genome, ensure the reliability of the reference genome, and provide an accurate basis for subsequent analysis.
[0074] S200. The resequencing data of at least one other strain of the first species are compared with the genome sequence of the target strain of the first species to obtain a second comparison result.
[0075] Other strains refer to strains that belong to the same species as the target strain but differ in their genetic background. For example, if the target strain is Furui Carp No. 2, then Jian Carp No. 2 would be considered an other strain.
[0076] In this step, the resequencing data from other strains serve as control data for comparison with the target strain.
[0077] Compared with the genome sequence of the target strain, the resequencing data of the other strains will have different alignment results (such as coverage and matching degree) due to the genetic variations between strains.
[0078] Specifically, the same alignment method and software (such as BWA, Hisat2, or Bowtie2) as in step S100 are used to ensure consistency in alignment standards and avoid systematic errors introduced by differences in software or parameters. For example, resequencing data from other strains are imported into BWA software for comparison with the genome sequence of the target strain, using the same reference genome file and core parameters as in S100 to obtain a second comparison result.
[0079] The second alignment result is a SAM or BAM format file generated for each other strain. This file records in detail the alignment position, alignment quality, and CIGAR string of each resequencing read on the target strain's genome sequence. It reflects the projection of the other strains' genome data onto the target strain's reference genome and can reveal which regions are conserved and which regions may have deletions or high variation.
[0080] S300. Compare the first comparison result of the target strain with the second comparison result of each other strain to obtain the result corresponding to each other strain.
[0081] The core of this step is cross-comparison, which involves analyzing the differences between the target strain and each other strain.
[0082] The comparison mainly refers to comparing the sequence coverage depth and coverage continuity of the two alignment results in different regions of the genome.
[0083] Specifically, this can be achieved using tools that calculate genomic coordinate coverage. For example, using the `genomecov` function in BEDTools software, the coverage depth of the first and second alignment results across the entire genome can be calculated. Then, using the `intersect` function in BEDTools or a custom script, genomic regions with non-zero coverage depth in the alignment results of the target strain, and genomic regions with zero coverage depth in the alignment results of another strain, can be identified.
[0084] The “results corresponding to each other strain” refers to one or more files. Each file records the coordinates of genomic regions of the target strain that are unique to a particular other strain and are not covered by reads from that other strain.
[0085] S400. Based on the results, determine the genomic region that exists in the genome sequence of the target strain and does not exist in at least one other strain, and define the genomic region as a specific sequence of the target strain relative to the at least one other strain.
[0086] This step is the result integration and final judgment stage. The "each result" refers to all the comparison results generated in step S300 for a single other strain.
[0087] The logic for this determination is that a genomic region can be identified as a specific sequence as long as it exists in the target strain and does not exist in any other strains involved in the comparison.
[0088] Specifically, this can be achieved through set operations. By systematically integrating all intermediate results generated from previous comparative analyses, and applying the principle of set difference operations, the unique genomic regions of the target strain can be mathematically and precisely defined.
[0089] The final identified specific sequences are continuous coordinate intervals within the genome of the target strain. These sequences represent insertion sequences unique to the target strain or regions where large-scale deletions have occurred in other strains. These sequences can be output in FASTA, BED, or other formats, serving as a direct basis for subsequent molecular marker development, functional studies, or strain identification.
[0090] like Figure 3 As shown, before steps S100 and S200 are executed, the following steps are also included:
[0091] S20. Use sequencing quality assessment tools to determine the sequencing quality of the resequencing data of the target strain and the sequencing quality of the resequencing data of at least one other strain.
[0092] Sequencing quality assessment tools refer to professional bioinformatics software used to systematically evaluate the quality of sequencing data. They can perform comprehensive quality checks on raw sequencing data and generate standardized quality reports. These reports are presented in the form of visual charts and statistical data, including key quality indicators such as sequencing error rate distribution, GC content anomaly detection, and adapter contamination level assessment.
[0093] Sequencing quality refers to the accuracy of sequencing each base, measured by Phred values. Q15 indicates an error rate of 3%, Q20 indicates an error rate of 1%, and Q30 indicates an error rate of 0.1%. High-quality data is fundamental to ensuring alignment accuracy.
[0094] Based on statistical probability models, the system evaluates the sequencing reliability of each base by calculating quality scores on the raw signal data generated during sequencing, thereby comprehensively understanding the quality status and error distribution characteristics of the sequencing data. For example, FastQC can be used for comprehensive quality assessment, generating a visual report containing various quality indicators; MultiQC tools can be used to integrate and analyze the quality results of multiple samples; or custom scripts in Bioinformatics Toolkits can be used for in-depth mining of specific indicators and quality trend analysis.
[0095] S40. Remove sequencing reads with sequencing quality below a preset value from the resequencing data of the target strain and the resequencing data of at least one other strain.
[0096] In the quality control phase, a sliding window algorithm is used to assess the quality of sequencing data. When the average quality value within a window falls below a threshold, that region is pruned. For example, in practice, Fastp or Trimmomatic software can be used for quality control.
[0097] The output after quality control processing is a high-quality file containing only sequencing reads that meet quality standards. This provides clean and reliable input data for subsequent sequence alignment analysis, ensuring the accuracy of the analysis results.
[0098] S60. Remove sequence adapters from the resequencing data of the target strain and the resequencing data of at least one other strain.
[0099] Sequence adapters are known sequence fragments added during the sequencing process, including sequencing primer binding sites and index sequences, such as the universal adapter sequence "AGATCGGAAGAGC" for the Illumina platform.
[0100] Specifically, in the connector removal stage, connector sequences in the read segment are identified and removed using an accurate sequence matching algorithm. For example, in practice, Fastp or Trimmomatic software can be used to perform connector removal.
[0101] The data preprocessing process includes two main steps: quality control and adapter removal. The preprocessed output file contains only high-quality sequencing reads, providing clean input data for subsequent alignment analysis.
[0102] In one possible implementation, such as Figure 4 As shown, step S300 includes:
[0103] S220. Calculate the overlap between the resequencing data of the target strain and the genome sequence of the target strain based on the first comparison result, and obtain a first calculation result; extract at least one first region with an overlap of 0 from the genome sequence of the target strain based on the first calculation result.
[0104] The first calculation result refers to the set of uncovered genomic regions identified by analyzing the alignment results of the target strain's own resequencing data with the reference genome. This file contains information on all regions in the target strain's reference genome that are not covered by its own resequencing data, specifically including: chromosome number, region start position, region end position, region length, and the coverage status (zero coverage) of the region in the target strain.
[0105] Overlap refers to the proportion of a resequencing read that covers a specific region of the genome. The formula is: Coverage = (Number of reads covering the region × Read length) / Region length. An overlap of 0 indicates that the region is completely uncovered by any resequencing read.
[0106] The first region refers to the continuous genomic fragments in the target strain's reference genome that are not covered by its own resequencing data, identified based on the first calculation results. These regions reflect potential assembly gaps, technical deviation regions, or highly polymorphic regions that may exist in the reference genome.
[0107] Based on the principle of sequence coverage depth statistics, regions in the reference genome not covered by the target strain's own resequencing data are identified by calculating the number of reads covered at each genomic location. These regions may be due to genome assembly errors, highly polymorphic regions, or technical biases.
[0108] The specific calculation process employs a sliding window-based statistical method. The reference genome is divided into fixed-size windows (typically 1 bp), the number of reads covered within each window is counted, and then adjacent zero-coverage windows are merged to form zero-coverage regions. For example, the genomecov function in BEDTools software, the mpileup function in SAMtools, or MOSdepth can be used for rapid coverage calculation.
[0109] S240. Calculate the overlap between the resequencing data of at least one other strain and the genome sequence of the target strain based on the second comparison result, and obtain a second calculation result; extract at least one second region with an overlap of 0 from the genome sequence of the target strain based on the second calculation result.
[0110] The second calculation result refers to the set of uncovered genomic regions identified by analyzing the alignment results of resequencing data from other strains onto the reference genome of the target strain. This file contains information on all continuous regions on the target reference genome of each other strain that are not covered by its resequencing data, specifically including: chromosome number, region start position, region end position, region length, and the coverage status (zero coverage) of the region in the corresponding strain.
[0111] The second region is a specific genomic fragment identified based on the second calculation result. It refers to a continuous DNA sequence region in the target strain's reference genome that is not covered by resequencing data from other comparison strains. Each second region has complete genomic coordinate information and a defined region length. These regions objectively reflect the potential missing or highly variable regions of each comparison strain relative to the target strain, providing a crucial comparative benchmark for subsequent specific sequence identification.
[0112] Also based on the statistical principle of coverage depth, but analyzing resequencing data from other strains. These uncovered regions reflect the true genetic differences between strains, including insertion / deletion variations and highly differentiated regions.
[0113] The specific calculation process uses the same analysis tools and parameter settings as step S220 to ensure the comparability of the results.
[0114] S260. Filter regions in each of the first and second regions whose sequence length is not less than a preset threshold.
[0115] The preset threshold refers to the minimum length standard set for screening candidate regions, determined based on biological significance and experimental feasibility. Preferably, the preset threshold is 80 bp. After screening, sequences with a length greater than or equal to 80 bp are obtained, which facilitates the subsequent design of specific PCR primers to achieve the detection effect of "the target strain has an amplification band while other strains do not" or "the amplification bands of the target strain and other strains are different in size".
[0116] The specific filtering process is based on the region length distribution characteristics. First, the length distribution of all candidate regions is statistically analyzed. Then, an appropriate length threshold is set according to subsequent experimental needs (such as PCR primer design, functional verification, etc.). Finally, regions with a length less than the threshold are filtered out.
[0117] Length filtering removes short fragments resulting from random sequencing errors or minor alignment biases, as these fragments typically lack biological significance or experimental validation value. For example, BEDTools' filter function can be used, or custom filters can be implemented using AWK, Python, or Perl scripts.
[0118] The output is a filtered region file after length selection, which has higher reliability and research value.
[0119] S280. Compare the selected first region with the selected second region to obtain the results corresponding to each of the other strains.
[0120] Based on the principle of set intersection, this method identifies genomic regions unique to the target strain and absent in other strains by comparing the distribution of zero-coverage regions between the target strain and other strains. For example, the `intersect` function in BEDTools can be used, along with the `GenomicRanges` package in R, or more complex region comparisons can be performed using Perl or Python scripts. In practice, a one-to-one comparison strategy can be adopted, where each candidate region of the target strain is independently compared with the corresponding region of each other strain, and the difference operation is used to filter out specific regions that exist only in the target strain.
[0121] The output results are multiple strain-specific comparison result files, each file recording the candidate specific regions of the target strain relative to a specific other strain.
[0122] This region comparison method based on set operations can effectively identify the true differences in genomic structure between strains, significantly reduce the probability of false positive results, and provide high-quality data support for strain-specific studies.
[0123] In one possible implementation, step S400 further includes:
[0124] S380 removes regions in the specific sequence where the proportion of repetitive sequences is greater than a preset ratio or the length of the repetitive sequences is greater than a preset length.
[0125] Repetitive sequences refer to sequence patterns in the genome that consist of repeated sequences of several base units, such as short tandem repeats like "ATCATCATCATC...", "CACACACA...", or "GAAGAAGAAGAA...". Due to their high sequence redundancy and structural simplicity, these sequences are not only difficult to amplify effectively using specific PCR primers, but they are also prone to sequence splicing errors during genome assembly, leading to reduced reliability of the assembly results.
[0126] The preset ratio refers to the upper limit threshold of the repetitive sequence content, usually set to 20% or 30%, which can be adjusted according to the genomic characteristics and analytical precision requirements of different species; the preset length refers to the maximum allowable length of a single repetitive sequence fragment, usually set to 50bp or 80bp, which can be optimized according to the needs of subsequent experimental verification.
[0127] The regions with many repetitive sequences refer to genomic regions where the content of repetitive sequences exceeds a preset threshold or contains extremely long repetitive fragments. These regions are prone to false positive comparison results during strain identification due to the multicopy characteristics of the sequences, affecting the accuracy of specific sequence identification.
[0128] Based on sequence similarity alignment and complexity analysis, this method systematically compares candidate specific regions with repetitive sequences in known repetitive sequence databases (such as RepBase, Dfam, or RepeatMasker Libraries). Simultaneously, it incorporates the complexity characteristics of the sequences themselves to accurately identify and filter regions with high repetition content, thereby significantly improving the reliability of the final specific sequence set. For example, IGV, RepeatMasker, WindowMasker, or DustMasker can be used for repetitive sequence annotation.
[0129] The output is a high-quality set of specific sequences filtered for repetitive sequences. These sequences have higher uniqueness and specificity, significantly improving the accuracy and reliability of strain identification and providing high-quality candidate sequence resources for subsequent molecular marker development, functional gene research, and breeding applications. This fine filtering step effectively reduces false positive results, increases the success rate of subsequent experimental verification, and lays a solid foundation for strain identification at the genome level.
[0130] This application also provides a computer program product, including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the methods for screening genome-specific sequences provided in this application.
[0131] This application also provides an electronic device, including at least one processor and a memory connected to the processor, wherein: the memory is used to store a computer program; the processor is used to execute the computer program, so that the electronic device can implement any of the methods for screening genome-specific sequences provided in this application. Reference Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing the genome-specific sequence screening method in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0132] like Figure 5 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0133] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0134] This application also provides a computer-readable storage medium carrying one or more computer programs. When these programs are executed by an electronic device, the device can implement any of the genome-specific sequence screening methods provided in this application. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0135] This application also provides specific implementation processes in its embodiments, such as... Figure 6 As shown, it includes the following steps:
[0136] 1. Quality control and alignment of resequencing data
[0137] The resequencing data of strains A and B were quality controlled using FastP 0.23.4 software. Reads with sequencing quality below Q15 and sequence adapters were removed. The quality-controlled resequencing reads of strains A / B were then aligned to the genome sequence of strain A using Hisat2 software. During alignment, the no-spliced-alignment parameter was set to prevent read truncation. A SAM file recording the alignment information was generated.
[0138] 2. Merging and processing of comparison data
[0139] Using the view function of Samtools 1.17, all SAM files were converted into binary BAM files. Then, the sort and index functions were used to sort and index the BAM files. All strain A alignments to strain A genomes were merged into a single file, A_to_A.bam, using the merge function of Samtools 1.17. The same method was used to obtain the merged file B_to_A.bam for strain B alignments to strain A genomes. Finally, the index function was used to index both BAM files.
[0140] 3. Extract the 0-read coverage region of the genome.
[0141] Using the genomecov module of Bedtools 2.29.1 software, the read coverage of each region in the A genome sequence in A_to_A.bam and B_to_A.bam was calculated. During the process, the bga parameter was set to find the region with 0 coverage, and the results were output in .bed file format. The final output files are A_to_A_zero_cov.bed and B_to_A_zero_cov.bed.
[0142] 4. Screening candidate differential fragments
[0143] Using the Perl script `filter_bed_gap_length.sh`, regions with a length ≥200bp in A_to_A_zero_cov.bed and B_to_A_zero_cov.bed are filtered to generate bed files: A_to_A_zero_cov_200.bed and B_to_A_zero_cov_200.bed. Using the Perl script `filter_no_overlap.pl`, regions of the A genome present in B_to_A_zero_cov_200.bed but not in A_to_A_zero_cov_200.bed are filtered; these are the regions specific to strain A relative to strain B, and this process generates the file A_special.bed.
[0144] 5. Remove duplicate sequences
[0145] Using IGV 2.18.2 software, visualize the A_to_A.bam and B_to_A.bam files, locate the regions shown in the A_special.bed file, remove regions containing many repetitive sequences, and retain the regions to form the file A_special_final.bed, which is the finally selected differential DNA fragment between strains A and B.
[0146] This application also provides process verification results in its embodiments, such as... Figure 7 As shown, the details are as follows:
[0147] Both Furui Carp No. 2 and Jianli Carp No. 2 are artificially bred varieties of carp (Cyprinus carpio). They are similar in appearance, especially in the juvenile stage, where there is almost no difference in appearance, making accurate differentiation difficult in actual aquaculture based solely on appearance. In the validation process, Furui Carp No. 2 was used as strain A, and Jianli Carp No. 2 as strain B. Twenty-six resequencing data points from Furui Carp No. 2 were downloaded from the NCBI database (project number: PRJNA684797), while the Furui Carp No. 2 genome and 60 Jianli Carp No. 2 resequencing data points came from previous research by the team. After data preparation, differentially expressed DNA sequences were screened according to the above analysis procedure, ultimately yielding seven sequences specific to Furui Carp No. 2 relative to Jianli Carp No. 2.
[0148] Primers were designed for PCR amplification of these sequences, and one differentially expressed DNA sequence was validated. A 241 bp band appeared in all 30 Furui Carp No. 2 individuals, while no band appeared in any of the 30 Jian Carp No. 2 samples. This DNA sequence can distinguish Furui Carp No. 2 and Jian Carp No. 2 with 100% accuracy. Figure 7 ).
[0149] Verification experiments have shown that the technical process described in this invention can indeed screen for DNA fragments that differ between different strains.
[0150] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0151] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0152] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs (Compact Disc-Read-Only Memory), or any other form of storage medium known in the art.
[0153] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0154] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for screening genome-specific sequences, characterized in that, include: The resequencing data of the target strain of the first species is compared with the genome sequence of the target strain of the first species to obtain the first comparison result; The resequencing data of at least one other strain of the first species are compared with the genome sequence of the target strain of the first species to obtain a second comparison result; The first comparison result of the target strain is compared with the second comparison result of each other strain to obtain the result corresponding to each other strain. Based on the results, a genomic region that exists in the genome sequence of the target strain and does not exist in at least one other strain is identified, and the genomic region is identified as a specific sequence of the target strain relative to the at least one other strain.
2. The method according to claim 1, characterized in that, Before comparing the resequencing data of the first species with the genome sequence of the first species, the method further includes: The sequencing quality of the resequencing data of the target strain and the sequencing quality of the resequencing data of at least one other strain were determined using sequencing quality assessment tools. Remove sequencing reads with sequencing quality below a preset value from the resequencing data of the target strain and the resequencing data of at least one other strain; Remove sequence adapters from the resequencing data of the target strain and the resequencing data of at least one other strain.
3. The method according to claim 1, characterized in that, The genome sequence of the target strain is obtained from third-generation sequencing data; The resequencing data for the other strains were obtained using next-generation sequencing.
4. The method according to claim 1, characterized in that, The step of comparing the first comparison result of the target strain with the second comparison result of each other strain to obtain the result corresponding to each other strain includes: The overlap between the resequencing data of the target strain and the genome sequence of the target strain is calculated based on the first comparison result to obtain a first calculation result; at least one first region with an overlap of 0 is extracted from the genome sequence of the target strain based on the first calculation result. Based on the second comparison result, the overlap between the resequencing data of at least one other strain and the genome sequence of the target strain is calculated to obtain a second calculation result; based on the second calculation result, at least one second region with an overlap of 0 is extracted from the genome sequence of the target strain. Filter regions in each of the first and second regions whose sequence length is not less than a preset threshold; The first selected region is compared with the second selected region to obtain the results corresponding to each of the other strains.
5. The method according to claim 4, characterized in that, The preset threshold is 80bp.
6. The method according to claim 1, characterized in that, Also includes: Remove regions in the specific sequence where the proportion of repetitive sequences is greater than a preset ratio or the length of the repetitive sequences is greater than a preset length.
7. The method according to claim 6, characterized in that, The preset ratio is 10%~40%, and the preset length is 40~90bp.
8. A computer program product, characterized in that, Includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to perform the method for screening genome-specific sequences as described in any one of claims 1 to 7.
9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the method for screening genome-specific sequences as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to perform the method for screening genome-specific sequences as described in any one of claims 1 to 7.