Screening method of genome high-specificity SNP (Single Nucleotide Polymorphism) probe, probe map and application of probe map in siniperca chuatsi breeding

Highly specific SNP probes were screened using a dual K-mer cross-validation process, a high-density probe map of mandarin fish was constructed, and a disease-resistant breeding chip was developed. This solved the problems of efficiency and accuracy in SNP marker screening in mandarin fish breeding, and improved breeding efficiency and the selection effect of disease resistance traits.

CN121780717APending Publication Date: 2026-04-03INST OF AQUATIC LIFE ACAD SINICA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently and accurately screen out SNP markers that can be used for disease resistance traits in mandarin fish farming, resulting in low breeding efficiency. Furthermore, existing methods cannot simultaneously achieve high detection capability and high identification accuracy.

Method used

Using a dual K-mer cross-validation process based on parental sequencing data, SNP probes with high specificity in both genomic sequence origin and physical coordinates were screened, a high-density probe map was constructed, and a disease-resistant breeding chip was developed.

Benefits of technology

It has enabled the accurate identification of SNP loci in the whole genome of mandarin fish and efficient typing of large-scale populations, which has significantly improved the efficiency and accuracy of genetic breeding and significantly enhanced the breeding effect of disease resistance traits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121780717A_ABST
    Figure CN121780717A_ABST
Patent Text Reader

Abstract

The invention discloses a screening method of a genome high-specificity SNP (Single Nucleotide Polymorphism) probe, a probe map and application of the probe map in siniperca chuatsi breeding. The screening method comprises the following steps: obtaining parents Clean reads and Mapped reads, and carrying out screening on the Clean reads and the Mapped reads, determining a two-state SNP (Single Nucleotide Polymorphism) site; a candidate probe sequence R1-Kmers is generated; verifying the uniqueness of the sequence on the basis of Clean reads, and screening out R2-Kmers; and the uniqueness of the coordinates is verified based on Mapped reads. A siniperca chuatsi genome high-density SNP probe map containing 3, 206 and 138 probes is constructed by using the method. Based on the atlas, 85 and 395 disease-resistant related SNP sites are screened out by performing genetic typing and correlation analysis on survival / death groups of virus counteracting families, and the siniperca chuatsi disease-resistant breeding liquid-phase chip is successfully constructed. Practical breeding verification shows that the viability attacking survival rate of the family screened by the chip is remarkably improved compared with that of a control group. The invention provides a key technical tool for efficient and accurate molecular breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of aquatic animal genetics and breeding and genomics, specifically to a method for screening highly specific SNP probes for the mandarin fish genome, a high-density SNP probe map constructed by this method, and its application in the selection of disease-resistant mandarin fish genomes and the construction of breeding chips. Background Technology

[0002] Mandarin fish is an important freshwater aquaculture species in my country. Since the mid-1980s, the scale of mandarin fish farming in my country has expanded rapidly, reaching 540,000 tons in 2024, with a value exceeding 40 billion yuan. However, the outbreak of infectious spleen and kidney necrosis virus (ISKNV) disease in mandarin fish has posed a persistent threat to its aquaculture industry, causing huge economic losses. Breeding new disease-resistant varieties is a key guarantee for the healthy development of the industry. Previous research by the inventors has found individual differences in disease resistance within the mandarin fish population, indicating the feasibility of discovering disease-resistant markers through genomic selection.

[0003] Precise genotyping of polymorphic sites in the genome is a core step in achieving genomic selection for superior varieties. Currently, there are two main technical approaches for large-scale SNP genotyping, each with significant limitations:

[0004] Firstly, chip-based technologies (such as SNP chips) are currently the gold standard for large-scale genotyping, offering stable data quality. However, their fundamental drawback lies in the limited number of detections and the fact that they can only detect known SNP sites pre-designed on the chip, failing to discover new sites and lacking flexibility and discovery capabilities.

[0005] Secondly, next-generation sequencing (NGS) technologies, especially whole-genome resequencing, can detect all SNPs, including rare variants and novel sites. However, this method generates a massive amount of data, and in the context of a complex genome, it is difficult to accurately determine whether the location of sequencing reads aligned to the reference genome is stable and reliable. This uncertainty in location directly leads to a decrease in the accuracy of subsequent SNP / Indel marker discovery and association analysis.

[0006] Specifically, in the process of whole-genome selection for fish, a common bottleneck faced by existing methods is how to stably and accurately identify marker loci that can be used for efficient genotyping from massive amounts of resequencing data. Both the inability of microarray technology to discover new loci and the difficulty in determining the stability of resequencing locations hinder the accurate discovery of trait-related marker loci, thus affecting breeding efficiency.

[0007] Therefore, there is an urgent need for a new technical solution that can balance high discovery capability with high identification accuracy in order to solve this fundamental technical problem in large-scale SNP marker screening and typing in complex plant and animal genomes. Summary of the Invention

[0008] In view of this, the purpose of this application is to provide a method for screening highly specific SNP probes of the genome, a probe map, and its application in mandarin fish breeding. This screening method uses a dual K-mer cross-validation process based on parental sequencing data to screen for SNP probes with high specificity in both genomic sequence origin and physical coordinates, thereby constructing a high-density probe map covering the entire genome. Based on this map, efficient disease resistance breeding chips can be further developed and successfully applied to the precise selection of disease resistance traits. The method established in this application not only achieves precise identification of SNP loci in the entire fish genome and efficient large-scale population typing, but also provides a reliable technical tool for genome breeding of other animals and plants, significantly improving the efficiency and accuracy of genetic breeding.

[0009] Therefore, this application provides the following technical solution:

[0010] Firstly, this application provides a method for screening genome-specific SNP probes, comprising the following steps: (1) obtaining genome resequencing data of a target biological individual, wherein the resequencing data includes clean sequencing reads.

[0011] (2) Based on the resequencing data, determine candidate dimorphic SNP sites from the reference genome; (3) Sequence generation: Using the dimorphic SNP site as the center, extract candidate sequences containing the site from the reference genome.

[0012] Probe sequences, denoted as R1-Kmers;

[0013] (4) Probe uniqueness verification based on clean reads: The clean reads are cut into short sequences to obtain Q1-Kmers; the Q1-Kmers are fully aligned with the R1-Kmers; based on the alignment results, sequences that meet the first condition are selected from the R1-Kmers and denoted as R2-Kmers; the first condition is: for an R1-Kmer, all Q1-Kmers that can be fully aligned with it are from the same clean read;

[0014] (5) Verification of probe coordinate uniqueness based on Mapped reads: The Mapped reads are truncated into short sequences to obtain Q2-Kmers; the Q2-Kmers are fully aligned with the R2-Kmers; based on the alignment results, sequences that meet the second condition are selected from the R2-Kmers; the second condition is: for an R2-Kmer, its own chromosome number (G) on the reference genome is... Kmer ), and the chromosome number (G) of all the mapped reads from which Q2-Kmers could be perfectly compared. ID (6) The R2-Kmers verified by steps (4) and (5) are identified as highly specific SNP probes.

[0015] Secondly, this application provides a high-density SNP probe map of the mandarin fish genome, the map containing multiple SNP probes, the SNP probes being R2-Kmers obtained using the screening method described in the first aspect.

[0016] Thirdly, this application provides a method for constructing a disease-resistant breeding chip for mandarin fish, comprising the following steps:

[0017] (a) Obtain survival and death samples from full-sibling families of mandarin fish challenged with the ISKNV virus;

[0018] (b) Genotyping of the surviving and dead samples using the high-density SNP probe map of the mandarin fish genome described in the second aspect;

[0019] (c) Compare and analyze the genotype data of the surviving and dead samples to screen out SNP loci associated with disease resistance traits;

[0020] (d) Based on the screened associated SNP sites, probes are synthesized and prepared into breeding chips.

[0021] Fourthly, this application provides a disease-resistant breeding chip for mandarin fish, which is obtained using the construction method described in the third aspect.

[0022] Compared with the prior art, this application has at least the following advantages and beneficial effects:

[0023] 1. This application pioneers a dual cross-validation process that combines "sequence origin uniqueness verification" based on parental clean reads with "genomic coordinate consistency verification" based on mapped reads. This process can empirically screen probe sequences that are non-repetitive in the original data and absolutely uniquely located on the reference genome, eliminating non-specific hybridization at the source and laying a methodological foundation for building a highly reliable detection platform.

[0024] 2. This application transforms the complex determination of probe specificity into deterministic algorithm steps based on precise K-mer alignment and coordinate calculation (such as "one-to-one correspondence", "chromosome number consistency", and region screening conditions). This process does not rely on human experience, is repeatable and efficient, and can further ensure the final detection sensitivity of the probe through steps such as "probe centering optimization", realizing the automated output of a ready-to-use high-quality probe set from massive amounts of data.

[0025] 3. This application constructs a high-density specific probe map for mandarin fish that can be directly used in breeding practices. Using the above method, a mandarin fish genome map containing over three million probes was successfully constructed. The probes in this map possess both high genome coverage and empirically proven high specificity, and can be directly used as a core database for genotyping and microarray design, solving the disconnect between marker "discovery" and "application".

[0026] 4. The mandarin fish disease resistance breeding chip developed in this application based on probe mapping has shown excellent performance in actual breeding screening. Disease-resistant families constructed using this chip, after viral challenge verification, exhibited significantly higher survival rates (e.g., 17.65%-44.22%) compared to the ordinary control population. This directly demonstrates the effectiveness of the screening markers used in this application and the significant practical value of the constructed technical system in accelerating the breeding of disease-resistant varieties and improving breeding efficiency.

[0027] 5. The core probe screening method of this application is universal and can be applied not only to the disease resistance trait of mandarin fish, but also to provide key technical support for the genomic selection breeding of other aquatic animals, livestock and poultry and crops, and has broad prospects for promotion and application. Attached Figure Description

[0028] Figure 1 This is a flowchart of the screening process for highly specific SNP probes for the genome in this application.

[0029] Figure 2 A high-density SNP probe map of the mandarin fish genome containing 3,206,138 SNP sites was constructed for this application.

[0030] Figure 3 This is a map showing the distribution of the 85K SNP locus in the genome of the breeding chip for disease resistance in mandarin fish (as described in this application). Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0032] The materials used in the following embodiments are not limited to those listed below, and other similar materials may be used instead. Unless otherwise specified, the instruments shall be used under conventional conditions or as recommended by the manufacturer. Those skilled in the art should have relevant knowledge of the use of conventional materials and instruments.

[0033] In this application, the terms "comprising," "including," "containing," or similar expressions are "open-ended," meaning that in addition to the listed elements, components, or steps, other elements, components, or steps not explicitly listed may be included. Furthermore, the ordinal numbers "first," "second," etc., used in this application are only used to distinguish between multiple similar objects or steps and are not intended to imply a specific order or priority of these objects or steps in time, space, importance, or any other aspect.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter of this application pertains. To enable those skilled in the art to more fully and accurately understand the technical solutions of this application, the following terms and definitions are explained:

[0035] 1. K-mer (or Kmer): refers to all possible short nucleotide sequence fragments of length K (K is a positive integer, usually an odd number) continuously extracted from a DNA or RNA sequence from beginning to end. In this application, K-mer, short sequence, and short sequence fragment have the same meaning.

[0036] 2. Clean reads: refers to high-quality sequencing reads obtained from raw data of high-throughput sequencing after quality control, removal of low-quality bases and adapter sequences, which can be used for subsequent bioinformatics analysis.

[0037] 3. Mapped reads: These are clean reads that have been successfully aligned to a specific location on the reference genome using sequence alignment software (such as Bowtie 2). Each mapped read records its aligned chromosome number, start coordinates, and end coordinates.

[0038] 4. Reference Genome: Refers to a known, complete genome sequence that serves as a standard for alignment and analysis. In one embodiment of this application, the genome of the mandarin fish (Siniperca chuatsi) is assembled at the chromosome level.

[0039] 5. Complete alignment (or 100% alignment): This means that during the alignment of two K-mer sequences, all corresponding bases from the 5' end to the 3' end of the two sequences must be completely identical, and no single base mismatch, insertion or deletion is allowed.

[0040] 6. One-to-one correspondence: In step (4) of this application, it specifically means that a candidate probe sequence (R1-Kmer) is completely aligned and covered by a short sequence (Q1-Kmer) derived from a single Clean read. That is, there are no two different Clean reads that can produce a Q1-Kmer that can be completely aligned with the same R1-Kmer.

[0041] 7. Dimorphic SNP sites: These are single nucleotide polymorphism sites located at a specific position in the genome within a population, where only two different base types are present. For example, at this position, some individuals in the population may have an A base, while others may have a G base, and no third base (such as C or T) may be present.

[0042] 8. Degenerate base symbols: Internationally recognized IUPAC codes are used to represent one of several possible bases. Specifically, in this application, these include: M (A or C), R (A or G), W (A or T), S (C or G), Y (C or T), and K (G or T). These are used to uniformly represent two possible bases at a dimorphic SNP site in the sequence.

[0043] 9. Coordinates (Z, K, X, Y): Z: refers to the physical location of the first base at the 5' end of a K-mer (such as R1-Kmer or R2-Kmer) on the reference genome; K: refers to the physical location of the SNP site on the reference genome; X: refers to the physical location of the alignment start position (leftmost coordinate) of a Mapped read on the reference genome; Y: refers to the physical location of the alignment end position (rightmost coordinate) of a Mapped read on the reference genome.

[0044] 10. Probe Centering: This refers to the position (K) of the target SNP site contained within a K-mer sequence used as a candidate probe being as close as possible to the midpoint of the K-mer sequence (Z+L / 2, where L is the K-mer length). In a preferred embodiment of this application, screening is achieved by minimizing the value of the expression |Z+25-K|.

[0045] The inventors of this application homogenize reference sequences and NGS sequences, and based on the alignment information of the NGS sequences, extract short sequences (Kmers) at specific coordinates. This gives the Kmers from the reference sequence both marker and localization characteristics, which are used to determine the stability of the NGS sequences in the genome. Building upon this, and based on the dimorphic characteristics of genetic markers, a novel genomic SNP marker screening and high-throughput genotyping detection technology is created. Furthermore, a breeding liquid-phase chip is developed to accurately identify complex trait-related markers starting from whole-genome marker sites in fish, thereby solving the bottleneck problem of difficulty in discovering trait-related marker sites.

[0046] Based on this, embodiments of this application provide a method for screening genome-specific SNP probes, such as... Figure 1 As shown.

[0047] The screening method includes the following steps: (1) obtaining the genome resequencing data of the target biological individual, wherein the resequencing data includes the original sequencing reads (Clean sequencing reads).

[0048] (2) Based on the resequencing data, determine candidate dimorphic SNP sites from the reference genome; (3) Sequence generation: Using the dimorphic SNP site as the center, extract candidate sequences containing the site from the reference genome.

[0049] Probe sequences, denoted as R1-Kmers;

[0050] (4) Probe uniqueness verification based on clean reads: The clean reads are cut into short sequences to obtain Q1-Kmers; the Q1-Kmers are fully aligned with the R1-Kmers; based on the alignment results, sequences that meet the first condition are selected from the R1-Kmers and denoted as R2-Kmers; the first condition is: for an R1-Kmer, all Q1-Kmers that can be fully aligned with it are from the same clean read;

[0051] (5) Verification of probe coordinate uniqueness based on Mapped reads: The Mapped reads are truncated into short sequences to obtain Q2-Kmers; the Q2-Kmers are fully aligned with the R2-Kmers; based on the alignment results, sequences that meet the second condition are selected from the R2-Kmers; the second condition is: for an R2-Kmer, its own chromosome number (G) on the reference genome is... Kmer ), and the chromosome number (G) of all the mapped reads from which Q2-Kmers could be perfectly compared. ID (6) The R2-Kmers verified by steps (4) and (5) are identified as highly specific SNP probes.

[0052] The above technical solution employs a "dual data source cross-validation" logic. First, the raw, unmapped clean reads (Q1-Kmers) are used to verify whether candidate probes (R1-Kmers) have a unique sequence origin in the real data, thus eliminating "false uniqueness" sequences caused by local duplication or sequencing coincidence. Second, precisely mapped reads (Q2-Kmers) are used to verify whether the probes that passed the first step (R2-Kmers) have unique and stable physical coordinates on the genome, thus resolving the localization ambiguity caused by genome assembly errors or unannotated homologous regions. This technical solution fundamentally solves the risk source of non-specific hybridization in high-throughput probe design, elevating probe screening from a "predictive" level to an "empirical" level, and providing a universal methodology for building any highly reliable genotyping detection platform.

[0053] In some preferred embodiments, step (3) includes: rewriting the dimorphic SNP site using degenerate base symbols, and then extracting R1-Kmers from the reference genome; the rewriting principle is: M=A / C, R=A / G, W=A / T, S=C / G, Y=C / T, K=G / T. For dimorphic SNPs (such as A / G), directly taking one base will lose the information of the other allele. Using IUPAC degenerate base symbols (such as R) for unified rewriting is essentially abstracting the two variant forms at the same physical location into a single "site" entity. This ensures that all subsequent K-mer cutting, alignment, and screening operations centered on this site can simultaneously accommodate both alleles, thereby ensuring that the finally screened probes can detect both genotypes at this site equally and effectively, avoiding allele bias in probe design.

[0054] In some embodiments, the truncation in step (3) is to slide and truncate short sequences of 50-70 bp at 1-base intervals to generate R1-Kmers; the cutting in steps (4) and (5) is to cut Clean reads and Mapped reads into short sequences of the same length to generate Q1-Kmers and Q2-Kmers, respectively.

[0055] In some preferred embodiments, the short sequence length of the truncation and trimming in step (3) is 51 bp. The "1 bp interval sliding truncation" ensures that every possible 51 bp window on the reference genome is traversed, achieving complete coverage. A "fixed length (e.g., 51 bp)" is a balanced choice: long enough to ensure its specificity in complex genomes (avoiding random matching), and short enough to accommodate sequencing read lengths and tolerate individual sequencing errors. Q1 / Q2-Kmers use the same length to ensure accurate alignment with R1-Kmers at the exact same scale. This approach provides an efficient, complete, and standardized sequence digitization method, transforming continuous genomes and sequencing reads into discrete, precisely alignable sets of K-mers, which forms the basis for all subsequent automated alignment and screening.

[0056] In some embodiments, in step (4), the unique identifier of the Clean read from which the Q1-Kmers originate is recorded; the "complete alignment" means that no mismatch, insertion, or deletion of any bases is allowed between the two short sequences being aligned; the "originating from the same Clean read" is determined based on the unique identifier. This step concretizes the abstract condition of "one-to-one correspondence" into a judgment based on the unique identifier of the Clean read, forcing the verification process to trace back to the most original sequencing data unit. The "complete alignment" is defined as 100% consistency, which sets an extremely high entry threshold, ensuring the absolute accuracy of the alignment results and avoiding fuzzy judgments introduced by allowing mismatches. Through this step, the key step of "sequence uniqueness verification" becomes strictly definable, automatically executed, and the results are indisputable, greatly enhancing the robustness and reproducibility of the method.

[0057] In some preferred embodiments, a region-specific screening step is further included after step (5): for the R2-Kmers, sequences located entirely within a continuous region of a single Mapped read are further screened; the screening is determined by the following conditions: let the 5' start coordinate of the R2-Kmer on the reference genome be Z, and its length be L; the start and end coordinates of a Mapped read used for alignment be X and Y, respectively; then X ≤ Z and Y ≥ Z + L must be satisfied. The reason for this step is that even if an R2-Kmer passes double verification, if it is located at the end of a Mapped read, its sequencing quality may be low, and its alignment coordinates (X, Y) itself may have boundary errors. The condition X ≤ Z and Y ≥ Z + L forces that the entire sequence of the R2-Kmer must be located within a high-quality continuous region of a Mapped read. This is equivalent to adding a "quality gate" to the probe template, screening out candidate sequences from the edge of the sequencing read with relatively low reliability, thereby further improving the overall reliability and signal stability of the final probe set.

[0058] In some preferred embodiments, a probe optimization step is further included after step (5) or step (6): for multiple R2-Kmers corresponding to the same SNP site, the R2-Kmer whose SNP site coordinate (K) is closest to the center of the R2-Kmer sequence is selected as the representative probe for that SNP site. The thermodynamic stability of DNA hybridization is significantly affected by the position of mismatched bases. When the SNP is located at the center of the probe sequence, a single base mismatch will cause the greatest damage to the stability of the entire double-stranded structure, resulting in a significant decrease in the melting temperature (Tm) of the hybrid. Therefore, selecting a probe that is in the center can maximize the signal difference of allele-specific hybridization. In subsequent microarray hybridization or liquid-phase capture experiments, this makes it easier and more accurate to clearly distinguish between homozygotes and heterozygotes, as well as between two different homozygotes, through strict elution conditions, thereby directly improving the sensitivity and resolution of genotyping.

[0059] Based on the above-mentioned SNP probe screening method, this application embodiment also provides a high-density SNP probe map of the mandarin fish genome, wherein the SNP probe map contains multiple SNP probes, and the SNP probes are R2-Kmers obtained by any of the aforementioned screening methods.

[0060] Based on this, the embodiments of this application also provide the application of the above-obtained SNP probe map in the preparation of a mandarin fish disease resistance breeding chip, that is, a method for constructing a mandarin fish disease resistance breeding chip, specifically including the following steps: (a) obtaining surviving and dead samples from a full-sib family population of mandarin fish challenged by ISKNV virus; (b) using the above-mentioned high-density SNP probe map of the mandarin fish genome to perform genotyping on the surviving and dead samples respectively; (c) comparing and analyzing the genotype data of the surviving and dead samples to screen out SNP sites associated with disease resistance traits; (d) based on the screened associated SNP sites, synthesizing probes and preparing a breeding chip.

[0061] High-throughput, high-accuracy genotyping was achieved using a high-density SNP probe map of the mandarin fish genome. Then, by comparing the genotype frequencies of disease-resistant (surviving) and susceptible (dead) populations, a subset of SNP loci significantly associated with disease resistance in mandarin fish was statistically identified. These loci are the "targets" that truly need attention in breeding practice. This achieves precise focusing from "massive markers across the entire genome" to "key trait-related markers." The breeding microarray constructed based on this is no longer a simple genotyping tool, but a dedicated and highly efficient tool for genomic selection targeting disease resistance. This microarray can effectively enrich disease resistance alleles, achieving predictable and verifiable breeding gains.

[0062] The technical solution of this application and the technical effects achieved will be described in detail below through more specific embodiments.

[0063] Example 1: Construction of a high-density SNP probe map of the mandarin fish genome

[0064] This embodiment provides a specific method for constructing a high-density SNP probe map of the mandarin fish genome containing 3,206,138 SNP sites. The specific steps are as follows.

[0065] S1. Construction of the whole sibling population and sample collection

[0066] One male and one female mandarin fish were collected from three independent river systems: the Yangtze River, the Pearl River, and the Xiang River, to construct three full-sib families. At three months of age, over 2000 individuals from each family were randomly selected and transferred to three sealed cement tanks indoors. After one week of temporary rearing, they were inoculated with ISKNV virus solution (2.47 × 10⁻⁶). 3Challenge was performed by immersing the fish in a solution of TCID50 / mL at 28 ℃ for 30 min under adequate dissolved oxygen conditions. At the peak of the disease outbreak and mortality, 100 dead individuals were randomly selected from each family, and their caudal fins were harvested. After the disease course ended (with no deaths for 2 weeks), the caudal fins of 100 surviving individuals were harvested. S2. Nucleic acid preparation and sequencing: Caudal fin DNA was extracted from all 600 surviving and dead samples using the Universal Genomic DNA Kit (CW2298M, CWBIO). Simultaneously, caudal fin DNA was extracted from the caudal fins of 6 parents from these 3 families. All DNA samples were quantitatively analyzed using Qubit and then re-sequencing using the BGI MGI platform at a PE150 depth of 50-60×. S3. Genome-wide candidate SNP locus discovery.

[0067] Bowtie 2 (v2.5.4) software was used to align the clean data obtained from parental resequencing with the reference genome (assembly version: ASM2008510v1, GenBank accession number: GCA_020085105.1, RefSeq accession number: GCF_020085105.1). (Alignment was performed using end-to-end mode, with the -no-unal parameter enabled to prevent unaligned reads from being output in the final SAM format results.) Mapped reads were obtained, ranging from 378,471,853 to 392,495,472. The chromosome numbers G corresponding to these mapped reads were recorded. ID The starting coordinates (X) and ending coordinates (Y) were determined. Then, the HaplotypeCaller module in the GATK toolkit (v.3.8) was used to identify SNPs and indel variations, and the BaseRecalibrator and IndelRealigner modules were used for base quality correction. The genomic variations of each material output by the HaplotypeCaller module were merged in GVCF format, and the original SNPs were further screened using GATK filtering expressions (parameters: 'QUAL<30.0||QD<2.0||FS>60.0||MQ<40.0||SOR>4.0', with a clustering window size of 5 and a cluster size of 2). The number of screened SNPs ranged from 1,027,295 to 1,083,285. After merging and deduplicating the SNPs detected from each of the six parents, a total of 3,479,873 usable SNPs were obtained, of which 3,411,492 were dimorphic and did not contain indels. S4. Probe discovery of candidate SNP sites across the entire mandarin fish genome

[0068] For the reference genome, the selected sites were rewritten according to the following SNP rules: M=A / C, R=A / G, W=A / T, S=C / G, Y=C / T, and K=G / T. After rewriting, the entire genome was truncated to 51bp Kmers at 1bp intervals. From these, 119,285,247 pairs of Kmers containing only one dimorphic SNP were selected, denoted as R1-Kmers, and their corresponding chromosome numbers G were recorded. Kmer The coordinates (Z at the 5' end and K at the SNP) correspond to 3,326,389 SNP sites. S5. Removal of repetitive probe sequences within the same read: Clean reads obtained from parental sequencing are also pruned with 51 bp Kmers to obtain Q1-Kmers (recording the source Clean read number). These are then completely aligned to R1-Kmers (requiring 100% sequence consistency). R1-Kmers covered by Q1-Kmers that correspond one-to-one with the Clean read numbers are retained. Each of the six parents is individually screened to obtain their own R2-Kmer candidate sets, ranging from 108,248,859 to 110,525,643 pairs. To obtain probe sequences that are consistently validated in all parents, Kmer sequences simultaneously present in the above six parental R2-Kmer candidate sets are selected, resulting in a total of 101,031,385 probe sequences, denoted as the unified R2-Kmer set. S6. Removal of repetitive probe sequences on different chromosomes

[0069] The mapped reads obtained in S3 are also pruned using Kmer to obtain Q2-Kmers, which are then fully aligned to R2-Kmers, leaving G. Kme r and G IDIdentical R2-Kmers. After individual testing of the six parents, 97,238,193-98,284,231 pairs of reference Kmers were retained. S7. Screening of probe sequences for specific regions of the same chromosome. Further, the above-mentioned region-specific screening (XZ≤0 and YZ≥50) was performed on the six parents individually, obtaining their respective sets of R2-Kmers that passed the screening, numbering 96,384,183-97,224,285 pairs. To obtain probe sequences located in stable genomic regions in all parents, the intersection of the above six sets was taken, resulting in a total of 93,295,248 pairs of R2-Kmers, thus completing the screening of probe sequences for specific genomic regions. S8. Construction of the SNP probe map of the mandarin fish genome: 3,206,138 R2-Kmers (with the smallest |Z+25-K| value) at the same SNP locus were selected. The remaining R2-Kmers (i.e., high-density SNP probes in the genome) were then constructed. Figure 2 As shown, a high-density probe map of the mandarin fish genome containing 3,206,138 highly specific SNP probes was obtained. Each probe (R2-Kmer) in this map underwent dual uniqueness verification in steps S5 and S6 to ensure its unique source and unique genomic coordinates in the actual sequencing data.

[0070] Example 2: Construction of a breeding chip for disease resistance in mandarin fish

[0071] This implementation provides a method for constructing a breeding chip for disease resistance traits in mandarin fish using the aforementioned probe maps. The specific steps are as follows:

[0072] (1) Classification of surviving and deceased subgroups in whole-sib family groups

[0073] Using a constructed high-density SNP probe map of the mandarin fish genome, genotyping was performed on the surviving and deceased subpopulations of three full-sib families after challenge. The genotyping criteria involved truncating the sequencing data of all genotyping samples to 51 bp Kmers at 1 bp intervals, and then performing a complete alignment with the whole-genome SNP probe map of the mandarin fish. If any probe locus matched one allele in the sample to be genotyped, and the number of alleles was ≥7, the sample was considered homozygous at that probe locus; if two alleles were detected at the genotyping locus, and the sum of the numbers of the two alleles was ≥7, it was considered heterozygous. Following this rule, 3,201,294, 3,205,194, and 3,204,229 genotypes were successfully identified in the three families, respectively.

[0074] (2) Analysis of common and segregating sites

[0075] A total of 3,200,714 common loci were detected in all three families. Then, loci that produced genotypic segregation in all three families were selected and redundancy was removed, resulting in a total of 3,194,024 loci. These were used for the next step of genotypic comparison analysis between the offspring (surviving and deceased samples) of the families and the surviving samples.

[0076] (3) Comparative analysis of genotypes of offspring and surviving samples from families

[0077] For the 3,194,024 loci whose genotypes were separated after offspring typing in step (2), the typing results of all loci in the offspring of each of the three family populations and their respective surviving samples were statistically analyzed and divided into 7 categories (AA, BB, AB, AA / BB, AA / AB, AB / BB, AA / AB / BB).

[0078] (4) Screening and summarizing candidate loci for breeding chips: Among the surviving samples of three families, the genotypes were AB and AB / BB, and the corresponding loci in the offspring population were genotyped as AA / AB, respectively.

[0079] A total of 139,394 loci were identified as AA / AB / BB. Additionally, 128,492 loci were identified in the surviving samples of three families with genotypes AB and AA / AB, and corresponding genotypes AB / BB and AA / AB / BB in their offspring populations. Redundancy was removed from the two sets of loci (with A and B as the recessive death alleles), resulting in 95,395 candidate loci.

[0080] (5) Identification of breeding chip sites

[0081] From the 95,395 candidate loci in step (4), three loci with a family phylogenetic death genotype of AA or BB were isolated, totaling 85,395 loci, which are the breeding liquid-phase chip loci. The distribution of these loci on the genome is as follows: Figure 3 As shown. (6) Synthesis of breeding liquid-phase chip

[0082] In step (5), among the 85,395 sites selected, 10 probes with a length of 120 bp were designed for each site to cover the target SNP. Based on the principle of DNA complementarity, the target probes were modified and synthesized using biotin labeling to complete the construction of the breeding liquid phase chip.

[0083] Example 3: Validation of the effect of mandarin fish genome breeding chip

[0084] This embodiment provides an effect verification of the mandarin fish genome breeding chip prepared in Example 2. The specific steps are as follows:

[0085] (1) Capture of chip probe sites

[0086] Ten mandarin fish samples were randomly selected to test the probe capture ability. After mixing and extracting tail fin DNA, a resequencing library was constructed. Biotin-modified probes were hybridized with the target region of the genome in liquid to form double strands.

[0087] (2) Detection rate test of chip sites

[0088] Streptavidin-coated magnetic beads were used to molecularly adsorb biotin-modified probes, thereby capturing target sites that hybridize with the probes. The captured target sequences were eluted, amplified, and re-sequencing, with a sequencing depth of approximately 10×, yielding 10.84 G clean reads. A total of 85,011 sites were successfully genotyped, with a chip site detection rate of 99.55%, indicating that the chip has a high detection rate.

[0089] (3) Test group construction

[0090] Using this breeding chip, a preliminary screening was conducted on 259 sexually mature mandarin fish parents collected by the inventors' team in the early stage. The proportion of resistance allele loci to the total detected loci was counted. Three pairs of parents with the highest abundance of resistance allele loci (58.63%, 62.51%, and 69.37%, respectively) were selected for one-to-one breeding. At the same time, a random population of wild-type was bred as a control.

[0091] (4) Testing the breeding effect of the chip

[0092] At 3 months of age, 273, 402, and 296 fish from three families and 416 fish from a randomized control group were randomly selected, and all fish were challenged with ISKNV virus by immersion. The results showed that the survival rates of the three candidate disease-resistant families were 25.92%, 37.18%, and 52.49%, respectively, significantly higher than the control group (8.27%) by 17.65-44.22 percentage points. These results validate the successful construction of the mandarin fish disease-resistant breeding chip and high-density SNP probe map in this application. With an increasing number of sexually mature mandarin fish parents, parents with higher abundance of resistant allele loci will be screened, potentially leading to breeding families with further enhanced disease resistance.

[0093] The present application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present application. The descriptions of the embodiments above are only for the purpose of helping to understand the present application and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for screening genome-specific SNP probes, characterized in that, The screening method includes the following steps: (1) obtaining the genome resequencing data of the target biological individual, wherein the resequencing data includes the original sequencing reads Clean (2) Based on the resequencing data, determine candidate dimorphic SNP sites from the reference genome; (3) Sequence generation: Using the dimorphic SNP site as the center, extract candidate sequences containing the site from the reference genome. Probe sequences, denoted as R1-Kmers; (4) Probe uniqueness verification based on clean reads: The clean reads are cut into short sequences to obtain Q1-Kmers; the Q1-Kmers are fully aligned with the R1-Kmers; based on the alignment results, sequences that meet the first condition are selected from the R1-Kmers and denoted as R2-Kmers; the first condition is: for an R1-Kmer, all Q1-Kmers that can be fully aligned with it are from the same clean read; (5) Uniqueness verification of probe coordinates based on Mapped reads: The Mapped reads are truncated into short sequences to obtain Q2-Kmers; the Q2-Kmers are fully aligned with the R2-Kmers; based on the alignment results, sequences that meet the second condition are selected from the R2-Kmers; the second condition is: for an R2-Kmer, its chromosome number G on the reference genome is G. Kmer Chromosome number G from all mapped reads derived from Q2-Kmers that can be perfectly compared with it. ID Same; (6) The R2-Kmers verified by steps (4) and (5) are identified as highly specific SNP probes.

2. The screening method according to claim 1, characterized in that, Step (3) includes: After degenerate base symbol rewriting of the dimorphic SNP sites, R1-Kmers are extracted from the reference genome; the rewriting principle is: M=A / C, R=A / G, W=A / T, S=C / G, Y=C / T, K=G / T.

3. The method according to claim 1, characterized in that, The truncation described in step (3) is to slide and truncate short sequences of 50-70 bp at intervals of 1 base to generate R1-Kmers; the cutting described in steps (4) and (5) is to cut the Clean reads and Mapped reads into short sequences of the same length to generate Q1-Kmers and Q2-Kmers, respectively.

4. The method according to claim 3, characterized in that, The short sequence length described in step (3) is 51 bp.

5. The method according to claim 1, characterized in that, In step (4), the unique identifier of the Clean read from which the Q1-Kmers originate is recorded; the complete alignment means that there are no mismatches, insertions or deletions of any bases between the two short sequences being aligned; the origin of the Clean read is determined based on the unique identifier.

6. The method according to claim 1, characterized in that, After step (5), a region-specific screening step is also included: for the R2-Kmers, sequences located entirely within a single Mapped read are further screened out; the screening is determined by the following conditions: let the starting coordinate of the 5' end of the R2-Kmer on the reference genome be Z, and its length be L; the starting and ending coordinates of a Mapped read used for alignment be X and Y, respectively; then X ≤ Z and Y ≥ Z + L must be satisfied.

7. The method according to claim 1, characterized in that, A probe optimization step is also included after step (5) or step (6): For multiple R2-Kmers corresponding to the same SNP site, the R2-Kmer whose SNP site coordinates (K) are closest to the center of the R2-Kmer sequence is selected as the representative probe for that SNP site.

8. A high-density SNP probe map of the mandarin fish genome, characterized in that, The SNP probes in the SNP probe map are obtained by any one of the screening methods in claims 1-7.

9. A method for constructing a disease-resistant breeding chip for mandarin fish, characterized in that, Includes the following steps: (a) Obtain survival and death samples from full-sibling families of mandarin fish challenged with the ISKNV virus; (b) Genotyping of the surviving and dead samples was performed using the above-mentioned high-density SNP probe map of the mandarin fish genome; (c) Compare and analyze the genotype data of the surviving and dead samples to screen out SNP loci associated with disease resistance traits; (d) Based on the screened associated SNP sites, probes are synthesized and prepared into breeding chips.

10. A disease-resistant breeding chip for mandarin fish, characterized in that, Obtained by the construction method described in claim 9.