A heijin 40k liquid chip based on single nucleotide polymorphism and application thereof

CN122104932APending Publication Date: 2026-05-29XINJIANG ACAD OF ANIMAL SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XINJIANG ACAD OF ANIMAL SCI
Filing Date
2026-02-05
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing sheep microchips have low coverage of specific functional sites for Turpan black sheep, resulting in low breeding efficiency, difficulty in ensuring breed purity, and inability to meet the needs of large-scale breeding.

Method used

A 40K liquid-phase chip for breeding Turpan black sheep was developed. By constructing a 40K SNP locus set for genotyping of Turpan black sheep, probe combinations were designed and probe kits were prepared. Genotyping was performed using the NGP liquid-phase chip system, and combined with bioinformatics analysis, genomic selection and kinship identification were achieved.

Benefits of technology

This has improved the precision and efficiency of breeding, shortened the breeding cycle, protected the unique genetic resources of Turpan black sheep, and enhanced the accuracy and economic value of breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present application relates to the field of genetic molecular breeding, and discloses a black sheep 40K liquid chip based on single nucleotide polymorphism and application thereof. The preparation method of the chip is as follows: phenotype data of Turpan black sheep is obtained, DNA of collected blood is extracted and sequenced; after quality control of original data obtained by sequencing of the DNA library, the original data is compared with a sheep reference genome, sequencing is performed after comparison, and BAM data is obtained; variation detection is performed on BAM data of all samples, all SNP sites are obtained and quality controlled; after quality control of the SNP data, chip background SNP site screening and breed-specific SNP site screening are performed; the selected chip background SNP sites and breed-specific SNP sites are combined, and a probe is designed to prepare a liquid chip. The chip can maximize the use of upstream and downstream SNP marker information of target sites, thereby improving breeding accuracy and efficiency and promoting the directional improvement of production performance of Turpan black sheep.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene molecular breeding, specifically to the Black Sheep 40K liquid phase chip based on single nucleotide polymorphism and its application, and more specifically to the Black Sheep 40K liquid phase chip. Background Technology

[0002] The rapid development of molecular breeding technology has driven the widespread application of genomics methods in animal genetic improvement, with genotype analysis becoming the core of improving breeding efficiency. Single nucleotide polymorphisms (SNPs), as the most widely distributed and highly stable genetic markers, have shown great potential in genome-wide association studies (GWAS) and genomic selection (GS), becoming a core tool in animal genetics research. In terms of microarray technology, liquid-phase microarrays, with their advantages of flexible customization, low cost, and high throughput, are gradually replacing the significantly limited solid-phase microarrays, becoming the mainstream technology in livestock and poultry breeding. Specialized liquid-phase microarrays for cattle and fine-wool sheep have been successfully developed in China, validating the maturity of this technology.

[0003] Turpan Black Sheep is a superior local sheep breed in my country and a national geographical indication product. It possesses excellent characteristics such as adaptability to extreme climates, tolerance to roughage, rapid growth, and tender meat. However, its breeding still mainly relies on traditional methods, which suffer from problems such as insufficient analysis of genetic background, unclear loci related to superior traits, and lack of precise tools for kinship identification. This leads to low breeding efficiency, difficulty in ensuring breed purity, and restricts the full realization of the breed's advantages.

[0004] Existing sheep microchips are mostly general-purpose, with low coverage of specific functional loci for Turpan Black sheep. The high cost and fixed-site characteristics of solid-phase microchips also fail to meet the needs of large-scale breeding. Therefore, designing a dedicated 40K liquid-phase breeding microchip for Turpan Black sheep is essential. This chip integrates significant effect SNP loci related to key traits such as growth, meat quality, and stress resistance in Black sheep. It can precisely support genetic breeding research, efficiently advance breeding improvement, achieve accurate identification of kinship, shorten the breeding cycle, improve selection accuracy, and protect its unique genetic resources, providing core technological support for the high-quality development of the Turpan Black sheep industry. Summary of the Invention

[0005] This invention develops a 40K liquid phase chip for breeding Turpan black sheep, and provides methods and strategies for developing chip target sites. Therefore, it provides a 40K liquid phase chip for genotyping in Turpan black sheep breeding and its applications.

[0006] The technical solution provided by this invention is as follows:

[0007] This invention first provides a method for constructing a 40K SNP locus set for genotyping in Turpan black sheep breeding, which includes the following steps:

[0008] Step 1: Obtain phenotypic data of Turpan black sheep and extract DNA from the collected blood;

[0009] Step 2: Construct DNA library and sequence. After quality control of the raw data, compare it with sheep reference genome. After the comparison is completed, sort the data to obtain BAM data.

[0010] Step 3: Perform variant detection on the BAM data of all samples to obtain all SNP sites;

[0011] Step 4: Use the obtained SNP data to screen for background SNP sites on the chip and for variety-specific SNP sites.

[0012] Specifically, in step one, venous blood is collected from Turpan black sheep to extract genomic DNA from the blood.

[0013] In step two, 0.1-1 μg of genomic DNA is taken and digested for the first time with the essential Mse I restriction endonuclease. After digestion, Solexa P1 and P2 adapters containing 6 bp barcodes are added to both ends of the fragment. Then, a second digestion is performed with a combination of second enzymes selected from Nla III, Hae II, Hae III, Msp I or EcoRI to adjust the number of tags. After that, the sequence containing the double adapters is amplified by PCR. The DNA fragments of the sample are separated by barcode and then pooled. The target region fragment is recovered by electrophoresis, and the PCR product is purified to obtain the GBS library.

[0014] Furthermore, during GBS library quality control, the library is first initially quantified and diluted to 1.5 ng / ul, then the library insert size is measured. Once the insert size meets expectations, the effective concentration of the library is accurately quantified by qRT-PCR (ensuring it is higher than 2 nmol / L) to ensure that the library quality is suitable for subsequent sequencing.

[0015] Libraries that pass the library inspection are subjected to paired-end 150 sequencing.

[0016] In step three, after preprocessing the raw data obtained from sequencing, the FastP software is used for data filtering.

[0017] The specific filtering rules are as follows: First, remove read pairs containing adapter sequences; second, for single-end sequencing reads, if the content of "N" bases exceeds 10% of the read length, or the number of low-quality bases (quality value Q≤5) exceeds 50% of the read length, then delete the corresponding paired reads. Through the above filtering steps, high-quality clean data is finally obtained. The filtered data is compared with the sheep reference genome (e.g., the sheep reference genome data published on the website https: / / www.ncbi.nlm.nih.gov / datasets / genome / GCF_016772045.2 / ) using BWA software to obtain high-quality SNP loci on autosomes.

[0018] More specifically, the population SNP detection was performed using samtools and bcftools software to compare the aligned data (specifically, after aligning and sorting the original data with the reference genome and generating a BAM format file, the detection and filtering analysis of SNPs were performed using samtools and bcftools software). After filtering (the parameters for filtering the obtained original SNP data using bcftools software were: dp:3; miss:0.1; maf:0.05), 455,143 high-quality autosomal SNP loci were obtained.

[0019] In step four, the simplified genome sequencing data of the Turpan black sheep breed is GCF_016772045.2 (sheep reference genome number on NCBI).

[0020] The principles for selecting target SNPs are: ① uniform distribution on each chromosome, with a denser distribution at both ends of each chromosome; ② polymorphism considerations: minor allele frequency (MAF) > 0.05 in Turpan black sheep; ③ heterozygosity of the loci ≤ 20%; ④ GWAS analysis of SNP loci and phenotypic data to obtain significant loci associated with traits such as body weight, body height, body length, chest width, chest depth, chest circumference, cannon bone circumference, tail width, and tail length in Turpan black sheep; ⑤ comparison with the QTLdb database to select SNP markers in the QTL regions of economic traits; ⑥ acquisition of known SNP loci containing genes associated with sheep body weight and body size.

[0021] In step five: the selected background SNP sites on the chip are merged with the breed-specific SNP sites to obtain the whole genome SNP site combination of Turpan black sheep, i.e., the SNP site set.

[0022] The present invention also provides a probe combination for Turpan black sheep genotyping, which is composed of a probe combination designed from the 40K SNP site set for Turpan black sheep genotyping obtained by the method described above.

[0023] The probe kit was further prepared into a genotyping kit for Turpan black sheep.

[0024] The present invention further provides a 40K liquid phase chip for black sheep genotyping, which is constructed using a probe combination with a probe length of 80-120 bp. Specifically, the chip structure adopts an NGP liquid phase chip using a solution-microbead system.

[0025] Specifically, the probe labeled with biotin at both ends is first incubated with streptavidin magnetic beads at a 1:1 molar ratio at room temperature to form a probe-magnetic bead complex; then the diluted complex suspension is injected into a standard PCR plate and freeze-dried under vacuum to generate a reversibly dried "probe-magnetic bead" film at the bottom of the well.

[0026] Preferably, the chip is in the form of a kit, comprising individually packaged probe mixtures and hybridization capture reagents.

[0027] This invention also provides the application of the probe combination or probe kit, or the 40K liquid phase chip for genotyping of Turpan black sheep breeding, in genotyping, genetic improvement, genetic diversity analysis, genomic selection breeding, and kinship identification of Turpan black sheep.

[0028] This invention also provides a method for genotyping genes in Turpan black sheep breeding, which uses the aforementioned 40K liquid phase chip for Turpan black sheep genotyping. Specifically, it includes the following steps:

[0029] I. Steps of Genotyping

[0030] (I) Experimental Operation Stage

[0031] S1. Sample Detection: After extracting genomic DNA, the concentration and purity of the DNA were detected by agarose gel electrophoresis and Qubit 4.0 fluorescence quantitative PCR instrument. Qualified samples were screened for subsequent experimental procedures.

[0032] S2. Library preparation and hybridization capture:

[0033] DNA fragmentation: Genomic DNA is randomly fragmented to an average length of 180-320 bp using enzyme digestion reagents;

[0034] End repair and A-tailing, adapter: End repair of fragmented DNA is performed, a single adenine A is added to the 3′ end, and a sequencing adapter containing an index is ligated;

[0035] Pre-PCR amplification: The ligation product is subjected to limited-cycle PCR amplification to obtain a pre-library library;

[0036] Hybrid capture: The pre-library was incubated with the 40K liquid phase chip at 48°C for 16 hours. The target fragments targeted by the 40K liquid phase chip were captured by magnetic beads coated with streptavidin and washed to remove non-specific binding products.

[0037] Post-capture PCR amplification: The enriched 40K liquid-phase chip targeted fragment product was subjected to 11 cycles of PCR amplification to obtain a sufficient sequencing library.

[0038] S3. Library quality control: The NGP Qubit dsDNA HS Assay Kit was used to accurately quantify the library on the Qubit 4.0 fluorometer to ensure that the library meets the requirements for use on the 40K liquid chromatography-mass sequencing platform;

[0039] S4. Sequencing: After pooling the qualified library according to the effective concentration and target data volume, perform paired-end sequencing on PE150. The sequence information of the target site on the 40K liquid chip is obtained by combining DNA nanosphere rolling circle amplification with joint probe anchoring ligation sequencing technology.

[0040] (II) Bioinformatics Analysis Stage

[0041] S1. Obtain raw data: Convert the raw image data from the sequencing platform into FASTQ format raw data (including sequence information and sequencing quality values).

[0042] S2. Data quality control: Use FASTP software to filter low-quality reads and retain qualified clean reads to ensure the accuracy of 40K liquid crystal chip targeted site analysis;

[0043] S3. Alignment analysis: Align clean reads to a reference genome (e.g., downloaded from a public database) using BWA software;

[0044] S4. Target site variant detection: Based on the alignment results, GATK HaplotypeCaller is used to perform genotyping on the target SNP sites targeted by the 40K liquid phase chip, generating VCF format variant files;

[0045] S5. Variance Annotation Analysis: ANNOVAR software was used to perform functional annotation of targeted variant sites on the 40K liquid-phase microarray to clarify their genomic location and amino acid level effects;

[0046] II. Criteria for Genotyping

[0047] (a) Sample quality inspection standards

[0048] DNA samples must meet the criteria of being "intact, without significant degradation, and with a concentration ≥50 ng / μL";

[0049] (II) Data Quality Control Standards

[0050] The following low-quality reads were filtered out, and only qualified data were retained for subsequent analysis of the target sites on the 40K liquid crystal chip:

[0051] S1. Reads containing sequencing adapter contamination;

[0052] S2. Reads with an N base ratio exceeding 10%;

[0053] S3. Reads in which more than 50% of the bases have a quality value ≤5;

[0054] (III) Document Quality Inspection Standards

[0055] The library concentration must meet the sequencing requirements of the 40K liquid-phase chip sequencing platform (DNBSEQ-T7) to be compatible with subsequent sequencing procedures;

[0056] (iv) Probe and Sequencing Standards

[0057] S1. Hybridization capture probe: a probe for 40K liquid phase chip, 80-120 bp in length, with GC content of 30%-70%;

[0058] S2. Sequencing Platform and Mode: The preferred platform is the BGI Genomics DNBSEQ-T7, with paired-end sequencing (PE150), to ensure efficient and accurate acquisition of target site sequence information from the 40K liquid-phase chip.

[0059] The advantages of the Turpan Black Sheep 40K liquid phase chip provided by this invention are:

[0060] 1. This SNP chip has a high content of site polymorphism information;

[0061] 2. The SNP chip is uniformly distributed across the genome;

[0062] 3. This chip can be used for genotyping, kinship identification, and screening of relevant economic trait variation sites in Turpan black sheep populations. Using the chip and analysis strategy designed in this invention, upstream and downstream SNP marker information at the target site can be utilized to the maximum extent, thereby improving breeding accuracy and efficiency, promoting targeted improvement of Turpan black sheep production performance, and increasing its economic value. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0064] Figure 1 This refers to the consistency rate of repeated sample detection in the chip evaluation process used to prepare the Turpan black sheep 40K liquid phase chip in Example 2.

[0065] Figure 2 This is a statistical analysis of the number and length distribution of 40,412 SNP loci on the chromosome used in Example 2 to prepare the 40K liquid phase chip for Turpan black sheep.

[0066] Figure 3 This is a statistical analysis of the density distribution of 40,412 SNP sites on chromosomes involved in the preparation of the 40K liquid phase chip for Turpan black sheep in Example 2.

[0067] Figure 4 This is a statistical analysis of the minimum allele frequency (MAF) distribution of 40,412 SNP loci on chromosomes involved in the preparation of the 40K liquid phase chip for Turpan black sheep in Example 2.

[0068] Figure 5 Example 2 is a statistical analysis of the variation types of 40,412 SNPs involved in the preparation of the Turpan Black Sheep 40K liquid phase chip.

[0069] Figure 6 The 3DPCA plot of 485 samples from the Turpan black sheep population structure analysis based on the black sheep 40K chip in Example 3;

[0070] Figure 7 This is a line graph of the K values ​​of 485 samples from the Turpan black sheep population structure analysis based on the black sheep 40K chip in Example 3.

[0071] Figure 8 This is a clustering diagram of the K values ​​of 485 samples from the Turpan black sheep population structure analysis based on the black sheep 40K chip in Example 3.

[0072] Figure 9 This is a heatmap of kinship among 485 samples from the Turpan black sheep population structure analysis based on the black sheep 40K chip in Example 3. Detailed Implementation

[0073] The following embodiments are intended to further illustrate the present invention, but not to limit it. Modifications or substitutions made to the embodiments without departing from the spirit and substance of the invention are all within the scope of the invention. Technical means not specifically specified in the embodiments are all conventional means well known to those skilled in the art. The technical solutions of some embodiments will be clearly and completely described below with reference to the accompanying drawings.

[0074] Example 1: GBS sequencing of Turpan black sheep

[0075] I. DNA Extraction and Quality Inspection

[0076] 129 Turpan black sheep of different breeds were selected from large-scale black sheep farms in the Turpan area. 5 ml of venous blood was collected from each sheep, and genomic DNA was extracted using the magnetic bead method universal genomic DNA extraction kit (catalog number: DP705) from the blood. The specific steps are as follows:

[0077] 1. Prepare in advance:

[0078] (1). Reagent pretreatment: Add anhydrous ethanol to 45 ml of buffer GDZ and 20 ml of wash buffer PWD in the DP705-01 kit respectively; if there is a precipitate in the GHL lysis buffer, it needs to be redissolved in a 37°C water bath and shaken before use.

[0079] (2) Premixed reagents: Mix the lysis buffer GHL in the DP705-01 kit with 1 ml of proteinase K in a certain proportion before use to improve the efficiency of operation.

[0080] (3) Sample thawing: Place the frozen blood samples of Turpan black sheep at room temperature until they are completely thawed, avoiding repeated freeze-thaw cycles.

[0081] 2. Take 250 μl of thawed blood sample from Turpan black sheep into a 2 ml centrifuge tube.

[0082] 3. Add 20 μl of pre-mixed proteinase K solution and 300 μl of lysis buffer GHL, vortex to mix, and lyse at 75 °C for 15 min, inverting the container 3 times during the lysing process, 3-5 times each time.

[0083] 4. Add 300 μl of isopropanol and shake to mix for 10 seconds.

[0084] 5. Add 15 μl of magnetic bead suspension GH. Before use, shake to mix to ensure that the magnetic beads are completely resuspended. Then shake to mix for 1 min, and let stand for a total of 9 min, shaking to mix for 1 min every 3 min.

[0085] 6. Place the centrifuge tube on the OSE-MF-01 magnetic rack and let it stand for 30 seconds until the magnetic beads are completely attracted. Then carefully remove the liquid.

[0086] 7. Add 900 μl of buffer GDZ and vortex to mix for 2 min; place the centrifuge tube on a magnetic rack and let it stand for 30 sec until the magnetic beads are completely adsorbed, then carefully remove the liquid.

[0087] 8. Add 500 μl of buffer GDZ and vortex to mix for 2 min; place the centrifuge tube on a magnetic rack and let it stand for 30 sec until the magnetic beads are completely adsorbed, then carefully aspirate the liquid.

[0088] 9. Remove the centrifuge tube from the magnetic rack, add 900 μl of PWD washing solution, and shake to mix for 2 min; place the centrifuge tube on the magnetic rack and let it stand for 30 sec until the magnetic beads are completely attracted, then carefully remove the liquid.

[0089] 10. Remove the centrifuge tube from the magnetic rack, add 300 μl of PWD washing solution, and shake to mix for 2 min; place the centrifuge tube on the magnetic rack and let it stand for 30 sec until the magnetic beads are completely attracted, then carefully remove the liquid.

[0090] 11. Leave the centrifuge tube on the magnetic rack and let it air dry at room temperature for 10-15 minutes (ensure that the ethanol has evaporated completely to avoid inhibiting subsequent enzyme reactions; do not dry for too long to prevent difficulty in eluting DNA).

[0091] 12. Remove the centrifuge tube from the magnetic rack, add 50-100 μl of elution buffer TB, vortex to mix, and incubate at 56°C for 10 min, inverting the tube 3 times during the incubation period, 3-5 times each time.

[0092] 13. Place the centrifuge tube, after shaking and mixing, on a magnetic rack and let it stand for 2 minutes until the magnetic beads are completely adsorbed. Then, carefully transfer the DNA solution to a new centrifuge tube to complete the DNA extraction.

[0093] 14. Integrity detection: Agarose gel electrophoresis was used to detect the fragment size and integrity of DNA in Turpan black sheep blood samples. The DNA fragment size was affected by factors such as sample storage time and shearing force during operation.

[0094] 15. Purity testing: DNA purity was determined using a UV spectrophotometer, expressed as OD. 260 nm / OD 280 A nm ratio in the range of 1.7 to 1.9 is considered acceptable.

[0095] 16. Concentration Detection: DNA concentration was determined using a UV spectrophotometer, OD... 260 When the nm value is 1, it is equivalent to approximately 50 μg / ml of double-stranded DNA.

[0096] 17. Sample preservation: Qualified blood DNA samples from Turpan black sheep will then undergo further library construction and sequencing.

[0097] II. DNA Library Construction

[0098] Library construction employed the GBS workflow adapted by Novogene. 0.1–1 μg of quantitatively qualified genomic DNA was used for the first digestion with the mandatory Mse I restriction endonuclease. After digestion, Solexa P1 and P2 adapters containing 6 bp barcodes were added to both ends of the fragments. A second digestion was then performed using a combination of second enzymes selected from Nla III, Hae II, Hae III, Msp I, or EcoRI to adjust the number of tags. PCR amplification of sequences containing dual adapters followed. DNA fragments from 129 samples were separated by barcode, pooled, and the target region fragments were recovered by electrophoresis. The PCR products were purified using AMPure XP beads to obtain the GBS library. For library quality control, the library was initially quantified using a Qubit 2.0 Fluorometer and diluted to 1.5 ng / µl. The insert size was then detected using an Agilent 2100 bioanalyzer. Once the insert size met expectations, the effective concentration of the library was accurately quantified by qRT-PCR (ensuring a concentration higher than 2 nmol / L) to guarantee the library quality for subsequent sequencing. Libraries that pass the library inspection are used with Illumina Novaseq. TM The sequencing platform was used to perform paired-end 150 sequencing.

[0099] III. Sequencing and Variance Detection

[0100] After preprocessing the raw data obtained from PE150 sequencing, data filtering was performed using FASTP software. The specific filtering rules were as follows: First, read pairs containing adapter sequences were removed; second, for single-end sequencing reads, if the "N" base content exceeded 10% of the read length, or the number of low-quality bases (quality value Q≤5) exceeded 50% of the read length, the corresponding paired reads were deleted. Through these filtering steps, high-quality clean data was finally obtained. The filtered data was compared with the sheep reference genome (https: / / www.ncbi.nlm.nih.gov / datasets / genome / GCF_016772045.2 / ) using BWA software. Population SNP detection was performed on the compared data using samtools and bcftools software. After filtering (parameters: dp:3; miss:0.1; maf:0.05), 455,143 high-quality autosomal SNPs were obtained for microarray design.

[0101] Example 2: Turpan Black Sheep Chip Design and Probe Preparation

[0102] I. Chip Site Design

[0103] Background site screening: After removing SNP sites on non-standard chromosomes and sex chromosomes from the obtained 455,143 SNP markers, further filtering was performed. Initial screening was conducted using bcftools and Plink software, with the following filtering conditions: ① sequencing depth DP ≥ 4X, ② site deletion rate (miss) ≤ 0.1, ③ minimum allele frequency (MAF) > 0.05, and ④ heterozygosity (het) < 0.2. This resulted in 100,000 high-quality SNPs, which were then used for probe design and sliding window screening to identify background SNP sites.

[0104] Functional site screening: ① Genome-wide association analysis (GWAS) was performed on the corrected phenotypes and GBS sequencing data of 129 Turpan black sheep individuals. A linear mixture model (LMM) using GEMMA software was employed to perform association analysis on nine traits: body weight, body height, body length, chest width, chest depth, chest circumference, cannon bone circumference, tail width, and tail length. The significance p-values ​​of each SNP were obtained using the Wald test. The significance threshold for association was set to −log[…]. 10 (P)>5 (i.e., P<1×10) -5 SNPs exceeding this threshold are considered candidate loci significantly associated with the trait. ② QTL data for sheep were downloaded from the AnimalQTLdb database. QTLs related to sheep weight and body size were compiled and compared with the quality-controlled SNP set of this project. The intersection was used to obtain the QTLs related to the weight and body size traits of Turpan black sheep. ③ Genes related to sheep weight and body size were retrieved from the Web of Science and PubMed databases. After screening, SNP loci for Turpan black sheep were extracted from these genes. Through the above steps, the functional loci SNP set for Turpan black sheep was obtained by integrating and removing duplicates. Table 1 shows the number, span, and total span of SNPs on each chromosome, and Table 2 shows some representative SNP loci.

[0105] Table 1. Statistics related to chip location

[0106]

[0107] Table 2. Representative sites on the chip

[0108]

[0109]

[0110] II. Chip Probe Fabrication

[0111] The probe design of the NGP liquid-phase chip is based on the principle of thermodynamic stability. It comprehensively considers the complexity of the target species' genome, the location of the target site, and the GC content in the vicinity. The design ensures that the number of repeats in the genome that strictly match the probe sequence does not exceed 200, the GC content is controlled between 30% and 70%, and the number of SNPs allowed does not exceed 5, ensuring 100% capture of common regions. For highly complex regions, the effective coverage is improved by adjusting the probe position and employing a multi-layer "shingled" probe placement scheme. The probe design length ranges from 80 to 120 bp, with an average length of approximately 100 bp, and probes are synthesized specifically for each SNP site.

[0112] The probe preparation process comprises two main steps: template synthesis and probe preparation. First, single-stranded DNA (oligo) is obtained through organic synthesis. After quality control (QC), pooling, and in vitro amplification, a DNA template (oligopool) is formed. Only after passing quality inspection can it be used for subsequent probe preparation. During preparation, biotin-labeled dNTPs / NTPs and nucleotide analogs (such as LNA and PNA) are used as synthetic raw materials to ultimately produce double-stranded DNA probes. Nucleotide analogs can increase the Tm value of probe-library binding, enhancing binding stability; while the biotin label carried by the probe enables specific binding to streptavidin-coated magnetic beads in subsequent experiments.

[0113] III. Chip Performance Evaluation Testing

[0114] NGP hybridization capture sequencing technology primarily utilizes probe hybridization to capture target regions, thereby detecting genetic and genome-related variations. Sample DNA is randomly fragmented, followed by WGS library construction, target region capture based on probe hybridization, adsorption using streptomycin beads, elution and enrichment of non-target molecules, PCR amplification, and sequencing to obtain sequencing information for the target region. In this project, 10 samples were selected, and two were randomly chosen as replicates. Microarray-based targeted capture sequencing was used to analyze the site detection rate and the genotyping concordance rate between the two replicates of the same sample. Table 2 summarizes the target region indicators. Figure 1 To improve the consistency rate of repeated sample testing, Figure 2 Chromosomal distribution at loci Figure 3 For chromosome density distribution, Figure 4 Site MAF distribution Figure 5 The variation type distribution is shown. The sample detection process mainly includes the following steps:

[0115] (I) Database Construction and Capture

[0116] (1) The DNA sample is subjected to sonication or enzyme digestion to break the DNA, end repair and 3' end addition of "A", adapter ligation and purification, and Pre-PCR library amplification to obtain the library required for hybridization capture.

[0117] (2) The capture method was performed according to “NGP Hyb & Wash Kit v2.0”. After hybridization of library and probe, binding of probe to magnetic beads, washing of non-specifically bound library, PCR amplification after capture, library quantification and quality control, sequencing was performed.

[0118] (II) Document Quality Inspection

[0119] Take 1 μL of the library and use the NGP Qubit dsDNA HS Assay Kit reagents to detect the library concentration on a Qubit 4.0 Fluorometer.

[0120] (III) Sequencing and Data Analysis

[0121] Library concentration was determined using the NGP Qubit dsDNA HS Assay Kit on a Qubit 4.0 Fluorometer, followed by sequencing using a high-throughput sequencing platform. After obtaining the raw data, quality control and alignment with a reference genome were performed. The alignment rate, coverage, and uniformity were statistically analyzed (Table 3) to assess whether the library construction and sequencing met the chip design standards.

[0122] Table 3. Summary of Indicators for the Target Area

[0123]

[0124] Example 3: Population Structure Analysis of Turpan Black Sheep Based on Black Sheep 40K Chip

[0125] The following analysis was performed on the 485 black sheep samples in this example using the black sheep 40K SNP chip designed in Example 2:

[0126] 1. Principal Component Analysis

[0127] Principal component analysis (PCA) was performed on the SNP dataset of all samples using PLINK software to calculate the variance explained by each principal component and the corresponding principal component score matrix. The three principal components (PC1, PC2, and PC3) with the strongest variance explained were selected for 3D PCA visualization, visually presenting the differences in genetic structure among the samples. The relevant analysis results are as follows: Figure 6 As shown.

[0128] 2. Group Structure Analysis

[0129] FastStructure software was used to infer the genetic structure of the population, assuming the number of clusters K ranged from 1 to 10. The optimal number of clusters was determined by the K value with the lowest cross-validation error (CV error). This was used to analyze the genetic stratification characteristics of the population, and a line graph of the K values ​​was plotted using R software. Simultaneously, R packages were used to plot population clustering diagrams corresponding to different K values. The results are shown below. Figure 7 , Figure 8 As shown.

[0130] 3. Kinship Analysis

[0131] Kinship analysis was performed using Tassel 5 software, resulting in a pairwise kinship matrix for 485 samples. The matrix was then visualized using the pheatmap package in R, creating a kinship clustering heatmap to visually reflect the degree of kinship among samples within the group. The results are as follows: Figure 9 As shown in the heatmap, most areas are light-colored, indicating that the overall genetic diversity of the population is high and inbreeding is less common, which can provide a basis for parent selection and avoiding inbreeding in breeding.

[0132] 4. Phylogenetic tree construction

[0133] Based on the genotyping data of 485 filtered samples, a phylogenetic tree was constructed using MEGA 12 software, and cluster analysis was performed using the neighbor-joining method. This method constructs the topology based on a distance matrix (p-distance). The p-distance model can directly quantify the degree of nucleotide differences between samples; simultaneously, 1000 bootstrap replicates were set to assess the confidence level of each branch node in the phylogenetic tree.

Claims

1. A method for constructing a 40K SNP locus set for black sheep genotyping, characterized in that, Includes the following steps: Step 1: Obtain phenotypic data of Turpan black sheep and extract DNA from the collected blood; Step 2: Construct DNA library and sequence. After quality control of the raw data, compare it with sheep reference genome. After the comparison is completed, sort the data to obtain BAM data. Step 3: Perform variant detection on the BAM data of all samples to obtain all SNP sites; Step 4: Use the obtained SNP data to screen for background SNP sites on the chip and for variety-specific SNP sites.

2. The method as described in claim 1, characterized in that, In step one, venous blood was collected from Turpan black sheep to extract genomic DNA from the blood. Preferably, in step two, 0.1-1 μg of genomic DNA is taken and digested for the first time with the essential Mse I restriction endonuclease. After digestion, Solexa P1 and P2 adapters containing 6 bp barcodes are added to both ends of the fragment. Then, a second digestion is performed with a combination of second enzymes selected from Nla III, Hae II, Hae III, Msp I or EcoRI to adjust the number of tags. After that, the sequence containing the double adapters is amplified by PCR. The DNA fragments of the sample are separated by barcode and then pooled. The target region fragment is recovered by electrophoresis, and the PCR product is purified to obtain the GBS library. Furthermore, during GBS library quality control, the library is first initially quantified and diluted to 1.5 ng / ul, then the library insert size is measured. Once the insert size meets expectations, the effective concentration of the library is accurately quantified by qRT-PCR (ensuring it is higher than 2 nmol / L) to ensure that the library quality is suitable for subsequent sequencing. Libraries that pass the library inspection are subjected to paired-end 150 sequencing.

3. The method as described in claim 1, characterized in that, In step three, after preprocessing the raw data obtained from sequencing, the FastP software is used for data filtering. The specific filtering rules are as follows: First, remove read pairs containing adapter sequences; second, for single-end sequencing reads, if the "N" base content exceeds 10% of the read length, or the number of low-quality bases (quality value Q≤5) exceeds 50% of the read length, then delete the corresponding paired reads. Through the above filtering steps, high-quality clean data is finally obtained. The filtered data is then compared with the sheep reference genome (e.g., the sheep reference genome data published on the website https: / / www.ncbi.nlm.nih.gov / datasets / genome / GCF_016772045.2 / ) using BWA software to obtain high-quality SNP loci on autosomes. Preferably, the population SNP detection is performed by comparing the aligned data with samtools and bcftools software (specifically, after aligning and sorting the original data with the reference genome and generating a BAM format file, the detection and filtering analysis of SNPs are performed using samtools and bcftools software). After filtering (the parameters for filtering the obtained original SNP data using bcftools software are: dp:3; miss:0.1; maf:0.05), 455,143 high-quality autosomal SNP loci are obtained.

4. The method as described in claim 1, characterized in that, In step four, the simplified genome sequencing data of the Turpan black sheep breed is GCF_016772045.2 (sheep reference genome number on NCBI). The principles for selecting target SNPs are: ① uniform distribution on each chromosome, with a denser distribution at both ends of each chromosome; ② polymorphism considerations: minor allele frequency (MAF) > 0.05 in Turpan black sheep; ③ heterozygosity of the loci ≤ 20%; ④ GWAS analysis of SNP loci and phenotypic data to obtain significant loci associated with traits such as body weight, body height, body length, chest width, chest depth, chest circumference, cannon bone circumference, tail width, and tail length in Turpan black sheep; ⑤ comparison with the QTLdb database to select SNP markers in the QTL regions of economic traits; ⑥ acquisition of known SNP loci containing genes associated with sheep body weight and body size.

5. The method as described in claim 1, characterized in that, In step five: the selected background SNP sites on the chip are merged with the breed-specific SNP sites to obtain the whole genome SNP site combination of Turpan black sheep, i.e., the SNP site set.

6. A probe combination for genotyping black sheep, characterized in that, It consists of a probe combination designed from the 40K SNP locus set for Turpan black sheep genotyping obtained by the method as described in any one of claims 1 to 5; Further, a probe kit for black sheep genotyping was prepared.

7. A 40K liquid phase chip for genotyping black sheep, characterized in that, It is constructed using the probe combination as described in claim 6, wherein the probe length is 80-120 bp. Specifically, the chip structure adopts an NGP liquid phase chip using a solution-microbead system.

8. The 40K liquid phase chip for black sheep genotyping as described in claim 7, characterized in that, The probe labeled with biotin at both ends was first incubated with streptavidin magnetic beads at a 1:1 molar ratio at room temperature to form a probe-magnetic bead complex. Then, the diluted complex suspension was injected into a standard PCR plate and freeze-dried under vacuum to generate a reversibly dried "probe-magnetic bead" film at the bottom of the well. Preferably, the chip is in the form of a kit, comprising individually packaged probe mixtures and hybridization capture reagents.

9. The application of the probe combination or probe kit as described in claim 6, or the 40K liquid phase chip for black sheep genotyping as described in claim 7 or 8, in black sheep genotyping, black sheep genetic breeding improvement, black sheep genetic diversity analysis, black sheep genome selection, and black sheep kinship identification.

10. A method for genotyping black sheep genes, characterized in that, Genotyping is performed using the 40K liquid phase chip for black sheep genotyping as described in claim 7 or 8, specifically including the following steps: I. Steps of genotyping: (I) Experimental Operation Stage: S1. Sample testing: After extracting genomic DNA, the concentration and purity of the DNA are tested, and qualified samples are screened to enter the subsequent experimental process; S2. Library preparation and hybridization capture: DNA fragmentation: Genomic DNA is randomly fragmented to an average length of 180-320 bp using enzyme digestion reagents; End repair and A-tailing, adapter: End repair of fragmented DNA is performed, a single adenine A is added to the 3′ end, and a sequencing adapter containing an index is ligated; Pre-PCR amplification: The ligation product is subjected to limited-cycle PCR amplification to obtain a pre-library library; Hybrid capture: The pre-library was incubated with a 40K liquid chip, and the target fragments targeted by the 40K liquid chip were captured by streptavidin-coated magnetic beads. Non-specific binding products were then washed away. Post-capture PCR amplification: The enriched 40K liquid-phase chip targeted fragment products were amplified by PCR to obtain a sufficient sequencing library; S3. Library quality control: Ensure that the library meets the requirements for use with the 40K liquid chromatography-mass spectrometry (LCS) sequencing platform; S4. Sequencing: After pooling the qualified library according to the effective concentration and target data volume, perform paired-end sequencing. Use DNA nanosphere rolling circle amplification combined with joint probe anchoring ligation sequencing technology to obtain the sequence information of the target site on the 40K liquid chip. (II) Bioinformatics Analysis Stage: S1. Obtain raw data: Convert the raw image data from the sequencing platform into FASTQ format raw data (including sequence information and sequencing quality values). S2. Data quality control: Use FASTP software to filter low-quality reads and retain qualified clean reads to ensure the accuracy of 40K liquid crystal chip targeted site analysis; S3. Alignment analysis: Align clean reads to a reference genome (e.g., downloaded from a public database) using BWA software; S4. Target site variant detection: Based on the alignment results, genotyping is performed on the target SNP sites targeted by the 40K liquid phase chip to generate VCF format variant files; S5. Variant Annotation Analysis: Functional annotation of targeted variant sites on the 40K liquid-phase chip was performed to clarify their genomic location and amino acid level effects; II. Criteria for Genotyping: (a) Sample quality control standards: DNA samples must meet the criteria of being "intact, without significant degradation, and with a concentration ≥50 ng / μL"; (II) Data Quality Control Standards: The following low-quality reads were filtered out, and only qualified data were retained for subsequent analysis of the target sites on the 40K liquid crystal chip: S1. Reads containing sequencing adapter contamination; S2. Reads with an N base ratio exceeding 10%; S3. Reads in which more than 50% of the bases have a quality value ≤5; (III) Document Quality Inspection Standards: The library concentration must meet the sequencing requirements of the 40K liquid-phase chip sequencing platform (DNBSEQ-T7) to be compatible with subsequent sequencing procedures; (iv) Probe and sequencing standards: S1. Hybridization capture probe: a probe for 40K liquid phase chip, 80-120 bp in length, with GC content of 30%-70%; S2. Sequencing Platform and Mode: The preferred platform is the BGI Genomics DNBSEQ-T7, with paired-end sequencing (PE150), to ensure efficient and accurate acquisition of target site sequence information from the 40K liquid-phase chip.