Evaluation method for purebred hu sheep marked with SNP molecules
High-throughput sequencing and genetic marker analysis establish a purebred evaluation model for Hu sheep, addressing the decline in breed purity by accurately identifying and protecting Hu sheep resources.
Patent Information
- Application Number
- JP2024014465
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-07
- Filing Date
- 2024-02-01
- Publication Date
- 2025-05-19
AI Technical Summary
The decline in the number of purebred Hu sheep due to crossbreeding with foreign breeds has led to a loss of breed characteristics, genetic degeneration, and disruption of family lines, necessitating a method to evaluate and protect the Hu sheep breed resources.
A method involving high-throughput sequencing to identify genetic mutations, using SNPs and INDELs, and creating a scoring model to evaluate the purity of Hu sheep by selecting genetic marker sites and applying multivariate linear regression analysis to establish a purebred evaluation model.
This method accurately identifies purebred Hu sheep, providing a scientific basis for their conservation and development, with a detection rate of genetic variations exceeding 90% and a model accuracy of 73.7% in predicting purebred flocks.
Smart Images

Figure 2025077938000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention is related to the technology in the fields of animal husbandry and molecular biology, and is a method for evaluating the pure breed of sheep. In particular, the present invention relates to a method for evaluating purebred Hu sheep labeled with SPN molecules. [Background technology]
[0002] Hu sheep is a famous local breed of white lamb skin sheep in China. In recent years, With the change in the market, the focus of the Hu sheep industry has shifted from sheepskin to mutton production, and some Hu sheep farming Breeders have gradually introduced other meat breeds to cross with Hu sheep, and to some extent, have improved the meat characteristics of Hu sheep. Although this has improved the breeding capacity, the purebred Hu sheep in the normal farming areas are being invaded by foreign breeds, and Many Hu sheep are experiencing the phenomenon of gene crowding, decline in breed characteristics, and degeneration of breed quality, and the number of purebred Hu sheep is declining. This has resulted in a rapid decline in the number of fish, disruption of family lines and bloodlines, and the loss of high-quality stocks.
[0003] In order to protect the Hu sheep breed resources, the breed protection effect of the Hu sheep breed protection station was monitored. The link is urgently needed to provide a foundation for the development and utilization of genetic resources of Hu sheep.
[0004] Single nucleotide polymorphisms (SPNs) and INDELs are the most common and most widely distributed of the gene groups. There are two types of genetic variations in gene groups. It is caused by mutations and only involves single nucleotide polymorphisms. However, the occurrence of INDELs is The length of the mutation varies, mainly related to the sequence environment and replication errors it encounters. The incidence rate varies among different tumors and is generally It is related to the size of the group. Using complete genome sequencing, SNPs and INDELs were considered. The current method is the most accurate and genome-wide detection method available. Using the gene subtyping information of P and INDEL groups, the similarity of genetic information between individuals was associated with sex, the genetic structure of the group, the selection signals received by special groups, and the complete genome You can get analysis etc. Summary of the Invention
[0005] Problems to be solved by the invention: In order to solve the above problems, the present invention provides a method for evaluating purebred Hu sheep tagged with SNPs. Using high-throughput data at the herd level, we have identified Hu sheep, Xiaotong Han sheep, and Dorper sheep. Detection of genetic variation across the entire genome in sheep and in a cross between Dorper and Hu sheep By using bioinformatics to identify the genetic characteristics of Hu sheep, the pure breed of Hu sheep was established. To establish an evaluation method to identify the breed of sheep and provide a scientific basis for the protection, development and utilization of the breed resources of Hu sheep. The objective is to provide Technical solutions to the problem:
[0006] The present invention relates to 1) Candidate genes with specific colony frequencies in Hu sheep colonies identified by high-throughput sequencing By identifying genetic mutations and comparing the frequency of genetic mutations between Hu sheep and other populations, We will develop a scoring method and perform sequencing to identify the genetic variation present in the sheep flock. A purebred evaluation model was created using various indicators, and genetic marker candidates were further evaluated based on the model results. The set is divided into 30 sets based on the site primer design and the site genome linkage situation. Selecting genetic marker sites for identification of Hu sheep pure breed; 2) The 30 genetic marker sites obtained from step 1) were subtyped using Sequenom SNP subtyping. Using the SPN data typed by the plink software, The IBS fraction is calculated by the above formula, where the IBS is the state homology, and there are 16 IBSs. obtaining relevant genetic marker sites; and 3) Based on the R language, a multivariate linear regression analysis model was used, of which 16 independent variables were included. are genetic marker sites of different genotypes, and the dependent variable is a specific IBS fraction. By creating a purebred evaluation model, a purebred evaluation model for the candidate Hu sheep purebred flocks can be established. The stage of making a decision; The present invention provides a method for evaluating purebred Hu sheep labeled with an SPN molecule comprising: Beneficial effects of the present invention compared to existing technologies:
[0007] 1) This invention uses high-throughput data at the colony level to identify Hu sheep and Xiaowei Han Genome-wide genetic diversity for sheep, Dorper sheep and Dorper x Hu sheep cross populations. By detecting mutations and using bioinformatics to identify the genetic characteristics of Hu sheep, To establish an evaluation method for identifying pure breeds of Hu sheep, and to promote the protection, development and utilization of Hu sheep breed resources. Provide scientific evidence. 2) The present invention identifies 30 genetic marker sites for use in the evaluation of Hu sheep purebredness. 3) The present invention further identifies a candidate set of genetic markers for use in the evaluation of the pure breeding of Hu sheep. do. [Brief description of the drawings]
[0008] [Figure 1] This is the technical route of Example 1.
[0009] [Diagram 2] 1 is an evolutionary tree of Example 1.
[0010] [Diagram 3] 1 is an experimental flow chart of Example 2.
[0011] [Figure 4] 1 is the Zscore conversion factor of IBS of Example 2.
[0012] [Diagram 5] FIG. 13 is a QQplot diagram of the model result evaluation of Example 2.
[0013] [Figure 6] 1 is a correlation between the IBS coefficients in Example 2 and the predicted IBS coefficients.
[0014] [Figure 7] This is the Zscore conversion coefficient of IBS for untested colonies. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] The present invention provides a method for evaluating the pure breed of Hu sheep tagged with SNP molecules, and the present invention is carried out at a herd level. Using high-throughput data, we have developed a method to measure the chromosome 1 (Hu sheep, Xian-tailed Han sheep, Dorper sheep and Doe sheep). We performed a genome-wide genetic mutation detection study on a cross between Per and Hu sheep, and Identifying the pure breed of Hu sheep by identifying the exclusive genetic characteristics of Hu sheep using informatics methods An evaluation method will be established to provide scientific basis for the conservation, development and utilization of Hu sheep breed resources.
[0016] The evaluation method includes: 1) Candidate genes with specific colony frequencies in Hu sheep colonies identified by high-throughput sequencing By identifying genetic mutations and comparing the frequency of genetic mutations in Hu sheep with those in other populations, We will develop a scoring method and perform sequencing to identify the genetic variation present in the sheep flock. A purebred evaluation model was created using various indicators, and genetic marker candidates were further evaluated based on the model results. The set is divided into 30 sets based on the site primer design and the site genome linkage situation. Selecting genetic marker sites for identification of Hu sheep pure breed; 2) The 30 genetic marker sites obtained from step 1) were subtyped using Sequenom SNP subtyping. Using the binning detection method, the SPN data subtyped by plink software were analyzed. The IBS fraction is calculated using the above method. The IBS is a state homology, and there are 16 IBSs. obtaining a genetic marker site that is closely related; and 3) Based on the R language, a multivariate linear regression analysis model was used, of which 16 independent variables were included. are genetic marker sites of different genotypes, and the dependent variable is a specific IBS fraction. By creating a purebred evaluation model, a purebred evaluation model for the candidate Hu sheep purebred flocks can be established. The method includes a step of determining
[0017] According to the genetic diversity analysis and conservation effect of Huzhou Hu sheep in the literature, the present invention 50 sheep, 10 Dorper sheep, 10 Han sheep, and a cross between Dorper and Hu sheep. Ten sheep are selected as the control group, and a total of 80 samples are selected. See Table 1. Hu sheep individuals are numbered with HY as the head, and Dorper sheep individuals are numbered with DB as the head. The small-tailed Han sheep are numbered as heads, the small-tailed Han sheep are numbered as heads, and the Dorper sheep are numbered as heads. The crossbred sheep between the Hu sheep and the DH sheep are numbered as the head. Some of the 1988 Hu sheep we obtained had horns, so we selected Hu sheep with some horns. Blood was collected from the neck of a sheep and placed into a disposable 2 mL vacuum blood collection tube (containing an anticoagulant). The blood samples were collected, brought back to the laboratory, stored at -20°C, and prepared for use using the imported product QIAGEN Blood Gen Extract DNA using the ome DNA Extraction Kit. Table 1. Individual samples used for high-throughput resequencing JPEG2025077938000002.jpg163115JPEG2025077938000003.jpg163115 Example 1
[0018] As shown in Figure 1, this example shows the construction of a library by resequencing performed on Hu sheep. There is sequencing.
[0019] The resequencing library creation involves total genomic DNA extraction, resequencing library creation, and Includes computerized high-throughput sequencing. Total genomic DNA extraction: for different populations DNA was extracted from the genomic samples and subjected to 0.8% agarose gel electrophoresis. The quality of DNA extraction was evaluated by extraction and at the same time, the DNA was analyzed by ultraviolet spectrophotometry. The DNA that passes the test is then further processed for end repair, A-tailing, and sequencing. The entire library was created through the steps of adding single adapters, purifying, and amplifying by PCR. do.
[0020] The created library is inspected for quality using the Agilent Bioanalyzer. The run library has a single peak, the adapter-free dimer, and is concentrated. The concentration is 2 nM or higher.
[0021] After the library was created, we performed preliminary quantification using Qubit2.0 and then used the qPCR method to determine the library. We provide accurate information on the effective concentration of the product to ensure the quality of the product. Afterwards, PE is sequenced by 2x150bp on an Illumina Hiseq.
[0022] Further quality control of the raw reads obtained The original sequence and adapter sequences are removed. Remove reads with a high proportion of low-quality bases (10%) and remove reads with a proportion of N > 10%. Remove ds. Single-end sequencing reads with fewer than 5 bases are longer than the read length. Reads that exceed 20% of the ratio are removed. In addition, the number of reads after filtration is further increased. Data such as amount, data yield, sequencing error rate, Q20 content, Q30 content, and GC content Perform statistics on the data.
[0023] The results are shown in Table 2. A total of 30.5 billion original sequencing sequences were obtained. On average, 380 million original sequencing sequences were detected per individual. The sequence contains approximately 57.2G of bases, and after filtration, a total of 29.8 billion quality-controlled strands are produced. High-quality sequence data was obtained, with an average of 373 million original sequences per individual. The sequencing sequence has a base length of about 56G, which satisfies the ~20X standard and is a high-quality variety. It can satisfy the appraisal needs. Table 2 Sequencing results for 80 sheep individuals JPEG2025077938000004.jpg163115JPEG2025077938000005.jpg163115 Comparison with the reference genome
[0024] For each individual sample, all quality control filtered data was collected using BWA software. The genome was compared to the sheep reference genome (GCA_016772045.1) with default parameters and then analyzed using Picar d. Using GATK and SAMtools, the compared sequences were recalibrated for overlap and basic quality. Then, duplicate data are removed and comparative statistics (i.e., deep coverage) are performed. All comparison documents (bam), including the detection of indels, are used for subsequent analysis processes.
[0025] As shown in Table 3, the total number of comparable reads of the reference genome is 30.2 billion, about 45.3 Gbp. It is a cardinal number and has a comparison rate of 99.76%. Table 3 Comparison results of 80 sheep JPEG2025077938000006.jpg163115JPEG2025077938000007.jpg163115 Detection of genetic variations (SNPs, Indels)
[0026] GATK soft process, HaplotypeCaller, GenotypeGVCFs and CombineGVCFs modules Using this method, we performed genetic mutation detection, genotyping, and colony merging for 80 individuals. Finally, the genotype file (VCF) of the detected mutation is obtained. The filtration module is used to perform forced filtering of genetic variations, including quality control. The criteria are as follows: QD < 2.0; QUAL < 30.0; SOR > 3.0; FS > 60.0; M Q < 40.0; MQRankSum < -12.5; ReadPosRankSum < -8.0.
[0027] At the same time, the sequencing depth is greater than 8, the coverage rate is greater than 30%, and there are sudden changes. The obtained InDels were selected by retaining the genetic variants with a minimum allele frequency of more than 0.05. Remove SNPs within 5 dp. Annotation of genetic variations
[0028] Using ANNOVA software, gene- or area-based annotations were performed for the detected variants. and download the genome gff annotation file (GCA_016772045.1) from NCBI. Notes on mutations: SNPs are divided into 8 types, and exon areas (synonyms, nonsynonyms) nim, stop gai and stop loss), splice sites, intron areas, 5' and 3 ' Includes UTRs, upstream and downstream areas and intergenic areas.
[0029] InDels further contain stop losses, stop gains, and frameshift mutations (3 bp insertions). or deletion). Genetic diversity assessment
[0030] Using SNP and InDels mutation information, the frequency distribution of the mutations in the different populations detected was analyzed. This is used to evaluate the diversity characteristics of the colony. Calculate the frequency of different SNP and InDels annotations based on the results of gene annotation above. Compare population frequency features to identify the preference for genomic feature distribution of genetic variation. InDels Information The length distribution of the chromosomes was further statistically analyzed to determine whether different length distributions exist in the InDels colonies. Compare the colony distribution frequencies.
[0031] On average, 23,972,025 SNPs and 2,441,291 InDels were detected per population. Among them, Hu sheep had the most genetic variations, with 26,262,209 SNPs and 2,693,430 InDels. Based on the s and InDels, a database of molecular genetic information of Huzhou Hu sheep was established. Sheep had the least genetic variation with 22,626,849 SNPs and 2,296,833 InDels. Colony structure analysis
[0032] Community structure analysis includes community principal component analysis, system evolutionary tree reconstruction, and community genetic structure analysis. Principal Component Analysis (PCA) is a method to analyze a data set. Analysis and simplified statistical methods.
[0033] A matrix is constructed from the SNP data of multiple individuals that have not been divided into colonies, and the feature vector of the matrix vector is Extract the main components (usually three) and create a scatter plot using two of the feature vectors. Depending on the distribution of the scatter plot, the unclassified individuals are divided into subpopulations. PCA is usually The method is mutually authenticated by genetic structure analysis and SNP-based system evolution analysis. The process was carried out by GCTA software using SNP data (SNPs with MAF less than 0.05) to identify the major A compositional analysis is carried out.
[0034] A system for reconstructing evolutionary trees based on SNPs. The evolutionary tree shows the closeness of kinship between species within a community. In the tree, each node has its nearest common ancestor on each branch. The length of the line segments between the nodes corresponds to the evolutionary distance (estimated evolutionary time). Each species is placed on a branched tree according to its relative kinship. Each leaf node represents one species, and the distance between two leaf nodes represents the distance between the corresponding species. This shows the degree of differentiation between the
[0035] Figure 2 shows the genetic relationships between sheep colonies, with Dorper and Short-tailed Han sheep being independent evolutionary groups. The Hu sheep colony has two relatively independent colonies, one with its own ranch and the other with its own lamb.
[0036] Admixture software is used to analyze the genetic structure of the population using SNP information. K = 2 to 10 ( In other words, the model with 2 to 10 ancestors was set as a mixed model, and the others were set using software. The default setting is as follows. The CV error value of different K is used to determine the K closest to the true value.
[0037] When K=2, Dorper sheep and Hu sheep have obvious similarity, and when K=3, Dorper sheep have obvious similarity. The phenomenon that Hu sheep colonies crossed with Parr sheep have a higher genetic relationship than Dorper sheep. was found. When K=4, the intracolony phenomenon present in the Hu sheep colony was also prominent. . Selection signal analysis for sheep
[0038] The selection signal of Hu sheep is calculated and evaluated by the differentiation index Fst of the Hu sheep flock. Fst is a group of one It indicates the degree of genetic diversity between subpopulations within the body, and the Fst value is generally between 0 and 1. The smaller the value, the less genetic differentiation of the subpopulations. If the value is 0, all This indicates that all individuals can mate freely with each other, and the degree of genetic differentiation is the lowest. The larger the value, the greater the genetic differentiation between the subpopulations. This indicates that subpopulations do not share any genetic diversity.
[0039] Using Vcftools, we performed FST detection within the entire genome, including windows. is set to 50Kb, the slippage window is set to 20Kb, and the top 5% is taken as the candidate area for Hu sheep. Creating an evaluation model for purebred Hu sheep
[0040] The evaluation of the Hu sheep pure breed is carried out by selecting representative genetic mutations () and detecting the Hu sheep pure breeds in the candidate sheep flocks. The purpose is to distinguish between bloodlines or hybrids. As shown in Table 4, there are six Select indicators and use them as the basis for the evaluation of genetic variation. A feature is classified according to the extent to which the mutation affects the functional genome (Table 5). Table 4. Evaluation parameters of the Hu sheep purebred evaluation model JPEG2025077938000008.jpg163115 Table 5. Classification of functional area characteristics JPEG2025077938000009.jpg163115
[0041] This invention creates a purebred identification model to identify the specific characteristics of crossbreed and purebred Hu sheep. A total of five indicators are used in the model to identify candidate genetic variants from different angles. The main indicators of genetic variation are the frequency of Hu sheep mutations and the frequency of mutations in other counties. The higher the mutation frequency of the body, the lower the mutation frequency of other groups, and the more the purebred characteristics of Hu sheep are captured. The detection rate of mutations in the Hu sheep colony is also an important indicator. It is used to indicate the reliability of the genetic variation, and the detection rate is usually over 90%, and the candidate features are Completely satisfy.
[0042] Based on the above indicators, an evaluation score can be calculated for a specific genetic mutation. Number = Mutation frequency of Hu sheep × Mutation frequency in other sheep flocks × Mutation detection rate in Hu sheep flocks; Random number of evaluation scores The ANO10 and PDLIM5 genes were evaluated in the Hu sheep population according to the results (see Table 6). These are taken as a set of candidate genetic markers. Table 6 Candidate genetic mutations for identification of Hu sheep pure breeds JPEG2025077938000010.jpg163115JPEG2025077938000011.jpg163115 Example 2
[0043] This example is based on Example 1, and includes two genes, ANO10 and PDLIM5, and 30 genetic markers. The site is used for PCR amplification to identify Hu sheep pure breeds.
[0044] Blood samples were collected from pedigree male sheep from the entire breeding colony at six Hu sheep breeding farms in Zhejiang Province. A total of 1902 blood samples were taken from the control group of Dorper sheep and crossbred Dorper and Hu sheep. Blood was collected from the neck vein of a sheep and placed in a disposable 2 mL vacuum blood collection tube (containing an anticoagulant). The blood samples were collected at 100 °C, taken to the laboratory, and stored at -20 °C until use. Extract DNA using the Genome DNA Extraction Kit. After purity testing, the DNA sample passed A total of 1898 samples were collected for further analysis. Some other Hu sheep blood samples were collected from pregnant ewes. As a result, the amount of blood collected was too low, resulting in insufficient purity and quantity of extracted DNA.
[0045] As shown in Figure 3, 30 SNP site subtyping was performed using Sequenom SNP subtyping detection. The specific detection method can be referred to patent ZL201811214821X, The title of the invention is "Eight SNP sites for identifying Tan and non-Tan sheep and their applications." be.
[0046] The 30 SNP site information and detection results are shown in Table 7. The total detection rate was 99.64. Only when this percentage standard is met can the accuracy of the calculation of subsequent models be guaranteed. Table 7. Location information and detection results JPEG2025077938000012.jpg163115
[0047] The present invention evaluates the genetic closeness between different individuals by calculating the IBS fraction. IBS stands for "identical by state" and refers to the same condition between individuals. Although they share many genetic mutations, these mutations have accumulated during genetic evolution. Therefore, it cannot be said that the two necessarily share a common ancestor. Since the relatives are close, it is used to evaluate the kinship between Hu sheep individuals. The IBS calculation is This is done using SNP data subtyped with plink software.
[0048] The Hu sheep purebred evaluation model is created using a multi-factor linear regression model, among which: The independent variables were 16 different genotypes of SNPs, and the dependent variable was a specific IBS fraction. At the same time, the variance and number of independent variables are determined by the stepwise method. At the same time, the model and model evaluation are optimized to produce the final purebred evaluation model. Finally, by creating a model, we can perform purebred identification of candidate Hu sheep flocks and purebred flocks. Both model creation and evaluation visualization are completed based on the R language.
[0049] Create genetic marker information for representative simulated individuals of the Hu sheep colony. Representative individual identification information for u sheep was obtained from the Hu sheep sequence colony (50 individuals) at the highest The allele types with the allele frequency of 1000 and 10000 were selected to obtain the genetic information of representative simulated individuals of Hu sheep. To ensure that the information has the characteristics of the Hu sheep herd and reduce the individual genetic deviation. Resequencing of IBS coefficients and Zscores for representative Hu sheep simulants The calculation results are shown in Table 8. Through IBS display and Zscore conversion, The latter IBS coefficient can be known and can better distinguish Hu sheep from other breeds. Zscore>0 is accepted as the discrimination criterion for Hu sheep pure breeds. Table 8. Calculation of IBS coefficients and Zscore conversion JPEG2025077938000013.jpg163115JPEG2025077938000014.jpg163115JPEG2025077938000015.jpg163115
[0050] As shown in Table 8 and Figure 4, a model was created using 30 gene expression sites and IBS fractions as parameters. At the same time, 30 genetic marker sites were labeled as independent variables and the IBS fraction was labeled as the dependent variable. A stepwise analysis was used to identify 16 genes significantly associated with IBS. The AIC value of the model was -773.8, and the model was highly evaluated. The detailed information of the model is shown in Table 9. Table 9. Model details JPEG2025077938000016.jpg163115
[0051] As shown in Figures 5 and 6, the QQplot evaluated by the model results is composed of sample values and predicted values. This reflects the relatively close results of both data sets. The correlation (r2) of 0.9846267 indicates that the model has excellent evaluation accuracy. expresses.
[0052] The IBS coefficient for the Hu sheep colony was predicted, and the prediction results were further converted into Zscore parameters. , where the calculation of the conversion factor is based on the mean and standard deviation of the resequencing samples. As shown in Figure 7, 7,219 Hu sheep samples with Zscore < 0 were classified as severely contaminated Hu sheep. Among the 80 Hu sheep samples, those with Zscore<0.3 and Zscore>0 were defined as Hu sheep with low contamination. Hu sheep with Zscore>0.3 can be classified as purebred Hu sheep.
[0053] The present invention uses a multi-element linear regression model to determine the 1-fold difference between 30 genetic marker sites. Six significant signature sites were selected to predict the IBS coefficients of representative individuals in the Hu sheep flock. The latter coefficients were used to successfully predict the breeding identification of untested Hu sheep flocks. Among them, 73.7% of the untested individuals were classified as purebred Hu sheep flocks, and 14.7% of the untested individuals were classified as less mixed H sheep flocks. The 11.53% of untested individuals were selectively selected as Hu sheep colonies with a high degree of heterogeneity. For colonies of the same classification, the IBS coefficient and the selection of the number of seeds to be left were significant. Thus, the individual can be removed.
[0054] The above description is merely a best mode for carrying out the present invention, and does not necessarily imply any formal or exemplary application of the present invention. The present invention does not impose any substantial limitations, and a person skilled in the art can easily understand the method of the present invention without departing from the scope of the present invention. In addition, some improvements and amendments can be made, but these improvements and amendments are within the scope of protection of the present invention. Those skilled in the art will be able to easily understand the invention without departing from the spirit and scope of the invention. , The above technical content is used to make equivalent changes such as minor variations, modifications and improvements. Any of these are considered to be embodiments of the present invention. Any equivalent changes, modifications and improvements made to the above embodiments are also included. It falls within the scope of protection of the technical means of the present invention.
Claims
1. 1) Candidate genes with specific colony frequencies in Hu sheep colonies using high-throughput sequencing methods By identifying genetic mutations and comparing the frequency of genetic mutations in Hu sheep with those in other populations, We will develop a scoring method and perform sequencing to identify the genetic variation present in the sheep flock. A purebred evaluation model was created using various indicators, and genetic marker candidates were further evaluated based on the model results. A set of 30 was provided based on the site primer design and the linkage situation of the site genome. selecting genetic marker sites for identification of Hu sheep pure breeds; 2) The 30 genetic marker sites obtained from step 1) were used for Sequenom SNP subtyping. Using the binning detection method, the SPN data subtyped by the plink software were The IBS fraction is calculated using the above method. The IBS is a state homology, and there are 16 IBSs. obtaining a closely related genetic marker site; and 3) Based on the R language, a multivariate linear regression analysis model was used, of which 16 independent variables were used. The genetic marker sites of different genotypes are the specific IBS fractions, and the dependent variable is By creating a purebred evaluation model, a purebred evaluation model for a candidate Hu sheep purebred flock can be obtained. determining whether the genetic marker candidate set includes the ANO10 gene and the PDLIM5 gene; In step 1), the 30 genetic marker sites are P64757122, P18235445, P83818961, P 83818969, P55276319, P55384071, P17642678, P17644657, P74169717, P19203813, P646 29274, P64629764, P64630486, P69523858, P15346217, P15347178, P15347396, P101963 616, P103941154, P50147220, P50194226, P6840723, P223703680, P37184275, P2994929 9, P30659550, P30778083, P44863176, P90333029 and P5313717, In step 2), the 16 genetic markers significantly associated with IBS were P30659550, P 15347396, P37184275, P223703680, P15346217, P55384071, P19203813, P64757122, P64 630486, P64629274, P50194226, P15347178, P74169717, P90333029, P69523858 and P1 Contains 8235445, In step 2), the larger the IBS, the closer the genetic relationship between the two individuals. Express the following: A method for evaluating purebred Hu sheep labeled with an SPN molecule, comprising:
2. In step 3), the purebred evaluation model is calculated by using the IBS coefficients. The genetic relationship between sheep individuals was evaluated, and the IBS coefficient and The calculation results of resequencing and Zscore conversion are shown in IBS display and Zscore conversion. We can know the IBS coefficient after Zscore transformation, and can distinguish Hu sheep from other breeds well. It is possible to distinguish between purebred Hu sheep when Zscore>0. A method for evaluating purebred Hu sheep labeled with the SPN molecule according to claim 1 .
3. In step 1), the multiple indicators used in the purebred evaluation model are Hu sheep herd mutation test Emergence rate, Hu sheep flock mutation frequency, other sheep flock mutation frequency, features located in functional areas, LD-S The number of NPs and the number of mutated bases are included, and the Hu sheep and the other groups are Hu sheep, Dorper sheep, Dorper sheep, and Dorper sheep.
2. The sheep according to claim 1, characterized in that it includes a cross between Parr sheep and Hu sheep and a small-tailed Han sheep. Method for evaluating purebred Hu sheep labeled with SPN molecules.
4. In step 1), multiple indicators are evaluated based on the calculated evaluation fraction for the candidate genetic variants. Fraction = number of mutations in Hu sheep colonies * mutation frequency in other sheep colonies * Hu sheep colony mutation detection rate, evaluation fraction is high The higher the probability of the candidate genetic mutation, the more likely the gene is to be selected as a genetic marker candidate set.
2. The application of the method for evaluating purebred Hu sheep labeled with the SPN molecule according to claim 1.
Citation Information
Patent Citations
Method for constructing tobacco core germplasm based on genomics and application thereof
CN112359102A
SNP (Single Nucleotide Polymorphism) molecular marker combination and application thereof in distinguishing Hebei small-tailed Han sheep
CN116732191A
Local ancestry inference using machine learning models
JP2023521893A
Cited By
Application of molecular marker influencing chest circumference character of Dumeng sheep
CN122235327A
Application of a molecular marker affecting chest circumference traits of duomun sheep
CN122235327B