Liquid chip for genome selection breeding of oncorhynchus nerka against edwardsiella tarda and application thereof
By developing a liquid-phase chip for genomic selection breeding of spotted sea bream resistant to Edwardsiella tarda containing 500 high-value SNP loci, and using MLE-rank, GWAS, and PVE analysis for screening, the problems of long breeding cycles and high costs were solved, and efficient and low-cost breeding of disease-resistant traits in spotted sea bream was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies for the prevention and control of Edwardsiella tarda disease in spotted sea bream have problems such as long breeding cycles, high costs, and difficulty in improving disease resistance traits. Traditional methods are also difficult to guarantee the accuracy of selection.
A liquid-phase microarray for genomic selection breeding of spotted sea bream resistant to Edwardsiella tarda was developed, containing 500 high-value SNP loci. These loci were screened using MLE-rank, GWAS, and PVE analyses. Capture probes were designed and synthesized, and a low-density liquid-phase microarray was constructed to assess disease resistance and estimate breeding value.
It significantly improved the accuracy of disease resistance prediction and selection efficiency, reduced genotyping costs, and enabled rapid and low-cost breeding of disease-resistant varieties of spotted sea bream.
Smart Images

Figure CN121472431B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of aquatic genomic selection breeding technology, specifically relating to liquid phase chip for genomic selection breeding of spotted sea bream resistant to Edwardsiella tarda and its application. Background Technology
[0002] Spotted sea bream ( Oplegnathus punctatus The spotted rock seabream is an important marine economic fish, prized for its delicate flesh, unique flavor, and high nutritional and medicinal value. Its aquaculture industry is rapidly developing due to its low farming costs, high market demand, and significant economic value. However, the presence of Edwardsiella tarda (…) has led to its decline. Edwardsiella tarda Bacterial diseases caused by pathogens such as [unspecified pathogens] are becoming increasingly frequent, especially during the high temperatures of summer. These diseases can cause skin ulcers, organ lesions, and other symptoms in farmed fish, resulting in extremely high mortality rates and severely hindering the healthy and sustainable development of the spotted sea bream industry. Traditional drug treatments pose risks of water pollution and drug resistance. Therefore, breeding disease-resistant strains is the fundamental way to achieve the green and sustainable development of the spotted sea bream industry.
[0003] In fish genetic breeding, traditional population selection, family selection, and hybridization breeding techniques are phenotype-dependent, have long breeding cycles, and struggle to guarantee selection accuracy for complex traits such as disease resistance, which are controlled by multiple genes and have low heritability, resulting in limited trait improvement. In recent years, Genomic Selection (GS) has utilized genome-wide genetic markers to construct a reference population and associate its genetic marker information with the target trait, estimating the breeding value (GEBV) of candidate populations without relying on their phenotypes. Compared to traditional methods, GS utilizes all SNP information to estimate the GEBV of candidate populations, significantly shortening generation intervals and greatly improving the accuracy and efficiency of selecting complex traits. Currently, this technology has demonstrated great potential in disease resistance breeding of various economically important fish species, such as large yellow croaker, turbot, and Atlantic salmon, providing a new technical pathway for improving disease resistance traits in fish.
[0004] The key reason limiting the application of genomic selection breeding is the cost of genotyping. Targeted SNP genotyping chip based on liquid capture (cGPS) can significantly reduce the cost of genotyping. This technology can reduce the cost of genotyping while ensuring detection accuracy by designing probes for high-depth sequencing of pre-selected SNP sites highly associated with target traits, which is a key path to promote the industrial application of genomic selection. In practice, although the low-density chip developed by this technology has been successful in species such as golden pompano and mandarin fish, the success of the low-density chip depends on the genetic information quality of the selected SNP sites. A high-quality SNP combination can significantly reduce the cost of genotyping while ensuring accuracy. However, the chip containing a large number of invalid SNP sites significantly reduces the prediction accuracy. Therefore, developing a low-cost genotyping chip for Oplegnathus punctatus containing high-value core SNP markers for the trait of resistance to Edwardsiella tarda is crucial for the sustainable development of the industry. SUMMARY
[0005] The purpose of the present application is to overcome the shortcomings of the prior art and provide a Oplegnathus punctatus genomic selection breeding liquid phase chip for resistance to Edwardsiella tarda and its application.
[0006] The technical solution of the present application is as follows:
[0007] A Oplegnathus punctatus genomic selection breeding liquid phase chip for resistance to Edwardsiella tarda comprises a probe combination capable of specifically capturing or detecting 500 SNP sites in the Oplegnathus punctatus reference genome (CNGBdb:CNA0019300), as shown in Table 1; the sequences of the 500 SNP sites are SEQ ID NO. 1-SEQ ID NO. 250, each sequence corresponds to two SNP sites, which are represented by the 36th and 107th bases of each sequence, and the 36th and 107th bases are indicated by the degenerate bases y / r.
[0008] Table 1: Information of 500 SNP sites
[0009]
[0010] The application of the liquid phase chip in any of the following:
[0011] (a) evaluating the ability of Oplegnathus punctatus individuals to resist Edwardsiella tarda;
[0012] (b) estimating the genomic breeding value of Oplegnathus punctatus individuals;
[0013] (c) screening the O. niloticus parents or families for resistance to E. tarda;
[0014] (d) preparing an O. niloticus disease-resistant breeding product.
[0015] A kit comprising the liquid chip and reagents for genotyping, the kit being used for O. niloticus genome selection breeding against E. tarda.
[0016] In order to achieve efficient and low-cost disease-resistant breeding, the application further provides a construction method of a low-density liquid chip for O. niloticus disease-resistant trait genome selection breeding, comprising the following steps:
[0017] (1) The application constructs Geno1 containing 2415773 SNPs by obtaining O. niloticus whole genome SNP markers; Geno2 containing 22757 SNPs is constructed by using a whole genome uniform distribution sampling method; the whole set Geno1 is screened by using maximum likelihood estimation ranking (MLE-rank), genome-wide association analysis (GWAS) and phenotype variance explanation rate (PVE) analysis, and a candidate core set containing 500 SNP sites highly associated with disease resistance traits is obtained.
[0018] (2) The 500 SNP sites screened by the method combining MLE-rank, GWAS and PVE analysis are used to perform genome selection by 10x10 cross-validation method on a reference population by using the GBLUP method, and the prediction accuracy is predicted by the area (AUC) under the ROC curve, and the prediction accuracy results are compared with those of Geno1 containing all SNPs and Geno2 containing uniformly distributed SNPs.
[0019] (3) In order to further verify that the 500 core SNP sites provided by the application can be efficiently applied to genome selection breeding, the prediction accuracy is compared with that of Geno1 containing all SNPs and Geno2 containing uniformly distributed SNPs in an independent verification population. And by constructing a GBLUP model with the phenotype and genotype data of the offspring, the parent GEBV of Geno1, Geno2 and Geno3 is obtained, and the Pearson correlation between the parent calculated family GEBV and the offspring survival rate is further verified, and it is verified that the 500 core SNP sites provided by the application are effective and can be efficiently applied to genome selection breeding.
[0020] (4) Based on the 500 SNP sites of Geno3 screened by MLE-rank, GWAS analysis and PVE analysis, corresponding capture probes are designed and synthesized to prepare a low-density liquid chip specially used for O. niloticus disease-resistant trait detection.
[0021] Compared with the prior art, the application has the following beneficial effects:
[0022] In view of the long breeding cycle, high cost, and difficulty in improving disease resistance of the traditional method, the application provides a liquid chip "Dream No. 1" for breeding of Oplegnathus punctatus against Edwardsiella tarda, which comprises 500 core SNP sites. In the reference population, the prediction accuracy of Geno3 of the 500 SNP sites screened by the method combining MLE-rank, GWAS and PVE analysis is higher than that of Geno1 of all SNP sites and Geno2 of evenly distributed SNP sites, so Geno3 is effective in the disease resistance of Oplegnathus punctatus. In the independent verification population, the prediction accuracy of the 500 core SNP sites provided by Geno3 is 0.87, which is significantly better than that of Geno1 (0.7) of all SNP sites and Geno2 (0.68) of conventional evenly distributed SNP sites. By calculating the family average GEBV of the parents and the survival rate of the offspring, the Pearson correlation between the breeding value predicted by the 500 core SNP sites and the actual survival rate of the offspring is 0.481, and the P value of the Pearson correlation coefficient significance test is 0.0046 < 0.01, and the correlation is extremely significant at the 0.01 level, so the 500 core SNP sites provided by the application are effective. The application realizes high prediction accuracy with extremely low marker density, successfully solves the industry problem that high accuracy and low cost are difficult to be balanced, and provides a key technical breakthrough for rapid and low-cost selection of Oplegnathus punctatus disease-resistant fine varieties. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 : GWAS diagram of 739 tail reference population;
[0024] Figure 2 : PVE diagram of 739 tail reference population;
[0025] Figure 3 : Prediction accuracy diagram of all SNP markers, 22757 evenly distributed SNP markers, 500 SNP markers selected by MLE-rank and GWAS of 739 tail reference population under 10x10 cross-validation method;
[0026] Figure 4 : Prediction accuracy diagram of 121 tail verification population evenly distributed SNP markers and 500 core SNP sites. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical scheme of the application clearer, the application will be further described in detail below in combination with the drawings. In the following test, the experimental methods described are conventional methods unless otherwise specified; the specific techniques or conditions not marked in the test are carried out according to the techniques or conditions described in the literature in the art or according to the product instructions; the reagents and materials described are commercially available unless otherwise specified.
[0028] Example 1: Phenotype and genotype and construction of different SNP marker sets
[0029] 1. Construction of phenotype data of O. marmoratus individuals
[0030] The O. marmoratus samples used in this example were collected in 2023 and 2024 by Huanghai Aquatic Products Co., Ltd. in Yantai City, Shandong Province, China. The infection experiment was performed by injecting the same concentration of Edwardsiella tarda into the abdominal cavity of O. marmoratus individuals. After the infection experiment, fin strips were collected from the dead individuals every 6 hours, and after no deaths for a week, the tail fins of the remaining O. marmoratus individuals were collected as surviving individuals. The recorded death and survival status data were used as phenotypes. A total of 860 O. marmoratus individuals were studied, including 739 individuals in the reference population and 121 individuals in the independent verification population.
[0031] 2. Screening of high-quality SNPs
[0032] DNA extraction was performed on 860 samples, and after strict quality detection, whole genome sequencing was performed using the Huada T7 sequencing platform. After quality control of the sequencing data, clean reads were aligned with the O. marmoratus reference genome (CNGBdb: CNA0019300), and after removing repetitive sequences, basic information statistics and mapping analysis were performed. Based on the alignment results of the sequencing sequences and the reference genome, SNP calling was performed using the Genome Analysis Toolkit (GATK), and only high-quality SNPs (QUAL>30, QD>2, MQ>40) were retained to generate the original vcf file. Quality control was performed using VCFtool, and the standards for removing low-quality SNPs were as follows: SNP sites with less than 15x average sequencing depth, sites with a deletion rate higher than 2%, SNP sites with a minor allele frequency (MAF) lower than 5%, and SNP sites not meeting Hardy Weinberg equilibrium (p-value set to 1x10 -6 ), only biallelic SNP sites were retained. After quality control, Beagle (v5.2) was used for genotype filling, and finally 2415773 high-quality SNP sites were obtained for subsequent research.
[0033] 3. Construction of SNP marker sets
[0034] The 2415773 high-quality SNP sites obtained by quality control are set as Geno1; 22757 SNP markers are screened out by the -bp-space command of Plink v1.9 software according to the physical position of the reference genome to set Geno2; 500 SNP sites with strong correlation with the phenotype are screened out by the MLE-rank model and GWAS analysis and PVE analysis to set Geno3, and the 500 sites are determined as the core marker set for constructing the breeding chip of the application.
[0035] The MLE-rank model is as follows:
[0036]
[0037] represents the transpose of the phenotype matrix of the individual with the phenotype of survival (recorded as 1), represents the phenotype matrix of the individual with the phenotype of death (recorded as 0), represents the genotype matrix of the individual with the phenotype of 1, represents the genotype matrix of the individual with the phenotype of 0, wherein AA / Aa / aa are marked as 0 / 1 / 2, p represents the proportion of the survival individuals in the population, and value represents the vector composed of the influence of each SNP site on the phenotype.
[0038] The MLE-rank screened SNP site is sorted from large to small by calculating the value of each SNP site in value, and the SNP site is screened.
[0039] GWAS is realized by the mlma instruction of gcta software, and the GWAS model is as follows:
[0040] y=P +Zu+
[0041] y represents the phenotype value, P represents the covariance matrix of the first three principal components, represents the effect value of the principal component, Z represents the SNP genotype matrix, and u represents the effect value of each SNP, represents the residual.
[0042] The results of the GWAS analysis are sorted from small to large by the p value, and the SNP site is screened.
[0043] The PVE calculation is realized by R language, and the PVE formula is as follows:
[0044] PVE=
[0045] beta is the effect value of the allele of SNP, MAF is the frequency of the minor allele, se is the standard error of the effect value, and N is the total number of individuals participating in the analysis.
[0046] The PVE values are sorted from large to small to screen the SNP sites.
[0047] The GWAS plot and PVE plot of the reference population of 739 tails are as follows Figure 1 and Figure 2 .
[0048] Example 2: Evaluation of genetic force of O. latus against E. tarda
[0049] 1. Purpose of the example
[0050] The purpose of this example is to compare the performance of the 500 SNP marker set (Geno3) based on MLE-rank, GWAS analysis and PVE screening, the whole genome marker set (Geno1) and the uniformly distributed SNP marker set (Geno2) in evaluating the genetic force of O. latus against E. tarda, and to verify the effectiveness of the SNP screening strategy of the present application.
[0051] 2. Evaluation method
[0052] The reference population (739 tails) constructed in Example 1 was used. Using the GBLUP model, combined with the probit connection function, the additive genetic variance and genetic force of Geno1, Geno2 and Geno3 were estimated by the Asreml-R software package.
[0053] 3. Results and analysis
[0054] The genetic parameter evaluation results of each data set are shown in Table 2.
[0055] Table 2: Evaluation of genetic force and variance components of three data sets
[0056]
[0057] The results of the evaluation of genetic force and variance components of the three data sets are shown in Table 2, in which the genetic force estimated by all SNPs (Geno1) is 0.182±0.062, and the additive genetic variance is 0.223±0.092; the genetic force estimated by uniform distribution (Geno2) is 0.163±0.055, and the additive genetic variance is 0.195±0.079; the genetic force estimated by 500 core SNP sites (Geno3) is 0.278±0.061, and the additive genetic variance is 0.384±0.117.
[0058] As shown by the results in Table 2, the estimated genetic parameters of the uniform distribution (Geno2) did not reach the level of all SNP markers (Genol), indicating that simple uniform distribution would result in loss of genetic information. The estimated genetic force and variance components based on the MLE-rank model, GWAS analysis and PVE analysis of the SNP dataset (Geno3) were significantly higher than those estimated by all SNP markers (Genol), indicating that Geno3 could reach 1.5 times the genetic force estimated by all SNP markers at 500 SNP sites, i.e., most of the genetic variation could be captured at low density, and the influence of "noise" unrelated to the target trait of all SNPs was reduced. Geno3 reduced the cost of genotyping while maintaining high accuracy of genetic parameters, providing a key technical foundation for establishing an economic and efficient genome selection breeding system for disease-resistant spotted grouper.
[0059] Example 3: Evaluation of the prediction accuracy of genomic estimated breeding values (GEBV) based on reference population cross-validation
[0060] 1. Purpose of the example
[0061] This example aims to compare the accuracy of the core SNP marker set (Geno3) screened by the present application, the whole genome SNP marker set (Genol) and the uniform distribution SNP marker set (Geno2) in predicting GEBV through 10-fold cross-validation, and to evaluate the performance of the screening strategy of the present application in the application of genomic selection breeding.
[0062] 2. Evaluation method
[0063] The reference population (739 individuals) constructed in Example 1 was used. Using the GBLUP model combined with the probit connection function, 10-fold cross-validation (10-fold cross-validation) was performed by the Asreml-R software package, with a total of 10 repetitions (100 iterations). In each iteration, 90% of the individuals were used as the training set to train the model, and 10% of the individuals were used as the validation set to test the model. The phenotypic data and GEBVs of the validation set were used to calculate the area under the ROC curve (AUC), and the prediction accuracy.
[0064] 3. Results and analysis
[0065] The cross-validation results of each repetition were averaged to obtain 10 prediction accuracy results to evaluate the performance of the model, and the results are shown in Table 3 and Figure 3 .
[0066] Table 3: Average prediction accuracy of 10-fold cross-validation
[0067]
[0068] As shown in Table 3, the average prediction accuracy of 100 cross-validation results shows that the prediction accuracy of Geno3 is 0.686, which is significantly higher than that of Geno1 (0.607) and Geno2 (0.647). This result shows that the 500 SNPs associated with the target trait screened by Geno3 not only effectively capture most of the genetic information, but also efficiently estimate GEBVs when the reference population is used for genomic selection, and the prediction accuracy. Therefore, Geno3 can be used for the production of liquid phase chips, which can effectively reduce the cost of genotyping and is conducive to the cultivation of disease-resistant good strains of O. lactatus.
[0069] Example 4: Application effect verification of core SNP marker set (Geno3) in independent verification population
[0070] 1. Purpose of the example
[0071] This example aims to evaluate the actual prediction performance of the 500 core SNP marker set (Geno3) finally screened by the application in an independent verification population (121 individuals), and compare it with Geno1 and Geno2, so as to verify the practical application value of the breeding chip developed based on Geno3 in the disease-resistant breeding of O. lactatus.
[0072] 2. Evaluation method
[0073] The independent verification population (121 individuals) defined in Example 1 was used. Using the GBLUP model, combined with the probit connection function, through the Asreml-R software package, taking the reference population (739 individuals) in Example 1 as the reference population and 121 individuals as the verification population, the reference population was used to train the independent verification population based on the SNP marker set of Geno1, Geno2 and Geno3. The area under the ROC curve (AUC) between the true phenotype (survival / death) of individuals in the independent verification population and the predicted GEBV of each marker set was calculated to compare the prediction accuracy of each marker set.
[0074] 3. Results and analysis
[0075] Figure 4 The results show that the prediction accuracy AUC of the 500 core SNP sites (Geno3) provided by the application is 0.87, which is significantly higher than the AUC of the uniform distribution (Geno2) of 0.68 and the AUC of all SNPs (Geno1) of 0.7. Therefore, the 500 core SNP sites finally screened based on the MLE-rank model and the GWAS model and the PVE analysis are effective for the breeding of O. lactatus against Edwardsiella tarda, and the produced liquid phase chip can be used for the breeding of O. lactatus against disease.
[0076] Example 5: Application verification of core marker set (Geno3) in family selection and parent evaluation
[0077] 1. Purpose of the example
[0078] The example aims to verify the practical effect and accuracy of the 500 core SNP marker set (Geno3) screened by the application in breeding by evaluating the GEBV of parents and families, and to compare with the whole genome SNP (Geno1) and uniformly distributed SNP (Geno2).
[0079] 2. Evaluation method
[0080] The parent data of 32 dams and 33 sires and 33 families composed of them were selected. Using the GBLUP model, the GEBV of each parent (sire, dam) was calculated based on the Geno1, Geno2 and Geno3 data sets respectively, and the average GEBV of each family corresponding to the GEBV of the parents was taken as the average GEBV of the family.
[0081] 3. Results and analysis
[0082] 3.1 Selection of resistant and susceptible families
[0083] By comparing the GEBV ranking of families calculated by the three different data sets, the accuracy of the 500 core SNP sites in screening resistant and susceptible families was tested.
[0084] Table 4: Average GEBV of families
[0085]
[0086] As shown in Table 4, the top 6 families of the average GEBV of the 500 core SNP data set and the total 2415773 SNP data set and the uniformly distributed 22757 SNP data set are the same, which are F2320, F2306, F2349, F2309, F2340, F2324, and the last 6 families are five of the same, which are F2326, F2336, F2321, F2325, F2317. The results show that although the Geno3 marker set of the application has only 500 SNPs, its decision-making ability in distinguishing resistant and susceptible families is almost the same as using whole genome data (Geno1) and uniformly distributed (Geno2), so the 500 core SNP sites provided by the application can be applied to practical breeding.
[0087] 3.2 Correlation analysis of prediction accuracy
[0088] The prediction accuracy is obtained by calculating the Pearson correlation coefficient between the estimated family average GEBV of each data set and the actual disease-resistant survival rate of the offspring of the family.
[0089] Table 5: Prediction accuracy of core SNP sites
[0090]
[0091] As shown in Table 5, the Pearson correlation coefficient is used to analyze the correlation between the family average GEBV and the family death survival rate of 500 core SNP sites, and the result shows that the correlation coefficient is 0.481. Using the Pearson correlation coefficient significance test, the P value is 0.0046 < 0.01, and the correlation is extremely significant at the 0.01 level, so the 500 core SNP sites provided by the present application are effective, which greatly reduces the cost of genotyping. Although the correlation of uniform distribution is slightly higher (0.536), the number of sites is 45.5 times that of the core SNP sites of the present application. Considering the cost of genotyping, the uniform distribution is not reasonable for cost investment in actual breeding. Therefore, the core SNP sites provided by the present application can ensure accuracy while reducing the cost of genotyping.
[0092] In summary, the 500 core SNP sites provided by the present application can accurately screen out disease-resistant families, which is consistent with the results of screening using whole genome SNP data, and can be directly used to guide breeding practice. The core SNP sites provided by the present application achieve 91.8% prediction accuracy with whole genome SNP using the lowest number of markers, and the required SNP number is only 0.02% of the total SNP number, which reduces the cost of genotyping. The low-cost chip developed based on this can promote the industrialization and scaling of the disease-resistant breeding of Oplegnathus punctatus.
Claims
1. A genomic selection breeding liquid chip for Oplegnathus fasciatus against Edwardsiella tarda, characterized in that, The chip comprises a probe combination capable of specifically capturing or detecting 500 SNP sites in the Oplegnathus fasciatus reference genome CNGBdb:CNA0019300, and the 500 SNP sites are shown in Table 1. Table 1: Information of 500 SNP sites 。 2. The liquid chip of claim 1, wherein, The screening method of the 500 SNP sites comprises the following steps: (a) obtaining whole genome resequencing data of an Oplegnathus fasciatus reference population, and obtaining a high-quality SNP marker set Geno1 after quality control and genotype filling; (b) respectively using a maximum likelihood estimation ranking model, whole genome association analysis and phenotype variance explanation rate analysis method to analyze and rank the SNP sites in Geno1; (c) comprehensively ranking the results of the three methods in step (b), and screening 500 SNP sites with the highest correlation with the Edwardsiella tarda resistance trait to form a core SNP marker set.
3. The use of the liquid chip of claim 1 in any one of the following: (a) preparing a product for evaluating the ability of Oplegnathus fasciatus individuals to resist Edwardsiella tarda; (b) preparing a product for estimating the genomic breeding value of Oplegnathus fasciatus individuals to resist Edwardsiella tarda; (c) preparing a product for screening Oplegnathus fasciatus parents or families resistant to Edwardsiella tarda.
4. A kit characterized in that, The kit comprises the liquid chip of claim 1 and reagents for genotyping.
5. The use of the kit of claim 4 in the preparation of an Oplegnathus fasciatus genomic selection breeding product resistant to Edwardsiella tarda.
Citation Information
Patent Citations
Method for breeding Edwardsiella tarda-resistant excellent flounder breed
CN104304096A
Preparation method and application of ironmodulatory peptide gene of oplegnathus punctatus and recombinant protein of ironmodulatory peptide gene
CN115010798A