Method for screening for high polyunsaturated fatty acid content superior parent in cyprinus carpio
By associating carp SNP locus genotypes with polyunsaturated fatty acid content, and using a ridge regression model to screen for superior carp parents, the problem of screening carp parents with high polyunsaturated fatty acid content in existing technologies has been solved, thereby improving the accuracy and efficiency of carp breeding and new germplasm identification.
Patent Information
- Application Number
- CN202510292442.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing technologies make it difficult to efficiently screen carp broodstock with high polyunsaturated fatty acid content, affecting the accuracy and efficiency of carp breeding and new germplasm identification.
By linking the SNP locus genotype of carp with the content of polyunsaturated fatty acids, the breeding value of carp was predicted using a ridge regression model, and superior parents with high polyunsaturated fatty acid content were screened out.
It enables accurate and efficient prediction of polyunsaturated fatty acid content in carp, provides excellent breeding materials, and improves the accuracy and efficiency of new carp germplasm identification.
Smart Images

Figure CN120319322B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of carp, specifically to a method for screening superior parent carp with high polyunsaturated fatty acid content. Background Technology
[0002] Carp (Cyprinus carpio) primarily inhabit still waters such as the middle and lower reaches of rivers, lakes, and reservoirs, especially favoring nutrient-rich, bottom-dwelling, or aquatic-plant-rich areas. Unsaturated fatty acids, particularly omega-3 and omega-6 fatty acids, are essential for human and animal health and life. Fish are one of the main sources of essential unsaturated fatty acids for human nutrition. The fatty acid composition of fish is highly plastic; the intake of exogenous fatty acids and the synthesis of endogenous fatty acids affect the deposition and metabolism of polyunsaturated fatty acids in fish. The intake of exogenous fatty acids mainly comes from feed rich in fish oil. However, fish oil has been recognized as a globally limited source of nutrients. Therefore, reducing the dependence of aquaculture feed on fish oil and enhancing the endogenous biosynthesis of unsaturated fatty acids in farmed fish is of great significance for increasing the supply of high-quality animal protein for humans. Carp, as an economically important aquatic product, can meet human consumption needs. Its unsaturated fatty acids, such as eicosapentaenoic acid (EPA) and docosahexaenoic acid (DHA), can effectively prevent cardiovascular disease, lower blood cholesterol levels, and reduce the probability of thrombosis. As a food product, especially fresh carp, quickly determining the EPA and DHA content in carp meat allows for rapid identification of the product's quality.
[0003] Therefore, it is particularly important to select carp parent fish with high polyunsaturated fatty acid content. Summary of the Invention
[0004] The inventors of this application have creatively correlated the SNP locus genotype of carp with the total content of polyunsaturated fatty acids, determining the polyunsaturated fatty acid content through SNP locus genotype and identifying dominant parents with high polyunsaturated fatty acid content. The method provided in this application can accurately and efficiently predict the total polyunsaturated fatty acid content of carp. The method provided in this application provides excellent materials for the identification and breeding of new carp germplasm.
[0005] Therefore, this application discloses at least the following technical solutions:
[0006] This embodiment discloses a method for screening superior carp parents with high polyunsaturated fatty acid content. The method includes: acquiring training feature data and training target data for a training set; the training feature data includes SNP genotype information of the training samples, and the training target data includes information on the total polyunsaturated fatty acid content of the training samples; training a ridge regression model using the training feature data and training target data to obtain a prediction model, wherein the prediction model characterizes the relationship between SNP genotype information and breeding value, and the breeding value characterizes the total polyunsaturated fatty acid content; acquiring the SNP genotype information of the sample to be tested; obtaining the breeding value of the sample to be tested based on the SNP genotype information and the prediction model; and screening superior carp parents with high polyunsaturated fatty acid content based on the breeding value.
[0007] By performing these steps, the electronic device obtains the breeding value directly based on the SNP genotype information and prediction model of the sample to be tested, and then screens out superior parents with high total polyunsaturated fatty acid content from the sample to be tested, providing excellent materials for the identification and breeding of new carp germplasm.
[0008] In some embodiments, the method further includes: providing a training set containing one or more samples, each training sample including training feature data and training target data of one or more carp.
[0009] In some embodiments, the SNP genotype information is the transformation value of one or more SNP loci genotype information of one or more carp.
[0010] In some embodiments, the SNP genotype information includes wild-type non-mutant genotype information, heterozygous genotype information, and purified mutant genotype information, wherein the wild-type non-mutant genotype information, heterozygous genotype information, and purified mutant genotype information are different.
[0011] In some embodiments, the SNP sites in the SNP genotype information include:
[0012] The carp genome NC_056572.1 contains a C>T mutation polymorphism at nucleotide position 35093078 on chromosome A1.
[0013] The carp genome NC_056572.1 contains a G>C mutation polymorphism at nucleotide position 171834 on chromosome A1.
[0014] The carp genome NC_056572.1 contains a C>T mutation polymorphism at nucleotide position 11053474 on chromosome A1.
[0015] Nucleotide position 26618712 on chromosome A1 of the carp genome NC_056572.1 shows a T>C mutation polymorphism; nucleotide position 20260097 on chromosome A12 of the carp genome NC_056583.1 shows a C>T mutation polymorphism; nucleotide position 12052513 on chromosome A12 of the carp genome NC_056583.1 shows a C>G mutation polymorphism; nucleotide position 20140419 on chromosome A12 of the carp genome NC_056583.1 shows a T>C mutation polymorphism; nucleotide position 24485947 on chromosome A14 of the carp genome NC_056585.1 shows a G>A mutation polymorphism; nucleotide position 24485947 on chromosome A14 of the carp genome NC_05658... Nucleotide position 18653036 of chromosome A14 in carp genome NC_056585.1 shows a G>A mutation polymorphism; nucleotide position 18984637 of chromosome A14 in carp genome NC_056585.1 shows a T>C mutation polymorphism; nucleotide position 18984627 of chromosome A14 in carp genome NC_056585.1 shows a T>C mutation polymorphism; nucleotide position 18984619 of chromosome A14 in carp genome NC_056585.1 shows a G>T mutation polymorphism; nucleotide position 23968026 of chromosome A14 in carp genome NC_056585.1 shows a G>C mutation polymorphism; nucleotide position 2 of chromosome A15 in carp genome NC_056586.1 shows a G>C mutation polymorphism; nucleotide position 2 of chromosome A15 in carp genome NC_056586.1 shows a G>A mutation polymorphism; nucleotide position 2 of chromosome A15 in carp genome NC_056586.1 shows a G>C mutation polymorphism; nucleotide position 2 of chromosome A15 in carp genome NC_056586.1 shows a G>C mutation polymorphism; nucleotide position 2 of chromosome A14 in carp genome NC_056585 ... Nucleotide position 1928001 shows a G>C mutation polymorphism; nucleotide position 802752 of chromosome A15 in carp genome NC_056586.1 shows an A>T mutation polymorphism; nucleotide position 802762 of chromosome A15 in carp genome NC_056586.1 shows an A>T mutation polymorphism; nucleotide position 802764 of chromosome A15 in carp genome NC_056586.1 shows an A>T mutation polymorphism; nucleotide position 6850717 of chromosome A16 in carp genome NC_056587.1 shows a T>C mutation polymorphism; nucleotide position 5355826 of chromosome A16 in carp genome NC_056587.1 shows a C>T mutation polymorphism. The following polymorphisms were observed in the carp genome NC_056587.1: Nucleotide position 5354912 of chromosome A16 shows an A>T mutation polymorphism; nucleotide position 5354913 of chromosome A16 shows a T>C mutation polymorphism; nucleotide position 8296639 of chromosome A16 shows a C>T mutation polymorphism; nucleotide position 4799340 of chromosome A16 shows a G>A mutation polymorphism; nucleotide position 4799336 of chromosome A16 shows an A>G mutation polymorphism; [The last part, "carp genome NC_056587.1," appears to be an error and is left untranslated.]Nucleotide position 3989726 of chromosome A16 in carp genome NC_056587.1 exhibits a T>C mutation polymorphism; nucleotide position 4939675 of chromosome A16 in carp genome NC_056587.1 exhibits a G>C mutation polymorphism; nucleotide position 11926439 of chromosome A18 in carp genome NC_056589.1 exhibits a G>A mutation polymorphism; nucleotide position 24946240 of chromosome A18 in carp genome NC_056589.1 exhibits a T>G mutation polymorphism; nucleotide position 2446713 of chromosome A2 in carp genome NC_056573.1 exhibits a C>A mutation polymorphism; nucleotide position 2111955 of chromosome A2 in carp genome NC_056573.1 exhibits a T>C mutation polymorphism. C>G mutation polymorphism; C>T mutation polymorphism exists at nucleotide 1842431 on chromosome A2 of the carp genome NC_056573.1; G>A mutation polymorphism exists at nucleotide 10731194 on chromosome A20 of the carp genome NC_056591.1; G>T mutation polymorphism exists at nucleotide 11906615 on chromosome A21 of the carp genome NC_056592.1; T>A mutation polymorphism exists at nucleotide 13620002 on chromosome A21 of the carp genome NC_056592.1; T>C mutation polymorphism exists at nucleotide 19643501 on chromosome A23 of the carp genome NC_056573.1; carp genome NC_056573.1... Nucleotide position 23931849 of chromosome A23 in carp genome NC_056574.1 shows a G>A mutation polymorphism; nucleotide position 18100115 of chromosome A3 in carp genome NC_056574.1 shows a T>C mutation polymorphism; nucleotide position 10859944 of chromosome A3 in carp genome NC_056574.1 shows a T>G mutation polymorphism; nucleotide position 10859946 of chromosome A3 in carp genome NC_056574.1 shows a T>G mutation polymorphism; nucleotide position 8426250 of chromosome A4 in carp genome NC_056575.1 shows a C>T mutation polymorphism; nucleotide position 8425648 of chromosome A4 in carp genome NC_056575.1 shows a... The following polymorphisms were observed in the carp genome NC_056575.1: T>C mutation polymorphism at position 20011914 of chromosome A4; T>A mutation polymorphism at position 39119153 of chromosome A5 of chromosome NC_056576.1; G>A mutation polymorphism at position 39119148 of chromosome A5 of chromosome NC_056576.1; T>C mutation polymorphism at position 6521311 of chromosome A5 of chromosome NC_056576.1; G>C mutation polymorphism at position 35333197 of chromosome A5 of chromosome NC_056577.1; and in the carp genome NC_056577.Nucleotide position 15309949 of chromosome A6 in carp genome NC_056577.1 shows a C>T mutation polymorphism; nucleotide position 5360090 of chromosome A6 in carp genome NC_056577.1 shows a T>C mutation polymorphism; nucleotide position 32004651 of chromosome A6 in carp genome NC_056577.1 shows a C>T mutation polymorphism; nucleotide position 40388036 of chromosome A7 in carp genome NC_056578.1 shows an A>T mutation polymorphism; nucleotide position 36226528 of chromosome A7 in carp genome NC_056578.1 shows a T>G mutation polymorphism; nucleotide position 36226538 of chromosome A7 in carp genome NC_056578.1 shows... The following polymorphisms were observed in the carp genome NC_056578.1: C>T mutation polymorphism; A>G mutation polymorphism at nucleotide position 3036102 on chromosome A7; T>C mutation polymorphism at nucleotide position 4467840 on chromosome A8; A>T mutation polymorphism at nucleotide position 16846973 on chromosome A8; G>A mutation polymorphism at nucleotide position 3296980 on chromosome A9; and T>C mutation polymorphism at nucleotide position 27008604 on chromosome A9. Nucleotide position 27008612 of chromosome A9 in carp genome NC_056580.1 shows a C>T mutation polymorphism; nucleotide position 1444937 of chromosome A9 in carp genome NC_056580.1 shows a C>A mutation polymorphism; nucleotide position 1737445 of chromosome A9 in carp genome NC_056580.1 shows a C>T mutation polymorphism; nucleotide position 4944998 of chromosome A9 in carp genome NC_056580.1 shows a G>A mutation polymorphism; nucleotide position 4944965 of chromosome A9 in carp genome NC_056580.1 shows a G>A mutation polymorphism; nucleotide position 4944960 of chromosome A9 in carp genome NC_056580.1 shows a T>T mutation polymorphism. C mutation polymorphism; T>C mutation polymorphism exists at nucleotide 4944980 on chromosome A9 of carp genome NC_056580.1; T>C mutation polymorphism exists at nucleotide 7101839 on chromosome B11 of carp genome NC_056607.1; T>G mutation polymorphism exists at nucleotide 9756853 on chromosome B11 of carp genome NC_056607.1; A>T mutation polymorphism exists at nucleotide 8439254 on chromosome B11 of carp genome NC_056607.1; T>A mutation polymorphism exists at nucleotide 26329356 on chromosome B12 of carp genome NC_056610.Nucleotide position 19109755 on chromosome B14 of the carp genome NC_056611.1 shows a C>G mutation polymorphism; nucleotide position 25384747 on chromosome B15 of the carp genome NC_056611.1 shows a T>A mutation polymorphism; nucleotide position 25384749 on chromosome B15 of the carp genome NC_056611.1 shows an A>T mutation polymorphism; nucleotide position 21342737 on chromosome B18 of the carp genome NC_056614.1 shows a T>A mutation polymorphism; nucleotide position N on chromosome B14 of the carp genome NC_056611.1 shows a T>A mutation polymorphism; Nucleotide position 26222460 of chromosome B19 in carp genome C_056615.1 exhibits a C>T mutation polymorphism; nucleotide position 5963093 of chromosome B19 in carp genome NC_056615.1 exhibits a T>A mutation polymorphism; nucleotide position 1978925 of chromosome B2 in carp genome NC_056598.1 exhibits a C>G mutation polymorphism; nucleotide position 31537582 of chromosome B2 in carp genome NC_056598.1 exhibits a T>A mutation polymorphism.
[0016] The carp genome NC_056616.1 contains an A>G mutation polymorphism at nucleotide position 30216632 on chromosome B20.
[0017] The carp genome NC_056616.1 contains an A>G mutation polymorphism at nucleotide position 25570376 on chromosome B20;
[0018] The carp genome NC_056618.1 contains an A>G mutation polymorphism at nucleotide position 18572947 on chromosome B22.
[0019] Nucleotide position 22921065 on chromosome B23 of the carp genome NC_056619.1 contains a C>A mutation polymorphism;
[0020] A G>C mutation polymorphism exists at nucleotide position 7761120 on chromosome B23 of the carp genome NC_056619.1;
[0021] The carp genome NC_056620.1 contains a T>C mutation polymorphism at nucleotide position 12425597 on chromosome B24;
[0022] The carp genome NC_056621.1 contains a G>C mutation polymorphism at nucleotide position 16283321 on chromosome B5.
[0023] The carp genome NC_056599.1 contains a G>C mutation polymorphism at nucleotide position 2626410 on chromosome B3.
[0024] The carp genome NC_056600.1 contains a T>A mutation polymorphism at nucleotide position 8600965 on chromosome B4.
[0025] The carp genome NC_056600.1 contains a T>C mutation polymorphism at nucleotide position 8325029 on chromosome B4.
[0026] The carp genome NC_056602.1 contains a T>A mutation polymorphism at nucleotide position 26046107 on chromosome B6.
[0027] The carp genome NC_056602.1 contains a G>A mutation polymorphism at nucleotide position 3369941 on chromosome B6.
[0028] The carp genome NC_056603.1 contains an A>G mutation polymorphism at nucleotide position 19808607 on chromosome B7.
[0029] The carp genome NC_056604.1 contains an A>C mutation polymorphism at nucleotide position 21307844 on chromosome B8.
[0030] Nucleotide position 12057742 on chromosome B8 of the carp genome NC_056604.1 contains a C>G mutation polymorphism;
[0031] The carp genome NC_056605.1 contains a T>A mutation polymorphism at nucleotide position 6208888 on chromosome B9.
[0032] Nucleotide position 491448 on chromosome NW_024879356.1 of the carp genome shows a C>G mutation polymorphism.
[0033] The NW_024879270.1 chromosome of the carp genome contains a C>T mutation polymorphism at nucleotide position 384685;
[0034] The NW_024879149.1 chromosome of the carp genome contains a G>A mutation polymorphism at nucleotide position 196646;
[0035] The NW_024872754.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 1182798;
[0036] The NW_024879270.1 chromosome of the carp genome contains a T>C mutation polymorphism at nucleotide position 611239.
[0037] The NW_024879270.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 616473;
[0038] The NW_024878853.1 chromosome of the carp genome contains a T>A mutation polymorphism at nucleotide position 9523;
[0039] The NW_024878853.1 chromosome of the carp genome contains an A>G mutation polymorphism at nucleotide position 9511;
[0040] The NW_024879195.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 762947;
[0041] The NW_024879195.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 762946;
[0042] The NW_024879211.1 chromosome of the carp genome contains a T>C mutation polymorphism at nucleotide position 519057;
[0043] The NW_024879211.1 chromosome of the carp genome contains a T>G mutation polymorphism at nucleotide position 519063;
[0044] The NW_024876713.1 chromosome of the carp genome contains a T>A mutation polymorphism at nucleotide position 1399.
[0045] In some embodiments, the steps for obtaining SNP genotype information include: filtering from at least one of carp resequencing data, reference genome sequence, and scaffold sequences not assembled into chromosomes to obtain a first SNP group; using the fixed effect of population structure and the random effect of kinship, screening from the first SNP group to obtain a second SNP group associated with the trait of total polyunsaturated fatty acid content in carp muscle; and performing conversion processing on the second SNP group to obtain the SNP gene information.
[0046] In some embodiments, a ridge regression model is trained using training feature data and training target data to obtain a prediction model. The prediction model characterizes the relationship between SNP genotype information and breeding values, whereby the breeding values characterize the total polyunsaturated fatty acid content information. This includes:
[0047] Establish the loss function for the ridge regression model;
[0048] The target ridge parameter and regression parameter of the ridge regression model are obtained based on the SNP genotype information and loss function of the training samples. The target ridge parameter and regression parameter include the ridge parameter and regression parameter.
[0049] The prediction model is obtained based on the target ridge parameters, regression parameters, and ridge regression model.
[0050] In some embodiments, a ridge regression model is trained using training feature data and training target data to obtain a prediction model. The prediction model characterizes the relationship between SNP genotype information and breeding values, whereby the breeding values characterize the total polyunsaturated fatty acid content information. This includes:
[0051] Establish the loss function for the ridge regression model;
[0052] Multiple sets of ridge parameters and regression parameters are obtained based on the loss function;
[0053] Based on the SNP genotype information, ridge parameters, regression parameters, and loss function of the training samples, the error of the SNP genotype information of the training samples for the ridge regression model under the ridge parameters and regression parameters is obtained. The error of the ridge regression model is the error between the predicted total polyunsaturated fatty acid content information obtained by the ridge regression model under the ridge parameters and regression parameters and the total polyunsaturated fatty acid content information of the training samples.
[0054] Based on the error corresponding to each set of ridge parameters and regression parameters, select the corresponding target ridge parameters and target regression parameters from multiple sets of ridge parameters and regression parameters;
[0055] The prediction model is obtained based on the target ridge parameters, target regression parameters, and ridge regression model.
[0056] In some embodiments, based on the SNP genotype information of the training samples, ridge parameters, regression parameters, and a loss function, the error of the SNP genotype information of the training samples for the ridge regression model under the ridge parameters and regression parameters is obtained. The error of the ridge regression model is the error between the predicted total polyunsaturated fatty acid content information obtained by the ridge regression model under the ridge parameters and regression parameters, and the total polyunsaturated fatty acid content information of the training samples; including:
[0057] Obtain one set of ridge parameters and regression parameters from multiple sets of ridge parameters and regression parameters;
[0058] The SNP genotype information of each training sample is calculated based on the loss function and applied to the sub-errors of the ridge parameters and regression parameters of that group in the ridge regression model.
[0059] Based on the sub-error of the training sample for the ridge regression model under the set of ridge parameters and regression parameters, obtain the SNP genotype information of the training sample for the mother error of the ridge regression model under the set of ridge parameters and regression parameters.
[0060] Obtain the parent error for each set of ridge parameters and regression parameters;
[0061] The error of the ridge regression model is determined based on the mother error of each set of ridge parameters and regression parameters.
[0062] In some embodiments, the prediction model is:
[0063]
[0064] in X is the breeding value of the i-th individual. ij : SNP genotype information of the i-th individual on the j-th marker, It is the effect value of the j-th label.
[0065] In some embodiments, when the breeding value of the sample to be tested is greater than or equal to 0.7, the sample to be tested originates from an individual or group of carp rich in unsaturated fatty acids.
[0066] In some embodiments, when the breeding value of the sample to be tested is less than 0.4, the sample to be tested originates from an individual or group of common carp.
[0067] In some embodiments, when the breeding value of the sample to be tested is between 0.4 and 0.7, the sample to be tested is derived from an individual or group of carp with moderate unsaturated fat content. Attached Figure Description
[0068] Figure 1 This is a flowchart illustrating the method for screening superior parents of carp with high polyunsaturated fatty acid content, as provided in the example.
[0069] Figure 2 This is a flowchart illustrating the SNP genotype information of the training samples provided in the example.
[0070] Figure 3 This is a schematic diagram of the steps for obtaining SNP genotype information provided in the example.
[0071] Figure 4 This is a flowchart illustrating step S20 provided in the embodiment.
[0072] Figure 5 A flowchart illustrating step S20 provided for other embodiments.
[0073] Figure 6 The flowchart of step S232 provided in the embodiment is shown. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. Reagents not specifically described in detail herein are all conventional reagents and are commercially available; methods not specifically described in detail are all conventional experimental methods and can be learned from the prior art.
[0075] The inventors of this application have creatively correlated the SNP locus genotype of carp with high polyunsaturated fatty acid content. By using the SNP locus genotype, the relative levels of high polyunsaturated fatty acid content among populations or individuals can be determined, thereby identifying dominant parents with high polyunsaturated fatty acid content. The method provided in this application can accurately and efficiently predict high polyunsaturated fatty acid content in carp. This method provides excellent materials for the identification and breeding of new carp germplasm.
[0076] In the application scenario of the method for screening carp with high polyunsaturated fatty acid content provided in this application embodiment, taking the device for screening carp with high polyunsaturated fatty acid content integrated into an electronic device as an example, the electronic device can acquire training feature data and training target data of the training set. The training feature data includes SNP genotype information of the training samples, and the training target data includes information on the total polyunsaturated fatty acid content of the training samples. The electronic device uses the training feature data and training target data to train a ridge regression model to obtain a prediction model. The prediction model represents the relationship between SNP genotype information and breeding value, and the breeding value represents information on the total polyunsaturated fatty acid content. The SNP genotype information of the sample to be tested is acquired. The breeding value of the sample to be tested is obtained based on the SNP genotype information and the prediction model. The carp with high polyunsaturated fatty acid content is screened based on the breeding value.
[0077] By performing these steps, the electronic device obtains the breeding value directly based on the SNP genotype information and prediction model of the sample to be tested, and then screens out superior parents with high total polyunsaturated fatty acid content from the sample to be tested, providing excellent materials for the identification and breeding of new carp germplasm.
[0078] Please see Figure 1 , Figure 1 This is a flowchart illustrating the method for screening carp parents with high polyunsaturated fatty acid content, as provided in this application embodiment. The specific process of the method for screening carp parents with high polyunsaturated fatty acid content provided in this application embodiment is as follows:
[0079] S10. Obtain the training feature data and training target data of the training set. The training feature data includes the SNP genotype information of the training samples, and the training target data includes the total content of polyunsaturated fatty acids of the training samples.
[0080] S20. The ridge regression model is trained using training feature data and training target data to obtain a prediction model. The prediction model represents the relationship between SNP genotype information and breeding value, and the breeding value represents the total content of polyunsaturated fatty acids.
[0081] S30. Obtain the SNP genotype information of the sample to be tested.
[0082] S40. Obtain the breeding value of the sample to be tested based on the SNP genotype information and prediction model.
[0083] S50. Select carp parents with high total polyunsaturated fatty acid content based on breeding values.
[0084] In step S10 of some embodiments, a training set containing one or more samples is provided for training the ridge regression model. Specifically, each training sample in this application includes training feature data and training target data of one or more carp.
[0085] The training feature data includes SNP genotype information of the training samples, where the SNP genotype information is the transformed value of the genotype information of significant SNP sites in one or more carp.
[0086] In one embodiment, such as Figure 2 As shown, the steps for obtaining SNP genotype information from training samples include:
[0087] S101. Filter from at least one of the following: resequencing data of carp, reference genome sequence, and scaffold sequence not assembled into chromosomes, to obtain the first SNP group;
[0088] S102. Transform the first SNP group to obtain SNP gene information.
[0089] S103. Using the fixed effect of population structure and the random effect of kinship, data of the second SNP group associated with the total content of polyunsaturated fatty acids in carp muscle were screened from the first SNP group.
[0090] In one specific embodiment, resequencing data of carp, a reference genome sequence (GCF_018340385.1), and some scaffold sequences not assembled into chromosomes were used. Genotype filtering was performed based on the criteria of sequencing depth 10X, integrity 0.1, minimum allele frequency ≥0.1, and deletion rate and heterozygosity <10%, to obtain the first SNP group within the genome. The first SNP group consists of multiple SNP loci. Using population structure as a fixed effect and kinship matrix as a random effect, an association analysis was performed on the SNP loci and the total polyunsaturated fatty acid content trait based on a mixed linear model to obtain the second SNP group. The second SNP group consists of multiple SNP loci significantly associated with the total polyunsaturated fatty acid content trait. There are three genotypes for SNP loci: A represents wild-type non-mutated nucleotides, and T represents mutated single nucleotides, such as AA, AT, and TT. The homozygous non-mutated genotype AA is converted to "-1", the heterozygous genotype AT is converted to "0", and the homozygous mutant genotype TT is converted to "1".
[0091] The basic form of the mixed linear model is: y = xβ + Zu + ∈. Here, y represents a vector of individual phenotypes, with a length equal to the number of individuals. x is the fixed effects matrix, including the intercept, SNP genotype, and covariates (e.g., population structure components). β is the fixed effects coefficient vector, calculated through statistical estimation. Z is the random effects matrix, usually a kinship matrix. u is the random effects vector, following a normal distribution. ∈ is the residual vector. The formula for calculating the kinship matrix is:
[0092]
[0093] K represents the kinship coefficient of an individual, m represents the total number of SNPs, g represents the encoding of the SNP allele, and p represents the frequency of the SNP allele.
[0094] In some embodiments, the S10 step includes wild-type non-mutant genotype information, heterozygous genotype information, and purified mutant genotype information, wherein the wild-type non-mutant genotype information, heterozygous genotype information, and purified mutant genotype information are different.
[0095] In some embodiments, as shown in Table 1, one test sample corresponds to 105 SNP genotypes. The genome (assembly number GCF_018340385.1) is the sequence number of the carp chromosome genome in GenBank, the position is the location of the SNP locus in that genome, the phenotypic explained variance rate is used to measure the degree of influence of gene variation on phenotypic variation, and polymorphism is the mutation polymorphism of the SNP locus. The part before the " / " is the reference genotype, and the part after the " / " is the mutant genotype. The dominant genotype represents the genotype with high total polyunsaturated fatty acid content, where the genotype value in parentheses is used, -1 is the homozygous non-mutant genotype, 0 is the heterozygous genotype, and 1 is the homozygous mutant genotype. The phenotypic mean of different genotypes in the population is the average D value of the genotype with high total polyunsaturated fatty acid content. The D value is the evaluation value of the total content of the seven target polyunsaturated fatty acids (linoleic acid, α-linolenic acid, γ-linolenic acid, eicosapentaenoic acid, dihomo-γ-linolenic acid, arachidonic acid, and docosahexaenoic acid).
[0096] Table 1. Information on SNP markers closely associated with high total polyunsaturated fatty acid content and phenotypic D values of their dominant genotypes.
[0097]
[0098]
[0099]
[0100]
[0101] The training target data includes information on the total polyunsaturated fatty acid content of the training samples. This information is the D-value used to evaluate the total polyunsaturated fatty acid content of one or more carp. Principal component analysis (PCA) of the polyunsaturated fatty acid content of each component in one or more carp is performed using the prcomp() function in R software to obtain a weighted composite value. This weighted composite value is then used to calculate the D-value for the total polyunsaturated fatty acid content of each individual carp.
[0102] The conversion process involves obtaining the total polyunsaturated fatty acid content of one or more carp and converting it into a total polyunsaturated fatty acid content evaluation value (D value), which ranges from 0 to 1. In some embodiments, each SNP locus in the test population corresponds to 2-3 genotypes (homozygous non-mutant genotype, heterozygous genotype, and homozygous mutant genotype), and the genotype information corresponding to each SNP locus is the average of the D values for the total polyunsaturated fatty acid content evaluation value in the test population.
[0103] Ridge regression is an improved linear regression estimation method widely used in data modeling scenarios with multicollinearity. The core purpose of ridge regression is to address the estimation instability problem encountered by ordinary least squares (OLS) when dealing with multicollinear data. When there is a high correlation (multicollinearity) among the independent variables, the variance of the regression coefficients estimated by OLS becomes very large, leading to unstable coefficient estimates that may deviate from the true values. Ridge regression addresses this by adding an L2 regularization term to the objective function of ordinary least squares, thus constraining the magnitude of the regression coefficients.
[0104] Specifically, in the ordinary linear regression model Y = βX + ε (where Y is the response variable, X is the independent variable matrix, β is the regression coefficient vector, and ε is the error term), the objective of ordinary least squares is to minimize the sum of squared residuals. The loss function of ridge regression is the sum of the sum of squared residuals and the regularization term. The sum of squared residuals is the sum of the squares of the differences between the model's brilliance values and the actual values, while the regularization term is the L2 norm of the model parameters.
[0105] In step S10, the training feature data is an n×m matrix, where n represents the number of samples and m represents the number of features. Each row in the matrix represents a training sample, and each column represents a feature. Additionally, a column of values can be added to the matrix as a bias term so that the ridge regression model includes an intercept term.
[0106] In step S10 of some embodiments, each row of the training feature data includes 106 features, of which 105 features are the D values of the phenotypic comprehensive evaluation values of the genotypes of the 105 SNP loci of the corresponding sample as shown in Table 1, and the features whose values are statistically obtained through library functions as bias terms. Correspondingly, the target vector is the total polyunsaturated fatty acid content information of the corresponding sample.
[0107] In step S10 provided in some embodiments, the training sample consists of genotype data and total polyunsaturated fatty acid content information of 280 carp samples as shown in Table 1.
[0108] In some embodiments, 280 carp individuals from different sources and without direct kinship were selected, and the content of fatty acids on the back was identified and quantitatively analyzed by GC-MS to obtain information on the total content of polyunsaturated fatty acids.
[0109] In one specific embodiment, the steps for obtaining the total polyunsaturated fatty acid content information of 280 individuals are as follows:
[0110] 1) Test materials
[0111] The carp population used in the experiment was selected from the Wudu Germplasm Resource Conservation Experimental Base of the Chinese Academy of Fishery Sciences.
[0112] To eliminate the influence of environmental and feed factors on the accumulation of unsaturated fatty acids in carp, the group was fed in the same manner. Muscle and liver samples were collected from 280 one-year-old carp and stored at -80°C.
[0113] Muscle samples from the same group were removed and stored in liquid nitrogen at low temperature. A mortar was pre-cooled with liquid nitrogen, maintaining a small amount of liquid nitrogen in the mortar at all times. A small amount of muscle tissue was taken into the mortar and ground into powder using a grinding stick. The samples were then gently ground with the grinding stick to ensure that all samples were thoroughly homogenized. The powder was then poured into sample tubes and stored at -80°C for later use in the determination of total unsaturated fatty acids.
[0114] 2) Detection of total polyunsaturated fatty acid content
[0115] Fatty acid methyl esterification: a) Add 2 mL of 0.5 mol / L sodium hydroxide methanol solution (1 g of solid sodium hydroxide dissolved in 50 mL of methanol, prepared and used in a fume hood) to a 15 mL centrifuge tube and weigh the total weight. Add 1 mL of powder (approximately 100 mg), shake to fully dissolve the sample, and weigh again. Calculate the difference between the two weighings to determine the actual powder weight. This helps determine the amount of hexane needed for final volume adjustment. (NaOH-methanol can dissolve fats, allowing fatty acids to undergo esterification with methanol); b) Heat the dissolved sample tube from step a at 100°C for 20 min in an oven, remove and cool (rapid cooling can be achieved in an ice box); c) Add 2 mL of boron trifluoride methanol solution (stored at 4°C, toxic, a catalyst for methyl esterification, used in a ventilated area), shake to mix, and heat at 100°C for 1 h in an oven; d) Cool to 30-40°C, add 1 mL of hexane, vortex for 30 s, immediately add 5 mL of saturated sodium chloride solution, and vortex again. e) After cooling and separating the layers, centrifuge at 3000g for 10 minutes, and transfer the upper layer of n-hexane to a clean 15mL centrifuge tube; f) Add another 1mL of n-hexane to wash the lower aqueous layer, vortex for 1 minute, centrifuge at 3000g for 10 minutes to separate the layers, and transfer the upper layer of n-hexane to the clean centrifuge tube from e, combining the supernatants; g) Blow with nitrogen until dry, and reconstitute with n-hexane (1mL n-hexane / 100mg powder) according to the ratio. Before placing the sample into the sample vial, the liquid should be filtered through a filter membrane (otherwise it will damage the machine); h) Add the reconstituted and quantified sample from g to a 2mL brown sample vial. For total volumes less than 250μl, an inner tube is required. Store at -80℃ before use.
[0116] The purified fatty acid methyl esters from the previous step were detected using a 7890 AGC System (Agilent Technologies, USA). Gas chromatography (GC) was employed to detect and analyze the fatty acid content in the muscle of the study population. The fatty acid components in the muscle were determined by comparing the retention time, peak area, and mass spectrometry fragments of the products with 37 fatty acid methyl ester standards. The relative content of each component was calculated using the normalization method based on the chromatographic peak area. The results showed that the total polyunsaturated fatty acid content per gram of muscle in the 280 carp population ranged from 12.82 to 134.49 mg, with an average of 64.73 ± 23.45 mg, and a D value of 0.12–0.94.
[0117] In some embodiments, such as Figure 3 As shown, the steps for obtaining SNP genotype information include:
[0118] S110, Obtain paired-terminal PE150 sequencing data of one carp population;
[0119] S120. Align the paired-end PE150 sequencing data with the reference genome to obtain SNP genotyping data;
[0120] S130. Filter the SNP genotyping data to obtain high-quality SNP loci;
[0121] S140. Principal component analysis and kinship analysis were performed on high-quality SNP loci to obtain the feature vectors (PCs) matrix of 280 carp individuals and the kinship coefficients (Kinship) matrix between each pair of individuals. These were used to control for false positives in GWAS analysis caused by population structure and to obtain filtered SNP loci.
[0122] S150. Correlation analysis was performed on the filtered SNP sites and total polyunsaturated fatty acid content information to obtain a threshold of <3.6×10. -5 The SNP sites, and the threshold <3.6×10 -5 SNP genotype information.
[0123] In one specific embodiment, the steps for obtaining SNP genotype information include: performing paired-end PE150 sequencing on 280 carp individuals using next-generation genome sequencing technology and the BGI DNBSEQ-T7 sequencing platform; performing sequence alignment (GCF_018340385.1) with the carp genome as a reference; using SAMtools software to detect population variation and obtain a VCF file storing SNP genotyping data; and using Plink and VCFtools software to filter genotypes according to the criteria of sequencing depth 10×, integrity 0.1, minimum allele frequency ≥0.1, and deletion rate and heterozygosity <10%, ultimately obtaining 40,965,117,262 high-quality SNP polymorphic sites. Principal component analysis and kinship analysis were performed on 40,965,117,262 high-quality SNP polymorphism loci using GCTA software. This yielded the eigenvector (PCs) matrix for all individuals and the kinship matrix (Kinship coefficients) between each pair of individuals. These were used to control for false positives in GWAS analysis due to population structure, resulting in filtered SNP loci. Based on these filtered SNP loci, total unsaturated fatty acids (TOFAs) were used as phenotypic data in carp. A single-environment model using GEMMA software was employed for genome-wide association analysis (GWAS) of TOFAs, with population structure and kinship as covariances, and a threshold set to <3.6 × 10⁻⁶. -5 We estimated the P-value and the phenotypic variance explained for each SNP locus.
[0124] In a specific step of obtaining SNP genotypic information, PLINK and SAMtools software were used to filter low-quality SNPs according to the following criteria: sequencing depth greater than 4, Q20 < 0.99, minimum allele frequency (MAF) < 5%, and deletion and heterozygosity rates < 10%. A total of 2,757,424 high-quality SNP polymorphism loci were obtained from 280 carp accessions for subsequent analysis. Principal component analysis (PCA) and kinship analysis were performed using PLINK and GEMMA software to obtain the top 5 principal component eigenvectors (PCs) matrix for all individuals and the kinship coefficients (Kinship matrices) between each pair of individuals. These were used to control for false positives in GWAS analysis caused by population structure.
[0125] In a specific step of obtaining SNP genotype information, based on the 2,757,424 SNP loci obtained through filtering, GWAS analysis was performed using GEMMA (https: / / github.com / genetics-statistics / GEMMA) software with a univariate linear mixed model (LMM) to conduct association analysis on the fatty acid content of total unsaturated fatty acids. Simultaneously, PCA and Kinshiap were used to control for population structure and kinship, with a significance detection threshold set at 3.3 × 10⁻⁶.-5 A total of 100 SNP sites significantly associated with total polyunsaturated fatty acids were detected (see [link to SNP analysis]). Figure 2 (As shown in the Manhattan plot), the phenotypic variation explained between 5.92% and 10.35%.
[0126] Some test cases also include step S20, which includes:
[0127] 1) Provide the total polyunsaturated fatty acid content information of 280 carp samples as target data, and provide the SNP genotype information of 280 carp samples as feature data. The SNP genotype information includes the genotype information of 105 SNP loci as shown in Table 1.
[0128] 2) Using 252 target data sets and 252 feature data sets as training sets, the following models were trained respectively: the optimal linear unbiased prediction model for the genome, the optimal linear unbiased prediction model for the extended genome, the ridge regression model, the elastic network, the support vector machine, the minimum absolute shrinkage and selection operator, the Bayesian ridge regression model, the random forest regression model, the Bayesian LASSO model, the LASSO model, the Bayesian A model, the Bayesian B model, and the Bayesian C model.
[0129] 3) Using 28 target data sets and 252 feature data sets as test sets, the trained genome best linear unbiased prediction model, extended genome best linear unbiased prediction model, ridge regression model, elastic network, support vector machine, minimum absolute shrinkage and selection operator, Bayesian ridge regression model, random forest regression model, Bayesian LASSO model, Bayesian A model, Bayesian B model and Bayesian C model were tested respectively.
[0130] 4) The ridge regression model was obtained by screening.
[0131] In step 3), the genome-wide breeding value (GEBV) is estimated in the validation population using genotype information and a prediction model. This step is repeated 10 times to eliminate sampling error, and the mean of the 10 GEBV replicates in the validation population is used as an indicator to evaluate the accuracy of genome-wide selection prediction. The total polyunsaturated fatty acid content of carp is analyzed using the aforementioned statistical model and SNP dataset.
[0132] Table 2 shows that, based on 105 SNP loci, the CV values for the predicted total polyunsaturated fatty acid content trait were all greater than 0.78. This CV represents the correlation coefficient between the predicted and actual values, where the predicted value is the breeding value. Most methods achieved CV values of 0.8 or higher for the breeding value, with only the random forest regression model showing an accuracy below 0.8. Among these ten methods, the ridge regression model was the best predictive model, outperforming the others. In Table 2, CV represents the accuracy of the breeding value based on genomic information, and SD_CV represents the standard deviation of the breeding value accuracy.
[0133] Table 2 shows the accuracy of genome-wide selection prediction using different models and SNP markers based on 280 individuals.
[0134]
[0135] In step S20 of some embodiments, the `ridge_model.fit(X_train, y_train)` method is called to input the SNP genotype information as feature data `X_train` and the total polyunsaturated fatty acid content information as target data `y_train` into the model for training. During training, the model obtains the ridge regression coefficients by minimizing the objective function (sum of squared residuals plus a regularization term) according to the set ridge parameters.
[0136] In some embodiments, such as Figure 4 As shown, step S20 includes:
[0137] S201. Establish the loss function for the ridge regression model;
[0138] S202. Obtain the target ridge parameter and regression parameter of the ridge regression model based on the SNP genotype information and loss function of the training samples. The target ridge parameter and regression parameter include the ridge parameter and regression parameter.
[0139] S203. Obtain the prediction model based on the target ridge parameters, regression parameters, and ridge regression model.
[0140] Among them, ridge parameters and regression parameters can include ridge parameters and regression parameters. Ridge regression adds a regularization term to the squared error. By determining the value of λ, a balance can be achieved between variance and bias: as λ increases, the model variance decreases while the bias increases.
[0141] In this embodiment, the loss function is the loss function of the ridge regression model, which is used to calculate the error between the output value of the ridge regression model on the sample and the true value.
[0142] In one embodiment, the loss function of the ridge regression model can be expressed as:
[0143]
[0144] Where n is the number of samples, m is the number of features, and y i x is the target value of the i-th sample. ij θ is the j-th feature value of the i-th sample. j λ is the regression parameter for the j-th feature, and λ is the ridge parameter that controls the strength of the regularization term. Adding this regularization term minimizes the sum of squared residuals while limiting the size of the regression coefficients, preventing them from becoming too large. This reduces model complexity and improves model stability and generalization ability.
[0145] In some embodiments, the loss function of the ridge regression model can be modified to obtain a regression parameter acquisition function, and then the ridge parameters and regression parameters can be obtained based on the regression parameter acquisition function.
[0146] For example, taking the derivative of this loss function yields:
[0147] 2X T (Y-XW)-2λW, where X is the matrix or vector of characteristic x, X T Let X be the transpose of y, and Y be a matrix or vector of y; then, let 2X T With (Y-XW)-2λW equal to zero, we can obtain the following regression parameter acquisition function:
[0148]
[0149] in, These are the regression parameters to be solved.
[0150] After obtaining the regression parameter acquisition function, the regression parameters can be calculated based on the formula and the training set feature data to finally obtain the ridge parameter and regression parameter.
[0151] In some embodiments, such as Figure 5 As shown, step S20 includes:
[0152] S211. Establish the loss function for the ridge regression model.
[0153] S222. Obtain multiple sets of ridge parameters and regression parameters based on the loss function.
[0154] S232. Based on the SNP genotype information, ridge parameters, regression parameters, and loss function of the training samples, obtain the error of the SNP genotype information of the training samples for the ridge regression model under the ridge parameters and regression parameters. The error of the ridge regression model is the error between the predicted total polyunsaturated fatty acid content information obtained by the ridge regression model under the ridge parameters and regression parameters and the total polyunsaturated fatty acid content information of the training samples.
[0155] S242. Based on the error corresponding to each set of ridge parameters and regression parameters, select the corresponding target ridge parameters and target regression parameters from multiple sets of ridge parameters and regression parameters.
[0156] S252. Obtain the prediction model based on the target ridge parameters, target regression parameters, and ridge regression model.
[0157] In some steps S232, such as Figure 6 As shown, it includes:
[0158] S233. Obtain a set of ridge parameters and regression parameters from multiple sets of ridge parameters and regression parameters.
[0159] S234. Calculate the SNP genotype information of each training sample based on the loss function and apply it to the sub-errors of the ridge parameters and regression parameters of that group in the ridge regression model.
[0160] S235. Based on the sub-error of the training sample for the ridge regression model under the set of ridge parameters and regression parameters, obtain the mother error of the ridge regression model for the SNP genotype information of the training sample under the set of ridge parameters and regression parameters.
[0161] S236. Obtain the mother error of each group of ridge parameters and regression parameters according to steps S233, S234 and S235.
[0162] S237. The error of the ridge regression model determined based on the mother error of each set of ridge parameters and regression parameters.
[0163] In step S234 provided in some embodiments, the sub-error of the ridge regression model corresponding to each genotype for each SNP locus is calculated.
[0164] In step S237 provided in some embodiments, in order to improve parameter accuracy and prediction precision, the average value of the mother error of each group of ridge parameters and regression parameters is calculated and used as the ridge regression model error.
[0165] In one embodiment, in order to improve parameter accuracy and prediction precision, after obtaining the mother error corresponding to each set of ridge parameters and regression parameters, the set of ridge parameters and regression parameters with the minimum mother error can be used as the target ridge parameters and target regression parameters of the ridge regression model, i.e., the target parameters.
[0166] The corresponding ridge regression model, or prediction model, can be mathematically expressed as follows:
[0167]
[0168] in X is the breeding value of the i-th individual. ijIt represents the SNP genotype information of the i-th individual at the j-th marker (usually encoded as 0, 1, 2, representing the number of alleles). It is the effect value of the j-th label.
[0169] In step S50 provided in some embodiments, when the breeding value of the sample to be tested is greater than or equal to 0.7, the sample to be tested originates from an individual or group of carp rich in unsaturated fatty acids; when the breeding value of the sample to be tested is less than 0.4, the sample to be tested originates from an individual or group of common carp; when the breeding value of the sample to be tested is between 0.4 and 0.7, the sample to be tested originates from an individual or group of carp with moderate unsaturated fat content.
[0170] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for screening superior paisley parents with high levels of polyunsaturated fatty acids, comprising: Obtain training feature data and training target data for the training set. The training feature data includes SNP genotype information of the training samples, and the SNP loci in the SNP genotype information include all of the following loci: The carp genome NC_056572.1 contains a C>T mutation polymorphism at nucleotide position 35093078 on chromosome A1. The carp genome NC_056572.1 contains a G>C mutation polymorphism at nucleotide position 171834 on chromosome A1. The carp genome NC_056572.1 contains a C>T mutation polymorphism at nucleotide position 11053474 on chromosome A1. The carp genome NC_056572.1 contains a T>C mutation polymorphism at nucleotide position 26618712 on chromosome A1. The carp genome NC_056583.1 contains a C>T mutation polymorphism at nucleotide position 20260097 on chromosome A12. Nucleotide position 12052513 on chromosome A12 of the carp genome NC_056583.1 contains a C>G mutation polymorphism; A T>C mutation polymorphism exists at nucleotide position 20140419 on chromosome A12 of the carp genome NC_056583.1; The carp genome NC_056585.1 contains a G>A mutation polymorphism at nucleotide position 24485947 on chromosome A14. The carp genome NC_056585.1 contains a G>A mutation polymorphism at nucleotide position 18653036 on chromosome A14. The carp genome NC_056585.1 contains a T>C mutation polymorphism at nucleotide position 18984637 on chromosome A14. The carp genome NC_056585.1 contains a T>C mutation polymorphism at nucleotide position 18984627 on chromosome A14. A G>T mutation polymorphism exists at nucleotide position 18984619 on chromosome A14 of the carp genome NC_056585.1; The carp genome NC_056585.1 contains a G>C mutation polymorphism at nucleotide position 23968026 on chromosome A14. A G>C mutation polymorphism exists at nucleotide position 21928001 on chromosome A15 of the carp genome NC_056586.1; The carp genome NC_056586.1 contains an A>T mutation polymorphism at nucleotide position 802752 on chromosome A15. The carp genome NC_056586.1 contains an A>T mutation polymorphism at nucleotide position 802762 on chromosome A15. The carp genome NC_056586.1 contains an A>T mutation polymorphism at nucleotide position 802764 on chromosome A15. A T>C mutation polymorphism exists at nucleotide position 6850717 on chromosome A16 of the carp genome NC_056587.1; The carp genome NC_056587.1 contains a C>T mutation polymorphism at nucleotide position 5355826 on chromosome A16; The carp genome NC_056587.1 contains an A>T mutation polymorphism at nucleotide position 5354912 on chromosome A16. The carp genome NC_056587.1 contains a T>C mutation polymorphism at nucleotide position 5354913 on chromosome A16. The carp genome NC_056587.1 contains a C>T mutation polymorphism at nucleotide position 8296639 on chromosome A16; A G>A mutation polymorphism exists at nucleotide position 4799340 on chromosome A16 of the carp genome NC_056587.1; The A>G mutation polymorphism exists at nucleotide position 4799336 on chromosome A16 of the carp genome NC_056587.1; The carp genome NC_056587.1 contains a T>C mutation polymorphism at nucleotide position 3989726 on chromosome A16; The carp genome NC_056587.1 contains a G>C mutation polymorphism at nucleotide position 4939675 on chromosome A16. The carp genome NC_056589.1 contains a G>A mutation polymorphism at nucleotide position 11926439 on chromosome A18; The carp genome NC_056589.1 contains a T>G mutation polymorphism at nucleotide position 24946240 on chromosome A18. The carp genome NC_056573.1 contains a C>A mutation polymorphism at nucleotide position 2446713 on chromosome A2. The carp genome NC_056573.1 contains a C>G mutation polymorphism at nucleotide position 2111955 on chromosome A2. Nucleotide position 1842431 on chromosome A2 of the carp genome NC_056573.1 contains a C>T mutation polymorphism; The carp genome NC_056591.1 contains a G>A mutation polymorphism at nucleotide position 10731194 on chromosome A20. A G>T mutation polymorphism exists at nucleotide position 11906615 on chromosome A21 of the carp genome NC_056592.1; A T>A mutation polymorphism exists at nucleotide position 13620002 on chromosome A21 of the carp genome NC_056592.1; A T>C mutation polymorphism exists at nucleotide position 19643501 on chromosome A23 of the carp genome NC_056594.1; The carp genome NC_056594.1 contains a G>A mutation polymorphism at nucleotide position 23931849 on chromosome A23; A T>C mutation polymorphism exists at nucleotide position 18100115 on chromosome A3 of the carp genome NC_056574.1; The carp genome NC_056574.1 contains a T>G mutation polymorphism at nucleotide position 10859944 on chromosome A3. The carp genome NC_056574.1 contains a T>G mutation polymorphism at nucleotide position 10859946 on chromosome A3. The carp genome NC_056575.1 contains a C>T mutation polymorphism at nucleotide position 8426250 on chromosome A4. The carp genome NC_056575.1 contains a T>C mutation polymorphism at nucleotide position 8425648 on chromosome A4. The carp genome NC_056575.1 contains an A>T mutation polymorphism at nucleotide position 20011914 on chromosome A4. The carp genome NC_056576.1 contains a T>A mutation polymorphism at nucleotide position 39119153 on chromosome A5; The carp genome NC_056576.1 contains a G>A mutation polymorphism at nucleotide position 39119148 on chromosome A5. The carp genome NC_056576.1 contains a T>C mutation polymorphism at nucleotide position 6521311 on chromosome A5. The carp genome NC_056576.1 contains a G>C mutation polymorphism at nucleotide position 35333197 on chromosome A5. The carp genome NC_056577.1 contains a C>T mutation polymorphism at nucleotide position 15309949 on chromosome A6. The carp genome NC_056577.1 contains a T>C mutation polymorphism at nucleotide position 5360090 on chromosome A6. The carp genome NC_056577.1 contains a C>T mutation polymorphism at nucleotide position 32004651 on chromosome A6. The carp genome NC_056578.1 contains an A>T mutation polymorphism at nucleotide position 40388036 on chromosome A7. The carp genome NC_056578.1 contains a T>G mutation polymorphism at nucleotide position 36226528 on chromosome A7. The carp genome NC_056578.1 contains a C>T mutation polymorphism at nucleotide position 36226538 on chromosome A7. The carp genome NC_056578.1 contains an A>G mutation polymorphism at nucleotide position 3036102 on chromosome A7. The carp genome NC_056579.1 contains a T>C mutation polymorphism at nucleotide position 4467840 on chromosome A8. The carp genome NC_056579.1 contains an A>T mutation polymorphism at nucleotide position 16846973 on chromosome A8; The carp genome NC_056580.1 contains a G>A mutation polymorphism at nucleotide position 3296980 on chromosome A9; The carp genome NC_056580.1 contains a T>C mutation polymorphism at nucleotide position 27008604 on chromosome A9. The carp genome NC_056580.1 contains a C>T mutation polymorphism at nucleotide position 27008612 on chromosome A9. The carp genome NC_056580.1 contains a C>A mutation polymorphism at nucleotide position 1444937 on chromosome A9; The carp genome NC_056580.1 contains a C>T mutation polymorphism at nucleotide position 1737445 on chromosome A9. The carp genome NC_056580.1 contains a G>A mutation polymorphism at nucleotide position 4944998 on chromosome A9; The carp genome NC_056580.1 contains a G>A mutation polymorphism at nucleotide position 4944965 on chromosome A9; The carp genome NC_056580.1 contains a T>C mutation polymorphism at nucleotide position 4944960 on chromosome A9; The carp genome NC_056580.1 contains a T>C mutation polymorphism at nucleotide position 4944980 on chromosome A9; A T>C mutation polymorphism exists at nucleotide position 7101839 on chromosome B11 of the carp genome NC_056607.1; The carp genome NC_056607.1 contains a T>G mutation polymorphism at nucleotide position 9756853 on chromosome B11; The carp genome NC_056607.1 contains an A>T mutation polymorphism at nucleotide position 8439254 on chromosome B11. The carp genome NC_056608.1 contains a T>A mutation polymorphism at nucleotide position 26329356 on chromosome B12. Nucleotide position 19109755 on chromosome B14 of the carp genome NC_056610.1 contains a C>G mutation polymorphism; The carp genome NC_056611.1 contains a T>A mutation polymorphism at nucleotide position 25384747 on chromosome B15; The carp genome NC_056611.1 contains an A>T mutation polymorphism at nucleotide position 25384749 on chromosome B15. A T>A mutation polymorphism exists at nucleotide position 21342737 on chromosome B18 of the carp genome NC_056614.1; The carp genome NC_056615.1 contains a C>T mutation polymorphism at nucleotide position 26222460 on chromosome B19. The carp genome NC_056615.1 contains a T>A mutation polymorphism at nucleotide position 5963093 on chromosome B19; Nucleotide position 1978925 on chromosome B2 of the carp genome NC_056598.1 contains a C>G mutation polymorphism; The carp genome NC_056598.1 contains a T>A mutation polymorphism at nucleotide position 31537582 on chromosome B2. The carp genome NC_056616.1 contains an A>G mutation polymorphism at nucleotide position 30216632 on chromosome B20. The carp genome NC_056616.1 contains an A>G mutation polymorphism at nucleotide position 25570376 on chromosome B20; The carp genome NC_056618.1 contains an A>G mutation polymorphism at nucleotide position 18572947 on chromosome B22. Nucleotide position 22921065 on chromosome B23 of the carp genome NC_056619.1 contains a C>A mutation polymorphism; A G>C mutation polymorphism exists at nucleotide position 7761120 on chromosome B23 of the carp genome NC_056619.1; The carp genome NC_056620.1 contains a T>C mutation polymorphism at nucleotide position 12425597 on chromosome B24; The carp genome NC_056621.1 contains a G>C mutation polymorphism at nucleotide position 16283321 on chromosome B5. The carp genome NC_056599.1 contains a G>C mutation polymorphism at nucleotide position 2626410 on chromosome B3. The carp genome NC_056600.1 contains a T>A mutation polymorphism at nucleotide position 8600965 on chromosome B4. The carp genome NC_056600.1 contains a T>C mutation polymorphism at nucleotide position 8325029 on chromosome B4. The carp genome NC_056602.1 contains a T>A mutation polymorphism at nucleotide position 26046107 on chromosome B6. The carp genome NC_056602.1 contains a G>A mutation polymorphism at nucleotide position 3369941 on chromosome B6. The carp genome NC_056603.1 contains an A>G mutation polymorphism at nucleotide position 19808607 on chromosome B7. The carp genome NC_056604.1 contains an A>C mutation polymorphism at nucleotide position 21307844 on chromosome B8. Nucleotide position 12057742 on chromosome B8 of the carp genome NC_056604.1 contains a C>G mutation polymorphism; The carp genome NC_056605.1 contains a T>A mutation polymorphism at nucleotide position 6208888 on chromosome B9. Nucleotide position 491448 on chromosome NW_024879356.1 of the carp genome shows a C>G mutation polymorphism. The NW_024879270.1 chromosome of the carp genome contains a C>T mutation polymorphism at nucleotide position 384685; The NW_024879149.1 chromosome of the carp genome contains a G>A mutation polymorphism at nucleotide position 196646; The NW_024872754.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 1182798; The NW_024879270.1 chromosome of the carp genome contains a T>C mutation polymorphism at nucleotide position 611239. The NW_024879270.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 616473; The NW_024878853.1 chromosome of the carp genome contains a T>A mutation polymorphism at nucleotide position 9523; The NW_024878853.1 chromosome of the carp genome contains an A>G mutation polymorphism at nucleotide position 9511; The NW_024879195.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 762947; The NW_024879195.1 chromosome of the carp genome contains an A>T mutation polymorphism at nucleotide position 762946; The NW_024879211.1 chromosome of the carp genome contains a T>C mutation polymorphism at nucleotide position 519057; The NW_024879211.1 chromosome of the carp genome contains a T>G mutation polymorphism at nucleotide position 519063; Nucleotide position 1399 of chromosome NW_024876713.1 in the carp genome exhibits a T>A mutation polymorphism; the training target data includes information on the total polyunsaturated fatty acid content of the training samples; The ridge regression model is trained using training feature data and training target data to obtain a prediction model. The prediction model represents the relationship between SNP genotype information and breeding value, and the breeding value represents the total content of polyunsaturated fatty acids. Obtain the SNP genotype information of the sample to be tested; The breeding value of the sample to be tested is obtained based on the SNP genotype information and the prediction model. Based on breeding values, select superior carp parents with high total polyunsaturated fatty acid content.
2. The method according to claim 1, further comprising: Provide a training set containing one or more samples, each training sample including training feature data and training target data of one or more carp.
3. The method according to claim 1, wherein the SNP genotype information is the transformation value of one or more SNP loci genotype information of one or more carp.
4. The method according to claim 1, wherein the SNP genotype information includes wild-type non-mutant genotype information, heterozygous genotype information, and purified mutant genotype information, wherein the wild-type non-mutant genotype information, heterozygous genotype information, and purified mutant genotype information are different.
5. The method according to claim 1, wherein a ridge regression model is trained using training feature data and training target data to obtain a prediction model, wherein the prediction model characterizes the relationship between SNP genotype information and breeding values, and the breeding values characterize the total content of polyunsaturated fatty acids; comprising: Establish the loss function for the ridge regression model; The target ridge parameters and regression parameters of the ridge regression model are obtained based on the SNP genotype information and loss function of the training samples. The prediction model is obtained based on the target ridge parameters, regression parameters, and ridge regression model.
6. The method according to claim 1, wherein a ridge regression model is trained using training feature data and training target data to obtain a prediction model, wherein the prediction model characterizes the relationship between SNP genotype information and breeding value, and the breeding value characterizes the total polyunsaturated fatty acid content information; comprising: Establish the loss function for the ridge regression model; Multiple sets of ridge parameters and regression parameters are obtained based on the loss function; Based on the SNP genotype information, ridge parameters, regression parameters, and loss function of the training samples, the error of the SNP genotype information of the training samples for the ridge regression model under the ridge parameters and regression parameters is obtained. The error of the ridge regression model is the error between the predicted total polyunsaturated fatty acid content information obtained by the ridge regression model under the ridge parameters and regression parameters and the total polyunsaturated fatty acid content information of the training samples. Based on the error corresponding to each set of ridge parameters and regression parameters, select the corresponding target ridge parameters and target regression parameters from multiple sets of ridge parameters and regression parameters; The prediction model is obtained based on the target ridge parameters, target regression parameters, and ridge regression model.
7. The method according to claim 6, wherein the error of the SNP genotype information of the training samples for the ridge regression model under the ridge parameters and regression parameters is obtained based on the SNP genotype information of the training samples, ridge parameters, regression parameters, and loss function, wherein the error of the ridge regression model is the error between the predicted total polyunsaturated fatty acid content information obtained by the ridge regression model under the ridge parameters and regression parameters and the total polyunsaturated fatty acid content information of the training samples; including: Obtain one set of ridge parameters and regression parameters from multiple sets of ridge parameters and regression parameters; The SNP genotype information of each training sample is calculated based on the loss function and applied to the sub-errors of the ridge parameters and regression parameters of that group in the ridge regression model. Based on the sub-error of the training sample for the ridge regression model under the set of ridge parameters and regression parameters, obtain the SNP genotype information of the training sample for the mother error of the ridge regression model under the set of ridge parameters and regression parameters. Obtain the parent error for each set of ridge parameters and regression parameters; The error of the ridge regression model is determined based on the mother error of each set of ridge parameters and regression parameters.
8. The method according to claim 1, wherein the prediction model is: in Xij represents the breeding value of the i-th individual, where Xij is the SNP genotype information of the i-th individual at the j-th marker. It is the effect value of the j-th label.
9. According to the method of claim 1, when the breeding value of the sample to be tested is greater than or equal to 0.7, the sample to be tested originates from an individual or group of carp rich in unsaturated fatty acids. When the breeding value of the sample to be tested is less than 0.4, the sample to be tested is derived from an individual or group of common carp. When the breeding value of the sample to be tested is between 0.4 and 0.7, the sample to be tested is derived from an individual or group of carp with moderate unsaturated fat content.
Citation Information
Patent Citations
Application clearing method and device, storage medium and electronic equipment
CN107807730A
Whole-genome selective breeding method and apparatus
CN111524545A