Genome information-based local pig optimal heterosis prediction method

By constructing a genomic information prediction method for local pigs, we can accurately identify purebred parents, simulate the hybridization process, separate dominant and recessive effects, and quantify non-additive effects. This solves the problem of bias in the assessment of genetic relationships in existing technologies, and enables efficient and accurate selection of hybrid combinations to stably breed ideal offspring.

CN121601040APending Publication Date: 2026-03-03CHONGQING ACAD OF ANIMAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511776720.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing methods for predicting the optimal heterosis of local pigs rely on pedigree records and limited genetic markers, which cannot fully reflect the complex genetic composition of the whole genome. This leads to biases in the assessment of genetic relationships, ignores the effects of dominance, overdominance, and epistasis, and results in blind selection of hybrid combinations. It is difficult to stably breed ideal offspring with highly aggregated traits, thus limiting the efficiency of the utilization of local pig germplasm resources.

Method used

Based on genomic information, we collected genomic data from Duroc pigs, Rongchang pigs, and Large White pigs. We constructed a purebred population using SNP genotype datasets, simulated the hybridization process, generated a phenotypic matching classification table, separated dominant and recessive weight parameters, quantified non-additive effects, constructed a dominant and recessive weight parameter table, and superimposed it on the traditional additive breeding value model to achieve direct quantitative prediction of heterosis.

Benefits of technology

Accurately identify purebred parents, simulate large-scale hybridization and propagation, separate dominant and recessive effects, improve the accuracy and efficiency of hybrid combination selection, and stably cultivate ideal offspring with highly integrated traits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601040A_ABST
    Figure CN121601040A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of genome information analysis, in particular to a local pig optimal heterosis prediction method based on genome information. According to the method, pure parents with clear genetic backgrounds are accurately identified and screened out by utilizing high-density SNP marker information covering a whole genome, a reference population is constructed, large-scale hybridization propagation is simulated, structured virtual population data is formed, and multi-character phenotype data of filial generations are standardized and weighted, so that the genetic backgrounds of the filial generations are optimized, and the genetic backgrounds of the filial generations are optimized. The method comprises the following steps of: dividing a plurality of characteristic groups such as a production performance type and a reproductive performance type, and analyzing the correlation between phenotypic deviation and genotype in each group, so as to separate and quantify dominant effect caused by heterozygous genotype and recessive effect caused by homozygous genotype, construct explicit and recessive weight parameters, and determine the genetic relationship between the dominant effect and the recessive effect; finally, the non-additive effect parameters are superposed to a traditional additive breeding value model, direct quantitative prediction of heterosis is achieved, and therefore the accuracy and efficiency of hybrid combination matching are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genomic information analysis technology, and in particular to a method for predicting optimal heterosis in local pigs based on genomic information. Background Technology

[0002] The field of genome information analysis technology involves the analysis and comparison of the whole genome sequence of an organism through high-throughput sequencing, genotyping, gene annotation, and bioinformatics algorithms, thereby revealing gene variation, structural polymorphism, and the intrinsic relationship between genetic information and phenotype. This includes the acquisition of genome sequencing data, sequence alignment, detection of single nucleotide polymorphisms, gene function annotation, and population genetic structure analysis.

[0003] Among them, the optimal heterosis prediction method for local pigs refers to the process of predicting and selecting the genetic potential of crosses between different pig breeds or lines based on phenotypic data of breeding populations and limited genetic marker information. Such methods typically rely on pedigree records and phenotypic determination, predicting the magnitude of heterosis by calculating genetic distance between parents or estimating genetic similarity.

[0004] Existing methods for predicting optimal heterosis in local pigs rely on pedigree records and limited genetic markers. They estimate the genetic distance between parents to predict the potential of hybrid offspring. This approach cannot fully reflect the complex genetic composition of the entire genome, and the accuracy of pedigree records is difficult to guarantee, leading to biases in the assessment of genetic relationships. More importantly, simply linking genetic distance with heterosis ignores non-additive effects such as dominance, overdominance, and epistasis. These effects are the biological basis of heterosis. Therefore, the prediction results cannot accurately reveal the genetic gains generated by interactions at specific gene loci, making the selection of hybrid combinations somewhat blind and making it difficult to stably breed ideal offspring with highly aggregated traits, thus limiting the utilization efficiency of local pig germplasm resources. Summary of the Invention

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for predicting optimal heterosis in local pigs based on genomic information, comprising the following steps: S1: Collect genomic information of Duroc pigs, Rongchang pigs and Large White pigs and perform quality control analysis. Verify the consistency and numbering of the genomic information after analysis and generate SNP genotype dataset; S2: Determine the bloodline origin of each type of pig based on the SNP genotype dataset and screen the corresponding purebred population data. Simulate the hybridization process based on the purebred population data to form the propagation population data. S3: Based on the propagation population data, construct a list of hybridization combinations for each type of pig, match it with the corresponding genomic information in the SNP genotype dataset, extract multiple phenotypic feature groups, and generate a phenotypic fitting classification table; S4: Determine the intra-group phenotypic bias contribution of each phenotypic feature in the phenotypic fitting classification table, and construct an explicit and implicit weight parameter table based on the magnitude of the bias contribution. S5: Combining the list of hybridization combinations for each type of pig with the dominant-recessive weight parameter table, calculate the predicted breeding value and the actual breeding value for each hybridization combination, analyze the correlation between the predicted breeding value and the actual breeding value, and output the prediction result of the optimal hybridization combination.

[0006] As a further embodiment of the present invention, the SNP genotype dataset includes SNP marker locus numbers, allele composition, and chromosome position parameters; the propagation population data includes purebred population genetic structure information, hybridization simulation recombination information, and offspring phenotypic parameters; the phenotypic fit classification table includes phenotypic feature group categories, phenotypic extended value ranges, and hybridization combination correspondences; the dominant-recessive weight parameter table includes dominant effect parameters, recessive effect parameters, and phenotypic deviation weight values; and the optimal hybridization combination prediction result includes hybridization combination number, predicted breeding value, actual breeding value, and correlation coefficient.

[0007] As a further aspect of the present invention, step S1 specifically comprises: S101: Collect genomic information data of Duroc pigs, Rongchang pigs and Large White pigs, perform signal recognition and data analysis on the genomic information data, extract the original probe signal intensity of each SNP marker site, determine the corresponding allele composition based on the probe signal intensity, obtain the allele type and chromosome position parameters of each SNP marker site, and establish the allele composition of the marker site. S102: For each of the marked sites, remove sites with missing probe signals or incorrect localization, calculate the minimum allele frequency and deletion rate for the remaining SNP marked sites, and compare the calculated frequency and deletion rate with the set screening threshold to obtain the SNP site set after quality control screening. S103: Call the chromosomal position parameters of each SNP marker locus in the SNP locus set after quality control screening, correct the numbers of all SNP marker loci that passed the screening, and integrate the corrected numbers with the corresponding allele composition to generate an SNP genotype dataset.

[0008] As a further aspect of the present invention, step S2 specifically comprises: S201: Based on the SNP genotype dataset, calculate the pedigree ratio parameters of Duroc pigs, Rongchang pigs and Large White pigs, and compare the allelic composition of each SNP marker locus with the allelic frequency distribution of the reference population to determine the pedigree. Then, compare the calculated pedigree ratio parameters with the preset pedigree threshold, delete samples below the pedigree threshold, and obtain purebred population data. S202: Based on the Bayesian genetic simulation method, the hybridization process between Duroc pigs, Rongchang pigs and Large White pigs is simulated. The virtual offspring gene structure is established by using the SNP marker locus allele composition and linkage recombination of the purebred population data. S203: Obtain the phenotypic information of the gene structure of the virtual offspring, calculate the real breeding value of the virtual offspring, and redistribute the purebred individuals of Duroc pigs, Rongchang pigs and Large White pigs as parents according to the real breeding value. Combine the redistributed parent and offspring data to generate propagation population data.

[0009] As a further aspect of the present invention, step S3 specifically comprises: S301: Based on the aforementioned propagation population data, select boars and sows from three breeds: Duroc, Rongchang, and Large White. According to the principle of matching bloodline ratio and breeding value, establish multiple hybridization combinations and integrate them into a hybridization combination list. S302: Based on the list of hybrid combinations, collect phenotypic observation data of each hybrid individual on the three traits of growth rate, backfat thickness and lean meat percentage, and match them with the SNP marker locus information of the same numbered hybrid individual in the SNP genotype dataset. Perform standardization and weighted summation operations on the phenotypic observation data of the matched hybrid individuals to obtain the phenotypic extended value. S303: Compare the phenotypic extended value with the phenotypic fit threshold, classify the hybrid individuals into three phenotypic characteristics: production performance type, reproductive performance type, and growth balance type, and generate a phenotypic fit classification table.

[0010] As a further aspect of the present invention, the principle for matching the pedigree ratio and breeding value is specifically as follows: Set the target pedigree ratio range for Duroc, Rongchang and Large White pigs in the crossbreeding combination; Obtain the breeding value of each boar and sow in the breeding population data. The breeding value includes the breeding value related to growth rate, backfat thickness and lean meat percentage. Select parental hybridization combinations that meet the requirements based on the target bloodline ratio range; For the selected parental combinations, calculate the expected breeding value of the offspring of each potential hybrid combination; The expected breeding value is compared with the preset optimal range of breeding values, and potential hybrid combinations whose expected breeding value falls within the optimal range of breeding values ​​are identified as hybrid combinations.

[0011] As a further aspect of the present invention, step S4 specifically comprises: S401: Calculate the intra-group phenotypic differences of the three characteristic groups (production performance type, reproductive performance type and growth balance type) in the phenotypic matching classification table respectively, determine the intra-group phenotypic bias contribution of each phenotypic characteristic, and obtain the intra-group phenotypic bias value. S402: Obtain the expression status of SNP marker sites for each hybrid offspring in the hybridization combination list, and calculate the offset contribution of heterozygous alleles to the phenotype and the independent offset contribution of recessive homozygous alleles in combination with the phenotypic deviation value within the group, and obtain the set of dominant effect parameters and recessive effect parameters. S403: Based on the independent bias contribution reflected by the set of dominant and recessive effect parameters, directional adjustments are made to the weights of the dominant and recessive effect parameters to represent the gene interaction characteristics of the three-variety hybrid population, and a dominant-recessive weight parameter table is constructed.

[0012] As a further aspect of the present invention, the process of calculating the shift contribution of the heterozygous allele to the phenotype and the independent bias contribution of the recessive homozygous allele specifically involves: For the three characteristic groups of production performance type, reproductive performance type and growth balance type, a linear regression model is constructed for each characteristic group. The linear regression model uses the intragroup phenotypic deviation as the dependent variable and the expression status of SNP marker sites in each hybrid offspring as the independent variable. The linear regression model is solved by joint error minimization. The regression coefficient associated with the heterozygous allele encoding is used as the shift contribution of the heterozygous allele to the phenotype, and the dominant effect parameter is obtained. The regression coefficients associated with the coding of the recessive homozygous alleles are used as the independent bias contribution of the recessive homozygous alleles to obtain the recessive effect parameters.

[0013] As a further aspect of the present invention, step S5 specifically comprises: S501: Combining the list of hybridization combinations and the table of dominant and recessive weight parameters, the predicted breeding value of each hybridization combination is calculated by superimposing the dominant effect parameter and the recessive effect parameter into the SNP locus effect of the hybrid individual. S502: Statistically calculate the actual breeding value corresponding to each hybrid combination in the propagation population data, and perform Pearson correlation coefficient and Spearman rank correlation coefficient calculation with the predicted breeding value to obtain the average predicted index of the hybrid combination; S503: Based on the average prediction index of the hybridization combination, sort the various mating types of Duroc pigs, Rongchang pigs and Large White pigs, and output the prediction result of the optimal hybridization combination.

[0014] As a further aspect of the present invention, the process of superimposing the dominant effect parameter and the recessive effect parameter into the SNP locus effect of the hybrid individual specifically involves: Obtain all SNP genotypes for each hybrid individual in the hybrid combination; For each SNP marker locus in each hybrid individual, determine whether the allelic composition of the SNP marker locus is heterozygous or recessive homozygous; If the alleles of the SNP marker site are heterozygous, then the dominant effect parameter corresponding to the dominant-recessive weight parameter table is summed and superimposed with the SNP site effect; If the allele of the SNP marker site becomes recessive homozygous, then the recessive effect parameter corresponding to the dominant-recessive weight parameter table is summed and superimposed with the SNP site effect.

[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, high-density SNP marker information covering the entire genome is used to accurately identify and screen purebred parents with clear genetic backgrounds. A reference population is constructed and large-scale hybridization is simulated to form structured virtual population data. By standardizing and weighting the multi-phenotypic data of the hybrid offspring, they are divided into multiple characteristic groups such as production performance type and reproductive performance type. Within each group, the correlation between phenotypic deviation and genotype is analyzed to separate and quantify the dominant effect of heterozygous genotype and the recessive effect of homozygous genotype. Dominant and recessive weight parameters are constructed. Finally, these non-additive effect parameters are superimposed on the traditional additive breeding value model to achieve direct quantitative prediction of heterosis, thereby greatly improving the accuracy and efficiency of hybrid combination selection. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a schematic diagram of the bloodline percentage results in the S2 process; Figure 5 A schematic diagram of the principal component analysis results for purebred population data in the S2 process; Figure 6 This is a detailed schematic diagram of S3 of the present invention; Figure 7 This is a schematic diagram showing the principal component analysis results for boars and sows of various breeds in the S3 process; Figure 8This is a detailed schematic diagram of S4 of the present invention; Figure 9 This is a detailed schematic diagram of S5 of the present invention; Figure 10 This diagram illustrates the prediction results using different genomic prediction models in the S5 workflow. Detailed Implementation

[0018] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0019] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0020] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0021] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0022] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0023] Please see Figures 1-10 This invention provides a method for predicting optimal heterosis in local pigs based on genomic information, comprising the following steps: S1: Collect genomic information of Duroc pigs, Rongchang pigs and Large White pigs and perform quality control analysis. Verify the consistency and numbering of the genomic information after analysis and generate SNP genotype dataset; S2: Determine the bloodline origin of each type of pig based on the SNP genotype dataset and screen the corresponding purebred population data. Simulate the hybridization process based on the purebred population data to form the propagation population data. S3: Construct a list of hybridization combinations for each type of pig based on the propagation population data, match it with the corresponding genomic information in the SNP genotype dataset, extract multiple phenotypic feature groups, and generate a phenotypic fitting classification table; S4: Determine the contribution of intragroup phenotypic bias to each phenotypic feature in the phenotypic fitting classification table, and construct an explicit and implicit weight parameter table based on the magnitude of the bias contribution. S5: Combining the list of hybridization combinations for each type of pig and the table of dominant and recessive weight parameters, calculate the predicted breeding value and the actual breeding value for each hybridization combination, analyze the correlation between the predicted breeding value and the actual breeding value, and output the prediction result of the optimal hybridization combination. The SNP genotype dataset includes SNP marker locus numbers, allele composition, and chromosome position parameters. The propagation population data includes genetic structure information of purebred populations, hybridization simulation recombination information, and offspring phenotypic parameters. The phenotypic fit classification table includes phenotypic feature group categories, phenotypic extended value ranges, and hybridization combination correspondences. The dominant-recessive weight parameter table includes dominant effect parameters, recessive effect parameters, and phenotypic deviation weight values. The optimal hybridization combination prediction results include hybridization combination number, predicted breeding value, actual breeding value, and correlation coefficient.

[0024] Please see Figure 2 Step S1 is as follows: S101: Collect genomic information data of Duroc pigs, Rongchang pigs and Large White pigs, perform signal recognition and data analysis on the genomic information data, extract the original probe signal intensity of each SNP marker site, determine the corresponding allele composition based on the probe signal intensity, obtain the allele type and chromosome position parameters of each SNP marker site, and establish the allele composition of the marker site. Biological samples, such as ear tissue or blood, were collected from Duroc, Rongchang, and Large White pigs, and genomic DNA was extracted from them. The extracted genomic DNA was genotyped using microarray technology. Specifically, the DNA samples were hybridized with a microarray containing hundreds of thousands of single nucleotide polymorphism (SNP) probes. After hybridization, the fluorescence signal intensity of each probe on the microarray was read using a laser scanner. Signal identification and data analysis of the genomic information data were performed as follows: for each SNP marker site, two probes were designed, corresponding to two possible alleles, such as allele A and allele B. The scanner recorded the fluorescence signal intensity values ​​emitted by these two probes. For example, for the pig individual "D001" at the "SNP-A01" site, the recorded probe signal intensity corresponding to allele A was 1.98, and the probe signal intensity corresponding to allele B was 0.21. The process of determining the allele composition based on probe signal intensity involves plotting the two signal intensity values ​​of all individuals at the same SNP locus into a scatter plot. A clustering algorithm automatically identifies three independent populations: one population with high intensity of allele A and low intensity of allele B, identified as homozygous AA; a second population with high intensity of allele B and low intensity of allele A, identified as homozygous BB; and a third population with moderate intensity of both signals, identified as heterozygous AB. Taking the "SNP-A01" locus as an example, the signal intensity combination (1.98, 0.21) of the individual "D001" falls into the cluster with high A signal and low B signal, therefore its allele composition is AA. Obtaining the allele type and chromosomal location parameters for each SNP marker locus involves comparing the unique identifier of each SNP probe on the chip with the pig reference genome sequence to determine the chromosome number and specific base pair position of each SNP locus. For example, "SNP-A01" was located at the 10,245,678th base pair on chromosome 7 through alignment. Finally, the determination results of all individuals at all loci were integrated to form a matrix, where the rows of the matrix represent individuals, the columns represent SNP loci, and the elements in the matrix are the allelic compositions of the corresponding individuals at the corresponding loci (such as AA, AB, BB), thus establishing the complete allelic composition of the marker loci.

[0025] S102: For the allelic composition of each marker site, remove sites with missing probe signals or incorrect localization, calculate the minimum allelic frequency and deletion rate of the retained SNP marker sites, and compare the calculated frequency and deletion rate with the set screening threshold to obtain the SNP site set after quality control screening. For the established allelic composition of marker sites, site removal is performed first. Sites lacking probe signals are removed if the genotyping success rate of a given SNP site across all tested individuals is below a specific criterion. Sites with incorrect localization are removed by comparing the physical location of each SNP site with the latest pig reference genome map. If a site's localization information does not match the reference map, or if it can be matched to multiple locations in the genome, it is considered incorrectly located and removed. When calculating the minimum allele frequency and deletion rate for the retained SNP marker sites, the deletion rate for each site is first calculated; this is the proportion of individuals who could not be successfully genotyped at that site out of the total number of individuals. For example, in a population of 1000 pigs, if the genotype of 60 pigs at the SNP-B02 site could not be determined, the deletion rate is 6%. Then, the minimum allele frequency is calculated; for a given SNP site, the frequency of occurrence of both alleles (e.g., G and T) across the entire population is counted. Assuming that in 940 successfully genotyped pigs, the G allele appeared 1200 times and the T allele appeared 680 times, for a total of 1880 alleles, then the frequency of G is 1200 / 1880, approximately 0.638, and the frequency of T is 680 / 1880, approximately 0.362. The minimum allele frequency is 0.362. The screening threshold is set as follows: the deletion rate screening threshold is determined by analyzing the deletion rate distribution of a batch of representative, high-quality historical genotyping data. This threshold is set at the inflection point of the distribution curve. After this inflection point, the rate of increase in the deletion rate accelerates sharply, usually corresponding to systematic failure of probe design or hybridization experiments. In this study, 5% of the value corresponding to this inflection point is used as the screening threshold. The minimum allele frequency (MAV) threshold for removal was set based on the definition of effective informative loci in population genetics. By analyzing the allele frequency spectrum of all loci in a reference population, the threshold was set at the lower 5% quantile of the frequency distribution, which was set to 1% in this case. Loci with frequencies below this level were removed because they carried too little population variation information and had a higher probability of being genotypic errors. The calculated frequencies and deletion rates were compared with the set removal threshold. Taking the aforementioned SNP-B02 locus as an example, its deletion rate of 6% was higher than the set 5% threshold, so this locus was removed. Another SNP-C03 locus had a calculated deletion rate of 2% and a minimum allele frequency of 0.25, both within the allowable range of the threshold, and was therefore retained. After performing this comparison and judgment on each locus, the final set of SNP loci after quality control screening was obtained.

[0026] S103: Call the chromosomal position parameters of each SNP marker locus in the SNP locus set after quality control screening, correct the numbers of all SNP marker loci that passed the screening, and integrate the corrected numbers with the corresponding allele composition to generate an SNP genotype dataset. First, the SNPs are sorted by chromosome number. Then, within each chromosome, they are sorted by their physical location from smallest to largest. After sorting, each SNP is reassigned a consecutive, unique numerical number. For example, the original SNP number "SNP-C03" is located at the 50,000th base pair on chromosome 1, while the original SNP number "SNP-A05" is located at the 40,000th base pair on chromosome 1. During the correction process, "SNP-A05" will be placed before "SNP-C03" and given a new, earlier sequential number. This operation ensures the logical continuity of all SNPs in their physical locations. Next, the corrected numbers are integrated with the corresponding allelic makeup. This integration process creates a new data structure where each record contains a corrected SNP number, its corresponding chromosome location information, and the allelic makeup of all individuals at that SNP locus. For example, for a newly numbered locus, its corresponding record will include the locus number, chromosome 1, location 40000, and the genotype "AA" for pig individual D001, the genotype "AG" for pig individual R001, and so on, until the data for all individuals are associated. By performing this numbering correction and data integration operation on all SNP loci that have passed quality control screening, a structured SNP genotype dataset is finally generated.

[0027] Please see Figure 3 Step S2 is as follows: S201: Based on the SNP genotype dataset, calculate the pedigree ratio parameters of Duroc pigs, Rongchang pigs and Large White pigs, and compare the allelic composition of each SNP marker locus with the allelic frequency distribution of the reference population to determine the pedigree. Then, compare the calculated pedigree ratio parameters with the preset pedigree threshold, delete samples below the pedigree threshold, and obtain purebred population data. First, a reference population needs to be constructed, consisting of SNP genotype data from a core group of individuals with known purebred lineage and no crossbreeding history. Based on the reference population data, the allele frequency distribution at each SNP locus for each breed is calculated. The allele composition of each individual to be tested at its SNP marker locus is compared with the allele frequency distribution of the reference population to determine its lineage. Specifically, for the individual to be tested, “DRW001,” its genotype at a certain SNP locus is AA, meaning its probability contribution from Duroc pig lineage is higher than that from Rongchang pig or Large White pig. This comparison is repeated for all SNP loci across the entire genome, and the probability information from all loci is integrated using an ancestry inference algorithm to finally calculate the proportion of Duroc, Rongchang, and Large White pig lineage in the individual's genome. For example, the calculated lineage proportion for individual “DRW001” is: Duroc 52%, Large White 28%, and Rongchang 20%. The lineage threshold is set based on a statistical analysis of the lineage proportion distribution of certified purebred individuals of each breed in the reference population. Specifically, the mean and standard deviation of the Duroc pedigree proportions of all purebred Duroc pigs in the reference population are calculated. The mean is then subtracted from three standard deviations to obtain a threshold. This method, based on statistical principles, ensures that only individuals with extremely high pedigree purity (located in the central region of the distribution) are considered purebred; in this case, the threshold is set at 95%. The calculated pedigree proportion parameter for each sample is compared to the preset pedigree threshold. Taking the individual "DRW001" as an example, its highest pedigree proportion is Duroc (52%), below the 95% threshold, therefore it is classified as a hybrid. Another individual, "D100," has a calculated pedigree proportion of 97% Duroc, 2% Large White, and 1% Rongchang pig; its Duroc pedigree is above 95%, therefore it is classified as a purebred Duroc sample. This judgment is applied to all individuals, and all samples below the pedigree threshold are deleted, resulting in the purebred population data. The `plotQ()` function in R is used to statistically analyze and plot the pedigree proportions of all individuals; please refer to [link to relevant documentation] for the results. Figure 4 For purebred population data, when performing principal component analysis, different populations should exhibit significant distances from different principal component perspectives. Please refer to the reference results. Figure 5 .

[0028] S202: Based on the Bayesian genetic simulation method, the hybridization process between Duroc pigs, Rongchang pigs and Large White pigs was simulated. The allelic composition of SNP marker sites and linkage recombination of purebred population data were used to establish the virtual offspring gene structure. First, based on the SNP genotype data of the purebred population, a haplotype library for each variety is compiled and constructed. The linkage disequilibrium strength and recombination rate between adjacent SNP loci are calculated; these parameters together constitute the genetic map. During the hybridization simulation, for two specified purebred parents, the process of gamete production is simulated first. During meiosis simulation, the program randomly simulates the location and frequency of genetic recombination events between homologous chromosomes of the parents, based on the constructed genetic map. For example, for a pair of chromosome 7s from the parents, depending on the recombination rate, a crossover may occur at a certain location, resulting in a new haplotype chromosome that mixes segments of the grandparents' chromosomes. Through multiple independent simulations, a large number of virtual gametes are generated for each parent, each gamete containing a complete set of recombinated haplotypes. After recombination using SNP marker allele composition and linkage relationships from purebred population data, a virtual gamete is randomly selected from a simulated Duroc boar gamete library and another from a Large White sow gamete library. Combining these two virtual gametes creates the genetic structure of a virtual first-generation hybrid (F1). The allele composition of each SNP locus in this virtual offspring's genome and its linkage relationships on chromosomes are determined by both the parental genetic structure and the simulated recombination process. By repeating this process, a large number of virtual offspring produced by different parental combinations can be simulated. The genetic structures of these offspring collectively constitute a virtual offspring genetic structure database.

[0029] S203: Obtain phenotypic information of the gene structure of virtual offspring, calculate the real breeding value of virtual offspring, and redistribute purebred individuals of Duroc pigs, Rongchang pigs and Large White pigs as parents based on the real breeding value. Combine the redistributed parent and offspring data to generate propagation population data. This is achieved by combining known SNP effect values ​​with the genotype data of virtual offspring. The specific method for calculating the true breeding value of a virtual offspring is as follows: For a given virtual offspring, all SNP loci in its genome are traversed. At each locus, based on its allele composition (e.g., AA, AB, BB), the value is converted to a numerical value (e.g., 0, 1, 2), and then multiplied by the known effect value of that SNP locus for the target trait. The sum of the calculation results for all SNP loci yields the true breeding value of the virtual offspring for that trait. For example, if a virtual offspring has genotypes of AA, AB, and BB at three SNP loci, corresponding to numerical values ​​of 0, 1, and 2, and the effect values ​​of these three loci on daily weight gain are +0.5 g, -0.2 g, and +1.1 g, respectively, then the breeding value contributed by these three loci is (0 × 0.5) + (1 × -0.2) + (2 × 1.1) = 2.0 g. The final true breeding value is obtained by summing the contribution values ​​of all loci across the entire genome. The redistribution of Duroc, Rongchang, and Large White purebred individuals as parents based on actual breeding values ​​refers to the reverse evaluation of the genetic value of the parents based on the breeding performance of simulated hybrid offspring. The selection criteria are set based on the principle of maximizing genetic progress while maintaining genetic diversity. Through multi-generational breeding simulations, the effects of different selection ratios (e.g., top 10%, 20%, 30%) on long-term genetic gain and inbreeding coefficient growth are tested. The selection ratio that yields the maximum genetic gain under a preset inbreeding growth rate constraint (e.g., less than 1% per generation) is selected; in this case, it is determined to be the top 20% of parents with the average breeding value of the simulated offspring. These selected individuals are redefined as core parents. Finally, the redistributed parent individual data are combined with all the virtual offspring individual data generated by them to generate a propagation population dataset.

[0030] Please see Figure 6 Step S3 is as follows: S301: Based on the propagation population data, boars and sows of three breeds—Duroc, Rongchang, and Large White—were selected. Multiple crossbreeding combinations were established according to the principle of matching pedigree ratios and breeding values, and integrated into a crossbreeding combination list. When performing principal component analysis on the selected Duroc, Rongchang, and Large White boars and sows, from different principal component perspectives, populations of different breeds should show significant distances, while boars and sows of the same breed should show significant clustering. Please refer to the reference results. Figure 7 .

[0031] The specific principles for matching pedigree ratios with breeding values ​​are as follows: Set the target pedigree ratio range for Duroc, Rongchang and Large White pigs in the crossbreeding combination; Obtain the breeding value of each boar and sow in the breeding population data. The breeding value includes the breeding value related to growth rate, backfat thickness and lean meat percentage. Select parental hybridization combinations that meet the requirements based on the target bloodline ratio range; For the selected parental combinations, calculate the expected breeding value of the offspring of each potential hybrid combination; The expected breeding value is compared with the preset optimal range of breeding values, and potential hybrid combinations whose expected breeding value falls within the optimal range of breeding values ​​are identified as hybrid combinations. A target pedigree ratio range for Duroc, Rongchang, and Large White pigs in crossbreeding combinations is established, based on market demand and production targets for the final commercial pigs. Breeding values ​​related to growth rate, backfat thickness, and lean meat percentage are obtained for each boar and sow in the breeding population data. Parental crossbreeding combinations meeting the target pedigree ratio range are selected. For the selected parental crossbreeds, the expected breeding value of each potential crossbreed's offspring is calculated. The expected breeding value of the offspring is calculated as the average of the corresponding trait breeding values ​​of its parents. For example, if a boar's daily weight gain breeding value is +100 grams and its paired sow's daily weight gain breeding value is +80 grams, then the expected daily weight gain breeding value of their offspring is (+100 + +80) / 2 = +90 grams. The expected breeding values ​​are compared with the preset optimal range of breeding values. This optimal range is determined by constructing a bioeconomic model that comprehensively considers economic parameters such as pork market prices, feed costs, veterinary drug and vaccine expenses, and fixed asset depreciation. The model calculates the net profit per unit carcass weight under different production performance levels. This optimal range is defined as the performance breeding value range that can generate the highest 10% net profit. For example, the model calculates that individuals with a daily weight gain breeding value between +85g and +110g, a backfat thickness breeding value between -1.5mm and -0.5mm, and a lean meat percentage breeding value between +1.0% and +2.0% have the highest market value; this is the optimal range of breeding values. Potential hybrid combinations whose expected breeding values ​​fall within this optimal range are determined as the final hybrid combinations. Taking the above offspring as an example, its expected daily weight gain breeding value of +90 grams falls within the optimal range [+85, +110]. If its expected breeding values ​​for backfat thickness and lean meat percentage also fall within their respective optimal ranges, then this parent pairing is confirmed as an effective hybrid combination and integrated into the hybrid combination list.

[0032] S302: Based on the list of hybrid combinations, collect phenotypic observation data of each hybrid individual on the three traits of growth rate, backfat thickness and lean meat percentage, and match them with the SNP marker locus information of the same numbered hybrid individual in the SNP genotype dataset. Perform standardization and weighted summation operations on the phenotypic observation data of the matched hybrid individuals to obtain the phenotypic extended value. Phenotypic data for growth rate, backfat thickness, and lean meat percentage were collected from each hybrid individual. These phenotypic data were then matched with the SNP marker loci information of the same-numbered hybrid individuals in the SNP genotype dataset. Standardization and weighted summation were performed on the matched hybrid individual phenotypic data to obtain phenotypic extended values. The standardization process involves subtracting the mean of that trait in the entire hybrid population from the original observed value of each trait, and then dividing by the standard deviation of that trait. For example, if the average daily weight gain of a hybrid population is 850 grams and the standard deviation is 50 grams, and an individual "H001" has a daily weight gain of 925 grams, its standardized daily weight gain value is (925-850) / 50 = 1.5. The same operation was performed on backfat thickness and lean meat percentage. The weights in the weighted summation were set according to the economic importance of different traits. The weights are determined by calculating the partial derivatives of each trait with respect to the total profit function. This profit function uses feed costs, meat sales revenue, etc., as variables, with each trait as the independent variable. The weight of each trait represents its marginal contribution to the total profit, and is normalized. For example, the economic weights for daily weight gain, backfat thickness, and lean meat percentage are calculated to be 0.4, -0.3, and 0.3, respectively. Therefore, the phenotypic expanded value of individual "H001" is: (1.5 × 0.4) + (standardized backfat value × -0.3) + (standardized lean meat percentage value × 0.3). This calculated composite value is the phenotypic expanded value of that individual.

[0033] S303: Compare the phenotypic extended values ​​with the phenotypic fit thresholds to classify hybrid individuals into three phenotypic characteristics: production performance type, reproductive performance type and growth balance type, and generate a phenotypic fit classification table. The calculated phenotypic expansion value of each hybrid individual is compared with a preset phenotypic fit threshold. The phenotypic fit threshold is set based on statistical analysis of the distribution of phenotypic expansion values ​​of all hybrid individuals in the propagation population data, combined with the preset demand ratios for different types of pigs in the breeding objectives. For example, if the breeding plan requires selecting 20% ​​of individuals as a specialized production performance group and 30% as a reserve population for reproductive performance, then all individuals' phenotypic expansion values ​​are sorted from highest to lowest. The values ​​in the top 20% are set as the lower threshold for the production performance type; the values ​​in the bottom 30% are set as the upper threshold for the reproductive performance type. The process of classifying hybrid individuals into three phenotypic categories—production performance type, reproductive performance type, and growth-balanced type—is as follows: If an individual's phenotypic expansion value is greater than the lower threshold for the production performance type, then the individual is classified as "production performance type." If an individual's phenotypic expansion value is less than the upper threshold for the reproductive performance type, then it is classified as "reproductive performance type." If an individual's phenotypic expansion value is between these two thresholds, then it is classified as "growth-balanced type." For example, the calculated lower threshold for the productive performance phenotype is 1.2, and the upper threshold for the reproductive performance phenotype is -0.5. An individual "H002" has a phenotypic expansion value of 1.5; because it is greater than 1.2, it is classified as a productive performance phenotype. Another individual "H003" has a phenotypic expansion value of -0.8; because it is less than -0.5, it is classified as a reproductive performance phenotype. By performing this classification operation on all hybrid individuals, a phenotypic fit classification table is finally generated.

[0034] Please see Figure 8 Step S4 is as follows: S401: Calculate the intra-group phenotypic differences of the three characteristic groups (production performance type, reproductive performance type, and growth balance type) in the phenotypic fitting classification table, determine the intra-group phenotypic bias contribution of each phenotypic characteristic, and obtain the intra-group phenotypic bias value. For each trait group, the raw phenotypic observation data of all individuals within the group are extracted, namely, the values ​​of daily weight gain, backfat thickness, and lean meat percentage. Then, the variance of the phenotypic values ​​of all individuals within that group is calculated for each trait. The contribution of intragroup phenotypic bias to each phenotypic trait is determined by comparing the variances of different groups on the same trait. The process of obtaining the intragroup phenotypic bias value is to subtract the average value of that trait within its trait group from the observed value of each individual for a given trait. For example, if the average daily weight gain of the "Productivity Performance" group is 950 grams, and the daily weight gain of individual "H002" in that group is 980 grams, then the intragroup phenotypic bias value for "H002" in daily weight gain is 980 - 950 = +30 grams. This calculation is performed for each individual and each trait, thus obtaining three intragroup phenotypic bias values ​​for each individual.

[0035] S402: Obtain the expression status of SNP marker sites for each hybrid offspring in the hybridization combination list, and calculate the offset contribution of heterozygous alleles to the phenotype and the independent offset contribution of recessive homozygous alleles in combination with the intragroup phenotypic deviation value, and obtain the set of dominant effect parameters and recessive effect parameters. The specific process for calculating the shift contribution of the heterozygous allele to the phenotype and the independent bias contribution of the recessive homozygous allele is as follows: For the three characteristic groups of production performance type, reproductive performance type and growth balance type, a linear regression model is constructed for each characteristic group. The linear regression model uses the intragroup phenotypic deviation as the dependent variable and the expression status of SNP marker sites in each hybrid offspring as the independent variable. The linear regression model is solved by joint error minimization. The regression coefficient associated with the heterozygous allele encoding is used as the shift contribution of the heterozygous allele to the phenotype, and the dominant effect parameter is obtained. The regression coefficients associated with the coding of the recessive homozygous alleles were used as the independent bias contribution of the recessive homozygous alleles to obtain the recessive effect parameters. The expression status of SNP marker loci for each offspring in the hybridization combination list is obtained. This status refers to the numerical encoding of the individual's genotype. To separate the dominant effect, a specific encoding scheme is used, representing the three genotypes AA, AG, and GG with two variables. The first variable represents the additive effect, encoded as 0, 1, and 2. The second variable represents the dominant deviation, encoded as 0 for the two homozygous individuals AA and GG, and as 1 for the heterozygous individual AG. The contribution of the heterozygous allele to the phenotypic shift and the independent deviation contribution of the recessive homozygous allele are calculated by combining the within-group phenotypic deviation values. This process involves constructing a linear regression model for each of the three trait groups: production performance, reproductive performance, and growth equilibrium. The model structure is: Individual within-group phenotypic deviation value = intercept + additive effect coefficient × additive encoding value + dominant effect coefficient × dominant deviation encoding value + residual. The dependent variable of the model is the within-group phenotypic deviation value for all individuals in the group on the specific trait. The independent variables are the numerical codes for each individual at all SNP loci. A linear regression model is solved by minimizing the joint error, an iterative optimization process aimed at finding a set of regression coefficients that minimizes the sum of squares of the differences between the predicted and actual dependent variable values. After solving, the regression coefficients associated with the heterozygous allele code (i.e., the dominant deviation code variable) are taken as the heterozygous allele's contribution to the phenotypic bias; this is the dominant effect parameter. The regression coefficients associated with the recessive homozygous allele code are taken as the independent bias contribution of the recessive homozygous allele, yielding the recessive effect parameter.

[0036] S403: Based on the independent bias contribution reflected by the sets of dominant and recessive effect parameters, the weights of the dominant and recessive effect parameters are adjusted directionally to represent the gene interaction characteristics of the three-variety hybrid population and to construct a dominant-recessive weight parameter table. Directional adjustment refers to determining whether an effect is beneficial or detrimental based on the breeding objective, and adjusting its weight in subsequent applications accordingly. The adjustment coefficients for the weight magnitudes are determined through cross-validation on independent validation datasets. Different combinations of adjustment coefficients are tested (e.g., enhancement coefficients of 1.1, 1.2, and 1.3 for consistent and significant effects; and weakening coefficients of 0, 0.25, and 0.5 for inconsistent effects), and the combination that maximizes the accuracy of the final predicted breeding values ​​is selected. The adjustment rule determined in this study is as follows: if an effect is statistically significant in at least two feature groups and its direction is consistent with the breeding objective, its weight magnitude is multiplied by 1.2; if it is significant only in one group, the magnitude remains unchanged; and if the direction is opposite in different groups, the magnitude is multiplied by 0. By performing this directional and magnitude adjustment on the dominant and recessive effect parameters for all SNP loci, a dominant-recessive weight parameter table is finally constructed.

[0037] Please see Figure 9 Step S5 is as follows: S501: Combining the list of hybrid combinations and the table of dominant and recessive weight parameters, the predicted breeding value of each hybrid combination is calculated by superimposing the dominant effect parameter and the recessive effect parameter into the SNP locus effect of the hybrid individual. The process of superimposing the dominant effect parameter and the recessive effect parameter into the SNP locus effect of the hybrid individual is as follows: Obtain all SNP genotypes for each hybrid individual in the hybrid combination; For each SNP marker locus in each hybrid individual, determine whether the allelic composition of the SNP marker locus is heterozygous or recessive homozygous; If the alleles of the SNP marker site are heterozygous, then the dominant effect parameter corresponding to the dominant-recessive weight parameter table is summed and superimposed with the SNP site effect; If the allele of the SNP marker site becomes recessive homozygous, then the recessive effect parameter corresponding to the dominant-recessive weight parameter table is summed and superimposed with the SNP site effect. First, obtain the genotypes of all SNPs for each potential offspring in the hybrid combination. For each SNP marker locus in each hybrid individual, determine its allelic composition. If the allelic composition of the SNP marker locus is heterozygous, then when calculating the contribution of that locus to the breeding value, the "dominant effect parameter" corresponding to that SNP locus and the target trait, found in the dominance-recessive weight parameter table, is summed and added to the base additive effect. For example, if the base additive effect estimate for a locus is +2 grams, and its adjusted dominant effect weight is +1.5 grams, then the total effect value of this heterozygous individual at this locus is 2 + 1.5 = 3.5 grams. If the allelic composition of the SNP marker locus is homozygous recessive, then the "recessive effect parameter" corresponding to the dominance-recessive weight parameter table is summed and added to the SNP locus effect. Repeat this operation for all SNP loci in the individual's genome, and sum the total effect values ​​of all loci to finally obtain the predicted breeding value of the hybrid individual containing dominant-recessive interaction effects. When calculating the predicted breeding value, multiple genome prediction models can be used for prediction. Please refer to the results for reference. Figure 10 .

[0038] S502: Calculate the actual breeding value corresponding to each hybrid combination in the statistical propagation population data and the predicted breeding value using Pearson correlation coefficient and Spearman rank correlation coefficient to obtain the average predicted index of the hybrid combination; The actual breeding values ​​for each hybrid combination in the propagation population data were statistically analyzed. The "predicted breeding values," which include dominant and recessive effects, were compared with these "actual breeding values" using Pearson correlation coefficients and Spearman rank correlation coefficients. The Pearson correlation coefficient assesses the strength and direction of the linear relationship between the two breeding value sequences. The Spearman rank correlation coefficient assesses the ordination consistency between the two breeding value sequences. These two correlation coefficients were then weighted and averaged to obtain the average prediction index for each hybrid combination. The weights were set based on the emphasis placed on prediction accuracy and ordination accuracy in the breeding objectives. Under a balanced breeding strategy that simultaneously prioritizes absolute value prediction accuracy and selection ordination correctness, equal weights were assigned to the two correlation coefficients. In this study, the weights for both the Pearson correlation coefficient and the Spearman rank correlation coefficient were set to 0.5. For example, if the Pearson correlation coefficient is 0.85 and the Spearman rank correlation coefficient is 0.80, then the average predictive index is (0.85×0.5)+(0.80×0.5)=0.825.

[0039] S503: Based on the average prediction index of hybrid combinations, sort the various mating types of Duroc pigs, Rongchang pigs and Large White pigs, and output the prediction result of the best hybrid combination; Based on the calculated average prediction index of the hybrid combinations, all pre-selected hybrid combinations of Duroc, Rongchang, and Large White pigs are ranked. The ranking rule is to arrange them from highest to lowest average prediction index value. A higher average prediction index indicates better consistency between the predicted breeding value (considering dominant and recessive interaction effects) and the actual breeding value based on the additive model. For example, if hybrid combination A has an average prediction index of 0.91 and hybrid combination B has an average prediction index of 0.88, then combination A will be ranked higher than combination B in the ranking list. By performing this ranking operation on all mating types in the hybrid combination list, a final ordered list is output, with the optimal hybrid combination prediction result at the top.

[0040] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for predicting optimal heterosis in local pigs based on genomic information, characterized in that, Includes the following steps: S1: Collect genomic information of Duroc pigs, Rongchang pigs and Large White pigs and perform quality control analysis. Verify the consistency and numbering of the genomic information after analysis and generate SNP genotype dataset; S2: Determine the bloodline origin of each type of pig based on the SNP genotype dataset and screen the corresponding purebred population data. Simulate the hybridization process based on the purebred population data to form the propagation population data. S3: Based on the propagation population data, construct a list of hybridization combinations for each type of pig, match it with the corresponding genomic information in the SNP genotype dataset, extract multiple phenotypic feature groups, and generate a phenotypic fitting classification table; S4: Determine the intra-group phenotypic bias contribution of each phenotypic feature in the phenotypic fitting classification table, and construct an explicit and implicit weight parameter table based on the magnitude of the bias contribution. S5: Combining the list of hybridization combinations for each type of pig with the dominant-recessive weight parameter table, calculate the predicted breeding value and the actual breeding value for each hybridization combination, analyze the correlation between the predicted breeding value and the actual breeding value, and output the prediction result of the optimal hybridization combination.

2. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 1, characterized in that, The SNP genotype dataset includes SNP marker locus numbers, allele composition, and chromosome position parameters. The propagation population data includes purebred population genetic structure information, hybridization simulation recombination information, and offspring phenotypic parameters. The phenotypic fit classification table includes phenotypic feature group categories, phenotypic extended value ranges, and hybridization combination correspondences. The dominant-recessive weight parameter table includes dominant effect parameters, recessive effect parameters, and phenotypic deviation weight values. The optimal hybridization combination prediction results include hybridization combination number, predicted breeding value, actual breeding value, and correlation coefficient.

3. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 1, characterized in that, Step S1 is as follows: S101: Collect genomic information data of Duroc pigs, Rongchang pigs and Large White pigs, perform signal recognition and data analysis on the genomic information data, extract the original probe signal intensity of each SNP marker site, determine the corresponding allele composition based on the probe signal intensity, obtain the allele type and chromosome position parameters of each SNP marker site, and establish the allele composition of the marker site. S102: For each of the marked sites, remove sites with missing probe signals or incorrect localization, calculate the minimum allele frequency and deletion rate for the remaining SNP marked sites, and compare the calculated frequency and deletion rate with the set screening threshold to obtain the SNP site set after quality control screening. S103: Call the chromosomal position parameters of each SNP marker locus in the SNP locus set after quality control screening, correct the numbers of all SNP marker loci that passed the screening, and integrate the corrected numbers with the corresponding allele composition to generate an SNP genotype dataset.

4. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 1, characterized in that, Step S2 is as follows: S201: Based on the SNP genotype dataset, calculate the pedigree ratio parameters of Duroc pigs, Rongchang pigs and Large White pigs, and compare the allelic composition of each SNP marker locus with the allelic frequency distribution of the reference population to determine the pedigree. Then, compare the calculated pedigree ratio parameters with the preset pedigree threshold, delete samples below the pedigree threshold, and obtain purebred population data. S202: Based on the Bayesian genetic simulation method, the hybridization process between Duroc pigs, Rongchang pigs and Large White pigs is simulated. The virtual offspring gene structure is established by using the SNP marker locus allele composition and linkage recombination of the purebred population data. S203: Obtain the phenotypic information of the gene structure of the virtual offspring, calculate the real breeding value of the virtual offspring, and redistribute the purebred individuals of Duroc pigs, Rongchang pigs and Large White pigs as parents according to the real breeding value. Combine the redistributed parent and offspring data to generate propagation population data.

5. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 1, characterized in that, Step S3 is as follows: S301: Based on the aforementioned propagation population data, select boars and sows from three breeds: Duroc, Rongchang, and Large White. According to the principle of matching bloodline ratio and breeding value, establish multiple hybridization combinations and integrate them into a hybridization combination list. S302: Based on the list of hybrid combinations, collect phenotypic observation data of each hybrid individual on the three traits of growth rate, backfat thickness and lean meat percentage, and match them with the SNP marker locus information of the same numbered hybrid individual in the SNP genotype dataset. Perform standardization and weighted summation operations on the phenotypic observation data of the matched hybrid individuals to obtain the phenotypic extended value. S303: Compare the phenotypic extended value with the phenotypic fit threshold, classify the hybrid individuals into three phenotypic characteristics: production performance type, reproductive performance type, and growth balance type, and generate a phenotypic fit classification table.

6. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 5, characterized in that, The specific principle for matching pedigree ratios with breeding values ​​is as follows: Set the target pedigree ratio range for Duroc, Rongchang and Large White pigs in the crossbreeding combination; Obtain the breeding value of each boar and sow in the breeding population data. The breeding value includes the breeding value related to growth rate, backfat thickness and lean meat percentage. Select parental hybridization combinations that meet the requirements based on the target bloodline ratio range; For the selected parental combinations, calculate the expected breeding value of the offspring of each potential hybrid combination; The expected breeding value is compared with the preset optimal range of breeding values, and potential hybrid combinations whose expected breeding value falls within the optimal range of breeding values ​​are identified as hybrid combinations.

7. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 1, characterized in that, Step S4 is as follows: S401: Calculate the intra-group phenotypic differences of the three characteristic groups (production performance type, reproductive performance type and growth balance type) in the phenotypic matching classification table respectively, determine the intra-group phenotypic bias contribution of each phenotypic characteristic, and obtain the intra-group phenotypic bias value. S402: Obtain the expression status of SNP marker sites for each hybrid offspring in the hybridization combination list, and calculate the offset contribution of heterozygous alleles to the phenotype and the independent offset contribution of recessive homozygous alleles in combination with the phenotypic deviation value within the group, and obtain the set of dominant effect parameters and recessive effect parameters. S403: Based on the independent bias contribution reflected by the set of dominant and recessive effect parameters, directional adjustments are made to the weights of the dominant and recessive effect parameters to represent the gene interaction characteristics of the three-variety hybrid population, and a dominant-recessive weight parameter table is constructed.

8. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 7, characterized in that, The process of calculating the shift contribution of heterozygous alleles to the phenotype and the independent bias contribution of recessive homozygous alleles is as follows: For the three characteristic groups of production performance type, reproductive performance type and growth balance type, a linear regression model is constructed for each characteristic group. The linear regression model uses the intragroup phenotypic deviation as the dependent variable and the expression status of SNP marker sites in each hybrid offspring as the independent variable. The linear regression model is solved by joint error minimization. The regression coefficient associated with the heterozygous allele encoding is used as the shift contribution of the heterozygous allele to the phenotype, and the dominant effect parameter is obtained. The regression coefficients associated with the coding of the recessive homozygous allele were used as independent bias contributions of the recessive homozygous allele to obtain the recessive effect parameters.

9. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 1, characterized in that, Step S5 is as follows: S501: Combining the list of hybridization combinations and the table of dominant and recessive weight parameters, the predicted breeding value of each hybridization combination is calculated by superimposing the dominant effect parameter and the recessive effect parameter into the SNP locus effect of the hybrid individual. S502: Statistically calculate the actual breeding value corresponding to each hybrid combination in the propagation population data, and perform Pearson correlation coefficient and Spearman rank correlation coefficient calculation with the predicted breeding value to obtain the average predicted index of the hybrid combination; S503: Based on the average prediction index of the hybridization combination, sort the various mating types of Duroc pigs, Rongchang pigs and Large White pigs, and output the prediction result of the optimal hybridization combination.

10. The method for predicting optimal heterosis in local pigs based on genomic information according to claim 9, characterized in that, The process of superimposing the dominant effect parameter and the recessive effect parameter into the SNP locus effect of the hybrid individual is as follows: Obtain all SNP genotypes for each hybrid individual in the hybrid combination; For each SNP marker locus in each hybrid individual, determine whether the allelic composition of the SNP marker locus is heterozygous or recessive homozygous; If the alleles of the SNP marker site are heterozygous, then the dominant effect parameter corresponding to the dominant-recessive weight parameter table is summed and superimposed with the SNP site effect; If the allele of the SNP marker site becomes recessive homozygous, then the recessive effect parameter corresponding to the dominant-recessive weight parameter table is summed and superimposed with the SNP site effect.