Pig variety identification method based on deep learning and genetic characteristics

Through the graph neural network model based on deep learning, learning the complex genetic relationships between pig breeds is solved, and the problem that the existing technology is difficult to accurately distinguish between purebred pigs and binary pigs is achieved, achieving higher-precision ancestry composition prediction and breed classification.

CN119993270APending Publication Date: 2025-05-13石家庄博瑞迪生物技术有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510100761.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing pig breed identification methods are difficult to accurately distinguish between purebred pigs and binary pigs, resulting in a decline in breeding efficiency and disorderly market order. Traditional methods have limited expression capabilities when facing complex genotype data, and cannot effectively capture genetic similarities and genotype subtle differences between samples.

Method used

Using a method based on deep learning and genetic characteristics, we learn the complex genetic relationships and ancestry structures between samples from genotype data through graph neural network models, construct the graph structure of genotype data, and use the graph convolutional layer to learn the interactions between nodes and the structural characteristics of the graph, and then make ancestry composition prediction and variety classification decisions.

Benefits of technology

It improves the accuracy of prediction of complex ancestry composition, can dynamically adjust the estimation of ancestry proportion, avoids the limitations of ancestry classification in traditional methods, and provides higher-precision variety identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993270A_ABST
    Figure CN119993270A_ABST
Patent Text Reader

Abstract

The invention discloses a pig variety identification method based on deep learning and genetic characteristics, and relates to the field of biological gene model prediction. According to the method, the graph neural network deep learning model is utilized, the complex genetic relationship and the lineage structure between the samples are learned from the genotype data, and the limitation of the prior art in lineage composition prediction is effectively overcome. By constructing a graph structure based on genotype data similarity and genetic relationship, the model can capture the interaction between gene loci and the relationship between samples, thereby providing more accurate bloodline proportion prediction. The method not only improves the prediction accuracy of the complex lineage composition, but also can dynamically adjust the estimation of the lineage proportion according to the specific genotype data of the sample, and avoids the limitation of lineage classification in a traditional method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological gene model prediction, and specifically to a method for pig breed identification based on deep learning and genetic characteristics. Background Art

[0002] At present, the application of genomic technology in animal breeding has been widely studied and practiced. Pig breed identification and pedigree analysis mainly rely on genomic typing data, especially the use of gene chips for large-scale genotype data detection. Common gene chips such as Dangkang 1K, Dangkang 50K, and Newgen 50K can provide high-quality genotype data through high-density marker site coverage. With the support of these technologies, genotype data has become an important basis for evaluating pig pedigree composition and breed identification. The current genotype data processing process usually includes steps such as data collection, quality control, genotype data filling, and breed label conversion, and finally makes breed classification decisions by establishing a classification model. For breed identification, mainstream methods usually rely on machine learning algorithms such as support vector machines (SVM) and random forests (RF), which can extract features from genotype data and classify samples based on these features. However, most of these traditional methods rely on artificially designed feature extraction, which may ignore some potential complex genetic relationships.

[0003] Pig breed identification and pedigree verification are key links in breeding and live pig trade. During the breeding process, some binary hybrid pigs are very similar to purebred pigs in body shape and appearance, which can easily lead to breeders misjudging the breed and mistaking binary hybrid pigs for purebred pigs for selection and matching. This misjudgment will not only affect breeding efficiency, but in severe cases may also lead to genetic degeneration of the entire group, thereby affecting the long-term development and economic benefits of the breeding project. In addition, in the live pig trade, there is a significant difference in the market value of purebred pigs and binary hybrid pigs, and the price of the former is often several times that of the latter. Therefore, in order to make high profits, some unscrupulous merchants sell binary hybrid pigs as purebred pigs. This behavior not only harms the rights and interests of consumers, but also seriously disrupts the market order and affects the healthy development of the industry.

[0004] Existing pig breed identification methods include appearance identification, pedigree identification, and linear regression algorithms based on genomic information. These methods have limitations to varying degrees in practical applications and are difficult to meet the high-precision requirements of modern breeding and live pig trade. For example, appearance identification requires extremely high professional skills from identification personnel, and some binary pigs and purebred pigs are extremely similar in body shape and appearance, so accurate identification cannot be performed by appearance. Pedigree identification is prone to recording errors, and unscrupulous merchants may tamper with the pedigree, resulting in unreliable identification results. Although the linear regression algorithm based on genomic information has a certain degree of accuracy, its precision is low and it is difficult to meet the needs of practical applications. With the development of the AI ​​field, academic researchers have begun to apply machine learning algorithms for breed identification, but these methods cannot provide pedigree composition information, and the accuracy cannot reach more than 99%. Most of them are only applicable to a single type of chip product and cannot adapt to a diverse market environment.

[0005] Secondly, most of the existing breed classification methods are based on simple pedigree ratio threshold judgments, and it is often difficult to provide accurate pedigree composition information for pig samples with multiple pedigrees. Especially when faced with complex three-way and four-way mating populations, the pedigree ratio predictions of traditional methods often have large errors. Existing models have limited expressive power when faced with complex interactions in genotype data and cannot effectively capture the genetic similarities between samples and subtle differences in genotypes. Since traditional methods are mostly shallow learning models, these models cannot learn deep graph structure information from the data, and therefore it is difficult to provide accurate pedigree composition predictions.

[0006] In addition, traditional genotype data analysis methods also have limitations in modeling the relationships between multiple genotypes. For example, most models cannot fully explore the complex genetic relationships between gene loci, and the interactions between gene loci are crucial for accurate ancestry composition prediction. Traditional genotype analysis methods usually only use single genotype information for analysis and lack in-depth modeling of the interactions between gene loci. This limitation leads to the fact that in the process of ancestry composition prediction, the model often ignores the structural relationship between samples, thus affecting the final prediction results. Summary of the invention

[0007] The present invention proposes a method for pig breed identification based on deep learning and genetic characteristics, with the aim of providing a breed identification algorithm suitable for mainstream gene chips on the market to solve the problem that it is difficult to distinguish between purebred pigs and dual-purpose pigs in actual production, which has a negative impact on breeding work and live pig trade.

[0008] Among them, a method for pig breed identification based on deep learning and genetic characteristics includes the following steps: S1. Collect pig breed information and genome typing test results. The genome typing test targets Landrace, Large White, Duroc, Changda and Duroc-Changda pigs, and uses one of the gene chips from Dangkang 1K, Dangkang 50K, Neogen 50K, SMIC-1 and Dangkang 100K for genome typing; S2. Perform quality control filling on the genotype data, convert the variety information into pedigree proportion according to the genetic law, and recode the genotype data using the One-Hot encoding method; S3. The recoded genotype data are used as nodes in the graph, and the genetic relationships between gene loci are used as edges to construct a graph structure of the genotype data, wherein the nodes are described by gene frequency and allele information, and the edges represent the interactions between genes; S4. Build a complete graph neural network model by learning the interactions between nodes and the structural features of the graph through the graph convolution layer; predict the lineage composition of the sample based on the graph neural network model; S5. Make a breed classification decision based on the predicted results output by the model according to the set pedigree ratio threshold.

[0009] Furthermore, in step S2, performing quality control filling on the genotype data specifically includes the following sub-steps: S201. Delete the minimum allele gene frequency less than 0.05; S202. Delete sites with a detection rate lower than 0.8; S203. Deletion sites and Indel sites are filled with NA.

[0010] Furthermore, in step S3, the genetic relationship between gene loci as an edge specifically includes the following sub-steps: S301. defining connections between nodes through a similarity matrix of genotype data; S302. According to the similarity matrix, a weighted adjacency matrix is ​​constructed, wherein each element represents the genotype similarity between nodes, and the adjacency matrix represents the structure of a graph defined based on genotype similarity.

[0011] Furthermore, in step S301, the construction of the similarity matrix of genotype data specifically includes the following steps: S3011. Use Hamming Distance to measure the number of different sites in the corresponding gene loci of two genotype samples; S3012. Construct similarity measurement based on Hamming Distance measurement results; S3013. Construct a similarity matrix based on the similarity measure, namely: ; Wherein, S represents the similarity matrix, and N represents the index of the sample.

[0012] Furthermore, in step S3011, the measurement process of Hamming Distance is: ; ; Among them, the represents the Hamming Distance from sample i to sample j, Represents the total number of sites in the genotype data, Represents the site index of the genotype, represents the indicator function, Indicates that the i-th sample is in The genotype at a gene locus takes a value of 0 or 1. Indicates that the jth sample is in The genotype at a gene locus has a value of 0 or 1.

[0013] Furthermore, in step S3012, the Hamming Distance measurement result is converted into a similarity measurement by the following formula: ; Among them, the Represents genotype sample i, the represents genotype sample j, Represents the Hamming Distance from sample i to sample j.

[0014] Furthermore, the step S4 specifically includes the following sub-steps: S401. taking the genotype data of each pig sample and the adjacency matrix of the constructed graph as inputs to the model; S402. Aggregating information of adjacent nodes in the graph according to the graph convolution layer; and gradually increasing the depth of node representation by stacking multiple layers of graph convolution layers; S403. Designing mean square error as a loss function, with the goal of minimizing the error between the predicted ancestry proportion and the actual label; S404. Predict the ancestry proportion of each sample through the fully connected layer and output the model prediction result.

[0015] Furthermore, in step S402, the information of adjacent nodes in the graph aggregated by the graph convolution layer is specifically expressed as: ; Among them, the Indicates The node representation of the layer represents the activation function, represents the normalized adjacency matrix, Indicates that Indicates The learning weight matrix of the layer, Indicates the number of hidden layers.

[0016] Furthermore, in step S404, the pedigree ratio of each sample is predicted by the fully connected layer as follows: ; Among them, the represents the ancestry proportion vector of the i-th sample, represents the weight matrix of the fully connected layer, represents the representation of the i-th sample in the last layer of the graph neural network. Represents the bias term.

[0017] Furthermore, in step S5, the variety classification decision is specifically as follows: When the proportion of a certain breed in the prediction result is ≥75%, it is judged to be a purebred pig of that breed; When the proportion of each breed pedigree in the prediction result is <75% and the proportion of Duroc pedigree is <20%, it is judged as a dual-purpose pig. In other cases, it is determined to be a three-yuan pig; Right now: ; Among them, the It represents the predicted pedigree percentage of sample i corresponding to variety j, and j=3 represents the Duroc pedigree percentage.

[0018] The beneficial effects of the invention are: The present invention uses a graph neural network deep learning model to learn the complex genetic relationships and pedigree structures between samples from genotype data, effectively overcoming the limitations of existing technologies in pedigree composition prediction. By constructing a graph structure based on genotype data similarity and genetic relationships, the model can capture the interactions between gene loci and the relationships between samples, thereby providing a more accurate prediction of pedigree proportion. The present invention not only improves the accuracy of predicting complex pedigree composition, but also can dynamically adjust the estimation of pedigree proportion according to the specific genotype data of the sample, avoiding the limitations of pedigree classification in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A method flow chart of a method for pig breed identification based on deep learning and genetic characteristics provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings, but the protection scope of the present invention is not limited to the following.

[0021] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention. It should be noted that relational terms such as the terms "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0023] Moreover, the terms "comprises," "comprising," or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or machine that includes a list of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article, or machine. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or machine that includes the element.

[0024] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.

[0025] Among them, Figure 1 , a method for pig breed identification based on deep learning and genetic characteristics, comprising the following steps: S1. Collect pig breed information and genome typing test results. The genome typing test targets Landrace, Large White, Duroc, Changda and Duroc-Changda pigs, and uses one of the gene chips from Dangkang 1K, Dangkang 50K, Neogen 50K, SMIC-1 and Dangkang 100K for genome typing; S2. Perform quality control filling on the genotype data, convert the variety information into pedigree proportion according to the genetic law, and recode the genotype data using the One-Hot encoding method; S3. The recoded genotype data are used as nodes in the graph, and the genetic relationships between gene loci are used as edges to construct a graph structure of the genotype data, wherein the nodes are described by gene frequency and allele information, and the edges represent the interactions between genes; S4. Build a complete graph neural network model by learning the interactions between nodes and the structural features of the graph through the graph convolution layer; predict the lineage composition of the sample based on the graph neural network model; S5. Make a breed classification decision based on the predicted results output by the model according to the set pedigree ratio threshold.

[0026] Specifically, a specific implementation plan of the above embodiment is given, pig breed information and genotype data are collected, and data quality control is performed; pig samples of different breeds including Landrace, Large White, Duroc, Changda binary pig, Duroc Changda triple pig, etc. are collected; each sample is genotyped using the following gene chips: Dangkang 1K, Dangkang 50K, Newgen 50K, SMIC No. 1, Dangkang 100K; the pedigree information of each breed is converted into a proportion vector, for example: Landrace is [1, 0, 0], Large White is [0, 1,0], Duroc is [0,0, 1], Changda binary is [0.5, 0.5, 0], Duroc Changda triple is [0.25, 0.25, 0.5].

[0027] Each gene locus is used as a node in the graph, and the genetic relationship between genes is used as an edge connection to construct the graph structure of the genotype; a graph convolution layer (GCN) or a graph attention layer (GAT) is used to learn the dependencies between nodes and the structural characteristics of the graph. Among them, the input: the genotype data after One-Hot encoding and the pedigree proportion of the variety information; the output: the pedigree composition prediction of each sample, including the proportion of different varieties (such as Changbai, Dabai, Duroc, etc.). Independent data sets are constructed for different gene chips (such as Dangkang 1K, Dangkang 50K, etc.). Using the cross-validation method, the hyperparameters of the model (such as learning rate, number of hidden layers, hidden layer dimension, batch size, etc.) are adjusted through the grid search algorithm.

[0028] Select a source domain (such as Landrace and Large White pig datasets) for training, and migrate the model to the target domain (such as Changda binary pig or Duchangda three-way pig dataset), and fine-tune the model parameters in the target domain to adapt it to the new dataset; fine-tune the genotype data in the target domain and adjust some parameters of the model (such as the weight of pedigree proportion) to adapt to new breeds or gene chip data; through fine-tuning, improve the performance of the model on new data, and optimize the threshold of pedigree proportion.

[0029] The graph neural network (GNN) model is used to output the pedigree ratio information of each sample; the pedigree ratio prediction results will be post-processed according to genetic laws; the breed identification threshold is set: when the pedigree ratio of a certain breed is ≥75%, it is judged as a purebred pig; when the pedigree ratio of all breeds is <75% and the Duroc pedigree ratio is <20%, it is judged as a two-way pig; in other cases, it is judged as a three-way pig; breeds are classified according to the pedigree ratio, and dynamically adjusted in combination with genetic laws to ensure the accuracy and flexibility of the classification results.

[0030] Furthermore, in step S2, performing quality control filling on the genotype data specifically includes the following sub-steps: S201. Delete the minimum allele gene frequency less than 0.05; S202. Delete sites with a detection rate lower than 0.8; S203. Deletion sites and Indel sites are filled with NA.

[0031] Furthermore, in step S3, the genetic relationship between gene loci as an edge specifically includes the following sub-steps: S301. defining connections between nodes through a similarity matrix of genotype data; S302. According to the similarity matrix, a weighted adjacency matrix is ​​constructed, wherein each element represents the genotype similarity between nodes, and the adjacency matrix represents the structure of a graph defined based on genotype similarity.

[0032] Furthermore, in step S301, the specific steps of constructing the similarity matrix of genotype data are as follows: S3011. Use Hamming Distance to measure the number of different sites in the corresponding gene loci of two genotype samples; S3012. Construct similarity measurement based on Hamming Distance measurement results; S3013. Construct a similarity matrix based on the similarity measure, namely: ; Wherein, S represents the similarity matrix, and N represents the index of the sample.

[0033] Furthermore, in step S3011, the measurement process of Hamming Distance is: ; ; Among them, the represents the Hamming Distance from sample i to sample j, Represents the total number of sites in the genotype data, Represents the site index of the genotype, represents the indicator function, Indicates that the i-th sample is in The genotype at a gene locus takes a value of 0 or 1. Indicates that the jth sample is in The genotype at a gene locus has a value of 0 or 1.

[0034] For example, suppose the Hamming distance between samples G1 and G2 is calculated: ; Calculate the difference at each site: Site 1: , , are not the same, so ; Site 2: , , are not the same, so ; Site 3: , , are not the same, so ; Site 4: , , are not the same, so .

[0035] so, ; It means that samples G1 and G2 have different genotypes at 3 sites, and the Hamming distance is 3.

[0036] Furthermore, in step S3012, the Hamming Distance measurement result is converted into a similarity measurement by the following formula: ; Among them, the Represents genotype sample i, the represents genotype sample j, Represents the Hamming Distance from sample i to sample j.

[0037] Specifically, through the above measurement, the calculated similarity or correlation value represents the genetic relationship between the samples, that is, the weight of the edge in the graph. When the genotypes of two samples are very similar, the edge weight between the two will be high, and vice versa. In addition, when the genotypes of two samples (pigs) are very similar at multiple sites (that is, they have the same or similar alleles), the edge weight between the two samples will be relatively large, indicating that they are genetically similar. If the genotypes of the two samples are very different at multiple sites, the edge weight between them will be low, indicating that the genetic relationship is distant and they may belong to different breeds or categories.

[0038] Furthermore, through the above adjacency matrix, the genetic similarity between samples can be used in the graph convolutional neural network (GNN) for information propagation and learning. In the subsequent network layers, the genotype information is convolved according to the similarity through the graph neural network, so that the model can effectively capture the relationship between genotype samples, thereby better predicting the ancestry composition.

[0039] Furthermore, the step S4 specifically includes the following sub-steps: S401. taking the genotype data of each pig sample and the adjacency matrix of the constructed graph as inputs to the model; S402. Aggregating information of adjacent nodes in the graph according to the graph convolution layer; and gradually increasing the depth of node representation by stacking multiple layers of graph convolution layers; S403. Designing mean square error as a loss function, with the goal of minimizing the error between the predicted ancestry proportion and the actual label; S404. Predict the ancestry proportion of each sample through the fully connected layer and output the model prediction result.

[0040] Furthermore, in step S402, the information of adjacent nodes in the graph aggregated by the graph convolution layer is specifically expressed as: ; Among them, the Indicates The node representation of the layer represents the activation function, represents the normalized adjacency matrix, Indicates that Indicates The learning weight matrix of the layer, Indicates the number of hidden layers.

[0041] Specifically, the adjacency matrix A is normalized by converting it into a similarity matrix through the inverse relationship of the distance.

[0042] Furthermore, in step S404, the pedigree ratio of each sample is predicted by the fully connected layer as follows: ; Among them, the represents the ancestry proportion vector of the i-th sample, represents the weight matrix of the fully connected layer, represents the representation of the i-th sample in the last layer of the graph neural network. Represents the bias term.

[0043] Furthermore, in step S5, the variety classification decision is specifically as follows: When the proportion of a certain breed in the prediction result is ≥75%, it is judged to be a purebred pig of that breed; When the proportion of each breed pedigree in the prediction result is <75% and the proportion of Duroc pedigree is <20%, it is judged as a dual-purpose pig. In other cases, it is determined to be a three-yuan pig; Right now: ; Among them, the It represents the predicted pedigree percentage of sample i corresponding to variety j, and j=3 represents the Duroc pedigree percentage.

[0044] Furthermore, as a preferred implementation scheme of the above embodiment, the Pearson correlation coefficient can be used instead of the Hamming distance. Specifically, using a more fine-grained correlation measure, the correlation of the genotypes can be expressed as: ; Among them, the and Respectively represent samples and The average value of the genotype at all sites, represents the similarity between sample i and sample j, M represents the total number of genotype sites, and m represents the index of the genotype site.

[0045] Furthermore, as a preferred implementation of the above embodiment, a test data verification step is also included: data that is not used for model training in step S3 and is from a different farm than the data set in step S3 is selected to verify the model prediction accuracy and generalization. Among the five chips, the prediction accuracy of four chips reached 100%, the prediction accuracy of SMIC-1 was 99.970%, and only one misjudgment occurred in 3311 individuals.

[0046] Furthermore, as a preferred implementation of the above embodiment, a system for pig breed identification based on deep learning and genetic characteristics is proposed, comprising: The data collection module is used to collect pig breed information and genome typing test results. The genome typing test targets Landrace, Large White, Duroc, Changda and Duroc, and uses one of the gene chips from Dangkang 1K, Dangkang 50K, Neogen 50K, SMIC-1 and Dangkang 100K for genome typing. The data quality control module is used to perform quality control filling on the genotype data, convert the variety information into pedigree proportion according to the genetic law, and recode the genotype data using the One-Hot encoding method; A model building module is used to construct a graph structure of genotype data by using the recoded genotype data as nodes in a graph and the genetic relationships between gene loci as edges, wherein the nodes are described by gene frequencies and allele information and the edges represent interactions between genes; The model prediction module is used to learn the interactions between nodes and the structural characteristics of the graph through the graph convolution layer to build a complete graph neural network model; based on the graph neural network model, the lineage composition of the sample is predicted; The classification decision module is used to make breed classification decisions on the prediction results output by the model according to the set pedigree ratio threshold.

[0047] In the data quality control module, quality control filling of genotype data specifically includes the following steps: deleting the minimum allele gene frequency less than 0.05; deleting sites with a detection rate lower than 0.8; and filling deletion sites and Indel sites with NA.

[0048] The model building module specifically includes: A node definition unit is used to define the connection between nodes through the similarity matrix of genotype data; The adjacency matrix unit is used to construct a weighted adjacency matrix according to the similarity matrix, wherein each element represents the genotype similarity between nodes, and the adjacency matrix represents the structure of a graph defined based on the genotype similarity.

[0049] In the node definition unit, the similarity matrix of genotype data is specifically constructed including: The Hamming Distance measurement subunit is used to measure the number of different sites of two genotype samples at corresponding gene sites through Hamming Distance; The similarity measurement subunit is used to construct a similarity measurement based on the Hamming Distance measurement result; The similarity matrix construction subunit is used to construct a similarity matrix according to the similarity measure.

[0050] The model prediction module specifically includes: A model input unit, used for taking the genotype data of each pig sample and the adjacency matrix of the constructed graph as inputs of the model; The convolution operation unit is used to aggregate the information of adjacent nodes in the graph according to the graph convolution layer; and gradually increase the depth of node representation by stacking multiple layers of graph convolution layers; The loss function definition unit is used to design the mean square error as the loss function, with the goal of minimizing the error between the predicted ancestry proportion and the actual label; The result prediction unit is used to predict the lineage proportion of each sample through the fully connected layer and output the model prediction result.

[0051] Specifically, the implementation principle process of the above embodiment is as follows: Collect pig breed information and genome typing test results; samples must include Landrace, Large White, Duroc, Changda and Duroc. Each pig must be genome typed using at least one gene chip from Dangkang 1K, Dangkang 50K, Neogen 50K, SMIC-1, and Dangkang 100K; Genotype data quality control filling, breed information and genotype data recoding, where genotype data quality control filling includes: deleting the minimum allele gene frequency less than 0.05; deleting sites with a detection rate lower than 0.8; and filling missing sites and Indel sites with NA. Breed information and genotype data recoding include: converting breed information into pedigree proportions according to genetic laws, such as the breed information of Changbai, Large White, Duroc, Changda Binyuan and Duroc Changda Sanyuan pigs are recoded as [1,0,0], 0,1,0], [0,0,1], [0.5,0.5,0], [0.25,0.25,0.5] respectively; using the One-Hot encoding method to recode genotype data.

[0052] Graph neural network model construction and parameter optimization; the model loss function is mean square error, and Adam is used as the optimizer to minimize the model loss function. Dangkang 1K, Dangkang 50K, Newgen 50K, SMIC-1, and Dangkang 100K each build a data set, divide their data sets into 4:1 ratios for model training, and determine the best hyperparameters for all chips through cross-validation and grid search algorithms. Among them, hyperparameters include: epoch, batchsize, dropout, learning rate, hidden layer, and hidden layer dimension.

[0053] The breed is determined based on the prediction results of the deep learning model and the genetic laws; if the sample's bloodline ratio in a certain breed is greater than or equal to 75%, the sample is judged as a purebred pig of that breed; otherwise, it is further determined whether the sample's bloodline ratio in Duroc is less than 20%, if so, it is judged as a two-way pig, otherwise, it is judged as a three-way pig.

[0054] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A method for pig breed identification based on deep learning and genetic characteristics, characterized in that: The following steps are involved: S1. Collect pig breed information and genome typing test results. The genome typing test targets Landrace, Large White, Duroc, Changda and Duroc-Changda pigs, and uses one of the gene chips from Dangkang 1K, Dangkang 50K, Neogen 50K, SMIC-1 and Dangkang 100K for genome typing; S2. Perform quality control filling on the genotype data, convert the variety information into pedigree proportion according to the genetic law, and recode the genotype data using the One-Hot encoding method; S3. The recoded genotype data are used as nodes in the graph, and the genetic relationships between gene loci are used as edges to construct a graph structure of the genotype data, wherein the nodes are described by gene frequency and allele information, and the edges represent the interactions between genes; S4. Build a complete graph neural network model by learning the interactions between nodes and the structural features of the graph through the graph convolution layer; predict the lineage composition of the sample based on the graph neural network model; S5. Make a breed classification decision based on the predicted results output by the model according to the set pedigree ratio threshold.

2. A method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 1, characterized in that: In step S2, performing quality control filling on genotype data specifically includes the following sub-steps: S201. Delete the minimum allele gene frequency less than 0.05; S202. Delete sites with a detection rate lower than 0.8; S203. Deletion sites and Indel sites are filled with NA.

3. The method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 1, characterized in that: In step S3, the genetic relationship between gene loci as an edge specifically includes the following sub-steps: S301. defining connections between nodes through a similarity matrix of genotype data; S302. According to the similarity matrix, a weighted adjacency matrix is ​​constructed, wherein each element represents the genotype similarity between nodes, and the adjacency matrix represents the structure of a graph defined based on genotype similarity.

4. A method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 3, characterized in that: In step S301, the construction of the similarity matrix of genotype data specifically includes the following steps: S3011. Use Hamming Distance to measure the number of different sites in the corresponding gene loci of two genotype samples; S3012. Construct similarity measurement based on Hamming Distance measurement results; S3013. Construct a similarity matrix based on the similarity measure, namely: ; Wherein, S represents the similarity matrix, and N represents the index of the sample.

5. The method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 4, characterized in that: In step S3011, the measurement process of Hamming Distance is: ; ; Among them, the represents the Hamming Distance from sample i to sample j, Represents the total number of sites in the genotype data, Represents the site index of the genotype, represents the indicator function, Indicates that the i-th sample is in The genotype at a gene locus takes a value of 0 or 1. Indicates that the jth sample is in The genotype at a gene locus has a value of 0 or 1.

6. A method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 4, characterized in that: In step S3012, the Hamming Distance measurement result is converted into a similarity measurement by the following formula: ; Among them, the Represents genotype sample i, the represents genotype sample j, Represents the Hamming Distance from sample i to sample j.

7. The method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 1, characterized in that: The step S4 specifically includes the following sub-steps: S401. taking the genotype data of each pig sample and the adjacency matrix of the constructed graph as inputs to the model; S402. Aggregating information of adjacent nodes in the graph according to the graph convolution layer; And gradually improve the depth of node representation by stacking multiple layers of graph convolutional layers; S403. Designing mean square error as a loss function, with the goal of minimizing the error between the predicted ancestry proportion and the actual label; S404. Predict the ancestry proportion of each sample through the fully connected layer and output the model prediction result.

8. The method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 6, characterized in that: In step S402, the information of adjacent nodes in the graph aggregated by the graph convolution layer is specifically expressed as: ; Among them, the Indicates The node representation of the layer represents the activation function, represents the normalized adjacency matrix, Indicates that Indicates The learning weight matrix of the layer, Indicates the number of hidden layers.

9. The method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 6, characterized in that: In step S404, the ancestry proportion of each sample is predicted by the fully connected layer as follows: ; Among them, the represents the ancestry proportion vector of the i-th sample, represents the weight matrix of the fully connected layer, represents the representation of the i-th sample in the last layer of the graph neural network. Represents the bias term.

10. The method for pig breed identification based on deep learning and genetic characteristics as claimed in claim 1, characterized in that: In step S5, the variety classification decision is specifically as follows: When the proportion of a certain breed in the prediction result is ≥75%, it is judged to be a purebred pig of that breed; When the proportion of each breed pedigree in the prediction result is <75% and the proportion of Duroc pedigree is <20%, it is judged as a dual-purpose pig. In other cases, it is determined to be a three-yuan pig; Right now: ; Among them, the It represents the predicted pedigree percentage of sample i corresponding to variety j, j=3, which represents the Duroc pedigree percentage.