A data management-based and cattle breeding information management method and system
By implementing quality control, standardization, stratified testing, and causal relationship mining of Wagyu breeding data, and combining it with a weighted genome prediction model, the problems of data chaos and environmental interference in traditional Wagyu breeding have been solved. This has enabled precise management and full life-cycle control of breeding data, thereby improving breeding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NIU ZHIGU HLDG (YANGXIN) CO LTD
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional Wagyu cattle breeding suffers from problems such as incomplete genomic information, difficulty in accurately capturing genetic characteristics, insufficient precision in genetic assessment, and scattered data storage and poor security, resulting in low breeding efficiency and difficulty in large-scale application of whole-genome selection technology.
By performing quality control verification and standardization on multi-source breeding pairing data, hierarchical classification and removal of environmentally unstable markers, using a hybrid constraint scoring network algorithm for causal relationship mining, combining a weighted single-step genome linear prediction model for genetic feature sorting and constraint matching, and implementing full-link traceability and access control, precise management of breeding data is achieved.
This has improved the accuracy and standardization of breeding data, ensured the precision and stability of breeding, realized the full life cycle management of Wagyu breeding data, and promoted the development of breeding towards standardization, precision and normalization.
Smart Images

Figure CN122493950A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data management technology, and more specifically, to a method and system for managing Wagyu cattle breeding information based on data management. Background Technology
[0002] Thoroughbred Wagyu cattle are the core breed for producing high-end Wagyu beef. Currently, the high-end Wagyu beef industry faces challenges such as low Wagyu cattle numbers, a lack of core breeds, reliance on crossbreeding for improvement, and an imperfect breeding system. Overall breeding efficiency is low, genetic progress is slow, and it is difficult to form a large-scale, high-quality industrial cluster. Genome-wide selection technology can achieve early and precise selection, shorten generation intervals, and accelerate genetic improvement, and has become an important development direction for beef cattle breeding.
[0003] However, there are still many technical bottlenecks in the current Wagyu breeding process, which restrict the implementation and industrialization of whole-genome selection technology:
[0004] The traditional Wagyu genome contains a large number of sequence gaps and complex repetitive regions, and the genomic information is not complete enough, making it difficult to accurately capture the unique genetic characteristics of Wagyu cattle. As a result, genome-based breeding decisions lack reliable data support and cannot achieve precision breeding.
[0005] Furthermore, the lack of mining and analysis mechanisms for massive sequencing data makes the stability of genetic markers susceptible to interference from environmental factors, and the determination of causal relationships among markers is not precise enough, making it difficult to distinguish between environmental influences and the role of genetics itself, which further reduces the accuracy of breeding assessment.
[0006] In addition, traditional Wagyu breeding relies heavily on human experience and judgment, resulting in insufficient accuracy in genetic assessment. Furthermore, breeding data is stored in a scattered manner, lacks unified standards, and suffers from poor data security and traceability, which affects breeding efficiency and restricts the large-scale application of whole-genome selection technology in Wagyu breeding.
[0007] There are currently no effective solutions to the problems in the relevant technologies. Summary of the Invention
[0008] In view of the problems in related technologies, this invention proposes a data management-based Wagyu cattle breeding information management method and system to overcome the aforementioned technical problems existing in the existing related technologies.
[0009] Therefore, the specific technical solution adopted by the present invention is as follows:
[0010] In a first aspect, the present invention proposes a method for managing Wagyu cattle breeding information based on data management, comprising:
[0011] S1. Perform quality control verification and standardization on the multi-source breeding pairing data to obtain a standardized breeding dataset;
[0012] S2. Based on the feeding environment factors, the standardized breeding dataset is hierarchically divided and labeled, and unstable environmental labels are removed to obtain environmental candidate labels and initial weight results.
[0013] S3. Using the hybrid constraint scoring network algorithm, causal relationship mining and site quantification are performed on environmental candidate markers and initial weight results to obtain environmental breeding causal markers and optimized site weight results.
[0014] S4. Using a weighted single-step genome linear prediction model, the genetic characteristics of the optimized site weights and the Wagyu whole genome data are sorted and constrained to obtain a breeding selection pairing dataset.
[0015] S5. Perform full-link traceability and access control on standardized breeding datasets, environmental breeding causal markers, and breeding selection and pairing datasets to obtain a full-link breeding data management archive, so as to realize the management of Wagyu breeding information and traceable breeding data.
[0016] Furthermore, the multi-source breeding pairing data underwent quality control verification and standardization processing to obtain a standardized breeding dataset, including:
[0017] S11. Obtain multi-source breeding pairing data, which includes Wagyu individual phylogenetic data, phenotypic determination data, genomic typing data, and pairing record data;
[0018] S12. Perform missing value detection on the multi-source breeding pairing data, and remove data fields with missing rates higher than a preset threshold to obtain the detected dataset.
[0019] S13. Using the three-standard-deviation principle, outlier detection is performed on the detected dataset, and a second verification is performed in combination with the threshold in the Wagyu breeding field to remove invalid outlier data and obtain the initial quality control screening dataset.
[0020] S14. Convert the unstructured paired record data in the initial screening dataset of quality control into structured data to obtain a standardized dataset;
[0021] S15. Using the maximum-minimum normalization method, the standardized dataset is unified in terms of dimensions, and the unified dataset is validated across all fields to generate a standardized breeding dataset.
[0022] Furthermore, based on feeding environment factors, the standardized breeding dataset was stratified and labeled, and environmentally unstable labels were removed, resulting in candidate environmental labels and initial weights, including:
[0023] S21. Obtain the feeding environment factors and divide the standardized breeding dataset into multiple environmental stratification subsets based on the feeding environment factors; wherein, the feeding environment factors include feeding density, dietary nutrient level, indoor temperature and humidity, and feeding area.
[0024] S22. Extract all genetic markers in multiple environmental stratified subsets respectively. Calculate the locus effect of each genetic marker in the corresponding environmental stratified subset using the breeding locus effect method, and record the causal edge direction corresponding to each marker to obtain a record table of marker effect and causal edge direction.
[0025] S23. Based on the preset locus effect size deviation threshold and the causal edge direction consistency judgment rule, compare the effect size deviation of the same locus genetic marker in different environmental stratified subsets, and verify the consistency of the causal edge direction of the same locus marker in different environmental stratified subsets.
[0026] S24. Based on the consistency test results, identify the markers whose site effect size deviation between different environmental stratified subsets exceeds the preset threshold, and the markers whose causal edge direction is inconsistent with that between different environmental stratified subsets, and remove all unstable markers.
[0027] S25. For the effective genetic markers after removing unstable markers, combine the mean locus effect size in each environmental stratification subset and assign initial locus weights.
[0028] S26. Integrate all retained valid genetic markers and initial site weights to generate a list of environmental candidate markers and initial weight results.
[0029] Furthermore, for the effective genetic markers after removing unstable markers, the mean locus effect size within each environmental stratification subset is combined, and initial locus weights are assigned, including:
[0030] S251. Based on the set of effective genetic markers after removing unstable markers, extract the locus effect size of each marker in the environmental stratification subset.
[0031] S252. Calculate the mean of the locus effect size of the same effective genetic marker in all environmental stratified subsets to obtain the uniform effect size mean.
[0032] S253. Use the mean effect size as the assigned value, and assign initial site weights to each effective genetic marker according to the weighting rules.
[0033] Furthermore, using a hybrid constraint scoring network algorithm, causal relationship mining and site quantification are performed on environmental candidate markers and initial weight results to obtain environmental breeding causal markers and optimized site weight results, including:
[0034] S31. Using the hybrid constraint scoring network algorithm, causal mining and site quantification are performed on the environmental candidate labels and initial weight results. Each label is set as a Bayesian network node, an initial node set of a directed acyclic graph is constructed, and the constraint conditions of the hybrid constraint scoring network algorithm are set.
[0035] S32. Calculate the candidate network structure score using the Bayesian information criterion structure scoring function.
[0036] S33. Based on the hybrid constraint scoring network algorithm, the constraint conditions and candidate network structure scores are obtained by traversing the network structure space using the classic hill-climbing search method. The initial weights are used as the starting point for iteration. The presence and direction of directed edges between nodes are adjusted, and the Bayesian information criterion structure scoring function is substituted to calculate the score of each candidate structure.
[0037] S34. Based on the optimal network structure, the intensity of the causal effect of each label is quantified using the Bayesian parameter estimation function, and weighted correction is completed by combining the initial weights.
[0038] S35. Extract the markers with causal relationships in the optimal structure as environmental breeding causal markers, integrate the quantified optimized site weights, and obtain the environmental breeding causal marker and optimized site weight results.
[0039] Furthermore, based on the constraints and candidate network structure scores of the hybrid constraint scoring network algorithm, the network structure space is traversed using a classic hill-climbing search method. Starting with the initial weights, the presence and direction of directed edges between nodes are adjusted, and the scores for each candidate structure are calculated using the Bayesian information criterion structure scoring function.
[0040] S331. The initial weights based on the environmental candidate labels are used as the starting point for the hill-climbing search method, and the current state and iteration number threshold of the network structure search are initialized.
[0041] S332. Under the constraints of the hybrid constraint scoring network algorithm, perform single-step local operations on the directed edges between Bayesian network nodes. The single-step local operations include adding directed edges, deleting directed edges, and flipping the direction of directed edges.
[0042] S333. After each directed edge local operation is completed, a candidate Bayesian network structure with a corresponding topology is generated.
[0043] S334. Input each generated candidate network structure into the Bayesian information criterion structure scoring function to complete the score calculation of the candidate structure, and record each candidate network structure and its score value in real time.
[0044] Furthermore, using a weighted single-step genome linear prediction model, the optimized site weights were matched with the Wagyu cattle whole genome data for genetic trait ranking and constraint matching, resulting in a breeding selection pairing dataset including:
[0045] S41. A weighted single-step genome prediction model composed of pedigree matrix construction technology, genome relation matrix construction technology, site weighted fusion technology, and hybrid linear model solution technology;
[0046] S42. Based on the whole genome data of Wagyu cattle, perform quality control on polymorphic genotype data and remove missing, abnormal, and non-compliant sample data from individual phylogenetic information;
[0047] S43. Based on the processed individual phylogenetic information, construct a phylogenetic additive genetic relationship matrix;
[0048] S44. Based on the quality-controlled Wagyu whole-genome polymorphism site data, a standard genome relation matrix is constructed; and the optimized site weights are assigned to the standard genome relation matrix to construct a weighted genome relation matrix, thereby completing the fusion of weight information and genome data;
[0049] S45. Merge the pedigree additive genetic relationship matrix with the weighted genomic relationship matrix to construct a weighted single-step genomic relationship matrix;
[0050] S46. Based on the weighted single-step genome relation matrix, the mixed linear model is solved using the constrained maximum likelihood method to obtain the estimated breeding value of each candidate individual;
[0051] S47. Based on the numerical value of the estimated breeding value from the genome, rank the candidate individuals according to their genetic characteristics.
[0052] S48. Based on the breeding selection constraints, perform constraint matching screening on the ranked individuals to eliminate individuals with inbreeding risk and those that do not meet the breeding objectives;
[0053] S49. Integrate the genetic trait ranking results with the constraint matching screening results to obtain the breeding selection pair dataset.
[0054] Furthermore, based on the weighted single-step genomic relation matrix, the constrained maximum likelihood method is used to solve the mixed linear model, yielding the estimated genomic breeding values for each candidate individual, including:
[0055] S461. Using the weighted single-step genome relation matrix as the basis of genetic covariance, construct a mixed linear model of fixed effects and random effects;
[0056] S462. Using the constrained maximum likelihood method, iteratively solve the variance components of the mixed linear model to obtain the unbiased variance components.
[0057] S463. Based on the variance component estimation results and the weighted single-step genome relationship matrix, calculate the random effect prediction value for each candidate individual;
[0058] S464. Integrate the fixed-effects estimates and random-effects predictions to obtain the genomic breeding values for each candidate individual.
[0059] Furthermore, full-chain traceability and access control are implemented on standardized breeding datasets, environmental breeding causal markers, and breeding selection and pairing datasets to obtain a full-chain breeding data management archive. This enables the management of Wagyu breeding information and ensures traceable breeding data, including:
[0060] S51. Data source nodes, processing nodes, and uniqueness traceability annotations are performed on the standardized breeding dataset, environmental breeding causal marker dataset, and breeding selection and pairing dataset to obtain basic information for full-link traceability.
[0061] S52. Verify the basic information of the whole-link traceability to obtain the verified whole-link traceability information;
[0062] S53. Based on the Wagyu breeding data management specifications, construct hierarchical permission control rules for data access permissions, data editing permissions, data export permissions, and data review permissions;
[0063] S54. Match and integrate the hierarchical permission control rules with the verified full-link traceability information, and perform complete verification, permission verification and traceability rationality verification on the integrated information results to obtain qualified Wagyu breeding control information, and standardize and structure it for archiving. Based on the sorting and archiving results, generate a full-link breeding data control archive.
[0064] Secondly, the present invention also provides a data management-based Wagyu cattle breeding information management system, comprising:
[0065] The data quality control and standardization module is used to perform quality control verification and standardization processing on multi-source breeding pairing data to obtain a standardized breeding dataset.
[0066] The environmental factor stratification test module is used to perform stratification and label testing on the standardized breeding dataset based on feeding environmental factors, and remove environmentally unstable labels to obtain environmental candidate labels and initial weight results.
[0067] The causal relationship mining and quantification module is used to mine causal relationships and quantify sites on environmental candidate labels and initial weight results using a hybrid constraint scoring network algorithm, so as to obtain environmental breeding causal labels and optimized site weight results.
[0068] The genetic feature ranking and matching module is used to rank and match the genetic features of the optimized site weights with the whole genome data of Wagyu cattle using a weighted single-step genome linear prediction model, so as to obtain a breeding selection and pairing dataset.
[0069] The end-to-end traceability and control module is used to perform end-to-end traceability and access control on standardized breeding datasets, environmental breeding causal markers, and breeding selection and pairing datasets to obtain a full-chain breeding data control archive, so as to realize the management of Wagyu breeding information and traceable breeding data.
[0070] The beneficial effects of this invention are as follows:
[0071] 1) This invention achieves precise control of Wagyu breeding data by using unified end-to-end data quality control and standardized processing, combined with a weighted single-step genome prediction model, thus solving the problems of chaotic and insufficient standardization in traditional breeding data and improving the accuracy and standardization of breeding data.
[0072] 2) This invention solves the problems of unstable markers and large genetic evaluation bias caused by environmental interference in traditional breeding by combining environmental stratification test and stable genetic marker screening with hybrid constraint scoring network algorithm, thereby ensuring the accuracy and stability of Wagyu breeding.
[0073] 3) This invention achieves full lifecycle management of Wagyu breeding data by combining full-process data traceability and access control with standardized data processing and precise genetic evaluation, thereby improving breeding efficiency and promoting the development of Wagyu breeding towards standardization, precision and normalization. It also solves the shortcomings of traditional breeding, such as data dispersion, difficulty in traceability and selection of pairs. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0075] Figure 1 This is a flowchart of a data management-based Wagyu breeding information management method according to an embodiment of the present invention.
[0076] Figure 2 This is a schematic diagram of a data management-based Wagyu cattle breeding information management system according to an embodiment of the present invention. Detailed Implementation
[0077] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.
[0078] According to an embodiment of the present invention, a method and system for managing Wagyu cattle breeding information based on data management is proposed.
[0079] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, a data management-based Wagyu breeding information management method according to an embodiment of the present invention includes:
[0080] Step S1: Perform quality control verification and standardization on the multi-source breeding pairing data to obtain a standardized breeding dataset;
[0081] Step S2: Based on the feeding environment factors, the standardized breeding dataset is hierarchically divided and labeled, and unstable environmental labels are removed to obtain environmental candidate labels and initial weight results.
[0082] Step S3: Using the hybrid constraint scoring network algorithm, perform causal relationship mining and site quantification on environmental candidate markers and initial weight results to obtain environmental breeding causal markers and optimized site weight results;
[0083] Step S4: Using a weighted single-step genome linear prediction model, the genetic characteristics of the optimized site weights are sorted and constraints are matched with the Wagyu whole genome data to obtain a breeding selection pairing dataset.
[0084] Step S5: Perform full-link traceability and access control on the standardized breeding dataset, environmental breeding causal markers, and breeding selection and pairing dataset to obtain a full-link breeding data control archive, so as to realize the management of Wagyu breeding information and traceable breeding data.
[0085] In this optional embodiment, the multi-source breeding pairing data undergoes quality control verification and standardization processing to obtain a standardized breeding dataset, including:
[0086] S11. Obtain multi-source breeding pairing data, which includes Wagyu individual phylogenetic data, phenotypic determination data, genomic typing data, and pairing record data;
[0087] S12. Perform missing value detection on the multi-source breeding pairing data, and remove data fields with missing rates higher than a preset threshold to obtain the detected dataset.
[0088] S13. Using the three-standard-deviation principle, outlier detection is performed on the detected dataset, and a second verification is performed in combination with the threshold in the Wagyu breeding field to remove invalid outlier data and obtain the initial quality control screening dataset.
[0089] S14. Convert the unstructured paired record data in the initial screening dataset of quality control into structured data to obtain a standardized dataset;
[0090] S15. Using the maximum-minimum normalization method, the standardized dataset is unified in terms of dimensions, and the unified dataset is validated across all fields to generate a standardized breeding dataset.
[0091] Specifically, the data was compiled using multi-source breeding pairing data from 1200 individuals at a core Wagyu cattle breeding farm. This included individual pedigree profiles, intramuscular fat content, marbling scores, phenotypic data, genomic typing data, and unstructured mating records. Missing values were statistically analyzed for each data field, and invalid fields exceeding the 10% threshold were removed to complete the initial screening of missing values. Simultaneously, outlier identification was performed on numerical phenotypic and genotypic data using the three-standard-deviation principle. Suspected outliers were further verified using established thresholds in Wagyu cattle breeding (e.g., intramuscular fat content ≥8%) to remove invalid data that did not meet breeding standards and generate the initial quality control screening dataset.
[0092] Unstructured mating records are converted into structured data format through field mapping and rule matching to achieve data format standardization. Max-min normalization is used to unify the dimensions of each feature dimension. Simultaneously, the consistency and integrity of all fields of data are verified, generating a standardized breeding dataset of 1080 qualified individuals.
[0093] In this optional embodiment, based on feeding environment factors, the standardized breeding dataset is stratified and labeled, and labels with unstable environments are removed to obtain candidate environmental labels and initial weights, including:
[0094] S21. Obtain the feeding environment factors and divide the standardized breeding dataset into multiple environmental stratification subsets based on the feeding environment factors; wherein, the feeding environment factors include feeding density, dietary nutrient level, indoor temperature and humidity, and feeding area.
[0095] S22. Extract all genetic markers in multiple environmental stratified subsets respectively. Calculate the locus effect of each genetic marker in the corresponding environmental stratified subset using the breeding locus effect method, and record the causal edge direction corresponding to each marker to obtain a record table of marker effect and causal edge direction.
[0096] S23. Based on the preset locus effect size deviation threshold and the causal edge direction consistency judgment rule, compare the effect size deviation of the same locus genetic marker in different environmental stratified subsets, and verify the consistency of the causal edge direction of the same locus marker in different environmental stratified subsets.
[0097] S24. Based on the consistency test results, identify the markers whose site effect size deviation between different environmental stratified subsets exceeds the preset threshold, and the markers whose causal edge direction is inconsistent with that between different environmental stratified subsets, and remove all unstable markers.
[0098] S25. For the effective genetic markers after removing unstable markers, combine the mean locus effect size in each environmental stratification subset and assign initial locus weights.
[0099] S26. Integrate all retained valid genetic markers and initial site weights to generate a list of environmental candidate markers and initial weight results.
[0100] In this optional embodiment, for the effective genetic markers after removing unstable markers, the mean locus effect size within each environmental stratification subset is combined, and initial locus weights are assigned, including:
[0101] S251. Based on the set of effective genetic markers after removing unstable markers, extract the locus effect size of each marker in the environmental stratification subset.
[0102] S252. Calculate the mean of the locus effect size of the same effective genetic marker in all environmental stratified subsets to obtain the uniform effect size mean.
[0103] S253. Use the mean effect size as the assigned value, and assign initial site weights to each effective genetic marker according to the weighting rules.
[0104] Specifically, using a standardized breeding dataset of 1080 purebred Wagyu cattle as input, real environmental factors for each individual were collected, covering four core indicators: stocking density, dietary nutrient level, indoor temperature and humidity, and feeding region. Stocking density was divided into three levels: 4 or more cattle per pen, 5 to 8 cattle per pen, and 9 or more cattle per pen. Dietary nutrient level was divided into three gradients: high concentrate, medium nutrient, and low energy. Indoor temperature and humidity were divided into three categories: 15 to 25 degrees Celsius and 50 to 70% suitable area, high temperature and high humidity area, and low temperature and low humidity area. Based on environmental factors, the standardized breeding dataset was stratified into 12 environmental stratification subsets, with each subset having a sample size of 85 to 95 cattle. Genetic markers of all genomic typing data in each environmental stratification subset were extracted. Using a quantitative genetics mixed linear method, the locus effect size of each genetic marker in the corresponding environmental stratification subset was recorded. Simultaneously with the Bayesian network structure learning results, the causal edge direction corresponding to each marker was recorded, generating a record table of marker effect size and causal edge direction.
[0105] Using a preset locus effect size deviation threshold of 0.3 and 80% causal edge direction consistency as the criterion, the effect size deviation of genetic markers at the same locus is compared across different environmental stratification subsets. The consistency of causal edge direction for the same locus marker within each subset is verified, and unstable genetic markers with effect size deviations exceeding the threshold or inconsistent causal edge directions are removed. For the remaining valid genetic markers after removing unstable markers, their mean locus effect size within each environmental stratification subset is calculated. Based on the mean effect size, initial locus weights are assigned according to normalization rules. All remaining valid genetic markers are integrated with their corresponding initial locus weights to generate an environmental candidate marker list and initial weight results.
[0106] In this optional embodiment, a hybrid constraint scoring network algorithm is used to mine causal relationships and quantify sites on environmental candidate markers and initial weight results, resulting in environmental breeding causal markers and optimized site weights, including:
[0107] S31. Using the hybrid constraint scoring network algorithm, causal mining and site quantification are performed on the environmental candidate labels and initial weight results. Each label is set as a Bayesian network node, an initial node set of a directed acyclic graph is constructed, and the constraint conditions of the hybrid constraint scoring network algorithm are set.
[0108] S32. Calculate the candidate network structure score using the Bayesian information criterion structure scoring function.
[0109] S33. Based on the hybrid constraint scoring network algorithm, the constraint conditions and candidate network structure scores are obtained by traversing the network structure space using the classic hill-climbing search method. The initial weights are used as the starting point for iteration. The presence and direction of directed edges between nodes are adjusted, and the Bayesian information criterion structure scoring function is substituted to calculate the score of each candidate structure.
[0110] S34. Based on the optimal network structure, the intensity of the causal effect of each label is quantified using the Bayesian parameter estimation function, and weighted correction is completed by combining the initial weights.
[0111] S35. Extract the markers with causal relationships in the optimal structure as environmental breeding causal markers, integrate the quantified optimized site weights, and obtain the environmental breeding causal marker and optimized site weight results.
[0112] In this optional embodiment, based on the constraints of the hybrid constraint scoring network algorithm and the scores of candidate network structures, the network structure space is traversed using a classic hill-climbing search method. Starting with the initial weights as the iteration starting point, the presence and direction of directed edges between nodes are adjusted, and the scores of each candidate structure are calculated using the Bayesian information criterion structure scoring function.
[0113] S331. The initial weights based on the environmental candidate labels are used as the starting point for the hill-climbing search method, and the current state and iteration number threshold of the network structure search are initialized.
[0114] S332. Under the constraints of the hybrid constraint scoring network algorithm, perform single-step local operations on the directed edges between Bayesian network nodes. The single-step local operations include adding directed edges, deleting directed edges, and flipping the direction of directed edges.
[0115] S333. After each directed edge local operation is completed, a candidate Bayesian network structure with a corresponding topology is generated.
[0116] S334. Input each generated candidate network structure into the Bayesian information criterion structure scoring function to complete the score calculation of the candidate structure, and record each candidate network structure and its score value in real time.
[0117] Specifically, using 8200 Wagyu SNP environmental candidate markers screened through environmental stability testing, their corresponding initial locus weights (i.e., weight distribution range of 0.10-0.90), and a standardized breeding dataset of 1080 Thoroughbred Wagyu cattle as input, a hybrid constraint scoring network algorithm was used for causal relationship mining and locus quantification. All candidate genetic markers were set as Bayesian network nodes, constructing an initial node set of a directed acyclic graph. Simultaneously, constraints were set for the hybrid constraint scoring network algorithm: prohibiting directed cycles, ensuring the linkage disequilibrium coefficient between genetic loci is less than or equal to 0.2, and retaining nodes with direct associations to feeding environment factors. The BIC score of the initial candidate network structure was calculated to be -12860 using the Bayesian Information Criterion Structure Scoring Function. The formula for calculating the Bayesian Information Criterion Structure Scoring Function is as follows:
[0118] ;
[0119] In the formula, Score candidate network structures G using the Bayesian information criterion for a given dataset K; n is the total number of network nodes; Let be the number of state combinations of the parent node set of the i-th node; Let be the number of discrete states of the i-th node; The number of samples in the dataset where the i-th node takes the k-th state and its parent node takes the j-th state; These are the conditional probability parameters under the corresponding conditions; This represents the total sample size.
[0120] Based on the constraints and initial structure scores of the hybrid constraint scoring network algorithm, the initial weights of each marker are used as the starting point for iteration. The network structure space is traversed using the classic hill-climbing search method. The presence and direction of directed edges between nodes are adjusted sequentially. After each structural adjustment, the Bayesian information criterion structure scoring function is substituted to calculate the BIC score of the corresponding candidate structure. After a maximum of 200 iterations, the optimal network structure with a BIC score of -8420 is obtained. Based on the optimal network structure, the causal effect strength of each marker node is quantified using the Bayesian parameter estimation function. The quantified value of the causal effect of each marker is in the range of 0.06-0.93. Combined with the initial site weights of each marker, a weighted correction is performed. 4582 genetic markers with stable causal associations in the optimal network structure are extracted as environmental breeding causal markers. At the same time, the optimized site weights after causal effect quantification correction are integrated to obtain the set of environmental breeding causal markers and the corresponding optimized site weight results.
[0121] The constraints of the hybrid constraint scoring network algorithm are as follows: the network is limited to a directed acyclic graph, causal loops are prohibited, hill-climbing search only adjusts the presence and direction of directed edges between nodes, without adding or deleting nodes or changing node types, and a maximum termination condition of 200 iterations is set. The Bayesian information criterion is used as the sole basis for structural scoring. Network nodes use genetic markers to be screened, irrelevant variables are not included, the initial site weights of each marker are used as the initial benchmark for iteration, the causal effect quantification value is constrained to be in the range of 0.06 to 0.93, and only structural relationships with stable causal associations in the network are retained.
[0122] In this optional embodiment, a weighted single-step genomic linear prediction model is used to rank and match the optimized site weights with the Wagyu whole genome data based on genetic characteristics, resulting in a breeding selection pairing dataset including:
[0123] S41. A weighted single-step genome prediction model composed of pedigree matrix construction technology, genome relation matrix construction technology, site weighted fusion technology, and hybrid linear model solution technology;
[0124] S42. Based on the whole genome data of Wagyu cattle, perform quality control on polymorphic genotype data and remove missing, abnormal, and non-compliant sample data from individual phylogenetic information;
[0125] S43. Based on the processed individual phylogenetic information, construct a phylogenetic additive genetic relationship matrix;
[0126] S44. Based on the quality-controlled Wagyu whole-genome polymorphism site data, a standard genome relation matrix is constructed; and the optimized site weights are assigned to the standard genome relation matrix to construct a weighted genome relation matrix, thereby completing the fusion of weight information and genome data;
[0127] S45. Merge the pedigree additive genetic relationship matrix with the weighted genomic relationship matrix to construct a weighted single-step genomic relationship matrix;
[0128] S46. Based on the weighted single-step genome relation matrix, the mixed linear model is solved using the constrained maximum likelihood method to obtain the estimated breeding value of each candidate individual;
[0129] S47. Based on the numerical value of the estimated breeding value from the genome, rank the candidate individuals according to their genetic characteristics.
[0130] S48. Based on the breeding selection constraints, perform constraint matching screening on the ranked individuals to eliminate individuals with inbreeding risk and those that do not meet the breeding objectives;
[0131] S49. Integrate the genetic trait ranking results with the constraint matching screening results to obtain the breeding selection pair dataset.
[0132] In this optional embodiment, based on the weighted single-step genomic relation matrix, the constrained maximum likelihood method is used to solve the mixed linear model to obtain the genomic estimated breeding values for each candidate individual, including:
[0133] S461. Using the weighted single-step genome relation matrix as the basis of genetic covariance, construct a mixed linear model of fixed effects and random effects;
[0134] S462. Using the constrained maximum likelihood method, iteratively solve the variance components of the mixed linear model to obtain the unbiased variance components.
[0135] S463. Based on the variance component estimation results and the weighted single-step genome relationship matrix, calculate the random effect prediction value for each candidate individual;
[0136] S464. Integrate the fixed-effects estimates and random-effects predictions to obtain the genomic breeding values for each candidate individual.
[0137] Specifically, using the optimized locus weights (i.e., weight distribution range of 0.12-0.92) corresponding to 4582 environmental breeding causal markers, whole-genome genotyping data of 1080 purebred Wagyu cattle, complete pedigree information of three generations of corresponding individuals, and meat quality phenotypic data such as intramuscular fat content and marbling score as inputs, quality control processing was performed on the whole-genome polymorphism genotype data and individual pedigree information of Wagyu cattle. The criteria included a locus detection rate greater than or equal to 95%, a minor allele frequency greater than or equal to 0.01, and a Hardy-Weinberg equilibrium test p-value greater than or equal to 1×10⁻⁶. −6 To meet quality control standards, missing, abnormal, and non-compliant loci and individuals were removed, retaining 1056 qualified individuals and valid SNP loci. Based on the complete pedigree information of the individuals after quality control, a pedigree additive genetic relationship matrix A covering more than 1200 related individuals was constructed. Simultaneously, a standard genome relationship matrix G was constructed based on the quality-controlled Wagyu whole-genome polymorphic locus data. The aforementioned optimized locus weights were assigned to the standard genome relationship matrix according to their corresponding locus positions, constructing a weighted genome relationship matrix, thus completing the deep integration of locus weight information and whole-genome data.
[0138] The pedigree additive genetic relationship matrix and the weighted genomic relationship matrix were fused to construct a weighted single-step genomic relationship matrix H. Based on this matrix, a mixed linear model was constructed, with year, sex, and initial weight as fixed effects and individual additive genetic effects as random effects. The variance components of the mixed linear model were iteratively solved using the constrained maximum likelihood method. After 35 iterations, the unbiased variance components were obtained. Based on the variance component estimation results and the weighted single-step genomic relationship matrix, the predicted random effects of each candidate individual were calculated. The fixed effects estimates and the predicted random effects were integrated to obtain the genomic estimated breeding value of the meat quality trait for each candidate individual. The breeding value distribution range was -2.35 to 3.68. The 1056 candidate individuals were ranked according to their genetic characteristics based on the estimated breeding values of the genome, from highest to lowest. Based on the breeding selection constraints, namely, the inbreeding coefficient of individuals is less than or equal to 6%, the estimated breeding value of the genome is in the top 30%, and the breeding value of the marbled trait is greater than or equal to 1.2, the ranked individuals were subjected to constraint matching screening. Individuals with inbreeding risk and those that do not meet the breeding goals of high-end meat quality were removed. The genetic characteristic ranking results and constraint matching screening results were integrated to generate a breeding selection pairing dataset of 286 core breeding candidate individuals.
[0139] In this optional embodiment, full-link traceability and access control are performed on the standardized breeding dataset, environmental breeding causal markers, and breeding selection and selection matching dataset to obtain a full-link breeding data management archive, thereby enabling the management of Wagyu breeding information and traceable breeding data, including:
[0140] S51. Data source nodes, processing nodes, and uniqueness traceability annotations are performed on the standardized breeding dataset, environmental breeding causal marker dataset, and breeding selection and pairing dataset to obtain basic information for full-link traceability.
[0141] S52. Verify the basic information of the whole-link traceability to obtain the verified whole-link traceability information;
[0142] S53. Based on the Wagyu breeding data management specifications, construct hierarchical permission control rules for data access permissions, data editing permissions, data export permissions, and data review permissions;
[0143] S54. Match and integrate the hierarchical permission control rules with the verified full-link traceability information, and perform complete verification, permission verification and traceability rationality verification on the integrated information results to obtain qualified Wagyu breeding control information, and standardize and structure it for archiving. Based on the sorting and archiving results, generate a full-link breeding data control archive.
[0144] Specifically, using a standardized breeding dataset of 1056 thoroughbred Wagyu cattle, an environmental breeding causal marker dataset including 4582 valid loci, and a breeding selection and pairing dataset covering 286 core breeding candidates as input, full-link traceability annotation was performed on the three core datasets. Each dataset was annotated with its corresponding data source nodes (i.e., unique ear tag identifier for each individual, sample collection breeding farm location, genotyping testing institution, and phenotyping time and personnel), full-process processing nodes, and unique codes for dataset and individual dimensions, generating basic information for full-link traceability. Simultaneously, the basic information for full-link traceability was verified for completeness, consistency, and uniqueness, eliminating abnormal content such as broken traceability links, duplicate identifiers, and mismatched information. After verification, full-link traceability information with 100% traceability link completeness and 100% individual identifier matching was obtained.
[0145] Based on industry standards for livestock and poultry genetic resource data management and data management requirements for Wagyu core breeding farms, a hierarchical access control rule was constructed, including data access permissions, data editing permissions, data export permissions, and data review permissions. The hierarchical access control rule was matched and integrated with the verified full-chain traceability information according to dataset and individual dimensions. The integrated information results were then subjected to integrity verification, permission conflict verification, and traceability rationality verification in sequence, with a 100% pass rate for all verifications. The verified Wagyu breeding control information was then standardized, organized, and structured for archiving. The organized and archived results were used to generate a full-chain breeding data control archive to achieve traceability and controllability of Wagyu breeding data throughout the entire process.
[0146] Table 1 Summary Table of Core Input / Output Datasets
[0147] Input data Standardized breeding dataset Covering 1056 thoroughbred Wagyu cattle, including pedigree, phenotype, and whole-genome typing data. Standardized quality control output Input data Environmental breeding causal marker dataset It contains 4582 valid sites and corresponding optimized site weights (i.e., distribution range 0.12-0.92). Causal mining output Input data Breeding selection and pairing dataset It covers 286 core breeding candidates, including information such as genome-estimated breeding values and inbreeding coefficients. Breeding value calculation output Output data Breeding end-to-end data management archive Includes full-chain traceability information, hierarchical permission configuration, and comprehensive archiving information of core breeding data. Full-process control output
[0148] The core input and output data, access control rules, and verification indicators of this invention are shown in Table 1. Through full-link traceability labeling and hierarchical access control, the traceability and controllability of the entire Wagyu breeding process data are realized, ensuring the security and integrity of the breeding data.
[0149] like Figure 2 As shown, according to another embodiment of the present invention, a data management-based Wagyu breeding information management system is also provided, comprising:
[0150] The data quality control and standardization module 101 is used to perform quality control verification and standardization processing on multi-source breeding pairing data to obtain a standardized breeding dataset.
[0151] The environmental factor stratification test module 102 is used to perform stratification and label test on the standardized breeding dataset based on feeding environmental factors, and remove environmentally unstable labels to obtain environmental candidate labels and initial weight results.
[0152] The causal relationship mining and quantification module 103 is used to mine causal relationships and quantify sites on environmental candidate labels and initial weight results using a hybrid constraint scoring network algorithm, so as to obtain environmental breeding causal labels and optimized site weight results.
[0153] The genetic feature sorting and matching module 104 is used to sort and match the genetic features of the optimized site weights with the whole genome data of Wagyu cattle using a weighted single-step genome linear prediction model, so as to obtain a breeding selection and pairing dataset.
[0154] The end-to-end traceability and control module 105 is used to perform end-to-end traceability and access control on standardized breeding datasets, environmental breeding causal markers, and breeding selection and pairing datasets to obtain a full-chain breeding data control archive, so as to realize the management of Wagyu breeding information and traceable breeding data.
[0155] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data-driven method for managing Wagyu cattle breeding information, characterized in that, include: S1. Perform quality control verification and standardization on the multi-source breeding pairing data to obtain a standardized breeding dataset; S2. Based on the feeding environment factors, the standardized breeding dataset is hierarchically divided and labeled, and unstable environmental labels are removed to obtain candidate environmental labels and initial weight results. S3. Using the hybrid constraint scoring network algorithm, causal relationship mining and site quantification are performed on environmental candidate markers and initial weight results to obtain environmental breeding causal markers and optimized site weight results. S4. Using a weighted single-step genome linear prediction model, the genetic characteristics of the optimized site weights and the Wagyu whole genome data are sorted and constrained to obtain a breeding selection pairing dataset. S5. Perform full-link traceability and access control on standardized breeding datasets, environmental breeding causal markers, and breeding selection and pairing datasets to obtain a full-link breeding data management archive, so as to realize the management of Wagyu breeding information and traceable breeding data.
2. The Wagyu breeding information management method based on data management according to claim 1, characterized in that, The process of quality control verification and standardization of multi-source breeding pairing data to obtain a standardized breeding dataset includes: S11. Obtain multi-source breeding pairing data, which includes Wagyu individual phylogenetic data, phenotypic determination data, genomic typing data, and pairing record data; S12. Perform missing value detection on the multi-source breeding pairing data, and remove data fields with missing rates higher than a preset threshold to obtain the detected dataset. S13. Using the three-standard-deviation principle, outlier detection is performed on the detected dataset, and a second verification is performed in combination with the threshold in the Wagyu breeding field to remove invalid outlier data and obtain the initial quality control screening dataset. S14. Convert the unstructured paired record data in the initial screening dataset of quality control into structured data to obtain a standardized dataset; S15. Using the maximum-minimum normalization method, the standardized dataset is unified in terms of dimensions, and the unified dataset is validated across all fields to generate a standardized breeding dataset.
3. The Wagyu breeding information management method based on data management according to claim 1, characterized in that, The process of hierarchically partitioning and labeling the standardized breeding dataset based on feeding environment factors, and removing environmentally unstable labels, yields the following environmental candidate labels and initial weights: S21. Obtain the feeding environment factors and divide the standardized breeding dataset into multiple environmental stratification subsets based on the feeding environment factors; wherein, the feeding environment factors include feeding density, dietary nutrient level, indoor temperature and humidity, and feeding area. S22. Extract all genetic markers in multiple environmental stratified subsets respectively. Calculate the locus effect of each genetic marker in the corresponding environmental stratified subset using the breeding locus effect method, and record the causal edge direction corresponding to each marker to obtain a record table of marker effect and causal edge direction. S23. Based on the preset locus effect size deviation threshold and the causal edge direction consistency judgment rule, compare the effect size deviation of the same locus genetic marker in different environmental stratified subsets, and verify the consistency of the causal edge direction of the same locus marker in different environmental stratified subsets. S24. Based on the consistency test results, identify the markers whose site effect size deviation between different environmental stratified subsets exceeds the preset threshold, and the markers whose causal edge direction is inconsistent with that between different environmental stratified subsets, and remove all unstable markers. S25. For the effective genetic markers after removing unstable markers, combine the mean locus effect size in each environmental stratification subset and assign initial locus weights. S26. Integrate all retained valid genetic markers and initial site weights to generate a list of environmental candidate markers and initial weight results.
4. The Wagyu breeding information management method based on data management according to claim 3, characterized in that, The process of assigning initial site weights to effective genetic markers after removing unstable markers, combined with the mean locus effect size within each environmental stratification subset, includes: S251. Based on the set of effective genetic markers after removing unstable markers, extract the locus effect size of each marker in the environmental stratification subset. S252. Calculate the mean of the locus effect size of the same effective genetic marker in all environmental stratified subsets to obtain the uniform effect size mean. S253. Use the mean effect size as the assigned value, and assign initial site weights to each effective genetic marker according to the weighting rules.
5. The Wagyu breeding information management method based on data management according to claim 1, characterized in that, The hybrid constraint scoring network algorithm is used to mine causal relationships and quantify sites on environmental candidate markers and initial weight results, resulting in environmental breeding causal markers and optimized site weights, including: S31. Using the hybrid constraint scoring network algorithm, causal mining and site quantification are performed on the environmental candidate labels and initial weight results. Each label is set as a Bayesian network node, an initial node set of a directed acyclic graph is constructed, and the constraint conditions of the hybrid constraint scoring network algorithm are set. S32. Calculate the candidate network structure score using the Bayesian information criterion structure scoring function. S33. Based on the hybrid constraint scoring network algorithm, the constraint conditions and candidate network structure scores are obtained by traversing the network structure space using the classic hill-climbing search method. The initial weights are used as the starting point for iteration. The presence and direction of directed edges between nodes are adjusted, and the Bayesian information criterion structure scoring function is substituted to calculate the score of each candidate structure. S34. Based on the optimal network structure, the intensity of the causal effect of each label is quantified using the Bayesian parameter estimation function, and weighted correction is completed by combining the initial weights. S35. Extract the markers with causal relationships in the optimal structure as environmental breeding causal markers, integrate the quantified optimized site weights, and obtain the environmental breeding causal marker and optimized site weight results.
6. The Wagyu breeding information management method based on data management according to claim 5, characterized in that, The constraint conditions and candidate network structure scores of the hybrid constraint scoring network algorithm are used to traverse the network structure space using a classic hill-climbing search method. Starting with the initial weights, the presence and direction of directed edges between nodes are adjusted, and the scores for each candidate structure are calculated using the Bayesian information criterion structure scoring function. S331. The initial weights based on the environmental candidate labels are used as the starting point for the hill-climbing search method, and the current state and iteration number threshold of the network structure search are initialized. S332. Under the constraints of the hybrid constraint scoring network algorithm, perform single-step local operations on the directed edges between Bayesian network nodes. The single-step local operations include adding directed edges, deleting directed edges, and flipping the direction of directed edges. S333. After each directed edge local operation is completed, a candidate Bayesian network structure with a corresponding topology is generated. S334. Input each generated candidate network structure into the Bayesian information criterion structure scoring function to complete the score calculation of the candidate structure, and record each candidate network structure and its score value in real time.
7. The Wagyu breeding information management method based on data management according to claim 1, characterized in that, The weighted single-step genome linear prediction model is used to rank and constrain the genetic characteristics of the optimized site weights and Wagyu whole genome data to obtain a breeding selection pairing dataset, which includes: S41. A weighted single-step genome prediction model composed of pedigree matrix construction technology, genome relation matrix construction technology, site weighted fusion technology, and hybrid linear model solution technology; S42. Based on the whole genome data of Wagyu cattle, perform quality control on polymorphic genotype data and remove missing, abnormal, and non-compliant sample data from individual phylogenetic information; S43. Based on the processed individual phylogenetic information, construct a pedigree additive genetic relationship matrix; S44. Based on the quality-controlled Wagyu whole-genome polymorphism site data, a standard genome relation matrix is constructed; and the optimized site weights are assigned to the standard genome relation matrix to construct a weighted genome relation matrix, thereby completing the fusion of weight information and genome data; S45. Merge the pedigree additive genetic relationship matrix with the weighted genomic relationship matrix to construct a weighted single-step genomic relationship matrix; S46. Based on the weighted single-step genome relation matrix, the mixed linear model is solved using the constrained maximum likelihood method to obtain the estimated breeding value of each candidate individual; S47. Based on the numerical value of the estimated breeding value from the genome, rank the candidate individuals according to their genetic characteristics. S48. Based on the breeding selection constraints, perform constraint matching screening on the ranked individuals to eliminate individuals with inbreeding risk and those that do not meet the breeding objectives; S49. Integrate the genetic trait ranking results with the constraint matching screening results to obtain the breeding selection pair dataset.
8. The Wagyu breeding information management method based on data management according to claim 7, characterized in that, The weighted single-step genomic relation matrix is used to solve the mixed linear model using the constrained maximum likelihood method to obtain the estimated genomic breeding values for each candidate individual, including: S461. Using the weighted single-step genome relation matrix as the basis of genetic covariance, construct a mixed linear model of fixed effects and random effects; S462. Using the constrained maximum likelihood method, iteratively solve the variance components of the mixed linear model to obtain the unbiased variance components. S463. Based on the variance component estimation results and the weighted single-step genome relationship matrix, calculate the random effect prediction value for each candidate individual; S464. Integrate the fixed-effects estimates and random-effects predictions to obtain the genomic breeding values for each candidate individual.
9. The Wagyu breeding information management method based on data management according to claim 1, characterized in that, The process involves full-chain traceability and access control of standardized breeding datasets, environmental breeding causal markers, and breeding selection and selection matching datasets to obtain a full-chain breeding data management archive. This enables the management of Wagyu breeding information and ensures traceable breeding data, including: S51. Data source nodes, processing nodes, and uniqueness traceability annotations are performed on the standardized breeding dataset, environmental breeding causal marker dataset, and breeding selection and pairing dataset to obtain basic information for full-link traceability. S52. Verify the basic information of the entire chain traceability to obtain the verified entire chain traceability information; S53. Based on the Wagyu breeding data management specifications, construct hierarchical permission control rules for data access permissions, data editing permissions, data export permissions, and data review permissions; S54. Match and integrate the hierarchical permission control rules with the verified full-link traceability information, and perform complete verification, permission verification and traceability rationality verification on the integrated information results to obtain qualified Wagyu breeding control information, and standardize and structure it for archiving. Based on the sorting and archiving results, generate a full-link breeding data control archive.
10. A data-management-based Wagyu breeding information management system, used to implement the data-management-based Wagyu breeding information management method according to any one of claims 1-9, characterized in that, include: The data quality control and standardization module is used to perform quality control verification and standardization processing on multi-source breeding pairing data to obtain a standardized breeding dataset. The environmental factor stratification test module is used to perform stratification and label testing on the standardized breeding dataset based on feeding environmental factors, and remove environmentally unstable labels to obtain environmental candidate labels and initial weight results. The causal relationship mining and quantification module is used to mine causal relationships and quantify sites on environmental candidate labels and initial weight results using a hybrid constraint scoring network algorithm, so as to obtain environmental breeding causal labels and optimized site weight results. The genetic feature ranking and matching module is used to rank and match the genetic features of the optimized site weights with the whole genome data of Wagyu cattle using a weighted single-step genome linear prediction model, so as to obtain a breeding selection and pairing dataset. The end-to-end traceability and control module is used to perform end-to-end traceability and access control on standardized breeding datasets, environmental breeding causal markers, and breeding selection and pairing datasets to obtain a full-chain breeding data control archive, so as to realize the management of Wagyu breeding information and traceable breeding data.