A method for crossbreeding improvement of a simmental

By using multimodal dynamic models and deep reinforcement learning techniques, combined with whole genome, phenotypic and environmental data of Simmental cattle, the recessive genetic load index was quantified, and the mating strategy was dynamically adjusted. This solved the problem of inaccurate mating in Simmental cattle crossbreeding improvement and achieved a balance between short-term improvement efficiency and long-term breeding safety.

CN122250424APending Publication Date: 2026-06-23INNER MONGOLIA XINGMU JIUYUAN ANIMAL HUSBANDRY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNER MONGOLIA XINGMU JIUYUAN ANIMAL HUSBANDRY CO LTD
Filing Date
2026-05-12
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

In the process of crossbreeding and improving Simmental cattle, existing technologies are unable to balance the risks of recessive inheritance across the entire genome, production performance, and environmental adaptability, resulting in inaccurate mating results. This can easily lead to an increase in the rate of deformed or weak offspring, affecting the long-term breeding safety of the population.

Method used

By acquiring whole genome sequence data, representative typographic data, and environmental and climatic time-series data, a multimodal dynamic model is used to quantify the recessive genetic load index. Based on a deep reinforcement learning model, mating schemes are generated, and mating strategies are dynamically adjusted to maximize production performance or genetic distance, avoiding the risks caused by single high-intensity breeding.

Benefits of technology

This approach achieves a dynamic balance between short-term improvement efficiency and long-term breeding safety in the crossbreeding improvement of Simmental cattle, reducing the rates of deformed and weak offspring, and improving the accuracy and safety of mating selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122250424A_ABST
    Figure CN122250424A_ABST
Patent Text Reader

Abstract

The present application relates to the field of animal genetic breeding and intelligent animal husbandry, in particular to a crossbreeding improvement method for Simmental cattle, comprising: obtaining whole genome sequence data, phenotypic data and target area environmental climate time series data of a target population; inputting the three types of data into a pre-trained multi-modal dynamic model, extracting features and quantifying the output of recessive genetic load index; judging the relationship between the index and the preset safety threshold; when the index is lower than the safety threshold, generating a first selection scheme based on a pre-constructed deep reinforcement learning model, with the goal of maximizing the intergenerational genetic progress rate of production performance; when the index is higher than or equal to the safety threshold, generating a second selection scheme with the goal of maximizing genetic distance; the present application realizes the dynamic balance between short-term improvement efficiency and long-term breeding safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of animal genetics and breeding and intelligent animal husbandry, specifically to a method for crossbreeding and improving Simmental cattle. Background Technology

[0002] Simmental cattle crossbreeding and breeding is an important part of beef cattle breeding and regional expansion. With the continuous improvement of gene testing, phenotyping and breeding environment monitoring methods, in order to ensure the coordinated improvement of crossbreeding efficiency and breeding safety, it is usually necessary to conduct a comprehensive evaluation of the genetic quality, production performance and environmental adaptability of the target population, and formulate corresponding selection and mating plans accordingly.

[0003] However, in existing Simmental cattle crossbreeding improvement processes, breeding decisions often focus on production performance indicators such as daily weight gain, carcass weight, and reproductive performance, or on static kinship and screening for single defective loci. This makes it difficult to simultaneously consider the coupled effects of whole-genome recessive genetic risks, current phenotypic performance, and regional environmental and climatic changes. To balance the speed of genetic improvement with breeding safety, empirical thresholds or single indicators are typically used to determine mating strategies. For example, bulls are directly selected based on breeding value, or pairings are made according to fixed inbreeding control standards. This decision-making approach lacks the ability to dynamically quantify the risk of amplified expression of recessive defective alleles under specific climatic stress conditions. Especially in regions with high temperature and humidity, and significant diurnal temperature fluctuations, recessive genetic risks are easily affected by climatic factors, making traditional mating methods susceptible to static judgments and human experience, resulting in inaccurate mating results. This limitation not only affects the rationality of crossbreeding improvement programs but may also lead to increased rates of deformities, weak offspring, or early failures in offspring, adversely affecting the long-term stable breeding of the population. Summary of the Invention

[0004] The purpose of this invention is to provide a method for crossbreeding and improving Simmental cattle, and to solve the following technical problems:

[0005] By unifying genomic risks, production performance improvement, and climate disturbance within the same technical process, we can avoid the sudden increase in deformity rate or weak fetus rate in subsequent two generations of calves caused by high-intensity selection based solely on normal phenotypes. This will achieve a dynamic balance between short-term improvement efficiency and long-term breeding safety.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] A method for crossbreeding and improving Simmental cattle, comprising the following steps:

[0008] Acquire whole genome sequence data, representative typological data, and environmental and climatic time-series data of the target population;

[0009] The whole genome sequence data, the current phenotypic data, and the environmental climate time series data are input into a pre-trained multimodal dynamic model to extract features and quantify and output the recessive genetic load index of the target population.

[0010] Determine the relationship between the recessive genetic burden index and a pre-set safety threshold;

[0011] When the recessive genetic load index is lower than the safety threshold, a first mating scheme for the combination of individuals within the target population is generated based on a pre-built deep reinforcement learning model with the goal of maximizing the rate of cross-generational genetic progress of production performance, and the first mating scheme is output.

[0012] When the recessive genetic load index is higher than or equal to the safety threshold, a second mating scheme is generated for the combination of individuals in the target population and individuals in the candidate population with the goal of maximizing genetic distance, and the second mating scheme is output.

[0013] Optionally, the step of inputting the whole genome sequence data, the current phenotypic data, and the environmental climate time-series data into a pre-trained multimodal dynamic model to extract features and quantify and output the recessive genetic load index of the target population specifically includes:

[0014] The distribution frequency of recessive defective alleles in the target population was identified based on the whole genome sequence data.

[0015] Environmental stress intensity characteristics are extracted based on the environmental climate time series data;

[0016] The distribution frequency and the environmental stress intensity features are fused using tensors to generate a feature vector, and the feature vector is input into a preset activation function to calculate the activation probability of the recessive defect allele in the current environment.

[0017] The recessive genetic load index is calculated by summing the product of the activation probability and the distribution frequency.

[0018] Optionally, the step of extracting environmental stress intensity features based on the environmental climate time-series data specifically includes:

[0019] Based on a pre-set climate safety threshold range, the data in the environmental climate time series data that are within or on the boundary of the climate safety threshold range are divided into normal climate data segments, and the data that are not within the climate safety threshold range are divided into extreme stress data segments.

[0020] Extract the duration and fluctuation amplitude of the extreme stress data segment;

[0021] After normalizing the duration and the fluctuation amplitude, the weighted sum of the two is calculated using preset weights to obtain the environmental stress intensity characteristics.

[0022] Optionally, the method further includes the step of predicting the failure risk rate of offspring, said failure risk rate being obtained through the following steps:

[0023] Obtain historical deformity rate data for the target population;

[0024] A regression model is constructed with the recessive genetic load index as the independent variable and the historical malformation rate data as the dependent variable as the mapping relationship;

[0025] Substitute the calculated recessive genetic load index into the mapping relationship to solve for the basic failure probability;

[0026] The calculated environmental stress intensity characteristics are mapped to environmental stress coefficients, and the environmental stress coefficients are used to perform product correction on the basic failure probability to obtain the progeny failure risk rate.

[0027] Optionally, the step of generating a first mating scheme based on a pre-built deep reinforcement learning model with the objective of maximizing the rate of genetic progress across generations of production performance specifically includes:

[0028] The whole genome sequence data and the current phenotypic data are input into the deep reinforcement learning model;

[0029] The deep reinforcement learning model is used to predict the intergenerational genetic progression rate of production performance for different combinations of individuals within the target population; the combination of individuals with the largest intergenerational genetic progression rate of production performance is selected to generate the first mating scheme.

[0030] Optionally, the step of generating a second mating scheme aimed at maximizing genetic distance specifically includes:

[0031] Obtain supplementary genomic data for candidate populations;

[0032] Calculate the genetic distance between the whole genome sequence data and the supplementary genome data;

[0033] The second mating scheme is generated by selecting the combination of individuals that maximizes the genetic distance and does not carry the same recessive defective allele.

[0034] Optionally, the step of determining the relationship between the recessive genetic burden index and a pre-set safety threshold may further include:

[0035] Collect failure sample data with defective phenotypes from historical breeding batches;

[0036] By retrospectively analyzing the multimodal data of the failed sample data in the corresponding historical breeding batches, the historical genetic load index is calculated, and a historical genetic load index distribution is constructed. Based on the historical genetic load index distribution, the inflection point of the historical genetic load index distribution is calculated, and the corresponding index value is taken as the critical mutation point.

[0037] The critical mutation point is set as the safety threshold.

[0038] Optionally, after the step of outputting the first configuration option or the second configuration option, the method further includes:

[0039] Collect Cenozoic genome data and Cenozoic phenotypic data after executing the first mating scheme or the second mating scheme;

[0040] The Cenozoic genome data and the Cenozoic phenotypic data are used as a feedback training set to update the parameters of the multimodal dynamic model.

[0041] The beneficial effects of this invention are:

[0042] 1. This invention uses a multimodal dynamic model to deeply integrate whole genome sequences, current genotypes, and environmental and climatic time-series data to quantitatively output the recessive genetic load index. It breaks through the limitations of traditional methods that rely solely on static indicators or single production performance, achieving a dynamic balance between short-term improvement efficiency and long-term breeding safety, and avoiding a significant increase in offspring deformity rates caused by maximizing a single production indicator.

[0043] 2. This invention extracts the duration and fluctuation range of extreme climate as environmental stress characteristics, and integrates them with the distribution frequency of recessive defective alleles to calculate the activation probability. This method transforms static genetic information into environmentally sensitive dynamic indicators, precisely quantifies the amplification effect of continuous climate pressure on dormant risk sites, and reduces the risk underestimation caused by rough assessment.

[0044] 3. This invention constructs a mapping model based on historical malformation rate data and combines it with the environmental stress coefficient to calculate the recessive genetic load index into the offspring failure risk rate. This mechanism transforms abstract genetic values ​​into intuitive probability indicators for grassroots production, providing a clear decision-making basis for whether to suspend aggressive breeding and avoiding the risk of calf malformation or weak offspring exceeding the safety boundary in advance.

[0045] 4. This invention automatically switches strategies based on the risk index; when the risk is under control, it uses deep reinforcement learning to maximize the rate of intergenerational progress in production performance; when the risk exceeds the limit, it pursues the maximization of genetic distance and avoids the overlap of the same latent defects; it can release the potential of breeding cattle during the safe period and actively expand genetic redundancy during the high-risk period to cut off the transmission of overlapping defects.

[0046] 5. This invention uses critical mutation points from historical failure sample data to set safety thresholds, thus eliminating reliance on subjective experience. At the same time, it introduces actual breeding data of the new generation as a feedback set to continuously update the model, giving the system intergenerational adaptive capabilities, enabling it to closely follow population evolution and climate change trends in the long term, and effectively avoiding decision drift during long-term operation. Attached Figure Description

[0047] The invention will now be further described with reference to the accompanying drawings.

[0048] Figure 1 This is a flowchart illustrating a method for crossbreeding and improving Simmental cattle, as provided in an embodiment of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] Please see Figure 1 A method for crossbreeding and improving Simmental cattle includes the following steps: acquiring whole genome sequence data, representative phenotypic data, and environmental and climatic time-series data of the target population; inputting the whole genome sequence data, representative phenotypic data, and environmental and climatic time-series data into a pre-trained multimodal dynamic model, extracting features, and quantifying and outputting the recessive genetic load index of the target population;

[0051] Determine the relationship between the recessive genetic load index and a pre-set safety threshold; when the recessive genetic load index is lower than the safety threshold, generate a first mating scheme for the combination of individuals within the target population based on a pre-built deep reinforcement learning model with the goal of maximizing the rate of cross-generational genetic progress of production performance, and output the first mating scheme.

[0052] When the recessive genetic load index is higher than or equal to the safety threshold, a second mating scheme is generated for the combination of individuals in the target population and individuals in the candidate population with the goal of maximizing genetic distance, and the second mating scheme is output.

[0053] This embodiment provides a dynamic breeding decision-making mechanism for Simmental cattle crossbreeding improvement scenarios. Specifically, the core cow herd, candidate bull herd, and climate time series information of the target pasture area of ​​the core beef cattle breeding base are uniformly modeled. Before mating each breeding batch, the recessive genetic risk of the population is quantitatively assessed, and then adaptively switched between high genetic progress scheme and high safety redundancy scheme according to the risk level.

[0054] In a spring breeding batch, the core breeding base collects whole genome sequence data of the target population; this data can be formed by splicing microarray genotyping data and key site resequencing data, for example, obtaining genotypes of several recessive defect-related sites for cow herds M1, M2, and M3 and candidate bull herds B1, B2, and B3 respectively.

[0055] At the same time, collect local representative data, including daily weight gain, carcass weight, feed conversion ratio, breeding interval, calf birth weight and heat resistance score; then obtain environmental climate time series data of the target area for the most recent 24 months, including at least daily maximum temperature, daily minimum temperature, relative humidity, number of consecutive high temperature days and diurnal temperature range;

[0056] During the model processing stage, the above three types of data are input into a pre-trained multimodal dynamic model. This model can be composed of a genome coding branch, a phenotypic coding branch, and a climate time-series coding branch. The genome coding branch is used to extract the distribution characteristics of latent defect sites, the phenotypic coding branch is used to extract production performance and adaptability characteristics, and the climate time-series coding branch is used to extract the trend of environmental stress changes.

[0057] Among them, the genome coding branch adopts a graph neural network structure to capture the epistatic effect between different gene loci; the phenotype coding branch adopts a fully connected feedforward neural network; the climate time series coding branch adopts a long short-term memory network or transformer architecture to extract the long-term dependence of meteorological sequences; in the graph neural network, nodes are defined as independent gene loci, and edges are defined as known epistatic interactions between genes.

[0058] Specifically, the known epistatic interactions were derived from the publicly available dataset of significant interaction sites from genome-wide association analysis of Simmental cattle in the AnimalQTLdb database; the weights of the edges between nodes were assigned as the linkage disequilibrium coefficients of the two loci in the corresponding GWAS analysis. Its value range is The message passing and update mechanism of the graph neural network adopts the GraphSAGE framework. It iteratively updates the feature vector of the current node by aggregating the features of neighboring nodes and combining them with its own features through a single-layer perceptron network. The output dimension of each encoding branch is uniformly set to 128 dimensions. The number of layers of the long short-term memory network is set to 2, and the number of hidden units is 64.

[0059] The multimodal data fusion method employs attention-based feature-level concatenation, where the three feature vectors are concatenated and input into a two-layer fully connected layer with ReLU activation to enhance model interpretability. During model training, the training set is constructed through supervised learning using whole-genome sequences, phenotypes, and corresponding climate data from core breeding farms over the past five years. The loss function is the mean squared error loss, calculated using the following formula: ,in This is the mean squared error loss value. This represents the total number of samples in the training set. The label is the normalized true deformity rate. This represents the recessive genetic load index output by the model.

[0060] The historical maximum true deformity rate record value is based on the highest measured deformity rate in batches subjected to extreme inbreeding or disease stress at this breeding base within the past ten years. The specific value used in this embodiment is... By dividing the actual batch defect rate by Complete the normalization process and transform it into a range of The true risk label between them is used as the true value. The Adam optimizer was selected, and the initial learning rate hyperparameter was set to 0.001.

[0061] The three features are fused to output a recessive genetic burden index; for ease of explanation, it is assumed that the index is standardized to between 0 and 1, with a larger value indicating a higher risk.

[0062] The following is a simulation calculation: Suppose that three key latent defect sites D1, D2, and D3 are identified in the target population, with distribution frequencies of 0.20, 0.10, and 0.05, respectively; After climate pathway extraction, the current environmental stress level in the target pastoral area is higher than the preset environmental stress threshold, and the activation degree of the three sites in the current environment is 0.80, 0.60, and 0.30, respectively;

[0063] After fusion calculation, the recessive genetic burden index of this batch is 0.20×0.80+0.10×0.60+0.05×0.30=0.235; if the preset safety threshold is 0.25, then this batch is judged to still be in the acceptable risk range.

[0064] When the index is below the threshold, the system calls a pre-built deep reinforcement learning model to generate the first mating scheme; the state input of the model includes at least the genome vector of the candidate individuals, the phenotypic vector, and the existing kinship structure information, and the action is to select which bull to pair with which group of cows, and the reward function is used to characterize the rate of genetic progress of production performance over multiple generations;

[0065] Specifically, the deep reinforcement learning model uses a proximal policy optimization algorithm for parameter updates; the reward function is constructed based on the weighted sum of the expected increment of the estimated breeding value of offspring in multi-generation simulations and the population average inbreeding coefficient penalty term, to ensure that inbreeding degradation is avoided while improving production performance; in the deep reinforcement learning model, the state space is defined as the concatenation of the genomic feature vectors and phenotypic feature vectors of the current cow herd and the candidate bull herd; the action space is defined as the specific pairing combination matrix of bulls and cows.

[0066] The policy network structure employs a three-layer feedforward neural network with 256 and 128 hidden layer nodes, respectively, and uses the Tanh activation function; the mathematical form of the reward function is... ,in For the return function value, Estimate the expected increment of breeding value for offspring. This represents the increase in the average inbreeding coefficient of the population. and The preset weight constant;

[0067] in The specific calculation method is as follows: The average breeding value of the parents to be selected is estimated based on the BLUP model, and the difference is obtained by subtracting the historical average breeding value of the current foundation herd. A preset weighting constant is determined through grid search combined with the experience of breeding experts in the field. To ensure a dynamic balance between performance improvement and inbreeding-controlled decline, this embodiment specifically sets it as follows: , ;

[0068] The value network and policy network have the same structure and the output dimension is 1; the pruning parameter of the near-end policy optimization algorithm is set to 0.2;

[0069] For example, the system predicts the cross-generational genetic progression rates for the three combinations as follows: , , If so, the corresponding optional configuration for G2 will be output first;

[0070] When the index is higher than or equal to the threshold, the system no longer pursues the highest genetic progress in the current generation, but switches to the second mating scheme, with the goal of maximizing genetic distance and avoiding overlap of the same recessive defect alleles;

[0071] At this point, candidate individuals from distantly related local yellow cattle or external supplementary populations can be introduced for crossbreeding and backcrossing to increase genetic diversity redundancy and reduce the risk of recessive homozygosity.

[0072] As an exception handling step, when there are missing input data, the system first checks the data integrity. If a cow lacks a key locus genotype but has high-density chip data from the same batch, then missing locus imputation is performed first. Specifically, missing locus imputation uses a genotype imputation algorithm based on a hidden Markov model. By comparing the haplotype reference panel of the breed, the recessive state is defined as the true haplotype and the observed state is defined as the genotype detected by the chip using the hidden Markov model. The forward-backward algorithm is used to calculate the state transition probability, infer and imput the allele status of the untyped key locus.

[0073] If the missing percentage exceeds the preset limit, for example, if it exceeds 20% of all key loci, then the individual will not be included in the automatic matching pool for this batch.

[0074] If there is a continuous interruption in climate time series data, such as a 7-day missing period, data from adjacent weather stations will be used to supplement the data; if the data cannot be supplemented, the environmental stress level will be raised to a warning level and the system will be forced to switch to a conservative mating mode; if the first mating option conflicts with the constraints of on-site breeding resources, such as a shortage of frozen semen in a certain bull, the alternative combination with the second highest return value will be used to supplement the data.

[0075] In a high-temperature and high-humidity pastoral area in Southwest China, the core farm plans to use Simmental cattle to improve the local yellow cattle herd in order to increase the daily weight gain of offspring and maintain their resistance. After the model conducted a joint evaluation of the current herd and candidate breeding bulls, it was found that the pastoral area experienced continuous alternation of high temperature and high humidity, which made some originally low-risk defect sites more easily activated in the actual feeding environment.

[0076] The system first outputs the recessive genetic load index, and then decides whether to implement the first mating scheme, which aims to maximize the rate of increase in meat yield, or the second mating scheme, which aims to increase genetic distance and reduce the homozygosity of defects.

[0077] The purpose of this step is to unify genomic risks, production performance improvement, and climate disturbance within the same technical process, and to avoid a sudden increase in the rate of deformed or weak calves in the subsequent two generations caused by high-intensity selection based solely on normal phenotypes, thereby achieving a dynamic balance between short-term improvement efficiency and long-term breeding safety.

[0078] In a preferred embodiment of the present invention, the steps of inputting whole genome sequence data, representative typographic data and environmental climate time series data into a pre-trained multimodal dynamic model, extracting features, and quantifying and outputting the recessive genetic load index of the target population specifically include: identifying the distribution frequency of recessive defective alleles in the target population based on whole genome sequence data;

[0079] Environmental stress intensity features are extracted from environmental and climate time series data; the distribution frequency and environmental stress intensity features are fused using tensors to generate feature vectors, and the feature vectors are input into a preset activation function to calculate the activation probability of recessive defective alleles in the current environment; the product of the activation probability and the distribution frequency is summed to calculate the recessive genetic load index.

[0080] This embodiment provides a quantitative mechanism for the recessive genetic burden index. Specifically, addressing the issue in the previous embodiment that only provided an overall judgment result without detailing the internal quantitative process, this embodiment further explicitly couples the frequency of occurrence of harmful recessive alleles in the population with the extent to which the current environment amplifies these risks, in order to output a comparable and traceable risk index.

[0081] The distribution frequency of recessive defective alleles was identified based on whole-genome sequence data. This distribution frequency can be statistically analyzed by locus or by defect type. For example, recessive defective loci D1, D2, and D3 were detected in the target population after merging the core cow herd and candidate bull herds. The specific extraction process involved using BWA alignment software to align the original whole-genome sequencing data to the Simmental reference genome, and using GATK software to detect SNP variations, screening for variant loci with a sequencing depth greater than 10-fold and a quality value greater than 30. Recessive defective alleles were defined as homozygous recessive mutation sites annotated on the reference genome that cause known genetic defects. The distribution frequency was calculated by locus, by dividing the number of individuals carrying the recessive defective allele in the target population by the total number of individuals in the target population.

[0082] If the total allele count is 200, the corresponding harmful allele counts are 40, 18, and 8, respectively, and the distribution frequencies are 0.20, 0.09, and 0.04. If a category merging method is used, multiple sites with similar mechanisms of action can be grouped into the same risk cluster before their frequencies are counted.

[0083] Environmental stress intensity characteristics are extracted from environmental and climate time series data; these characteristics can be derived from the comprehensive coding of factors such as high temperature, high humidity, and drastic diurnal fluctuations within a time window.

[0084] For ease of deduction, let the model output the current environmental stress vector of the pastoral area as... These correspond to the susceptibility of D1, D2, and D3 to unfavorable phenotypic expression induced in the current environment, respectively.

[0085] Tensor fusion is performed on the distribution frequency and environmental stress intensity features to generate feature vectors. Tensor fusion can be understood as pairing each locus frequency with the corresponding environmental stress feature and mapping it together with the phenotypic adaptive features to the same feature space, so that the calculation process of the recessive genetic load index has the traceability of intermediate variables.

[0086] In a specific implementation scenario, if the distribution frequency vector is The environmental stress vector is This will generate three sets of local fusion features: [0.20, 0.80], [0.09, 0.60], and [0.04, 0.30]. These features are then further processed through a linear mapping layer and a preset activation function.

[0087] In computational scenarios at the micro-level, tensor fusion employs linear weighted fusion of multi-dimensional features to reduce computational overhead. Specifically, the intermediate fusion response values... The calculation logic is to weight the distribution frequency. With the Distribution frequency of recessive defective alleles Multiply by, plus environmental stress weight With the Environmental stress characteristics corresponding to each site Multiply, then add the bias term. The calculation formula is as follows:

[0088]

[0089] in, This is the identification number for the recessive defective allele. , The intermediate fusion response value represents the total number of recessive defective alleles identified in the target population. Input a preset activation function, such as the sigmoid function, combined with the natural constant. Calculate the activation probability :

[0090]

[0091] Through the structured computational logic described above, an explicit coupling mapping between distribution frequency and environmental stress is achieved; for ease of deduction, it is assumed that... , , After training and assigning fixed values, and following the mapping and activation calculations described above, three activation probabilities are directly output. ;

[0092] Summing the products of activation probability and distribution frequency yields the recessive genetic burden index; corresponding to the example above, 0.20×0.78+0.09×0.55+0.04×0.22=0.2143 can be calculated.

[0093] If different defect categories have different degrees of harm, a hazard weight can be introduced before summing. For example, a first preset weight can be assigned to lethal defects, and a second preset weight can be assigned to subclinical defects. The first preset weight is greater than the second preset weight to form a weighted index.

[0094] As an anomaly handling step, if a harmful allele is not detected at a certain locus in the current population, the distribution frequency of that locus is recorded as 0 and does not contribute directly to the sum; if a component of the environmental stress vector is abnormally high, such as exceeding the upper limit of the training phase, truncation or normalization is performed first to avoid exponential distortion caused by a single extreme weather data.

[0095] If all features are zero after fusion, it indicates that no significant genetic risk has been identified or the data quality is insufficient. In this case, the system marks the batch as low-risk and awaiting review, rather than directly considering it as absolutely safe.

[0096] In the aforementioned spring breeding batch in the southwestern pastoral area, the base found that the local area experienced two consecutive weeks of high daytime temperatures and high nighttime humidity, and the proportion of calves with weak pregnancies exceeded the preset warning threshold for weak pregnancies.

[0097] The system not only counts the frequency of defective alleles, but also encodes this continuous climate pressure into environmental stress characteristics; thus, even if some defective sites are not at significant risk under normal conditions, they will have a higher activation probability in the current batch, ultimately leading to an increase in the recessive genetic load index.

[0098] The purpose of this step is to transform static genetic information into dynamic risk indicators with environmental sensitivity, thereby enabling the early quantification of risk sites that are not activated under normal conditions but are amplified under abnormal climate conditions.

[0099] In a preferred embodiment of the present invention, the step of extracting environmental stress intensity characteristics based on environmental climate time series data specifically includes:

[0100] Based on the pre-set climate safety threshold range, the data in the environmental climate time series data that are within or on the boundary of the climate safety threshold range are divided into normal climate data segments, and the data that are not within the climate safety threshold range are divided into extreme stress data segments.

[0101] Extract the duration and fluctuation range of extreme stress data segments; after normalizing the duration and fluctuation range, calculate the weighted sum of the two using preset weights to obtain the environmental stress intensity characteristics.

[0102] This embodiment provides a refined extraction step for the intensity features of environmental stress. Specifically, in the previous embodiment, if the environmental climate is simply encoded as a single stress value, it is easy to mask the different effects of short-term peaks and long-term persistence risks on the expression of latent defects. Therefore, this embodiment further introduces two dimensions: climate safety threshold range, duration, and fluctuation amplitude, to quantify environmental stress in segments.

[0103] The environmental climate time series data were segmented according to a pre-set climate safety threshold interval. This pre-set climate safety threshold interval was determined through statistical intersection analysis of historical meteorological data from the target area over the past five years and the temperature and humidity index ranges of the target cattle herd during the same period, excluding heat stress. Specifically, the statistical determination method involved taking the 95% confidence interval of the distribution of each meteorological indicator within the heat stress-free period. The calculation of the 95% confidence interval was based on the total data from national meteorological observation stations in the target area over the past five years. The daily meteorological sample data were rigorously calculated using the parameter estimation method under the statistical premise that the temperature and humidity indices approximately follow a normal distribution.

[0104] As a specific application example, for the southern hot and humid pastoral areas where Simmental improved cattle are introduced, the normal range can be defined as an average daily temperature of 18°C ​​to 28°C and a relative humidity of 40% to 75%.

[0105] If any key climate indicator exceeds the range on a given day, that time slice is marked as an extreme stress data segment; to reflect continuity, adjacent days exceeding the limit can be merged into the same segment.

[0106] As a specific implementation scenario, suppose that in a continuous 10-day temperature and humidity combination, the first to third days are within a safe range, the average daily temperature from the fourth to the seventh day is 31℃, 33℃, 32℃, and 30℃ respectively, and the humidity exceeds 80% in all of them, the eighth day returns to normal, and the ninth and tenth days are again characterized by high temperature and high humidity; then two extreme stress data segments can be obtained, with the first segment lasting for 4 days and the second segment lasting for 2 days.

[0107] Extract the duration and fluctuation range of extreme stress data segments; the duration can be measured in days; the fluctuation range can be defined as the deviation of each out-of-limit indicator from the safety boundary; for example, in the first segment, the humidity deviations from the 75% upper limit are 6, 8, 7, and 5 respectively; to eliminate the dimensional difference between temperature and humidity, before comparison, the absolute deviation of each indicator must be divided by its corresponding preset maximum tolerable deviation to convert it into a dimensionless relative deviation coefficient; the maximum value among multiple relative deviation coefficients at each moment can be taken as the representative value, where... This represents the function that takes the maximum value, and then the average value over the entire segment is used to obtain the fluctuation range of that segment;

[0108] Corresponding to the first paragraph, assuming that the absolute deviations of the above temperature and humidity, after conversion, correspond to dimensionless relative deviation coefficients that are proportional to the original values, the maximum relative deviations for the four days are 6, i.e. 8 7 and 5 i.e. The average value of the entire segment is given by the following: For ease of deduction, assume that the system calculates the fluctuation amplitude of the second segment to be 4.0.

[0109] The duration and fluctuation amplitude are normalized; specifically, the normalization process uses the minimum-maximum normalization method, and the calculation logic is to normalize the duration or fluctuation amplitude to be normalized. Subtract the default minimum value that is usually set to 0 Then divide by the preset maximum reference value Compared with the preset minimum value The difference is used to obtain the normalized duration or fluctuation range. :

[0110]

[0111] It should be noted that, since duration and fluctuation amplitude belong to different characteristic dimensions, the formula above... , , and These are general calculation symbols; when normalizing the duration, the above parameters represent the current value of the duration, the preset maximum reference value, the preset minimum value, and the normalization result, respectively; when normalizing the fluctuation amplitude, the above same parameters independently represent the current value of the fluctuation amplitude, the preset maximum reference value, the preset minimum value, and the normalization result; in actual execution, both are substituted with specific values ​​of their respective dimensions for independent calculation.

[0112] If the system sets the maximum duration within the watch window to 10 days, that is... , The maximum fluctuation range is 10, that is , The normalized value for the first segment, which lasts for 4 days, is... Normalized fluctuation range is ;

[0113] The second segment uses values ​​of 0.2 and 0.4 respectively; by clearly defining the normalization boundaries and calculation rules, both dimensions of the extreme stress data segment are mapped to between 0 and 1, eliminating dimensional differences; the mathematical definition of duration is the number of consecutive natural days within the extreme stress data segment; the mathematical definition of fluctuation amplitude is the maximum value of the ratio of the absolute value of each exceeding indicator deviating from the climate safety threshold boundary to the preset maximum tolerance deviation within the data segment, and its calculation formula is as follows: ,in This indicates the current value of the exceeded indicator. This indicates the climate safety threshold boundary corresponding to this indicator. This indicates the preset maximum tolerance deviation for this indicator.

[0114] The pre-set weights are determined using the analytic hierarchy process (AHP), based on historical expert evaluations to construct a judgment matrix and calculate eigenvectors. The maximum tolerance deviation is determined by the maximum absolute deviation of the corresponding meteorological indicator recorded in historical extreme disaster years. For example, for the hot and humid pastoral areas in southern China, the data for this historical extreme disaster year comes from the continuous extreme heat wave event recorded by the national meteorological station in the region in 2013, where the maximum absolute deviation of temperature... Values The maximum absolute deviation of humidity Values The original data of the judgment matrix is ​​set to 1 and 1.5 in the first row and 0.67 and 1 in the second row. The calculated consistency ratio is less than 0.1.

[0115] The judgment matrix consists of two evaluation indicators: the duration and amplitude of extreme weather events. These indicators were constructed by three senior livestock breeding experts based on their scoring of the induction importance of genetic defect expression. The original data is set as the first row. and This indicates that duration is slightly more important than fluctuation range; the second line... and Based on this The matrix is ​​calculated to have its largest eigenvalue. The consistency index is calculated as follows:

[0116]

[0117] Due to this embodiment Substituting into the above formula, we can calculate the result. ;because Random consistency index of order matrix Therefore, the consistency ratio It fully meets the consistency test requirements of the judgment matrix;

[0118] Then, the weighted sum of the two is calculated using preset weights; if the duration weight is 0.6 and the fluctuation amplitude weight is 0.4, then the stress intensity of the first segment is 0.6×0.4+0.4×0.65=0.50, and the stress intensity of the second segment is 0.6×0.2+0.4×0.4=0.28.

[0119] If the total stress intensity of the entire window is required, the maximum value, average value, or time-weighted sum of each segment can be used. In this embodiment, the time-weighted average can be used, that is, the weighted value of the first segment of 4 days and the second segment of 2 days is calculated as (0.50×4+0.28×2) / 6, and the total stress intensity is approximately 0.427.

[0120] The reason for introducing this treatment is that if the previous scheme only looks at whether the limit is exceeded on a certain day, it will regard 32℃ for 7 consecutive days and 32℃ for only half a day as similar situations, which will easily underestimate the cumulative effect of long-term high-pressure environment. By adding the duration and fluctuation range, the environmental pressure can be more stably correlated with the degree of exposure of calves to the development window.

[0121] As an anomaly handling step, if there are no extreme stress data segments within a statistical window, the environmental stress intensity characteristic is recorded as 0, indicating that the current climate factors do not amplify hidden risks.

[0122] If the denominator of the duration is 0 or the maximum fluctuation range reference value is not configured, the system will fall back to the default standardized range and write a configuration missing alarm to the model log; if data from different weather stations conflict, for example, if the temperature deviation in the same area exceeds the set limit, the data source that is closer to the breeding farm and has a more complete timestamp will be used first, and the data from conflicting stations will be downgraded.

[0123] In the aforementioned pastoral areas, the base does not simply use the recent temperature reaching the preset high temperature threshold as an empirical judgment, but quantifies the number of consecutive days of high temperature and high humidity and the degree of deviation from the safe range.

[0124] In this way, when the first four consecutive days of high temperature and humidity approach the sensitive window for embryo survival, the model will automatically give an environmental stress intensity higher than that of the second period, thereby increasing the activation probability of the corresponding defect sites.

[0125] The purpose of this step is to transform climate disturbances from qualitative descriptions into computable stress characteristics, thereby enabling precise identification of persistent environmental stresses and reducing risk underestimation caused by overly coarse granularity in climate characteristic coding.

[0126] In a preferred embodiment of the present invention, the method further includes the step of predicting the offspring failure risk rate, which is obtained through the following steps: obtaining historical malformation rate data of the target population; constructing a regression model with the recessive genetic load index as the independent variable and the historical malformation rate data as the dependent variable as the mapping relationship;

[0127] The calculated recessive genetic load index is substituted into the mapping relationship to solve for the basic failure probability; the calculated environmental stress intensity characteristics are mapped to the environmental stress coefficient, and the basic failure probability is corrected by multiplying the environmental stress coefficient to obtain the failure risk rate of offspring.

[0128] This embodiment provides a prediction mechanism for the failure risk rate of offspring. Specifically, although the aforementioned process can output a recessive genetic load index, in actual production, grassroots propagation farms are more concerned with the approximate proportion of deformed, weak, or prematurely deceased offspring in the next generation. Therefore, this embodiment further maps the genetic load index to a failure risk rate that can be directly used for production early warning.

[0129] Obtain historical deformity rate data for the target population; this data can come from multiple consecutive breeding batches, including the total number of calves in each batch, the number of visible deformities, the number of weak calves, and the number of short-term deaths after birth;

[0130] To correspond with the genetic burden assessment, the recessive genetic burden index at that time can be calculated for each historical batch, forming a set of historical sample pairs; for example, the indices and malformation rates of five historical batches are as follows: 0.10 corresponds to 2%, 0.15 corresponds to 4%, 0.20 corresponds to 7%, 0.28 corresponds to 13%, and 0.35 corresponds to 18%;

[0131] A regression model is constructed with the genetic load index as the independent variable and the historical malformation rate as the dependent variable. The regression model can be linear regression, piecewise regression, or nonlinear regression. This embodiment uses piecewise regression to better reflect the actual characteristics of the change rate in the low-risk interval being lower than the first preset change rate threshold and the change rate in the high-risk interval being higher than or equal to the second preset change rate threshold.

[0132] As a specific implementation scenario: when the index is below 0.20, the basic failure probability is estimated as 0.3 × index + 0.01; when the index is above or equal to 0.20, the basic failure probability is estimated as 0.7 × index - 0.07. The breakpoint selection is based on: using the Chow test to traverse all historical sample data within the region, the historical sample data covers the past eight consecutive spring / autumn breeding batches of this breeding base, and the effective data sample size for each batch. Head; the step size for traversal search is strictly set to And the statistical significance level of the Chow test was set at 1. Provided that statistical significance is satisfied, select the option that minimizes the sum of squared residuals of the piecewise regression. As the optimal breakpoint; if the current batch index is 0.24, then the basic failure probability is 0.7 × 0.24 - 0.07 = 0.098, or 9.8%;

[0133] Secondly, an environmental stress coefficient is extracted based on environmental and climate time series data. This coefficient is used to reflect the differences in actual harm caused by the same genetic load in different environments. For ease of application, the environmental stress intensity in the previous embodiment can be mapped to a product correction coefficient, for example, 1.0 under normal conditions, 1.2 under moderate stress, and 1.5 under severe stress.

[0134] If the current batch is under continuous high temperature and high humidity, and the system gives an environmental stress coefficient of 1.4, then the failure risk rate of the offspring is 9.8% × 1.4 = 13.72%.

[0135] The reason for adding this layer of prediction is that simply providing the index value lacks intuitive quantitative indicators to support production applications; for example, the numerical difference between the index 0.24 and 0.28 is less than the preset difference threshold, but if the environmental stress coefficient is added, it may correspond to a rapid jump from around 10% to a failure boundary close to 15%; by explicitly predicting the failure risk rate, the system can more easily form the basis for whether to suspend the operation of a certain batch of high-progress selection.

[0136] As an anomaly handling step, if the sample size of historical abnormality rate is insufficient, such as less than 3 complete batches, batch-level regression is not directly fitted. Instead, a prior basic model of the same variety and region is called and a low confidence label is added to the output result.

[0137] If a certain historical batch has obvious anomalies, such as an outbreak of disease causing the deformity rate to deviate from the normal distribution range and exceed the preset judgment threshold, and is unrelated to genetic risk, it will be removed or downweighted before modeling.

[0138] If the calculated progeny failure risk rate after correction exceeds 100%, it will be truncated to 100%; if it is less than 0, it will be truncated to 0; if the environmental stress coefficient is missing, it will be set to 1.0 by default, and a prompt will be made that climate data needs to be collected.

[0139] In the summer batch at the aforementioned core base, the system calculated that the recessive genetic load index of a certain maternal population combined with a candidate breeding bull population was 0.26; based on the historical sample mapping relationship, its basic failure probability was approximately 11.2%.

[0140] However, the area where this batch was deployed was experiencing continuous high temperatures and humid nights, increasing the environmental stress coefficient to 1.35. As a result, the failure risk rate of the offspring was raised to 15.12%, which is close to the failure threshold set by the base. At this point, the system will suggest abandoning the original high-intensity configuration option.

[0141] The purpose of this step is to transform abstract genetic risks into failure probability indicators for production decisions, thereby enabling early warning of whether the failure boundary of calf deformity rate or weak fetus rate will exceed 15% for two consecutive generations.

[0142] In a preferred embodiment of the present invention, the step of generating a first mating scheme based on a pre-built deep reinforcement learning model with the goal of maximizing the rate of cross-generational genetic progress in production performance specifically includes: inputting whole genome sequence data and current phenotypic data into the deep reinforcement learning model;

[0143] A deep reinforcement learning model is used to predict the intergenerational genetic progression rate of production performance for different combinations of individuals within the target population; the combination of individuals with the highest intergenerational genetic progression rate of production performance is selected to generate the first mating scheme.

[0144] This embodiment provides a mechanism for generating a first mating scheme in a low-risk range. Specifically, the aforementioned process has given the general idea of ​​entering the radical improvement mode when the risk is acceptable. However, without a clear explanation of how candidate combinations are compared, it is still difficult to reproduce in engineering. Therefore, this embodiment further elaborates on how a deep reinforcement learning model works in mating scheme generation.

[0145] The whole genome sequence data and phenotypic data of the target population are input into a deep reinforcement learning model; in a breeding cycle, each candidate cow and bull is encoded as a state vector;

[0146] This state vector can simultaneously include genomic breeding values, key production performance phenotypes, kinship constraints, and defect locus carrier status; the model's action space can be defined as selecting a breeding bull to pair with a group of cows.

[0147] The reward function is set as the improvement of the comprehensive production performance index of offspring in the next generation or multiple generations, which can be obtained by linear or nonlinear combination of indicators such as daily weight gain improvement, carcass quality improvement and reproductive performance maintenance;

[0148] As a specific implementation scenario, assuming there are two individuals, A1 and A2, in the cow herd and two individuals, S1 and S2, in the bull herd, the candidate combinations can be simplified to four groups: A1-S1, A1-S2, A2-S1, and A2-S2. The model predicts the intergenerational genetic progress rate for each of the four combinations, obtaining scores of 2.3, 1.9, 2.1, and 2.5. These values ​​can be interpreted as comprehensive production performance improvement scores; the higher the score, the higher the efficiency of offspring improvement. Since the genetic load index of the current batch is below the safety threshold, the system selects the A2-S2 combination with the highest score as the first mating option.

[0149] If the actual scenario involves one bull corresponding to multiple cows, the system can perform batch allocation based on resource constraints. For example, although a superior breeding bull S2 is predicted to have the highest growth rate, its annual frozen semen supply limit can only cover 30 cows. In this case, the model will continue to select the second-best combination from the remaining candidates to fill the remaining cow sites, instead of forcibly allocating all cows to the same breeding bull.

[0150] The reason for introducing this approach is that if an overly conservative genetic distance-first strategy is still adopted in the low-risk range, it will reduce the efficiency of rapidly improving local yellow cattle in the core farm and prolong the time for breeding new strains. By using a reinforcement learning model to search for rewards across multiple generations and output auditable intermediate decision logic, short-term phenotypic improvement and long-term genomic trends can be incorporated into the decision-making process simultaneously.

[0151] As an exception handling step, if the difference between the predicted values ​​of different combinations is too small, for example, the difference between the highest and second highest values ​​is less than the preset threshold of 0.05, then the combination with a resource consumption coefficient lower than the preset coefficient threshold and an available inventory resource quantity greater than the preset inventory threshold is selected to avoid the instability of the scheme due to small fluctuations in the model.

[0152] If the highest-scoring combination contains a marked breeding forbidden pair, such as two carriers of the same defect site, it is directly removed and reselected from the remaining combinations; if the model output is abnormal, such as all combinations having negative scores, it indicates that the current candidate population is not suitable for high-progression mating, and the system automatically reverts to conservative mode and prompts for supplementary seed sources.

[0153] In the early spring batch at the aforementioned core base, the system calculated the recessive genetic load index of the target population to be 0.21, which is below the safety threshold of 0.25.

[0154] Therefore, the model predicted several combinations of Simmental bulls and local yellow cattle herds and found that the combination of S7 and herd G3 had the highest comprehensive improvement rate in daily weight gain, carcass rate and subsequent reproductive stability, and did not trigger the prohibition rule. Finally, it was listed as the priority implementation plan.

[0155] The purpose of this step is to fully unleash the genetic potential of superior breeding cattle while risks remain under control, thereby maximizing the rate of cross-generational genetic progress of core production performance.

[0156] In a preferred embodiment of the present invention, the step of generating a second mating scheme with the goal of maximizing genetic distance specifically includes: obtaining supplementary genome data of the candidate population; and calculating the genetic distance between the whole genome sequence data and the supplementary genome data.

[0157] The second mating scheme is generated by selecting the combination of individuals that maximizes genetic distance and does not carry the same recessive defective allele.

[0158] This embodiment provides a second mating scheme generation mechanism for high-risk intervals. Specifically, the previous embodiment is suitable for pursuing improvement efficiency when the risk is low. However, when the genetic load index reaches the danger interval, if mating is still carried out in a concentrated manner according to the highest genetic progress, it is easy to cause homozygous enrichment of the same recessive defect alleles in the offspring. Therefore, this embodiment introduces the principle of supplementary population and maximizing genetic distance to hedge the risks of the existing core population.

[0159] Obtain supplementary genomic data from candidate populations; these supplementary populations may be distant Simmental strains, local yellow cattle conservation populations, or externally introduced breeding cattle populations that have been verified not to carry specific recessive defect loci; to ensure compatibility, the supplementary data and the core population selected should meet the requirement of completing genotyping on the same set of key loci.

[0160] Calculate the genetic distance between the target population and the supplementary population; the genetic distance can be calculated based on the proportion of loci differences across the entire genome, principal component spatial distance, or phylogenetic coefficient transformation values;

[0161] In this preferred embodiment, state consistency distance is used to quantify genetic distance. The specific calculation steps are as follows: the proportion of genotypes of the target cow and the supplementary candidate bull that do not share the same alleles at all detection sites is counted. This proportion is the genetic distance between the two, and the larger the value, the more distant the genetic background.

[0162] As a specific implementation scenario, the genetic distances between the current high-risk cow M8 and the three supplementary candidate bulls C1, C2, and C3 are set to 0.35, 0.52, and 0.47, respectively; C2 has the largest genetic distance value and is selected first, with the highest priority.

[0163] However, genetic distance alone is not enough; it is also necessary to screen for combinations that do not carry the same recessive defect alleles. For example, if M8 carries D1 and D3, C2 also carries D3, C3 carries only D2, and C1 does not carry any known identical defects, then although C2 has the largest genetic distance, it is not suitable for direct pairing with M8 because it carries D3 in common.

[0164] The system will select the scheme that maximizes the genetic distance from the remaining combinations, provided that the constraints such as not carrying the same defect are met. In this case, C3 is preferred over C1 because the distance is 0.47 and it does not have the same defect as M8.

[0165] Furthermore, if it is necessary to generate schemes in batches for a population, the screening process can be divided into two levels; the first level is to remove all combinations that have the same pathogenic allele overlap.

[0166] The second level sorts the remaining combinations by genetic distance and outputs several candidate solutions based on the minimum production performance baseline and resource availability; this way, high-risk batches will not continue to accumulate defective homozygosities, nor will they completely lose the direction of improvement.

[0167] The reason for introducing this mechanism is that while simply stopping mating or introducing species randomly can reduce local risks, it will cause an imbalance in the breeding program. By introducing species from distant relatives and avoiding defective species, we can make room for the population to restore diversity and for subsequent improvement while enduring a certain short-term performance decline.

[0168] As an exception handling step, if no individual in the supplementary population meets the conditions such as not carrying the same defect, the system can perform stratified processing according to the defect hazard level: first, prohibit combinations that overlap with lethal or high-risk defects, then select the candidate individual with the largest genetic distance from the low-risk defect overlap, and mark it as for use only in risk-mitigation batches.

[0169] If the supplementary population genotype data is incomplete, key loci testing will be performed first; if the supplementary testing period cannot meet the current mating window, the individual will not participate for the time being; if multiple candidate combinations have the same genetic distance, the combination with a stress resistance score higher than the preset stress resistance standard will be selected.

[0170] After the aforementioned pastoral area entered a sustained high-temperature season, the system found that the recessive genetic load index of the core maternal population rose to 0.29, which exceeded the threshold of 0.25. The base originally planned to continue using a high-yield Simmental bull, but this bull and several cows had overlapping carriers at the D2 locus.

[0171] The system then compared candidate bulls from the local yellow cattle breeding population and ultimately selected individuals with a large genetic distance from the core maternal population and no overlap with high-risk defect sites to perform a backcrossing and cleaning program.

[0172] The purpose of this step is to actively expand genetic redundancy and cut off the overlapping transmission of the same defective sites during high-risk phases, thereby achieving dynamic hedging against the outbreak of recessive genetic load.

[0173] In a preferred embodiment of the present invention, the step of determining the relationship between the recessive genetic load index and the preset safety threshold further includes, before: collecting failure sample data of defective phenotypes in historical breeding batches;

[0174] By retrospectively analyzing the multimodal data of failed sample data in the corresponding historical breeding batches, the historical genetic load index is calculated, and a historical genetic load index distribution is constructed. Based on the constructed historical genetic load index distribution, the inflection point of the historical genetic load index distribution is calculated, and the corresponding index value is used as the critical mutation point. The critical mutation point is set as a safety threshold.

[0175] This embodiment provides a safety threshold adaptive setting mechanism; specifically, if the safety threshold in the aforementioned process relies entirely on experience and manual specification, it may not reflect the true failure boundary of a specific base, a specific pastoral area, and a specific hybrid base population.

[0176] Especially when the defect background is complex due to historical disordered hybridization, the fixed threshold is easy to be too high or too low; therefore, this embodiment uses failure sample data to infer the critical mutation point and establishes the safety threshold on the basis of the actual historical consequences.

[0177] Collect data on failed samples with defective phenotypes from historical breeding batches; these failed samples may include deformed calves, weak fetuses, short-term deaths after birth, and batches with declining reproductive quality for two consecutive generations.

[0178] For each failed batch, the historical genetic load index corresponding to its mating time was traced back; assuming that 8 historical batches were collected, their indices were 0.12, 0.15, 0.18, 0.21, 0.24, 0.26, 0.31, and 0.34, respectively, corresponding to a gradual increase in failure manifestations from slight increase to a continuous high proportion of deformities;

[0179] The inflection point is calculated based on these historical index distributions; this inflection point can be understood as the point where the risk growth trend begins to accelerate significantly.

[0180] In a specific implementation scenario, historical samples are sorted by exponent from smallest to largest, and the increment of failure severity between adjacent batches is calculated. If, starting around 0.24, the failure severity rapidly increases from slight fluctuations to batch anomalies, then 0.24 can be considered a trend inflection point. Alternatively, a smooth curve fitting can be used to find the median of the interval where curvature changes significantly as a critical value. The specific algorithm for the critical inflection point is as follows: the Kneedle algorithm is used to calculate the maximum curvature point after differentiating the quadratic smoothed fitting curve. Specifically, a spline smoothing function is used to fit the discrete historical data, and the smoothing sensitivity parameter is set to... This effectively filters out local data noise interference; then, it extracts the point of maximum curvature after differentiation, where the local curvature exceeds the global average curvature. The mutation was confirmed to occur at the time of the index change, and the corresponding index value was extracted as the critical mutation point.

[0181] Set this critical mutation point as a safety threshold; for example, historical data from a certain base shows that when the index is below 0.23, the deformity rate mostly remains below 5%;

[0182] When the index reaches 0.25 or higher, some batches quickly jump to 12% or even higher; then the midpoint between 0.24 and 0.25, 0.245, can be determined as the threshold. In actual deployment, the system can take 0.25 as a control line for easy management.

[0183] The reason for introducing this treatment is that if the aforementioned scheme directly applies a uniform threshold, the error may be large in some basic populations. Because different bases have different historical accumulation of latent defects, environmental exposure and management levels, the same genetic load index does not necessarily correspond to the same actual damage. By fitting the critical mutation point through historical failure samples and generating a verifiable threshold determination process, the threshold can be made to fit the actual boundary of the field more closely.

[0184] As an anomaly handling step, if there are too few historical failure samples to reliably identify inflection points, the system adopts a two-layer threshold strategy: first, the industry reference threshold is used as the default value, and then it is gradually corrected as new batches of data are added.

[0185] If historical data shows obvious clusters, such as high-altitude and hot and humid regions being mixed together, making the inflection point unclear, then calculate by region or climate zone first, and then set safety thresholds for each.

[0186] If multiple local inflection points are calculated, the lower one should be selected as the control line to avoid underestimating the risk.

[0187] When reviewing the breeding records of the past seven years at the aforementioned core base, it was found that once the recessive genetic load index exceeded approximately 0.25, and the batch was released into a hot and humid pastoral area, the proportion of weak calves and atypical degeneration exceeded the preset alarm threshold. Based on this, the base set 0.25 as the farm's control red line, so that the system no longer relies on fixed experience judgment in subsequent batches, but automatically switches strategies based on historical failure boundaries.

[0188] The purpose of this step is to transform the safety threshold from subjective experience into an objective control point based on historical loss data, thereby achieving more robust risk classification and strategy switching.

[0189] In a preferred embodiment of the present invention, after the step of outputting the first mating scheme or the second mating scheme, the method further includes: collecting Cenozoic genome data and Cenozoic phenotypic data after executing the first mating scheme or the second mating scheme; using the Cenozoic genome data and Cenozoic phenotypic data as a feedback training set to update the parameters of the multimodal dynamic model.

[0190] This embodiment provides a closed-loop update mechanism for the model; specifically, the aforementioned scheme can make decisions before mating, but if the model relies solely on historical samples without incorporating results from the new generation for a long period of time, the early model may gradually deviate from reality as climate patterns change and population structure evolves.

[0191] Especially in hot and humid pastoral areas, the environmental stress response patterns will change in stages, so it is necessary to feed the real results back into the model after each generation is born;

[0192] After implementing the first or second mating scheme, collect genomic data and phenotypic data of the newborn calves. Genomic data may include genotypes of key loci in newborn calves and information on parental origin confirmation. Phenotypic data may include birth weight, survival rate, early growth rate, malformation records, heat resistance, etc. To ensure temporal consistency, the preferred data should be stored in a batch according to the corresponding mating scheme.

[0193] The new generation data is organized into a feedback training set; specifically, the parent combination characteristics and the actual results of the new generation can be combined to form a feedback sample; for example, in a certain batch, the original predicted genetic burden risk of combination P1 was low, but only 1 weak fetus was born out of 20 calves. This sample strengthens the model's safety judgment on this combination in this regional environment.

[0194] If the original prediction of the P2 combination was controllable, but three animals actually died early, the model will treat this sample as a correction signal of high risk.

[0195] As a specific implementation scenario, suppose there are three feedback samples R1, R2, and R3 in a certain training set; R1 corresponds to a predicted risk of 0.18 and an actual failure rate of 5%; R2 corresponds to a predicted risk of 0.22 and an actual failure rate of 6%; R3 corresponds to a predicted risk of 0.21 and an actual failure rate of 14%.

[0196] If the batch containing R3 is also accompanied by continuous high temperature and high humidity, the model will increase the weight of this type of environmental feature on the latent risk when updating the parameters, making the output more conservative under similar conditions in the future.

[0197] When the updated model encounters similar input again, it may output 0.21 instead of 0.21, which may now increase to 0.26, thus prompting the system to switch to the second option.

[0198] The reason for introducing closed-loop updates is that the previous layer method assumes that the model parameters are fixed, while the actual breeding process has obvious intergenerational delay characteristics; the real consequences of a certain batch of mating often only become apparent after several months or even longer. If these posterior results are not included in the model, the system may repeatedly underestimate the risks under specific climatic backgrounds; the introduction of a feedback mechanism enables the model to have the technical capabilities of feedback correction and continuous state tracking.

[0199] As an exception handling step, if the data of a certain new generation batch is incomplete, such as having only birth weight but no subsequent survival information, then the sample of that batch will only participate in the update of some parameters and will not participate in the failure risk mapping correction; if the parent identity is not confirmed, then the sample will not be used for pair-level updates, but only for population-level statistics.

[0200] If a batch is impacted by exogenous diseases or other non-genetic factors, an anomaly marker is added before the data is fed into the model, and the relevant samples are downweighted to prevent non-genetic accidents from being mistaken for genetic laws. If multiple batches of data have poor quality, the system maintains the old parameters and prompts for re-sampling. The above anomaly handling process is automatically executed according to the set rules and is fully recorded in the system log. The real data in the specific embodiment of this invention comes from the Simmental cattle breeding database of a core breeding base over the past five years, and the data scale includes multimodal records of 8,500 core cows and 120 bulls.

[0201] In the label definition, the deformity rate label is the percentage of calves with significant physical deformities or weak fetuses in a historical breeding batch out of the total number of newborn calves in that batch; the model validation adopts the ten-fold cross-validation method, randomly dividing the total dataset into ten equal parts, and selecting nine parts in turn as the training set to verify the Pearson correlation coefficient between the predicted value of the recessive genetic load index and the actual label; other evaluation indicators include root mean square error and mean absolute error; the overfitting monitoring method is to introduce an early stopping mechanism, separating 10% of the training set as the validation set in each round of training, and terminating the training early if the loss of the validation set does not decrease for 10 consecutive cycles.

[0202] After completing one round of breeding during the high-temperature season and one round during the normal season at the aforementioned core base, the base conducted genotyping and phenotypic verification on the newborn calves. The results showed that some combinations that were judged to be low-risk by the original model during the high-temperature season actually had a higher proportion of weak pregnancies than similar combinations during the normal season.

[0203] After the system added these results from the new generation to the feedback training set, the multimodal dynamic model increased the weight of the effect of the duration of high temperature and humidity on the activation of specific defect sites. This adjustment process was completed autonomously by the model based on the input data. When similar climate is encountered again in the next breeding season, the system will identify the target population into the high-risk range earlier and actively provide a conservative solution with a greater genetic distance.

[0204] The purpose of this step is to enable the model to have intergenerational self-correction capabilities, thereby achieving continuous adaptation to changes in population structure and environmental stress, and reducing judgment drift during long-term operation.

[0205] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A method for crossbreeding and improving Simmental cattle, characterized in that, Includes the following steps: Acquire whole genome sequence data, representative typological data, and environmental and climatic time-series data of the target population; The whole genome sequence data, the current phenotypic data, and the environmental climate time series data are input into a pre-trained multimodal dynamic model to extract features and quantify and output the recessive genetic load index of the target population. Determine the relationship between the recessive genetic burden index and a pre-set safety threshold; When the recessive genetic load index is lower than the safety threshold, a first mating scheme for the combination of individuals within the target population is generated based on a pre-built deep reinforcement learning model with the goal of maximizing the rate of cross-generational genetic progress of production performance, and the first mating scheme is output. When the recessive genetic load index is higher than or equal to the safety threshold, a second mating scheme is generated for the combination of individuals in the target population and individuals in the candidate population with the goal of maximizing genetic distance, and the second mating scheme is output.

2. The method for crossbreeding and improving Simmental cattle according to claim 1, characterized in that, The steps of inputting the whole genome sequence data, the current phenotypic data, and the environmental climate time-series data into a pre-trained multimodal dynamic model, extracting features, and quantifying and outputting the recessive genetic load index of the target population specifically include: The distribution frequency of recessive defective alleles in the target population was identified based on the whole genome sequence data. Environmental stress intensity characteristics are extracted based on the environmental climate time series data; The distribution frequency and the environmental stress intensity features are fused using tensors to generate a feature vector, and the feature vector is input into a preset activation function to calculate the activation probability of the recessive defect allele in the current environment. The recessive genetic load index is calculated by summing the product of the activation probability and the distribution frequency.

3. The method for crossbreeding and improving Simmental cattle according to claim 2, characterized in that, The steps for extracting environmental stress intensity characteristics based on the environmental climate time series data specifically include: Based on a pre-set climate safety threshold range, the data in the environmental climate time series data that are within or on the boundary of the climate safety threshold range are divided into normal climate data segments, and the data that are not within the climate safety threshold range are divided into extreme stress data segments. Extract the duration and fluctuation amplitude of the extreme stress data segment; After normalizing the duration and the fluctuation amplitude, the weighted sum of the two is calculated using preset weights to obtain the environmental stress intensity characteristics.

4. The method for crossbreeding and improving Simmental cattle according to claim 3, characterized in that, The method further includes a step of predicting the progeny failure risk rate, which is obtained through the following steps: Obtain historical deformity rate data for the target population; A regression model is constructed with the recessive genetic load index as the independent variable and the historical malformation rate data as the dependent variable as the mapping relationship; Substitute the calculated recessive genetic load index into the mapping relationship to solve for the basic failure probability; The calculated environmental stress intensity characteristics are mapped to environmental stress coefficients, and the environmental stress coefficients are used to perform product correction on the basic failure probability to obtain the progeny failure risk rate.

5. The method for crossbreeding and improving Simmental cattle according to claim 1, characterized in that, The steps for generating a first mating scheme based on a pre-built deep reinforcement learning model, aimed at maximizing the rate of genetic progress in production performance across generations, specifically include: The whole genome sequence data and the current phenotypic data are input into the deep reinforcement learning model; The deep reinforcement learning model is used to predict the intergenerational genetic progression rate of production performance for different combinations of individuals within the target population; the combination of individuals with the largest intergenerational genetic progression rate of production performance is selected to generate the first mating scheme.

6. The method for crossbreeding and improving Simmental cattle according to claim 2, characterized in that, The steps for generating a second mating scheme aimed at maximizing genetic distance specifically include: Obtain supplementary genomic data for candidate populations; Calculate the genetic distance between the whole genome sequence data and the supplementary genome data; The second mating scheme is generated by selecting the combination of individuals that maximizes the genetic distance and does not carry the same recessive defective allele.

7. The method for crossbreeding and improving Simmental cattle according to claim 1, characterized in that, The step of determining the relationship between the recessive genetic burden index and a pre-set safety threshold further includes, prior to: Collect failure sample data with defective phenotypes from historical breeding batches; By retrospectively analyzing the multimodal data of the failed sample data in the corresponding historical breeding batches, the historical genetic load index is calculated, and a historical genetic load index distribution is constructed. Based on the historical genetic load index distribution, the inflection point of the historical genetic load index distribution is calculated, and the corresponding index value is taken as the critical mutation point. The critical mutation point is set as the safety threshold.

8. The method for crossbreeding and improving Simmental cattle according to claim 1, characterized in that, After the step of outputting the first configuration option or the second configuration option, the method further includes: Collect Cenozoic genome data and Cenozoic phenotypic data after executing the first mating scheme or the second mating scheme; The Cenozoic genome data and the Cenozoic phenotypic data are used as a feedback training set to update the parameters of the multimodal dynamic model.