Artificial intelligence-based beef cattle carbon emission data intelligent analysis and breeding decision-making system
By analyzing the waveform characteristics of methane release from beef cattle and combining them with an epigenetic marker database, a differentiated gene editing scheme was designed. This solved the problems of low gene target localization and editing efficiency in traditional beef cattle carbon emission analysis, and improved the accuracy and safety of low-carbon breeding of beef cattle.
Patent Information
- Application Number
- CN202511100964.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional carbon emission analysis and breeding techniques for beef cattle cannot accurately analyze the dynamic characteristics and genetic regulatory mechanisms of methane release waveforms, resulting in the inability to locate key gene targets. Furthermore, the lack of differentiated editing strategies leads to low efficiency and may cause non-target trait abnormalities, making it difficult to meet the refined requirements of low-carbon farming.
The methane release waveform characteristics were analyzed by gene editing target prediction unit, and gene targets that regulate carbon efficiency were located by combining epigenetic marker database. Differentiated gene editing schemes were designed, including guiding ribonucleic acid sequences, homologous recombination repair templates and base editing. A deep convolution-attention hybrid model was used for target prediction and feedback optimization.
It has enabled the precise targeting of key genes for carbon efficiency regulation, improved the safety and efficiency of gene editing, reduced non-target trait abnormalities, enhanced the reliability and adaptability of breeding results, and contributed to the low-carbon transformation of animal husbandry.
Smart Images

Figure CN120998300A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-carbon livestock farming and precision breeding technology, specifically to an artificial intelligence-based intelligent analysis and breeding decision-making system for beef cattle carbon emission data. Background Technology
[0002] Low-carbon livestock farming and precision breeding is an important technology. In the context of low-carbon agricultural development, methane emissions from beef cattle farming are a significant source of greenhouse gases. Precisely monitoring carbon emission data and combining it with genetic selection to reduce the carbon footprint is of great significance for achieving goals such as sustainable development of livestock farming.
[0003] This technology breaks through the limitations of traditional breeding that relies on experience, providing data-driven decision support for the efficient breeding of low-carbon beef cattle breeds. Traditional breeding methods that rely solely on phenotypic observation or single indicators are no longer sufficient to meet the refined needs of low-carbon farming. However, traditional beef cattle carbon emission analysis and breeding technologies suffer from a core problem of insufficient precision in carbon efficiency regulation. Existing solutions only assess carbon efficiency through overall methane emissions, without deeply analyzing the dynamic characteristics of methane release waveforms and their correlation with genetic regulatory mechanisms. This makes it impossible to locate key gene targets that affect carbon efficiency. At the same time, there is a lack of differentiated editing strategies for targets with different regulatory efficiencies. Using a uniform editing method for high-impact and low-impact genes is not only inefficient but may also lead to abnormalities in non-target traits. These problems result in limited improvement in the carbon efficiency of the bred beef cattle, making it difficult to promote and apply on a large scale and meet the urgent needs of the low-carbon transformation of animal husbandry. To solve this technical problem, we provide an AI-based intelligent analysis and breeding decision system for beef cattle carbon emission data. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence, so as to solve the problems mentioned in the background art.
[0005] 1. Since traditional methods only assess overall methane emissions without analyzing the relationship between waveform dynamics and genetic mechanisms, they cannot locate key gene targets. Therefore, this case uses a gene editing target prediction unit to analyze methane release waveform characteristics and match them with an epigenetic marker database, which can accurately locate gene targets that regulate carbon efficiency.
[0006] 2. Since traditional methods use a uniform editing approach for targets with different efficiencies, which is inefficient and may cause abnormalities, this case uses a gene editing instruction generation unit to design different editing schemes for high-, medium- and low-efficiency targets, which can improve editing efficiency and reduce abnormalities in non-target traits.
[0007] To achieve the above objectives, an AI-based intelligent analysis and breeding decision-making system for beef cattle carbon emission data is provided, including a phenotypic and carbon efficiency fusion unit. This unit acquires dynamic methane release waveforms and rumination rhythm time-series data of individual beef cattle via an implanted rumen capsule, and calculates a carbon efficiency ratio correction value based on the dynamic methane release waveforms. The system is characterized by further comprising:
[0008] The gene editing target prediction unit is used to input the carbon efficiency ratio correction value into the pre-trained deep convolutional-attention hybrid model and perform the following operations:
[0009] Analyze the key abrupt changes in the methane release waveform, including the steep slope of the waveform and the duration of the plateau phase;
[0010] Match the key mutation features with the epigenetic marker database to locate editable target genes that regulate methane metabolism. Based on the regulatory efficacy of the target genes to the carbon efficiency ratio correction value, output a list of high regulatory efficacy targets and a list of medium and low regulatory efficacy targets.
[0011] The gene editing instruction generation unit is used to design guide RNA sequences and homologous recombination repair templates for genes in the list of high regulatory efficacy targets, generate base editing schemes to inhibit translation initiation for genes in the list of medium and low regulatory efficacy targets, synchronize the instruction set to the bovine gene editing platform, and trigger micromanipulation of in vitro fertilized eggs.
[0012] The phenotypic feedback optimization unit is used to collect methane release waveforms in real time after the gene-edited offspring are born. When the deviation between the mutation characteristics of the methane release waveform and the expected value exceeds a set threshold, it automatically feeds back a target efficacy correction signal to the gene editing target prediction unit.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0014] 1. The gene editing target prediction unit uses a deep convolution-attention hybrid model to analyze key mutational features such as the steep slope of the methane release waveform and the duration of the plateau phase. Combined with an epigenetic marker database, it locates easily editable target genes through cosine similarity matching. Based on the level of regulatory efficacy on carbon efficiency, it accurately distinguishes high, medium and low regulatory efficacy targets. This mechanism breaks through the limitations of traditional methods that rely solely on the overall methane emissions, and achieves precise targeting of key genes for carbon efficiency regulation, providing clear targets for subsequent editing.
[0015] 2. The gene editing instruction generation unit designs guide ribonucleic acid sequences with disulfide bond modifications and photosensitive switches, along with dynamically responsive homologous recombination repair templates, targeting high-regulatory-efficiency targets. It reduces off-target risks through quantum annealing algorithms and activates editing only in specific regions. For medium- and low-regulatory-efficiency targets, it uses base editing to inhibit translation initiation, avoiding the inefficiency caused by uniform editing, reducing the possibility of non-target trait abnormalities, and improving the safety and effectiveness of editing.
[0016] 3. The methane release waveform data of gene-edited offspring is fed back to the target prediction unit in real time to correct the target efficacy parameters and ensure that the breeding effect continues to meet expectations. The deep convolution-attention hybrid model eliminates the interference of breed differences through multimodal data training and transfer learning, and enhances the adaptability to different beef cattle breeds. At the same time, the biological basis of the waveform features is verified by combining rumen microbiome data, which further ensures the editability and regulatory effectiveness of the target, providing reliable support for the continuous breeding of low-carbon beef cattle breeds and helping the low-carbon transformation of animal husbandry. Attached Figure Description
[0017] Figure 1 This is an overall block diagram of the present invention.
[0018] The meanings of the labels in the diagram are as follows:
[0019] 1. Phenotype and carbon efficiency fusion unit; 2. Gene editing target prediction unit; 3. Gene editing instruction generation unit; 4. Phenotype feedback optimization unit. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] This invention provides an AI-based intelligent analysis and breeding decision-making system for beef cattle carbon emission data. Please refer to [link / reference]. Figure 1 As shown, it includes a phenotypic and carbon efficiency fusion unit 1, used to acquire dynamic methane release waveforms and rumination rhythm time-series data of individual beef cattle through an implanted rumen capsule, and calculate carbon efficiency ratio correction values based on the dynamic methane release waveforms. Its characteristic is that it further includes:
[0022] Gene editing target prediction unit 2 is used to input the carbon efficiency ratio correction value into the pre-trained deep convolutional-attention hybrid model and perform the following operations:
[0023] Analyze the key abrupt changes in the methane release waveform, including the steep slope of the waveform and the duration of the plateau phase;
[0024] After acquiring dynamic methane release waveforms and rumination rhythm time-series data of individual beef cattle in the phenotypic and carbon efficiency fusion unit 1 and calculating the carbon efficiency ratio correction value, the gene editing target prediction unit 2 performs operations to analyze key mutation features in the methane release waveforms in order to accurately locate key genetic characteristics affecting carbon efficiency. Specifically, this includes:
[0025] Specifically, a one-dimensional convolution kernel is used to perform a sliding scan on the waveform. The kernel size can be set according to the waveform's fluctuation period to extract local pulse signals. These pulse signals correspond to the instantaneous methane release event triggered by rumen contraction. This method can separate discrete release features directly related to physiological activities from continuous waveforms, providing a basic unit for subsequent analysis and ensuring that no key instantaneous release information is missed. Next, to highlight the rapid metabolic features related to genetic variations, attention weights are used to amplify the correlation between the amplitude and time span of the methane release waveform during the steep rise phase. The steep rise phase is defined as the waveform interval where the methane concentration increases by more than a set threshold per unit time. This threshold is determined based on the population average increase. The time span refers to the length of time from the start to the peak of this interval. Attention mechanisms assign higher weight to intervals with large amplitudes and short time spans, indicating rapid metabolic responses. This strengthens rapid metabolic characteristics determined by genetic factors, facilitating subsequent association with specific gene targets and enhancing the correlation between characteristics and genetic basis. Finally, to assess the energy metabolism efficiency of beef cattle, an adaptive sliding window is used to capture the periodicity of the plateau period duration. The plateau period refers to the interval where methane concentration remains at a high level with fluctuations less than a set value. The size of the sliding window automatically adjusts according to the waveform period of the previous stage to fully capture its duration and repetitive patterns. This operation reflects the ability to maintain methane production homeostasis, which is directly related to energy metabolism efficiency and can provide a basis for locating energy metabolism-related genes. By using multi-scale decomposition to extract transient release events, attention mechanisms to enhance rapid metabolic characteristics, and adaptive windows to capture homeostasis patterns, the key mutation features synergistically analyzed in these three steps not only comprehensively reflect the dynamic characteristics of methane release but also accurately associate them with genetic regulatory mechanisms. This lays a reliable foundation for subsequent matching of epigenetic markers and locating editing targets, effectively improving the specificity and accuracy of target prediction.
[0026] After resolving the key mutational features of the methane release waveform using a deep convolution-attention hybrid model, further validation using rumen microbiome metagenomic data is needed to ensure that these features are not random fluctuations but have a clear biological basis.
[0027] The biological basis of waveform features is validated based on rumen microbiome metagenomic data. This data, obtained through gene sequencing and bioinformatics analysis of rumen contents samples collected from beef cattle, includes genetic information and species abundance data of all microorganisms in the rumen. Correlation analysis with the resolved waveform features provides a microbiological explanation for these features, avoiding the inclusion of biologically meaningless noise features in subsequent analyses and ensuring feature reliability. When a steep slope is detected—the inclination of the rapid rise in methane concentration—and a strong positive correlation is found between the abundance of specific methanogens (correlation coefficient exceeding 0.7), it indicates that the steep slope is triggered by the metabolic activity of this type of methanogen. At this point, the system automatically associates the metabolic pathway gene clusters of this type of methanogen, such as the methyl-CoM reductase gene cluster involved in methane synthesis, with an epigenetic marker database. This database stores epigenetic markers related to gene expression regulation. This links microbial metabolic features with host epigenetic regulation, providing a theoretical basis for subsequent targeting of editable targets and enhancing the biological relevance of the targets. Meanwhile, to verify the biological significance of the plateau duration, i.e. the duration of stable methane concentration, and to reverse-validate the causal relationship between the plateau duration and the expression level of key enzyme genes using gene editing interference technology, this study ensures that key mutant features have editable biological targets. Specifically, for key enzyme genes involved in methane metabolism, their expression was temporarily inhibited through gene editing, and changes in the plateau duration were observed. If the plateau duration was significantly shortened when the enzyme gene expression level decreased, and prolonged when the expression level recovered, it indicates a causal relationship between the two. This operation can clarify the direct link between the plateau feature and the function of a specific gene, avoiding erroneous conclusions based solely on correlation. Through the above steps, the association between the steep slope and the abundance of methanogens and the underlying metabolic gene basis can be verified, and the causal relationship between the plateau duration and key enzyme genes can be confirmed. Ultimately, this ensures that the key mutant features identified correspond to editable biological targets, providing a solid biological basis for the subsequent precise location of gene editing targets. This effectively reduces the risk of ineffective editing due to the lack of actual regulatory basis for the features, and improves the scientificity and reliability of the entire breeding decision-making system.
[0028] Match key mutation features with epigenetic marker databases to locate editable target genes that regulate methane metabolism. Based on the regulatory efficacy of target genes to carbon efficiency ratio correction values, output lists of high-efficiency and medium-to-low-efficiency target genes.
[0029] After analyzing and validating the biological basis of key mutational features in methane release waveforms, it is necessary to match these features with epigenetic marker databases to associate them with specific gene targets. The specific steps involved in matching key mutational features with epigenetic marker databases include:
[0030] First, it is necessary to clarify the source of the pre-constructed epigenetic association map. This map is a gene regulatory relationship network formed by integrating epigenetic data and methane release characteristic data of beef cattle populations through machine learning modeling. Nodes represent genes or epigenetic markers, and edges represent the regulatory relationships between them. Inputting the steep slope and plateau duration into the pre-constructed epigenetic association map can quickly locate potential targets by leveraging the existing association relationships in the map, avoiding the inefficiency caused by searching from scratch. Next, the nodes of the epigenetic association map are traversed through a bidirectional recurrent neural network. This network can simultaneously trace the association path forward (from feature to gene) and backward (from gene to feature), ensuring that no possible regulatory relationships are missed. Through this bidirectional traversal, gene nodes related to the steep slope and plateau duration can be comprehensively mined, improving the comprehensiveness of target localization. During the traversal, the similarity between the current methane release waveform feature vector (composed of steep slope and plateau duration) and the epigenetic marker vector of the gene regulatory region (methylation level vector of the promoter region) needs to be calculated. Cosine similarity is used as the calculation method. This method judges the degree of similarity by measuring the angle between the two vectors. The closer the value is to 1, the higher the similarity. When the similarity exceeds a set threshold (0.8, which is determined based on the accuracy of historical matching data), the corresponding gene is marked as an easily editable target gene. This can screen out the genes most closely associated with the feature through quantitative similarity indicators, thereby improving the accuracy of the target. Furthermore, since gene promoter or enhancer regions are key areas for regulating gene expression, editing these regions is more likely to affect gene function. Therefore, among the easily editable target genes selected, targets located in promoter or enhancer regions are preferentially retained. This operation can improve the efficiency of subsequent gene editing and ensure that the edited gene can effectively regulate the metabolic processes related to methane release. By inputting features into the epigenetic association map, traversing the bidirectional recursive neural network, screening with cosine similarity calculation, and prioritizing key region targets, a series of operations work together to accelerate target localization by utilizing existing epigenetic association knowledge and ensuring the relevance and editability of targets through quantitative indicators and regional priorities. This lays a reliable foundation for subsequent target efficacy grading and editing scheme design, effectively improving the efficiency and accuracy of gene editing target prediction.
[0031] After matching key mutational features with epigenetic marker databases and locating easily editable target genes, differentiated editing strategies for targets with varying impacts need to be developed. This requires grading the regulatory efficacy of the target gene on the carbon efficiency ratio correction value, specifically including:
[0032] First, it is necessary to clarify the source of the historical gene editing experimental dataset. This dataset contains records of changes in the carbon efficiency ratio correction values of beef cattle after gene editing of the same or similar targets in the past. This data is obtained through the system's built-in database retrieval function, providing a historical reference for assessing the regulatory efficacy of the current target and avoiding grading based solely on theoretical speculation, thus ensuring the objectivity of the grading criteria. For each easily editable target gene, the fluctuation range of the carbon efficiency ratio correction value of beef cattle after editing is extracted from the historical dataset, i.e., the difference between the maximum and minimum values of the edited carbon efficiency ratio correction value. This range directly reflects the actual impact of the target on carbon efficiency; the greater the fluctuation, the stronger the regulatory effect, providing a quantitative indicator for grading. Simultaneously, it is necessary to determine the population standard deviation of the carbon efficiency ratio correction value. This standard deviation is obtained by calculating the dispersion of the carbon efficiency ratio correction value of the same breed of beef cattle that has not undergone gene editing, reflecting the level of carbon efficiency fluctuation under natural conditions. This serves as a benchmark for measuring regulatory efficacy. By comparing with natural fluctuations, the additional effects brought by gene editing can be clearly identified, avoiding misjudging natural fluctuations as editing effects. The specific grading rules are as follows:
[0033] If the fluctuation range after target editing is greater than twice the population standard deviation, it indicates that the target can significantly change carbon efficiency and is classified into the list of high-efficiency targets. If the fluctuation range is less than the population standard deviation but greater than a preset fold standard deviation (usually set to 0.5, depending on the editing precision requirements), it indicates that the target has some but limited impact on carbon efficiency and is classified into the list of medium-to-low-efficiency targets. Other targets with fluctuation ranges less than or equal to the preset fold standard deviation are automatically filtered out due to their weak regulatory effects. This rule can clearly distinguish the regulatory capabilities of different targets, providing a clear basis for subsequent differentiated editing. By retrieving fluctuation ranges from historical data, quantifying and classifying them based on the population standard deviation, and filtering out inefficient targets, the scientific validity and reproducibility of the classification results are ensured. This also allows subsequent gene editing work to focus on high-efficiency targets, avoiding resource waste. At the same time, it provides a classification basis for designing appropriate editing schemes for medium-to-low-efficiency targets, effectively improving the targeting and overall efficiency of gene editing.
[0034] Training methods for deep convolutional-attention hybrid models include:
[0035] To ensure that the deep convolution-attention hybrid model can accurately analyze the key mutation features of the methane release waveform, a multi-step training method is required. The training of this model is the foundation for gene editing target prediction and can provide reliable algorithmic support for subsequent feature analysis.
[0036] First, since a single data type cannot fully reflect the relationship between carbon metabolism and genetic regulation in beef cattle, it is necessary to integrate multimodal data to construct a training set. Dynamic methane release waveforms (collected via implanted rumen capsules) and rumination rhythm time-series data (time series of beef cattle rumination behavior recorded by sensors) need to undergo wavelet transform for noise reduction. Wavelet transform effectively separates signal from noise, retains key waveform features, and removes noise caused by environmental interference and equipment errors. These are then input into a 3D convolutional layer to extract spatiotemporal features. The 3D convolutional layer can simultaneously capture the changing trends in the time dimension and the correlation in the feature dimension, such as the synchronicity of methane concentration fluctuations over time and rumination behavior. This transforms the raw data into a more representative feature vector, improving the model's ability to capture key features. Simultaneously, the whole-genome methylation map of the corresponding beef cattle individual is imported as a supervision label. The whole-genome methylation map, obtained through gene sequencing technology, reflects the regulatory state of gene expression and is directly related to the activity of carbon metabolism-related genes. This provides a clear learning objective for model training, enabling the model to learn the potential correlation between waveform features and epigenetic regulation, ensuring the accuracy of the training direction. Given the relatively mature research on human epigenetic regulation and the limited data on beef cattle, a transfer learning approach was adopted to adapt the attention weight matrix pre-trained on human epigenetic regulation to the beef cattle data domain. The adaptation process involved adjusting matrix parameters to match the regulatory patterns in human data with the characteristics of beef cattle data. This approach leverages existing knowledge to accelerate model convergence, reduces reliance on large amounts of labeled beef cattle data, and improves training efficiency. Since different beef cattle breeds have varying genetic backgrounds, potentially leading to inconsistent model interpretation results for different breeds, an adversarial domain adaptation technique was employed to eliminate breed-specific interference. This technique constructs a domain discriminator, enabling the model to learn universal features unaffected by breed. This keeps the model's feature interpretation error for steep waveform slopes within a preset level, set according to detection accuracy requirements. This operation ensures stable interpretation results across different beef cattle breeds, enhancing the model's versatility. By integrating multimodal data to construct a training set, using whole-genome methylation maps as supervisory labels, accelerating training through transfer learning, and eliminating variety differences through adversarial techniques, a series of synergistic operations not only enriched the model's learning materials but also optimized the training process by leveraging existing knowledge and technology. The resulting model can accurately analyze the key mutational features of methane release waveforms, providing reliable algorithmic support for subsequent target prediction and effectively improving the analytical accuracy and applicability of the entire system.
[0037] Gene editing instruction generation unit 3 is used to design guide RNA sequences and homologous recombination repair templates for genes in the list of high regulatory efficacy targets, and to generate base editing schemes to inhibit translation initiation for genes in the list of medium and low regulatory efficacy targets. The instruction set is synchronized to the bovine gene editing platform to trigger micromanipulation of in vitro fertilized eggs.
[0038] Gene editing instruction generation unit 3 designs and guides RNA sequences for a list of high-regulatory-efficiency targets, specifically including:
[0039] After completing the classification of the regulatory efficacy of target genes and identifying a list of high-efficiency targets, to ensure the accuracy and safety of gene editing, gene editing instruction generation unit 3 needs to design guide RNA sequences for these high-impact targets. First, since the guide RNA sequence may cause off-target effects if it binds to non-target genes, an off-target effect prediction engine based on the quantum annealing algorithm needs to be invoked. This engine efficiently searches for possible binding sites through the quantum annealing algorithm, and can find potential off-target regions more quickly than traditional algorithms. At the same time, the chromatin open region data of the target gene is input. The chromatin open region refers to the loose region in the gene that is easy to bind to regulatory factors. It is obtained through chromatin immunoprecipitation sequencing technology and reflects the editable active region of the gene. This approach can focus on designing sequences for the active region of the gene, reduce non-specific binding to inactive regions, and provide a precise range of regions for predicting off-target risks. Next, the risk of potential binding sites was assessed by simulating the unwinding barrier of DNA double strands. The unwinding barrier refers to the energy required for the double strands to open; higher energy indicates a tighter binding and less susceptibility to misbinding. Sequences were designed preferentially from regions with low unwinding barriers and off-target risk prediction values below a set threshold, thus avoiding high off-target risk sites. This process allows for the selection of more specific sequences from a molecular mechanics perspective, significantly reducing interference with non-target genes during editing. In terms of sequence structure design, a guide ribonucleic acid backbone structure with disulfide bond modifications was generated. Disulfide bonds enhance backbone stability and reduce the probability of degradation by nucleases. Simultaneously, a light-sensitive switch peptide chain was attached to the ribonucleic acid chain at a specific end, the 3' end. This peptide chain inhibits the activity of the guide ribonucleic acid in the absence of near-infrared light and releases this inhibition upon light exposure, ensuring that editing activity is activated only in the fertilized egg region exposed to near-infrared light. This allows for precise spatiotemporal control of editing through external light, preventing accidental editing in non-target cells or developmental stages. By using quantum annealing algorithms to predict off-target risks, simulating unspinning barriers to screen safe regions, modifying the backbone and attaching photosensitive switches, a series of operations were conducted to ensure the specificity and stability of the guiding RNA sequence and to achieve controllable activation of editing activity. This provided a key sequence basis for subsequent efficient and safe gene editing operations, effectively improving the accuracy and safety of gene editing and reducing the risk of unexpected traits.
[0040] The design method for homologous recombination repair templates specifically includes:
[0041] Based on the design of guide RNA sequences for high regulatory efficacy targets, in order to ensure that the carbon efficiency ratio can be precisely regulated after gene editing, it is necessary to design a suitable homologous recombination repair template. This template is a key tool for achieving precise repair and functional regulation of target genes.
[0042] First, a dynamic response repair template library is constructed. This library contains various modular repair units that can be adjusted according to gene regulation needs. Its construction is based on regulatory network data of genes related to carbon metabolism in beef cattle. By integrating gene interaction relationships and expression regulation data, it can provide diverse repair solutions for different targets, avoiding the problem that a single template cannot adapt to complex regulatory needs and improving the template's versatility. Next, based on the topological weight of the target gene in the carbon efficiency ratio regulation network—which reflects the gene's core position in the regulatory network—and calculated using network analysis algorithms, higher weights indicate a more critical role for the gene in the overall regulation of carbon efficiency ratio. Selective insertion of promoter elements regulated by methylation levels is performed. The activity of these promoters changes with gene methylation levels; high methylation levels result in low promoter activity, and vice versa. These elements can be obtained from a gene regulatory element database. This approach allows the expression of the repaired gene to be dynamically regulated by its own methylation state, avoiding excessively strong or weak gene expression and achieving steady-state regulation of carbon efficiency ratio. When the target is located at a core node of the regulatory network, in genes ranking within the top 20% of topological weight, a bidirectional promoter-driven fluorescently labeled protein, such as green fluorescent protein (GFP), is used. Its expression directly reflects the success of gene editing through fluorescence signals. A specific codon repressor transport RNA is also used. This transport RNA recognizes and inhibits stop codons in genes, prolonging protein synthesis and enhancing gene co-expression. This setup allows for monitoring the editing effect through fluorescence signals and enhancing the function of core node genes through regulatory factors, ensuring strong regulation of carbon efficiency. When the target is located at a peripheral node, in genes ranking within the bottom 50% of topological weight, the coding sequence of a gene editing activation protein is linked to the repair template. This protein activates the expression of other carbon metabolism-related genes, forming a negative feedback loop for carbon efficiency. Specifically, when carbon efficiency is too high, the expression of the activation protein increases, promoting the function of carbon metabolism-related genes and reducing carbon efficiency; when carbon efficiency is too low, the expression of the activation protein decreases, maintaining carbon metabolism balance. This design enables fine-tuning of carbon efficiency through the synergistic effect of peripheral node genes driving the entire regulatory network. By constructing a dynamic repair template library, selecting regulatory elements based on topological weights, and designing differentiated expression units for core and edge nodes, homologous recombination repair templates can not only accurately repair target genes, but also dynamically regulate their functions according to the role of genes in the regulatory network. This provides a stable molecular basis for the efficient regulation of carbon efficiency after gene editing, effectively improving the targeting and persistence of gene editing in improving carbon efficiency in beef cattle.
[0043] The generation of base editing schemes to suppress translation initiation includes the following operations:
[0044] After designing gene editing protocols for high-efficiency targets, for genes in the list of medium- and low-efficiency targets, since their impact on carbon efficiency is relatively weak, targeted base editing protocols are needed to inhibit translation initiation, thereby moderately regulating their function.
[0045] First, to accurately locate the editing site, the secondary structure of the gene's messenger RNA (MRNA) needs to be analyzed. MRNA carries genetic information and guides protein synthesis; its secondary structure, such as stem-loop structures, affects the translation process. Single-molecule real-time sequencing technology is used. This technology can observe the nucleic acid chain synthesis process in real time, accurately capture the structural information formed by base pairing, and obtain the spatial folding state of MRNA. This allows direct observation of the position and stability of key structures such as stem-loops, providing a structural basis for selecting editing sites and ensuring that the selection of editing sites is based on molecular structural characteristics. Next, among the analyzed secondary structures, stem-loop regions with free energy below a preset threshold are selected as editing sites. The lower the free energy, the more stable the stem-loop structure, and the greater its influence on translation initiation. The preset threshold is set based on the structural stability data of similar genes. Structural changes in these regions are more likely to interfere with the formation of the translation initiation complex. This improves the efficiency of editing in inhibiting the translation process, ensuring a significant reduction in gene expression after editing. Subsequently, by fusing cytosine deaminase with a transcription activator-like effector protein, the transcription activator-like effector protein can precisely recognize specific deoxyribonucleic acid sequences, guiding the cytosine deaminase to act on the target site. The cytosine deaminase can catalyze the base transversion of cytosine to uracil, directionally catalyzing the target base transversion, causing the start codon (AUG, the signal to start translation) to mutate into a stop codon (UAG, UAA, which will terminate protein synthesis). This operation can directly block the initiation of translation, fundamentally inhibiting gene expression, and the directional catalysis ensures the specificity of the editing, reducing the impact on other sequences. Simultaneously, to further enhance the inhibitory effect, guide RNA targeting adjacent intron splicing enhancers was designed. Intron splicing enhancers are sequences that promote intron splicing and ensure the correct processing of messenger RNA; their spatial conformation is crucial to function. This guide RNA can bind to the enhancer region, blocking its spatial conformation folding, preventing it from functioning properly, leading to abnormal messenger RNA processing and indirectly inhibiting the translation process. This dual mechanism strengthens the inhibitory effect on genes with low to medium regulatory efficacy, ensuring a moderate reduction in their function to match carbon efficiency ratio regulation. Through single-molecule sequencing to resolve the structure, selecting stable stem-loop regions as sites, directional catalytic mutation of the start codon, and simultaneous blocking of splicing enhancer function, a series of synergistic operations achieve precise inhibition of genes with low to medium regulatory efficacy while avoiding over-editing that interferes with the function of other genes. This provides a mild and controllable editing scheme for moderate carbon efficiency ratio regulation, effectively balancing editing effectiveness and biosafety, and improving the adaptability of the entire gene editing system to different efficacy targets.
[0046] The specific procedures for triggering micromanipulation of in vitro fertilized eggs include:
[0047] After completing the design of guide RNA sequences for high-regulatory-efficiency targets and the construction of homologous recombination repair templates, as well as the generation of base editing schemes for medium- and low-regulatory-efficiency targets, gene editing needs to be achieved through in vitro fertilized egg micromanipulation. This process is a key step in transforming the editing scheme into actual gene modification.
[0048] First, to ensure the editing complex can precisely target the fertilized egg and avoid degradation, it is encapsulated in a liposome carrier. Liposomes are vesicles composed of a lipid bilayer that can encapsulate biomolecules and fuse with the cell membrane. They are obtained from cell banks or reagent suppliers, and the surface of the liposomes is modified with peptide chains that target the zona pellucida protein of the fertilized egg. The zona pellucida is the outer glycoprotein layer of the fertilized egg. This peptide chain is extracted from antibodies targeting the zona pellucida protein by immunoprecipitation and can specifically bind to the zona pellucida, ensuring that the liposomes precisely target the fertilized egg. This improves the delivery efficiency of the editing complex to the fertilized egg, reduces non-specific effects on other cells, and enhances the targeting of the operation. Next, the fertilized egg was observed and positioned using a micromanipulator, an instrument capable of highly precise positioning and manipulation of tiny biological samples. This ensured that the editing tool accurately acted on the target location, providing a precise spatial reference for subsequent operations and preventing editing failure due to positioning deviations. Subsequently, an alternating magnetic field was applied, with the magnetic field strength and frequency set according to the magnetic response characteristics of the liposomes. This caused the liposome carrier to rupture due to magnetostriction, releasing the internal pH-responsive editing component. This component is only activated under specific pH conditions within the fertilized egg, preventing premature activation during delivery. This control ensures the editing component activates at the correct time and location, enhancing the spatiotemporal precision of the editing. Simultaneously, near-infrared light of a specific wavelength was applied through an optical fiber array. This wavelength of light causes minimal cell damage and can be recognized by the photosensitive switch peptide chains pre-coupled to it, activating the photosensitive switch in the editing complex and further initiating the editing process. This dual activation mechanism minimizes erroneous editing and improves operational safety. After the operation, spectral analysis technology is used to determine the success of the editing by detecting the characteristic spectral changes of the gene sequences before and after editing. Non-destructive testing is performed on the fertilized eggs to screen out embryos with high editing efficiency, i.e., those whose spectral signals show that the target gene modification has achieved the expected results. Only these embryos are transferred into the mother's uterus. This ensures that the transferred embryos have the expected gene editing effect, improves the success rate of breeding, and reduces the waste of resources caused by ineffective transplantation. Through liposome targeted delivery, microscopic positioning, synergistic activation by alternating magnetic fields and near-infrared light, and spectral detection screening, a series of operations not only ensure the accuracy and safety of gene editing, but also improve the quality of embryo transfer through efficient screening. This lays a solid foundation for breeding beef cattle offspring with low carbon efficiency and effectively promotes the transformation from gene editing programs to actual breeding results.
[0049] Phenotypic feedback optimization unit 4 is used to collect methane release waveforms in real time after the gene-edited offspring are born. When the deviation between the mutation characteristics of the methane release waveform and the expected value exceeds a set threshold, it automatically feeds back a target efficacy correction signal to gene editing target prediction unit 2.
[0050] In this invention, the phenotype and carbon efficiency fusion unit 1 acquires dynamic methane release waveforms and rumination rhythm data through an implanted rumen capsule, and calculates the carbon efficiency ratio correction value. The gene editing target prediction unit 2 uses a deep convolution-attention hybrid model to analyze key mutation features of the waveform, matches the epigenetic marker database, locates and hierarchically regulates the target genes of carbon efficiency. The gene editing instruction generation unit 3 designs guiding ribonucleic acid sequences, homologous recombination repair templates, and base editing schemes for high, medium, and low regulatory efficacy targets, respectively, and synchronizes them to the breeding cattle gene editing platform. The phenotype feedback optimization unit 4 collects the methane release waveforms of offspring in real time, corrects the target efficacy parameters, and realizes precise location and differentiated editing of carbon efficiency-related genes, thereby improving the efficiency of low-carbon breeding of beef cattle.
[0051] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-based intelligent analysis and breeding decision-making system for beef cattle carbon emission data, comprising a phenotypic and carbon efficiency fusion unit (1), used to acquire dynamic methane release waveforms and rumination rhythm time-series data of individual beef cattle via an implanted rumen capsule, and to calculate a carbon efficiency ratio correction value based on the dynamic methane release waveforms, characterized in that, Also includes: The gene editing target prediction unit (2) is used to input the carbon efficiency ratio correction value into the pre-trained deep convolutional-attention hybrid model and perform the following operations: Analyze the key abrupt changes in the methane release waveform, including the steep slope of the waveform and the duration of the plateau phase; Match the key mutation features with the epigenetic marker database to locate editable target genes that regulate methane metabolism. Based on the regulatory efficacy of the target genes to the carbon efficiency ratio correction value, output a list of high regulatory efficacy targets and a list of medium and low regulatory efficacy targets. The gene editing instruction generation unit (3) is used to design a guide ribonucleic acid sequence and a homologous recombination repair template for genes in the list of high regulatory efficacy targets, generate a base editing scheme to inhibit translation initiation for genes in the list of medium and low regulatory efficacy targets, synchronize the instruction set to the bovine gene editing platform, and trigger micromanipulation of in vitro fertilized eggs. Phenotypic feedback optimization unit (4) is used to collect methane release waveforms in real time after the gene-edited offspring are born. When the deviation between the mutation characteristics of the methane release waveform and the expected value exceeds the set threshold, it automatically feeds back the target efficacy correction signal to the gene editing target prediction unit (2).
2. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 1, characterized in that, The operation of parsing key mutation features in the methane release waveform in the gene editing target prediction unit (2) specifically includes: The dynamic methane release waveform is decomposed into multiple scales using a deep convolution-attention hybrid model. Local pulse signals in the methane release waveform are extracted using a one-dimensional convolution kernel to capture the instantaneous methane release event triggered by rumen contraction. The correlation between the amplitude of methane release waveform and time span during the steep rise phase is amplified by attention weighting. The steep rise phase is the waveform interval in which the methane concentration increases by more than a set threshold per unit time. The time span refers to the length of time from the start to the peak of this interval. This is used to enhance the rapid metabolic characteristics related to genetic variation. Finally, an adaptive sliding window was used to capture the periodic pattern of the plateau duration, reflecting the ability to maintain methane production steady state, and was used to evaluate the energy metabolism efficiency of beef cattle.
3. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 1, characterized in that, The operation of matching key mutation features with epigenetic marker databases specifically includes: The steep slope and plateau duration are input into a pre-constructed epigenetic association map. The nodes of the epigenetic association map are traversed through a bidirectional recursive neural network to locate easily editable target genes whose similarity to the key mutation features of the current methane release waveform exceeds a set threshold. The similarity is calculated based on the cosine similarity between the methane release waveform feature vector and the epigenetic marker vector of the gene regulatory region, and targets located in the gene promoter or enhancer region are preferentially screened.
4. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 3, characterized in that, The classification based on the regulatory efficacy of target genes on carbon efficiency ratio correction values specifically includes: For each easily editable target gene, its historical gene editing experimental dataset is retrieved, and the fluctuation range of the carbon efficiency ratio correction value of beef cattle after editing the target is extracted. Targets with a fluctuation range greater than twice the population standard deviation of the carbon efficiency ratio correction value are classified into the list of high regulatory efficacy targets, targets with a fluctuation range less than the standard deviation but greater than a preset multiple of the standard deviation are classified into the list of medium and low regulatory efficacy targets, and the remaining targets are automatically filtered.
5. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 1, characterized in that, The training method for the deep convolutional-attention hybrid model specifically includes: A training set was constructed by fusing multimodal data. The dynamic methane release waveform and rumination rhythm time series data were denoising by wavelet transform and then input into a three-dimensional convolutional layer to extract spatiotemporal features. The whole genome methylation map of the corresponding beef cattle individual was simultaneously imported as a supervision label. The attention weight matrix pre-trained in human epigenetic regulation was adapted to the beef cattle data domain through transfer learning. Finally, adversarial domain adaptation technology was used to eliminate the interference of breed differences, so that the feature parsing error of the model for the steep slope of the waveform was controlled within the preset level.
6. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 2, characterized in that, The operation of analyzing the key abrupt changes in the methane release waveform further includes: Based on the biological basis of waveform features validated by metagenomic data of the rumen microbiome, when a strong positive correlation is detected between the steep slope and the abundance of a specific methanogen, the metabolic pathway gene clusters are automatically linked to an epigenetic marker database. Furthermore, the gene editing system interference technology is used to reverse-validate the causal relationship between the duration of the plateau phase and the expression level of key enzyme genes, ensuring that key mutation features have editable biological targets.
7. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 1, characterized in that, The gene editing instruction generation unit (3) designs and guides ribonucleic acid sequences for a list of high-regulatory-efficiency targets, specifically including: The off-target effect prediction engine based on quantum annealing algorithm is invoked. The chromatin open region data of the target gene is input. By simulating the deoxyribonucleic acid double-strand unwinding energy barrier, high off-target risk sites are avoided. A guide ribonucleic acid backbone structure with disulfide bond modification is generated, and a photosensitive switch peptide chain is attached to a specific end of it, so that the editing activity is activated only in the fertilized egg region exposed to near-infrared light.
8. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 7, characterized in that, The design method of the homologous recombination repair template specifically includes: A dynamic response repair template library was constructed. Based on the topological weight of the target gene in the carbon efficiency ratio regulatory network, promoter elements regulated by methylation level were selectively inserted. When the target is located at the core node of the regulatory network, a bidirectional promoter was used to drive the co-expression of fluorescently labeled protein and special codon repressor transport RNA. When the target is located at the edge node, the coding sequence of gene editing activation protein was connected to form a carbon efficiency ratio negative feedback loop.
9. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 1, characterized in that, The base editing scheme for generating bases to suppress translation initiation specifically includes: For genes in the list of targets with medium to low regulatory efficacy, single-molecule real-time sequencing technology was used to analyze their messenger ribonucleic acid secondary structures. Stem-loop regions with free energy below a preset threshold were selected as editing sites. Cytosine deaminase, which is fused with transcription activator-like effector protein, was used to catalyze base transversion, causing the start codon to mutate into a stop codon. Simultaneously, guide ribonucleic acid targeting the splicing enhancer of adjacent introns was designed to block its spatial conformation folding.
10. The intelligent analysis and breeding decision-making system for beef cattle carbon emission data based on artificial intelligence according to claim 1, characterized in that, The specific procedure for triggering micromanipulation of in vitro fertilized eggs includes: The editing complex is encapsulated in a liposome carrier, and its surface is modified with peptide chains targeting the zona pellucida protein of the fertilized egg. After the fertilized egg is positioned in a micromanipulator, an alternating magnetic field is applied to cause the liposome to release an acid-base responsive editing component. At the same time, a photosensitive switch is activated by applying near-infrared light of a specific wavelength through an optical fiber array. After the operation is completed, the editing efficiency is non-destructively detected using spectral analysis technology, and only high-success-rate embryos are transferred to the mother.