A method for targeted nutrient source mining based on beef cattle enterotype-host gene interaction

By constructing a targeted nutrient source mining method based on the interaction between the intestinal type and host genes in beef cattle, and utilizing the co-expression network of microecology and host genes and molecular conformation search, the instability problem of feed nutrient source development in existing technologies has been solved, achieving precise nutrient regulation and efficient development of new nutrient sources.

CN121884961BActive Publication Date: 2026-05-15内蒙古元牛繁育科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610344111.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-20
Publication Date
2026-05-15
Estimated Expiration
2046-03-20

AI Technical Summary

Technical Problem

Existing methods for developing feed nutrient sources rely on large-scale animal trials and lack molecular-level validation, resulting in inconsistent effects of screened products in cattle herds with different genetic backgrounds, and failing to achieve precise regulation of beef cattle growth performance.

Method used

By constructing a targeted nutrient source mining method based on the interaction between the beef cattle gut type and the host gene, and utilizing the co-expression network of the microecology and the host gene, combined with molecular conformation search and Gibbs free energy calculation, targeted effector factors are screened out and matched with natural raw materials to generate precise nutritional regulation schemes.

Benefits of technology

This enables a deep mapping from macroscopic phenotypic classification to microscopic gene regulation, ensuring that metabolic targets have a biological interaction basis, and improving the success rate of developing novel nutrient sources and their practical application effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884961B_ABST
    Figure CN121884961B_ABST
Patent Text Reader

Abstract

The application discloses a targeted nutrient source mining method based on beef cattle intestinal type-host gene interaction, and particularly relates to the field of biological information data processing, and is used for solving the problems of lack of molecular mechanism verification and poor targeting of existing nutrient source development. First, based on the microbiome and transcriptome data, the nutrient-related intestinal type is divided, and the co-expression network is constructed to screen the host co-expression specific genes driven by the specific intestinal type. Then, the model of the group metabolite ligand set and the host protein receptor is established, the molecular conformation search and the Gibbs free energy calculation are performed, and the high-activity targeted effect factor is screened based on the physical affinity. Finally, the targeted effect factor is used for traversing matching and quantitative screening of natural raw material liquid chromatography-mass spectrometry data, and a targeted nutrient source recommendation list for the specific intestinal type is generated. The application combines multi-omics correlation analysis and molecular thermodynamic verification, and constructs a precise development closed loop from micro mechanism analysis to macro raw material matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics data processing, specifically to a method for targeted nutrient source mining based on the interaction between beef cattle gut-host genes. Background Technology

[0002] Beef cattle farming is an important component of modern animal husbandry and one of the main sources of high-quality animal protein for humans. In large-scale farming, feed conversion efficiency is a core indicator determining the economic benefits and resource utilization rate of the operation. As the livestock industry develops towards precision farming, traditional herd-based, average feeding methods ignore the differences in digestive and metabolic characteristics among individual beef cattle, making it difficult to meet the demands of modern farming for maximizing production performance. Matching appropriate nutritional programs to the physiological characteristics of different subgroups and precisely controlling these programs to improve the growth performance of beef cattle has become an important development direction in feed engineering and animal nutrition.

[0003] However, existing methods for developing feed nutrients have significant technical limitations in practical applications. Traditional development models primarily rely on large-scale animal feeding trials, screening for effective components based on phenotypic data. This validation model is not only time-consuming and costly, but also suffers from inconsistent efficacy in cattle herds with different genetic backgrounds due to a lack of understanding of the microscopic mechanisms of action. Although high-throughput sequencing technology has been used to analyze gut microbiota or host genes, existing analytical methods are mainly based on single-dimensional statistical association analysis, i.e., calculating the linear correlation coefficient between changes in microbial abundance and phenotypic data. This statistical association-based method struggles to distinguish between biological causal relationships and numerical co-occurrences between microorganisms and the host, easily introducing false positive results. More critically, existing technologies fail to verify at the molecular physical level whether microbial metabolites possess the spatial structural basis for binding to host cell receptors. This results in many nutrients screened based on statistical correlation failing in practical applications because they cannot complete key molecular signal transduction, thus hindering precise regulation of host growth performance. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a targeted nutrient source mining method based on the interaction between beef cattle gut type and host genes, thus solving the problems mentioned above.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a targeted nutrient source mining method based on beef cattle intestinal type-host gene interaction, comprising the following steps: S1. Obtaining intestinal microbiome sequencing data and blood transcriptome sequencing data of beef cattle populations; constructing a microbiome abundance matrix based on the intestinal microbiome sequencing data and calculating the microecological distance between samples; clustering the beef cattle population into different nutrient-associated intestinal types and extracting the core indicator bacteria genera for each nutrient-associated intestinal type; using the nutrient-associated intestinal type as a grouping label to construct a co-expression network for the blood transcriptome sequencing data, and screening for co-expressed specific genes highly associated with specific nutrient-associated intestinal types; S2. Retrieving secondary metabolites that the core indicator bacteria genera can synthesize from the microbial metabolic pathway database, and establishing ligand molecules containing the chemical structures of the secondary metabolites. Dataset; Obtain the three-dimensional crystal structure data of the protein encoded by the co-expressed specific gene, and extract the spatial coordinate parameters of the protein active binding site; S3. Map the chemical structure in the ligand molecule dataset to the spatial coordinate parameter range of the protein active binding site, perform molecular conformation search and Gibbs free energy calculation, and output the binding energy score; Sort the ligand molecule dataset by affinity according to the binding energy score, and select metabolite molecules with binding energy scores better than a preset threshold as the targeting effector factors for targeting the co-expressed specific gene; S4. Use the targeting effector factors as search probes, perform traversal matching in the feature library containing natural feed ingredient components and liquid chromatography-mass spectrometry data, screen natural raw materials containing the targeting effector factors and whose content meets the preset concentration, and generate a recommended list of targeted nutrient sources for nutritionally related intestinal types.

[0006] Further, the following process was performed to obtain gut microbiome sequencing data and blood transcriptome sequencing data of beef cattle. Based on the gut microbiome sequencing data, a microbial abundance matrix was constructed, and the microecological distance between samples was calculated. The beef cattle population was then clustered into different nutrient-associated enterotypes, and the core indicator genera for each nutrient-associated enterotype were extracted. The specific steps are as follows: Sequence denoising and feature clustering were performed on the gut microbiome sequencing data to generate a feature sequence list containing species classification information. The feature sequence list was accumulated and normalized to construct a microbial abundance matrix reflecting the proportion of species in the samples. Based on the microbial abundance matrix, the probability distribution difference between sample pairs in the beef cattle population was calculated using the Jensen-Shannon divergence algorithm to generate a microecological distance matrix. This matrix was then input into a clustering algorithm around a central point. The optimal number of clusters was determined by calculating the silhouette coefficient, thus classifying the beef cattle population into nutrient-associated enterotypes. The sample data within each nutrient-associated enterotype were traversed, and the frequency and average relative abundance of bacterial genera within the nutrient-associated enterotype were calculated. Genera with both frequency and average relative abundance exceeding a preset threshold were identified as core indicator genera.

[0007] Furthermore, using nutrient-associated gut type as a grouping label, a co-expression network was constructed from blood transcriptome sequencing data. The specific process for screening co-expression-specific genes highly associated with specific nutrient-associated gut types is as follows: The blood transcriptome sequencing data was standardized, and the correlation coefficient matrix of gene expression levels was calculated. A soft threshold power parameter was introduced to transform the correlation coefficient matrix into a weighted adjacency matrix. The topological overlap matrix reflecting the tightness of gene connections was further calculated. Based on the topological overlap matrix, a hierarchical clustering algorithm was used to identify gene co-expression modules. The first principal component of the gene co-expression module was calculated as the module feature gene. The correlation between the module feature gene and the nutrient-associated gut type, which is a binary feature vector, was calculated. Specific modules whose correlation meets the preset significance criteria were screened. Within the specific module, the association significance between the gene and the nutrient-associated gut type was calculated. Genes with association significance better than the screening threshold were locked as co-expression-specific genes.

[0008] Furthermore, the specific process of retrieving secondary metabolites that core indicator bacteria genera can synthesize from the microbial metabolic pathway database and establishing a ligand molecule dataset containing the chemical structures of secondary metabolites is as follows: The core indicator bacteria genera are mapped to the microbial genome database, functional genomic information is extracted, and the enzyme-catalyzed reaction network is reconstructed using the metabolic pathway database. The substrate-product transformation pathway is traced within the enzyme-catalyzed reaction network, secondary metabolites located at the end of the metabolic pathway are identified, primary metabolites of basic life activities are removed, and a potential metabolite list is generated. The chemical structure encoding files of secondary metabolites in the potential metabolite list are obtained by retrieving chemical small molecule structure databases. These chemical structure encoding files are filtered according to the upper limit of molecular weight and drug-likeness rules to establish a ligand molecule dataset.

[0009] Further, the specific process for obtaining the three-dimensional crystal structure data of the protein encoded by the co-expression-specific gene and extracting the spatial coordinate parameters of the protein active binding site is as follows: A biomolecular structure database is searched to find protein crystal structure files that match the coding sequence of the co-expression-specific gene. The protein crystal structure files are preprocessed by dehydration and charge balancing. A geometric probe algorithm is used to scan the protein surface to identify concave regions with specific volume characteristics as potential active binding sites. The geometric center coordinates of the potential active binding sites are calculated. Three-dimensional spatial constraint parameters are set according to the spatial extension range of the potential active binding sites. The geometric center coordinates and the three-dimensional spatial constraint parameters are combined and output as the spatial coordinate parameters of the protein active binding site.

[0010] Furthermore, the chemical structures in the ligand molecule dataset are mapped to the spatial coordinate parameters of the protein active binding site. A molecular conformation search and Gibbs free energy calculation are performed to output the binding energy score. The specific process is as follows: A three-dimensional affinity potential energy grid matrix is ​​constructed based on the spatial coordinate parameters. Rotatable chemical bonds in the chemical structures of the ligand molecule dataset are set to maintain molecular flexibility. Through a global optimization algorithm, a semi-flexible conformation search is performed on the chemical structures in the ligand molecule dataset within the three-dimensional affinity potential energy grid matrix. The van der Waals forces, electrostatic interactions, and hydrogen bond potential energy between the ligand and receptor are iteratively calculated. The minimum energy value after iterative convergence is output as the binding energy score.

[0011] Furthermore, the specific process of ranking the ligand molecule dataset by affinity based on binding energy scores and selecting metabolite molecules with binding energy scores better than a preset threshold as targeting effectors for co-expressing specific genes is as follows: Root mean square bias clustering analysis is performed on the ligand conformations generated by conformation search to extract representative binding patterns of dominant conformation clusters; the ligand molecule dataset is ranked by affinity from low to high binding energy scores, and candidate molecules with high rankings and binding energy scores lower than the preset energy threshold are selected; the contact between candidate molecules and protein active binding site residues is checked, molecules with steric hindrance conflicts are eliminated, and the remaining molecules are identified as targeting effectors.

[0012] Furthermore, the specific process of performing a traversal matching in a feature library containing natural feed ingredient components and liquid chromatography-mass spectrometry (LC-MS) data, using the targeting effect factor as a search probe, is as follows: The chemical structure of the targeting effect factor is analyzed, and the theoretical mass-to-charge ratio and characteristic ion peak information of the targeting effect factor under LC-MS detection conditions are predicted; the feature library containing natural feed ingredient components and LC-MS data is traversed, and the secondary mass spectrometry information of known components of each natural raw material in the feature library is retrieved; the cosine similarity between the characteristic ion peak information and the secondary mass spectrometry data is calculated, and natural raw materials with a cosine similarity higher than the preset matching standard are marked as initially matched raw materials.

[0013] Furthermore, the specific process for selecting natural raw materials containing target effect factors at preset concentrations and generating a recommended list of targeted nutrient sources for nutrition-related gut types is as follows: Extract the ion current peak area at the chromatographic retention time of the corresponding target effect factor in the initially targeted raw materials; perform integral calculation on the ion current peak area using a preset standard curve equation to calculate the relative content of the target effect factor in the natural raw materials; compare the relative content with the preset concentration, retain natural raw materials with a relative content higher than the preset concentration, obtain the name, origin, and target effect factor content data of the retained natural raw materials, and generate a recommended list of targeted nutrient sources in descending order of content.

[0014] The present invention has the following beneficial effects:

[0015] (1) A targeted nutrient source mining method based on the interaction between beef cattle gut type and host genes solves the problem of inaccurate target identification caused by the neglect of sample heterogeneity and reliance on only a single-dimensional statistical association in existing technologies by constructing a co-expression coupling network of microecology and host genes. This invention first calculates the microecological distance based on microbiome data and performs clustering and typing to divide the complex beef cattle population into nutrient-related gut types with biological commonalities, reducing the background noise of microecological data; at the same time, a host gene co-expression network is constructed using nutrient-related gut types as group labels, which can accurately identify host-specific response genes driven by core bacteria in specific gut type environments. This method systematically couples the microbial community structure with the host gene expression network, extracts core indicator bacteria genera and co-expressed specific genes, and realizes a deep mapping from macroscopic phenotypic classification to microscopic gene regulatory networks, ensuring that the metabolic targets mined subsequently have a clear biological interaction basis.

[0016] (2) A targeted nutrient source mining method based on the interaction between beef cattle intestinal type and host genes solves the problem of high false positives and inapplicability caused by the lack of physical-level validity verification in existing technologies by introducing molecular conformation search and thermodynamic binding energy calculation. This invention no longer relies solely on statistical predictions from bioinformatics, but extracts the spatial coordinate parameters of the host protein active binding sites and uses computer simulation technology to perform molecular conformation search and Gibbs free energy calculation at the atomic level. This process verifies the spatial matching degree and binding affinity between metabolites and host receptors from a physicochemical perspective, effectively eliminating invalid molecules that are statistically relevant but physically unable to bind. In addition, by traversing and matching the screened targeted effect factors with natural raw material liquid chromatography-mass spectrometry data, a natural raw material scheme containing effective components and meeting the concentration requirements is directly output, opening up the transformation path from theoretical calculation to feed production application and significantly improving the success rate of developing new nutrient sources.

[0017] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0018] Figure 1 This is a flowchart of a targeted nutrient source mining method based on the interaction between the intestinal type and the host gene in beef cattle, according to the present invention. Detailed Implementation

[0019] This application provides a targeted nutrient source mining method based on the interaction between the intestinal type of beef cattle and the host gene. This method solves the problems of low screening accuracy due to reliance on weak statistical associations and unstable practical application effects due to the lack of molecular-level physical binding verification in the development of existing feed nutrient sources.

[0020] The overall concept of the solution in this application embodiment is as follows:

[0021] First, multi-omics sequencing data of beef cattle populations were acquired. Clustering algorithms were used to identify different nutrient-associated gut types, and a host gene co-expression network was constructed based on this to identify host-specific co-expressed genes driven by specific gut microbiota. Subsequently, databases were searched to obtain the structures of potential metabolites of core gut microbiota and the protein structure parameters of host-specific genes, establishing a ligand and receptor structure dataset. Next, molecular docking technology was used to simulate the binding process between metabolites and host proteins. By calculating Gibbs free energy, high-affinity targeting effectors were screened to verify the effectiveness of the interaction from a physical perspective. Finally, the targeting effectors were used as search probes to traverse the feature library of natural feed ingredients, matching natural ingredients rich in effective components to generate a precise nutritional regulation program for specific nutrient-associated gut types.

[0022] Please see Figure 1 This invention provides a technical solution: a targeted nutrient source mining method based on the interaction between beef cattle gut type and host genes, comprising the following steps: S1. Obtaining gut microbiome sequencing data and blood transcriptome sequencing data of beef cattle populations; constructing a microbiome abundance matrix based on gut microbiome sequencing data and calculating the microecological distance between samples; clustering the beef cattle population into different nutrient-associated gut types and extracting the core indicator bacteria genera for each nutrient-associated gut type; using the nutrient-associated gut type as a grouping label to construct a co-expression network for the blood transcriptome sequencing data, and screening for co-expression specific genes highly associated with specific nutrient-associated gut types; S2. Retrieving secondary metabolites that the core indicator bacteria genera can synthesize from the microbial metabolic pathway database, and establishing a ligand molecule dataset containing the chemical structures of secondary metabolites; S3. Obtain the three-dimensional crystal structure data of the protein encoded by the co-expressed specific gene and extract the spatial coordinate parameters of the protein active binding site; S4. Map the chemical structure in the ligand molecule dataset to the spatial coordinate parameter range of the protein active binding site, perform molecular conformation search and Gibbs free energy calculation, and output the binding energy score; Sort the ligand molecule dataset by affinity according to the binding energy score, and select metabolite molecules with binding energy scores better than a preset threshold as the targeting effector factors for targeting the co-expressed specific gene; S5. Use the targeting effector factors as search probes to perform traversal matching in a feature library containing natural feed ingredient components and liquid chromatography-mass spectrometry data, screen natural raw materials containing the targeting effector factors and whose content meets the preset concentration, and generate a recommended list of targeted nutrient sources for nutritionally related intestinal types.

[0023] In this implementation scheme, step S1 primarily involves dimensionality reduction, classification, and feature extraction of biological big data to address sample heterogeneity interference and pinpoint biological targets. Specifically, the system first standardizes the acquired high-throughput sequencing data, constructs a microbial abundance matrix to digitally represent the microbial community structure, and calculates the microecological distance between samples using algorithms such as the Jensen-Shannon divergence. This distance reflects the degree of difference in the probability distribution of gut microbiota composition among different individuals. Based on this distance, clustering is performed to classify the complex beef cattle population into several nutritionally associated gut types with similar metabolic characteristics. This concept refers to a microbial community pattern exhibiting a stable structure under specific feeding or physiological conditions. Furthermore, using this gut type classification as a supervisory label, a weighted gene co-expression network analysis method is employed to construct a host gene co-expression network, identifying co-expression-specific genes whose expression patterns highly synergistically change with specific gut types. This establishes a preliminary statistical association between the microecological environment and host physiological responses, providing precise input for subsequent mechanism analysis. In step S2, the system primarily completes the cross-omics mapping from biological information to physicochemical models, providing a structured data foundation for computer simulation screening. The system retrieves a dedicated database and maps the core indicator bacterial genera identified in S1 to their potential secreted secondary metabolites, establishing a ligand molecule dataset. Here, secondary metabolites refer to small molecule compounds synthesized by microorganisms at specific growth stages that have regulatory functions on the host but are not related to microbial growth. Simultaneously, for the host-specific co-expressed genes screened in S1, the system obtains the three-dimensional crystal structures of their encoded proteins and uses geometric algorithms to extract the spatial coordinate parameters of the protein's active binding sites. These parameters precisely describe the geometry and physicochemical properties of the depressions or pockets on the protein surface that can specifically bind to small molecules. This step transforms abstract gene and microbial community names into computer-recognizable ligand chemical structures and receptor spatial models, realizing the materialization of biological information. In step S3, computer-aided drug design technology is mainly used to perform virtual screening and physical verification, aiming to eliminate spurious correlations and identify effectors with real physical forces. The system places the ligand small molecules within the active pocket coordinate range of the receptor protein, performs molecular conformation search to simulate the dynamic recognition process between the two at the microscopic level, and calculates the Gibbs free energy score. Gibbs free energy, a state function in thermodynamics describing the stability of a system, is used here to quantify the stability of the system after a metabolite binds to a host protein. A lower value indicates a tighter binding and a more stable complex. Based on this score, all candidate molecules are ranked by affinity, and a threshold is set for selection. This allows for the screening of metabolite molecules that can physically bind with high affinity and potentially activate host genes—i.e., targeted effector factors. This process confirms the effectiveness of the microbial-host interaction from a physicochemical perspective.In step S4, the main focus is on the reverse matching from microscopic molecular discovery to macroscopic feed application, aiming to generate a practical and precise nutritional regulation solution. The system uses the screened target effect factors as fingerprint features and searches through a feature library containing natural raw material components and liquid chromatography-mass spectrometry (LC-MS) data. LC-MS data refers to the chemical fingerprint spectrum of substances obtained using liquid chromatography separation and mass spectrometry detection techniques, which can achieve qualitative and quantitative analysis of trace components in complex mixtures. By comparing the similarity between the characteristic ion peaks of the target effect factors and the raw material spectrum, and combining this with standard curves for content calculation, the system can accurately locate natural plant or agricultural by-product raw materials rich in the effective component, and finally generate a recommendation list matching specific nutritionally associated intestinal patterns, realizing the transformation from theoretical calculation to actual feed formulation development.

[0024] Specifically, the process of acquiring gut microbiome sequencing data and blood transcriptome sequencing data of beef cattle populations, constructing a microbial abundance matrix based on the gut microbiome sequencing data, calculating the microecological distance between samples, and clustering the beef cattle population into different nutrient-associated enterotypes and extracting the core indicator genera for each nutrient-associated enterotype is as follows: Sequence denoising and feature clustering are performed on the gut microbiome sequencing data to generate a feature sequence list containing species classification information. The feature sequence list is then accumulated and normalized to construct a microbial abundance matrix reflecting the proportion of species in the samples. Based on the microbial abundance matrix, the probability distribution difference between sample pairs in the beef cattle population is calculated using the Jensen-Shannon divergence algorithm to generate a microecological distance matrix. This matrix is ​​then input into a clustering algorithm around a central point. The optimal number of clusters is determined by calculating the silhouette coefficient, thus classifying the beef cattle population into nutrient-associated enterotypes. The sample data within each nutrient-associated enterotype are traversed, and the frequency and average relative abundance of bacterial genera within the nutrient-associated enterotype are calculated. Genera with both frequency and average relative abundance exceeding a preset threshold are identified as core indicator genera.

[0025] In this implementation plan, the raw sequencing data is first subjected to sequence denoising and normalization to eliminate systematic errors caused by uneven sequencing depth and ensure that different samples are compared on the same baseline. After constructing the microbial abundance matrix, this plan uses the Jensen-Shannon divergence algorithm to measure the microecological distance between samples. Jensen-Shannon divergence is a similarity measure based on probability distribution. Compared to traditional Euclidean distance, it can more accurately handle the sparsity and zero-value problems common in microbiome data, truly reflecting the structural differences in the gut microbiota composition of different beef cattle individuals. The formula for calculating microecological distance is as follows: In the formula, : Indicates beef cattle sample Compared with beef cattle samples The Jansen-Shannon microecological distance between them; : Indicates the total number of microbial species detected; : indicates the first Microorganisms in the sample The relative abundance in; : indicates the first Microorganisms in the sample The relative abundance in; : Represents the natural logarithm. After calculating the distance matrix, the system executes a clustering algorithm around the centroid. This algorithm iteratively finds the most representative centroid (i.e., the core sample) such that the sum of the dissimilarity of all samples within a cluster to that centroid is minimized, thus ensuring the robustness of the gut-like pattern. The objective function for clustering is as follows: In the formula, : Represents the total cost function value of the cluster analysis; : Indicates the number of nutritionally associated gut types, which is determined by the principle of maximizing the silhouette coefficient; : indicates that it belongs to the first A sample set of intestinal types; : indicates the first A central point sample of an intestinal pattern; : Indicates a sample With the center point of the corresponding intestinal type The Jansen-Shannon microecological distance between them. Through the above steps, the system classifies the population into structurally stable enterotypes, and further identifies the core indicator genera of each enterotype through dual threshold screening of frequency and abundance, providing clear microbial targets for subsequent analysis.

[0026] Specifically, the process of constructing a co-expression network using nutrient-associated gut type as a grouping label for blood transcriptome sequencing data and screening for co-expression-specific genes highly associated with specific nutrient-associated gut types is as follows: The blood transcriptome sequencing data is standardized, and the correlation coefficient matrix of gene expression levels is calculated. A soft threshold power parameter is introduced to transform the correlation coefficient matrix into a weighted adjacency matrix, and a topological overlap matrix reflecting gene connectivity is further calculated. Based on the topological overlap matrix, a hierarchical clustering algorithm is used to identify gene co-expression modules. The first principal component of each gene co-expression module is calculated as the module's feature gene. The correlation between the module's feature gene and the nutrient-associated gut type, which is used as a binary feature vector, is calculated, and specific modules whose correlation meets a preset significance standard are screened. Within each specific module, the association significance between the gene and the nutrient-associated gut type is calculated, and genes with association significance better than the screening threshold are identified as co-expression-specific genes.

[0027] In this implementation scheme, a weighted gene co-expression network analysis strategy is used to process transcriptome data, aiming to identify synergistic regulatory patterns among genes from a systems biology perspective, rather than independent changes in single genes. After calculating the correlation coefficient matrix, a soft-threshold power parameter is introduced to transform the correlation values ​​into a weighted adjacency matrix. This step ensures that the gene network follows a scale-free network distribution, i.e., a few nodes have a large number of connections, which conforms to the true topology of biological networks. To further eliminate noise and enhance the robustness of network connections, the system calculates a topological overlap matrix. This matrix considers not only the direct correlation between two genes but also the degree of overlap of their common neighbor nodes. The formula for calculating the topological overlap matrix is ​​as follows: In the formula, : Represents gene With genes The topological overlap measure between them; : Represents gene With genes Connection strength in a weighted adjacency matrix; : Represents gene With the third node gene The connection strength; : Represents gene With the third node gene The connection strength; : Represents a minimum value function. After identifying gene modules based on the topological overlap matrix, the system calculates the characteristic genes of each module and screens for specific modules. To accurately pinpoint key driver genes within a module, the system calculates a co-expression specificity index within that specific module. This index comprehensively measures the gene's position within the module and its association with gut type. The formula for calculating gene association significance is as follows: In the formula, : Represents gene The gene significance value is used to measure the degree of correlation between the gene and the nutritionally associated gut type; : Represents gene Expression vectors in each sample; : Represents the binary feature vector of nutrient-associated gut type. It takes a value of 1 when the sample belongs to this gut type, and 0 otherwise. : Represents the Pearson correlation coefficient calculation function. High correlation coefficients are filtered by setting a threshold. By identifying the most significant host-specific genes affected by specific gut microbiota environments, the system successfully pinpointed the host-specific genes most significantly influenced by specific gut microbiota environments, achieving effective coupling between the gut microbiota and host genetic information.

[0028] Specifically, the process of retrieving secondary metabolites synthesized by core indicator bacteria genera from the microbial metabolic pathway database and establishing a ligand molecule dataset containing the chemical structures of these secondary metabolites is as follows: The core indicator bacteria genera are mapped to a microbial genome database; functional genomic information is extracted; and the enzyme-catalyzed reaction network is reconstructed using the metabolic pathway database. The substrate-product transformation pathway is traced within the enzyme-catalyzed reaction network, and secondary metabolites located at the end of the metabolic pathway are identified. Primary metabolites of basic life activities are removed, generating a potential metabolite list. Chemical structure encoding files of secondary metabolites in the potential metabolite list are obtained by searching a small molecule structure database. These files are then filtered according to molecular weight limits and drug-likeness rules to establish a ligand molecule dataset.

[0029] In this implementation scheme, the system first performs microbial genome mapping and enzyme-catalyzed reaction network reconstruction to explore the potential metabolic capabilities of the microbial community. The core indicator genera identified in step S1 are compared with publicly available microbial genome databases to extract functional gene annotation information from their whole genome sequences, particularly gene clusters encoding metabolic enzymes. Using a metabolic pathway database as a reference template, the identified enzymes are connected into a complex enzyme-catalyzed reaction network according to substrate-product transformation relationships. Within this network, the system identifies compounds at the end of metabolic pathways that no longer participate in the construction of basic cell structures or energy maintenance through topological structure analysis, defining them as secondary metabolites. To ensure that the screened molecules have the potential to become feed additives or drugs, the system establishes a ligand molecule dataset and introduces cheminformatics filtering rules. These rules not only consider the chemical stability of the molecules but also quantitatively assess their drug-likeness, i.e., the likelihood of the molecules being absorbed and utilized in vivo. Drug-likeness assessment is achieved by calculating a metabolite suitability score, the formula of which is as follows: In the formula, The metabolite suitability score represents the m-th secondary metabolite, which is used to measure the physicochemical properties of the molecule as a ligand. , , : These represent the weighting coefficients for molecular weight limitation, lipid-water partition coefficient, and hydrogen bond formation ability, respectively. Their values ​​are determined through expert scoring or training with historical drug data. : Represents the molecular weight of the m-th secondary metabolite; This indicates a preset upper limit threshold for molecular weight, such as 500 Daltons; : Represents an indicator function, which takes the value 1 when the condition is met, and 0 otherwise; : Represents the lipid-water partition coefficient of the m-th secondary metabolite; : Represents the target value of the optimal fat-water partition coefficient; : Indicates the allowable range of fluctuation in the lipid-water partition coefficient; : Indicates the number of hydrogen bond acceptors in a molecule; : Indicates the number of hydrogen bond donors in the molecule; : Represents the total number of heavy atoms in the molecule. By setting a scoring threshold, the system eliminates molecules that are too large, too polar, or too weak, which are not conducive to permeation of biological membranes, thus establishing a high-quality ligand molecule dataset.

[0030] Specifically, the process of obtaining the three-dimensional crystal structure data of the protein encoded by the co-expression-specific gene and extracting the spatial coordinate parameters of the protein active binding site is as follows: A biomolecular structure database is searched to find protein crystal structure files that match the coding sequence of the co-expression-specific gene. The protein crystal structure files are preprocessed by dehydration and charge balancing. A geometric probe algorithm is used to scan the protein surface to identify recessed regions with specific volume characteristics as potential active binding sites. The geometric center coordinates of the potential active binding sites are calculated. Three-dimensional spatial constraint parameters are set according to the spatial extension range of the potential active binding sites. The geometric center coordinates and the three-dimensional spatial constraint parameters are combined and output as the spatial coordinate parameters of the protein active binding site.

[0031] In this implementation scheme, the system performs protein 3D structure acquisition and binding site identification for co-expression-specific genes, aiming to construct an accurate receptor physical model. For proteins with resolved crystal structures, the system directly calls the structure file; for unresolved proteins, homology modeling methods are used to predict their folding conformation. Preprocessing steps include removing water of crystallization molecules, adding polar hydrogen atoms, and balancing the side chain charge of amino acids according to the physiological pH environment, which is crucial for accurately calculating electrostatic interactions. Subsequently, a geometric probe algorithm is used to scan the protein surface. This algorithm simulates a probe sphere with a specific radius rolling on a van der Waals surface to detect recessed regions that can accommodate small molecules. To define the search range for molecular docking, the system needs to convert irregular recessed regions into regular spatial parameters that can be recognized by the computer, i.e., calculating the geometric center and 3D spatial constraint box of potential active binding sites. The formula for calculating the spatial coordinate parameters is as follows: ; In the formula, : Indicates the potential active binding site at The coordinates of the geometric center in the dimension. : Represents the dimension direction of the Cartesian coordinate system, with values ​​of x, y, or z. : Indicates the total number of amino acid residue atoms that constitute potential active binding sites. : indicates that the a-th atom is in Spatial coordinates in a dimension. : Indicates the three-dimensional spatial constraint parameters in Side length in the dimension. This parameter defines the size of the search box during molecular docking. : Represents the set of atoms that constitute potential active binding sites. max: Represents the function that takes the maximum value. min: Represents the function that takes the minimum value. : This represents the preset spatial buffer distance value to prevent edge effects. Through the above calculations, the system outputs accurate receptor grid parameters, ensuring that subsequent molecular docking processes can be concentrated in the most biologically active regions, greatly improving computational efficiency and the accuracy of results.

[0032] Specifically, the process of mapping the chemical structures in the ligand molecule dataset to the spatial coordinate parameters of the protein active binding site, performing molecular conformation search and Gibbs free energy calculation, and outputting the binding energy score is as follows: A three-dimensional affinity potential energy grid matrix is ​​constructed based on the spatial coordinate parameters, and rotatable chemical bonds in the chemical structures of the ligand molecule dataset are set to maintain molecular flexibility; through a global optimization algorithm, a semi-flexible conformation search is performed on the chemical structures in the ligand molecule dataset within the three-dimensional affinity potential energy grid matrix, and the van der Waals forces, electrostatic interactions, and hydrogen bond potential energies between the ligand and the receptor are iteratively calculated. The minimum energy value after iterative convergence is output as the binding energy score.

[0033] In this implementation scheme, the system first constructs a three-dimensional affinity potential energy lattice matrix based on the geometric characteristics of the receptor protein. This is a data structure that discretizes and stores a continuous spatial energy distribution, aiming to accelerate the subsequent energy calculation process. The system defines a spatial box around the receptor active site, dividing it into regularly arranged tiny cubic lattice points. The interaction potential energy between each probe atom and the receptor at each lattice point is pre-calculated and stored. Subsequently, the system sets the rotatable chemical bonds of the ligand molecule, allowing it to change its torsion angle during docking, thereby simulating the dynamic adaptation process of small molecules in a microscopic environment, i.e., a semi-flexible conformation search. To find the binding posture with the lowest energy, the system employs a global optimization algorithm to search within the potential energy lattice matrix, avoiding getting trapped in local energy minima. In each iteration of the search, the system calculates the Gibbs free energy binding score between the ligand and receptor based on an empirical force field function. This score comprehensively reflects the key interactions that contribute to the thermodynamic stability of the complex, and its calculation formula is as follows: In the formula, : Represents the Gibbs free energy binding score between the ligand and the receptor; : Indicates the total number of heavy atoms in a ligand molecule; : Indicates the total number of heavy atoms within the active site range of the receptor protein; This represents the p-th atom in the ligand molecule; This represents the q-th atom in the receptor protein; : Represents the Euclidean distance between atoms p and q; : Represents the van der Waals repulsion coefficient between atom p and atom q; : Represents the van der Waals attraction coefficient between atom p and atom q; : Represents the partial charge value of atom p; : Represents the partial charge value of atom q; : Represents the dielectric constant of the simulated environment, used to adjust the strength of electrostatic interactions; Represents the set of atom pairs that form hydrogen bonds; This indicates the maximum potential depth of a hydrogen bond. The value represents the deviation of the hydrogen bond angle from the ideal angle. Using the above formula, the system quantifies the contributions of van der Waals forces, electrostatic interactions, and hydrogen bonds to the binding affinity. The minimum energy value after iterative convergence represents the optimal binding state of this conformation.

[0034] Specifically, the process of ranking the ligand molecule dataset by affinity based on binding energy scores and selecting metabolite molecules with binding energy scores above a preset threshold as targeting effectors for co-expressing specific genes is as follows: Root mean square bias clustering analysis is performed on the ligand conformations generated by conformation search to extract representative binding patterns from the dominant conformation clusters; the ligand molecule dataset is ranked by affinity from low to high binding energy scores, and candidate molecules with high rankings and binding energy scores below a preset energy threshold are selected; the contact between candidate molecules and protein active binding site residues is checked, molecules with steric hindrance conflicts are eliminated, and the remaining molecules are identified as targeting effectors.

[0035] In this implementation scheme, to extract statistically significant binding patterns from a massive number of conformational search results, the system first performs root mean square deviation (RMSD) cluster analysis. Because molecular docking algorithms are random, a single lowest-energy conformation may be a chance result. Cluster analysis aims to group conformations with highly similar spatial structures into a single cluster, identifying the most frequently occurring dominant conformation cluster, which represents the most thermodynamically probable binding state. RMSD is used to quantify the spatial positional differences between two ligand conformations, and its calculation formula is as follows: In the formula, : Represents the root mean square deviation between the u-th ligand conformation and the v-th ligand conformation; : Indicates the number of non-hydrogen heavy atoms in the ligand molecule; : Represents the spatial position vector of the k-th heavy atom in the u-th conformation; : Represents the spatial position vector of the k-th heavy atom in the v-th conformation; : Represents the calculation of the vector's modulus. After completing clustering and extracting representative conformations, the system ranks all candidate molecules by affinity based on their binding energy scores. A lower binding energy score (usually negative) indicates more free energy released between molecules and a tighter binding. The system selects molecules with high rankings and scores below a preset energy threshold, which can be determined by the known statistical distribution of binding energies of positive controls, for example, by subtracting the standard deviation from the average binding energy of the controls. Finally, the system performs a steric hindrance check, eliminating unreasonable conformations that, while having low energies in the mathematical model, result in atomic overlap or collisions in physical space. This ensures that the selected targeting effectors can successfully embed into the receptor pocket and exert their regulatory effects in the real biological environment.

[0036] Specifically, the process of using the target effect factor as a search probe to perform a traversal matching in a feature library containing natural feed ingredient components and liquid chromatography-mass spectrometry (LC-MS) data is as follows: The chemical structure of the target effect factor is analyzed, and the theoretical mass-to-charge ratio and characteristic ion peak information of the target effect factor under LC-MS detection conditions are predicted; the feature library containing natural feed ingredient components and LC-MS data is traversed, and the secondary mass spectrometry information of known components of each natural raw material in the feature library is retrieved; the cosine similarity between the characteristic ion peak information and the secondary mass spectrometry data is calculated, and natural raw materials with a cosine similarity higher than the preset matching standard are marked as initially matched raw materials.

[0037] In this implementation scheme, the system first performs virtual spectrum prediction and matching based on the principle of liquid chromatography-mass spectrometry (LC-MS) to solve the cross-modal retrieval problem between chemical structures and measured spectra. The system analyzes the chemical structure of the target effect factor determined in step S3, calculates its theoretical mass-to-charge ratio (MTBR) based on a preset ionization mode (e.g., positive or negative ion mode), and predicts its characteristic ion peak information based on chemical bond breaking rules. Subsequently, the system traverses the secondary mass spectrometry data of natural raw materials in the feature library. Secondary mass spectrometry data reflects the distribution of fragment ions generated by collision-induced dissociation of molecules, serving as a fingerprint for substance identification. To quantify the similarity between the target effect factor and a certain component in the raw material, the system calculates the cosine similarity of their spectra. This algorithm treats the mass spectrum as a high-dimensional vector and evaluates the matching degree by calculating the cosine value of the angle between the vectors. The formula for calculating the cosine similarity is as follows: In the formula, : Represents the cosine similarity between the theoretically predicted spectral vector T of the targeting effect factor and the experimentally measured spectral vector E of a certain component in a natural raw material; : This represents the total number of intervals into which the mass-to-charge ratio range is discretized, that is, dividing the continuous mass spectrum axis into several tiny mass windows; : indicates the h-th quality interval; : Represents the theoretical relative abundance of the characteristic ion of the targeting effect factor in the h-th mass interval; : Indicates the experimental relative abundance of ions in the h-th mass interval of the measured spectrum of natural raw materials; This indicates a summation operation. By setting a cosine similarity threshold (e.g., greater than 0.9), the system can quickly identify natural raw materials containing molecules with structures highly consistent with the target effector from complex plant extract data, achieving precise traceability from single compounds to mixtures of raw materials.

[0038] Specifically, the process of selecting natural raw materials containing target effect factors at preset concentrations and generating a recommended list of targeted nutrient sources for nutritionally linked gut microbiota is as follows: Extract the ion current peak area at the chromatographic retention time of the corresponding target effect factor in the initially targeted raw material; perform integral calculation on the ion current peak area using a preset standard curve equation to calculate the relative content of the target effect factor in the natural raw material; compare the relative content with the preset concentration, retain natural raw materials with a relative content higher than the preset concentration, obtain the name, origin, and target effect factor content data of the retained natural raw materials, and generate a recommended list of targeted nutrient sources in descending order of content.

[0039] In this implementation scheme, the system performs quantitative analysis and list generation based on chromatographic elution curves to assess whether the content of active ingredients in natural raw materials has reached the effective concentration. For initially identified natural raw materials, the system extracts the extract ion chromatogram at a specific chromatographic retention time and integrates the peak area of ​​the ion flow within that retention time window. Retention time is the time a substance resides in the chromatographic column, used for qualitative analysis; peak area is proportional to the substance concentration and used for quantitative calculation. The system uses a pre-set standard curve equation for this type of compound to convert the integrated signal response value into the actual physical concentration. The formula for calculating the relative content is as follows: In the formula, : Indicates the calculated relative content of the targeting effect factor in the natural raw material, usually expressed in milligrams per kilogram; This represents the integrated area value of the target characteristic ion peak in the extracted ion chromatogram. The slope coefficient of the standard curve equation reflects the sensitivity of the detection instrument to the compound. : Represents the intercept coefficient of the standard curve equation, reflecting the level of background noise signal; This represents the dilution factor during sample pretreatment. After obtaining the quantitative results, the system performs threshold filtering to remove raw materials that, although containing effective components, are present in extremely low concentrations and cannot meet the nutritional regulation needs of animals. Finally, the system extracts and retains the metadata (name, origin) of the raw materials and the calculated content data, and generates a recommended list by sorting them in descending order of content, providing a direct material basis for feed formulation design.

[0040] In summary, this application has at least the following effects:

[0041] A targeted nutrient source mining method based on the interaction between the intestinal type and host genes in beef cattle is proposed. By constructing a microecology-host co-expression coupling network, the microbial community structure and the host gene regulatory network are deeply mapped, effectively overcoming the shortcomings of traditional single-dimensional statistical analysis that ignores individual heterogeneity and biological causal mechanisms, and achieving precise target identification. Furthermore, molecular conformation search and Gibbs free energy calculation are introduced to verify the effectiveness of the interaction pathway between microbial metabolites and host receptors at the atomic level, significantly reducing the false positive rate caused by relying solely on statistical association. Finally, the method combines liquid chromatography-mass spectrometry (LC-MS) feature library for reverse matching and quantitative screening of natural raw materials, opening up a transformation path from microscopic molecular mechanism analysis to macroscopic feed formulation development, greatly shortening the research and development cycle and improving the practical application effect of precision nutrition regulation.

[0042] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0043] This invention is described with reference to flowchart illustrations and / or block diagrams of systems, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0044] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0045] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0046] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0047] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for targeted nutrient source mining based on beef cattle intestinal-host gene interaction, characterized in that, Includes the following steps: S1. Obtain gut microbiome sequencing data and blood transcriptome sequencing data of beef cattle population, construct a microbiome abundance matrix based on gut microbiome sequencing data and calculate the microecological distance between samples, cluster the beef cattle population into different nutrient-associated enterotypes and extract the core indicator bacteria genus of each nutrient-associated enterotype; Using nutrient-associated gut type as a grouping tag, a co-expression network was constructed from blood transcriptome sequencing data to screen for co-expression-specific genes that are highly associated with specific nutrient-associated gut types. S2. Retrieve secondary metabolites that core indicator bacteria genera can synthesize from the microbial metabolic pathway database, and establish a ligand molecule dataset containing the chemical structures of secondary metabolites; obtain the three-dimensional crystal structure data of proteins encoded by co-expressed specific genes, and extract the spatial coordinate parameters of protein active binding sites; S3. Map the chemical structures in the ligand molecule dataset to the spatial coordinate parameters of the protein active binding site, perform molecular conformation search and Gibbs free energy calculation, and output the binding energy score; sort the ligand molecule dataset by affinity based on the binding energy score, and select metabolite molecules with binding energy scores better than a preset threshold as targeted effector factors for targeted regulation of co-expressed specific genes. S4. Using the target effect factor as a search probe, perform a traversal matching in the feature library containing natural feed ingredient components and liquid chromatography-mass spectrometry data, screen natural raw materials containing the target effect factor and whose content meets the preset concentration, and generate a recommended list of targeted nutrient sources for nutrition-related gut types.

2. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 1, characterized in that: The specific process of obtaining gut microbiome sequencing data and blood transcriptome sequencing data of beef cattle, constructing a microbiome abundance matrix based on the gut microbiome sequencing data, calculating the microecological distance between samples, clustering the beef cattle population into different nutrient-associated gut types, and extracting the core indicator bacteria genera for each nutrient-associated gut type is as follows: Sequence denoising and feature clustering were performed on the gut microbiome sequencing data to generate a feature sequence list containing species classification information. The feature sequence list was accumulated and scaled normalized to construct a microbial abundance matrix reflecting the proportion of species in the sample. Based on the microbial abundance matrix, the probability distribution difference between sample pairs in the beef cattle population is calculated using the Jensen-Shannon divergence algorithm to generate a microecological distance matrix. The microecological distance matrix is ​​then input into a clustering algorithm around the center point. The optimal number of clusters is determined by calculating the silhouette coefficient, and the beef cattle population is divided into nutritionally associated intestinal types. The sample data within the nutrient-associated enterotype are traversed, and the frequency and average relative abundance of bacterial genera in the nutrient-associated enterotype are calculated. Genera whose frequency and average relative abundance are both higher than a preset threshold are identified as core indicator genera.

3. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 1, characterized in that: The specific process of constructing a co-expression network from blood transcriptome sequencing data using nutrient-associated gut type as a grouping label, and screening for co-expression-specific genes highly associated with specific nutrient-associated gut types, is as follows: Blood transcriptome sequencing data were standardized, and the correlation coefficient matrix of gene expression was calculated. A soft threshold power parameter was introduced to transform the correlation coefficient matrix into a weighted adjacency matrix, and the topological overlap matrix reflecting the tightness of gene connections was further calculated. Gene co-expression modules are identified using a hierarchical clustering algorithm based on the topological overlap matrix. The first principal component of the gene co-expression module is calculated as the module feature gene. The correlation between the module feature gene and the nutritional association gut type, which is a binary feature vector, is calculated. Specific modules whose correlation meets the preset significance criteria are screened out. Within the specificity module, the significance of the association between genes and nutrient-associated gut types is calculated, and genes with association significance exceeding the screening threshold are identified as co-expressed specific genes.

4. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 1, characterized in that: The specific process of retrieving secondary metabolites that core indicator bacteria genera can synthesize from the microbial metabolic pathway database and establishing a ligand molecule dataset containing the chemical structures of these secondary metabolites is as follows: The core indicator bacteria genera are mapped to the microbial genome database, functional genomic information is extracted, and the enzyme reaction network is reconstructed using the metabolic pathway database. The substrate and product transformation pathways are tracked in the enzyme reaction network, secondary metabolites located at the end of the metabolic pathway are identified, primary metabolites of basic life activities are removed, and a list of potential metabolites is generated. The chemical structure coding files of secondary metabolites in the potential metabolite list are obtained by searching the chemical small molecule structure database. The chemical structure coding files are filtered according to the upper limit of molecular weight and drug-likeness rules to establish a ligand molecule dataset.

5. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 1, characterized in that: The specific process for obtaining the three-dimensional crystal structure data of the protein encoded by the co-expression-specific gene and extracting the spatial coordinate parameters of the protein's active binding site is as follows: Search biomacromolecule structure databases to find protein crystal structure files that match the coding sequences of co-expressed specific genes. Perform dehydration and charge balance preprocessing on the protein crystal structure files. Use a geometric probe algorithm to scan the protein surface and identify concave regions with specific volume characteristics as potential active binding sites. Calculate the geometric center coordinates of the potential active binding site, set the three-dimensional spatial constraint parameters according to the spatial extension range of the potential active binding site, and output the geometric center coordinates and the three-dimensional spatial constraint parameters as the spatial coordinate parameters of the protein active binding site.

6. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 1, characterized in that: The specific process of mapping the chemical structures in the ligand molecule dataset to the spatial coordinate parameters of the protein active binding site, performing molecular conformation search and Gibbs free energy calculation, and outputting the binding energy score is as follows: A three-dimensional affinity potential energy lattice matrix is ​​constructed based on spatial coordinate parameters, and rotatable chemical bonds of the centralized chemical structure of ligand molecules are set to maintain molecular flexibility. Using a global optimization algorithm, a semi-flexible conformation search is performed on the chemical structures in the ligand molecule dataset within the three-dimensional affinity potential energy lattice matrix. The van der Waals forces, electrostatic interactions, and hydrogen bond potential energy between the ligand and the receptor are calculated iteratively, and the minimum energy value after iterative convergence is output as the binding energy score.

7. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 6, characterized in that: The specific process of ranking ligand molecule datasets by affinity based on binding energy scores and selecting metabolite molecules with binding energy scores better than a preset threshold as target effectors for regulating co-expressed specific genes is as follows: Root mean square bias clustering analysis was performed on the ligand conformations generated by conformation search to extract representative binding patterns of dominant conformation clusters; The ligand molecule dataset is sorted by affinity from low to high according to the binding energy score, and candidate molecules with high ranking and binding energy scores below a preset energy threshold are selected. The contact between candidate molecules and protein active binding site residues is examined, molecules with steric hindrance conflicts are eliminated, and the remaining molecules are identified as targeting effectors.

8. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 1, characterized in that: The specific process of using targeted effect factors as retrieval probes to perform traversal matching in a feature library containing natural feed ingredient components and liquid chromatography-mass spectrometry data is as follows: The chemical structure of the targeting effect factor was analyzed, and the theoretical mass-to-charge ratio and characteristic ion peak information of the targeting effect factor were predicted under the detection environment of liquid chromatography-mass spectrometry. Traverse the feature library containing natural feed ingredient components and liquid chromatography-mass spectrometry data, and retrieve the secondary mass spectrometry information of known components of each natural ingredient in the feature library; The cosine similarity between characteristic ion peak information and secondary mass spectrometry data is calculated, and natural raw materials with a cosine similarity higher than the preset matching standard are marked as preliminary hit raw materials.

9. The method for targeted nutrient source mining based on the interaction between the intestinal type and the host gene in beef cattle according to claim 8, characterized in that: The specific process of selecting natural raw materials containing targeted effect factors at preset concentrations and generating a recommended list of targeted nutrient sources for nutrition-related gut types is as follows: Extract the ion current peak area of ​​the corresponding target effect factor in the raw material at the chromatographic retention time, and use the preset standard curve equation to perform integral calculation on the ion current peak area to calculate the relative content of the target effect factor in the natural raw material. The relative content is compared with the preset concentration. Natural raw materials with a relative content higher than the preset concentration are retained. The name, origin and content of the targeted effect factor of the retained natural raw materials are obtained. A recommended list of targeted nutrient sources is generated in descending order of content.