Feed intelligent formula optimization method based on effectiveness of background trace elements
By using machine learning models and multi-objective optimization algorithms, the problem of predicting the release rate of trace elements across raw materials and enzymes has been solved, enabling precise feed formulation optimization, reducing the amount of exogenous additives, reducing heavy metal emissions, and promoting the green and sustainable development of animal husbandry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-06-25
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies cannot predict the release efficiency of background trace elements in feed across raw materials and enzymes, making it impossible to accurately determine the amount of exogenous addition, which affects breeding costs and environmental pollution.
By acquiring baseline data and enzyme characteristics of feed ingredients, a machine learning model is trained to predict the release rate of trace elements. A multi-objective optimization algorithm is then used to determine the optimal ratio of raw materials and the amount of enzyme added, thereby achieving precision nutrition and green farming.
It enables the prediction of the release rate of unknown raw materials and enzyme combinations, reduces the amount of exogenous additives, reduces heavy metal emissions, and meets the requirements of green aquaculture.
Smart Images

Figure CN122494086A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of feed processing technology, and more specifically to a method for intelligent feed formulation optimization based on the availability of background trace elements. Background Technology
[0002] In livestock farming, the addition and utilization of trace elements are key factors affecting production performance. Essential elements such as copper, iron, zinc, and manganese have been widely proven to play indispensable roles in animal production. For example, high-zinc feeding in early piglets can effectively reduce diarrhea rates, high-manganese feeding can improve the reproductive performance of pregnant sows, and high-copper feeding can significantly improve the growth performance of early piglets. However, current livestock farming advocates green and sustainable development, and the widespread addition of high concentrations of trace elements leads to excessive emissions from livestock farming, causing difficulties in downstream manure treatment and environmental pollution. Although conventional feed ingredients (such as corn, soybean meal, and sorghum) contain considerable background trace elements, anti-nutritional factors such as phytic acid in plants have a strong chelating effect on metal ions, resulting in almost zero absorption of these trace elements by animals, further exacerbating the emission of trace elements in manure. Researching how to improve the release efficiency of background trace elements in feed and reduce the amount of exogenous addition through intelligent methods is not only related to the control of farming costs but also directly affects the environmental protection level and sustainable development capacity of livestock farming. The importance of this field lies in its ability to reduce heavy metal emissions at the source through technological means, achieving synergy between precision nutrition and green farming.
[0003] However, existing methods often face the challenge of generalizing predictions across different raw materials and enzymes when assessing and optimizing the release efficiency of background trace elements. Many approaches rely on in vitro digestion experiments of specific raw material-enzyme combinations, neglecting the differences in phytic acid structure among different raw materials and the diversity of release characteristics among different enzyme preparations. This necessitates repeating the entire experiment every time a batch of raw materials or a new enzyme is used, resulting in high costs and hindering large-scale application. Furthermore, current technologies are completely incapable of handling commercially available enzymes whose composition information is not publicly available, creating a long-standing technological blind spot. While some approaches attempt to utilize enzyme preparations to promote release, they are limited to qualitative conclusions based on a single enzyme or raw material, lacking the ability to model the complex nonlinear relationship between "raw material characteristics—enzyme characteristics—release rate," and failing to consider the expected release rate as a multi-objective optimization variable in conjunction with raw material ratios, enzyme types, and dosages. This prevents the system from accurately predicting the release efficiency of background trace elements under new raw material combinations or new enzyme treatments, thus affecting the precise decision-making for minimizing exogenous additions and the effective achievement of emission reduction effects. In the context of intelligent feed formulation optimization, the core technical challenge lies in how to accurately quantify and predict the trace element release rate under the combination of multiple raw materials and multiple enzymes (including enzymes with unknown components) without relying on repeated experiments, and dynamically integrate it into the formulation optimization model. Summary of the Invention
[0004] The purpose of this invention is to provide a method for intelligent feed formulation optimization based on the availability of background trace elements, in order to address the shortcomings of the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligent feed formulation optimization based on the availability of background trace elements, comprising: S1. Obtain baseline data for various feed ingredients, including the content of at least one trace element in each feed ingredient and the physicochemical characteristic parameters of each feed ingredient. S2. For each of the multiple candidate enzymes, in an in vitro digestion system simulating the animal digestive tract environment, the release rate of at least one trace element of each feed ingredient after treatment with the candidate enzyme is measured, and a feed ingredient-enzyme-release rate dataset is constructed. S3. Extract the physicochemical characteristic parameters and the enzyme feature vectors of the candidate enzymes as multidimensional input variables, and use the release rate as the output variable to train a machine learning model to obtain a trace element release rate prediction model; wherein, the enzyme feature vectors contain at least the identification information of the candidate enzymes or the functional category parameters of the enzymes. S4. Input the target raw material combination to be optimized and its corresponding candidate enzyme into the release rate prediction model, and output the expected trace element release rate of the target raw material combination under the treatment of the candidate enzyme. S5. Based on the expected trace element release rate and the preset animal nutritional requirements standard, the optimal raw material ratio, enzyme type and enzyme addition amount are determined by a multi-objective optimization algorithm to minimize the amount of exogenous trace elements added.
[0006] In a preferred embodiment, the feed ingredients are selected from at least two of corn, wheat, barley, sorghum, rice, wheat bran, soybean meal, cottonseed meal, rapeseed meal, peanut meal, and DDGS; the trace elements are selected from at least two of iron, copper, zinc, and manganese.
[0007] In a preferred embodiment, the physicochemical characteristic parameters in S1 include at least one of the following: phytic acid content, phytic acid phosphorus content, total phosphorus content, crude fiber content, crude protein content, crude ash content, and near-infrared spectral characteristics.
[0008] In a preferred embodiment, the simulated animal digestive tract environment in S2 includes two continuous simulated digestive systems: a simulated gastric juice environment, a simulated small intestinal juice environment, or both; wherein the pH of the simulated gastric juice environment is 2.0–4.0, and the pH of the simulated small intestinal juice environment is 6.0–7.5.
[0009] In a preferred embodiment, the candidate enzyme includes at least one of phytase, xylanase, β-glucanase, cellulase, and a complex enzyme preparation; The candidate enzymes include enzymes with known component information and commercial enzymes with undisclosed component information; For commercial enzymes whose component information is not publicly available, the enzyme feature vector uses the release rate data of the commercial enzyme on a variety of predetermined raw materials in S2 as its functional characterization. The functional characterization reflects the distribution of the release effect of the commercial enzyme on the at least one trace element on the variety of predetermined raw materials.
[0010] In a preferred embodiment, the release rate prediction model can predict the corresponding trace element release rate for new raw material combinations and / or new candidate enzyme combinations that have not been measured in the S2 experiment, without having to re-perform the entire S2 measurement process for the new raw material combination and / or new candidate enzyme combination.
[0011] In a preferred embodiment, the multi-objective optimization algorithm in S5 includes: The expected trace element release rate output by S4 is used as the input variable of the first objective function. The first objective is to maximize the expected trace element release rate, and the second objective is to minimize the amount of exogenous trace element added. The NRC nutritional requirements of animals in the growth stage are used as constraints. The solution is obtained by using genetic algorithm, particle swarm optimization or Bayesian optimization. The decision variables of the multi-objective optimization algorithm include: the mass percentage of each raw material, the type of enzyme selected, and the amount of enzyme added.
[0012] In a preferred embodiment, the method for constructing the enzyme feature vector includes: The release rate data of each candidate enzyme measured against multiple predetermined raw materials in S2 are used to construct a high-dimensional functional fingerprint. The high-dimensional functional fingerprint is then mapped into a low-dimensional embedding vector using a contrastive learning or metric learning neural network, which serves as the feature input of the candidate enzyme in the prediction model. In the comparative learning, the positive sample pairs are the release rate data of the same enzyme in different batches or under different experimental conditions, and the negative sample pairs are the release rate data of different enzymes.
[0013] In a preferred embodiment, in step S2, the release rate of multiple trace elements under coexisting conditions is also measured to construct an extended dataset containing element interaction effects. The machine learning model in S3 is a graph neural network or a multi-task learning model, whose hidden layer contains edge connections that represent the interdependencies between trace elements.
[0014] In a preferred embodiment, the machine learning model is a time-series prediction model; In step S2, the release rate of trace elements is measured multiple times at different time points during the simulated digestion process to construct a release kinetic curve for each raw material-enzyme combination. The release kinetic curve contains release rate data at at least three time points. The output of the time-series prediction model is the predicted release kinetic curve.
[0015] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention extracts the physicochemical characteristic parameters of feed ingredients and the enzyme feature vectors of candidate enzymes to train a machine learning model, resulting in a microelement release rate prediction model. This model can learn mapping patterns from known feed ingredient-enzyme combination release rate data and directly output the expected release rate for new feed ingredient combinations and / or new candidate enzyme combinations that have not been tested in vitro, eliminating the need for repeated, time-consuming simulated digestion measurements. This represents a paradigm shift from one-time training to unlimited prediction, significantly reducing the technical application threshold and formulation optimization costs for feed companies, and providing a feasible solution for large-scale promotion.
[0016] This invention addresses the widespread secrecy surrounding the composition of commercial compound enzyme preparations in the feed industry by proposing a method for constructing enzyme feature vectors based on functional characterization. For commercial enzymes with undisclosed composition information, the release rate data measured on various predetermined raw materials is directly used as a functional fingerprint and incorporated into the input features of the prediction model. This method is completely independent of the enzyme's chemical composition information, enabling the intelligent optimization framework to handle a large number of commercially available enzymes with undisclosed composition in actual production. It overcomes a long-standing industrial bottleneck that existing technologies have been unable to overcome, demonstrating strong industrial adaptability and practical application value.
[0017] This invention directly uses the expected trace element release rate output by the prediction model as the first objective function of a multi-objective optimization algorithm, with maximizing the utilization rate of background trace elements as the core driving force. Simultaneously, minimizing the amount of exogenous trace elements added is the second objective, constrained by animal nutritional requirements. Genetic algorithms, particle swarm optimization, or Bayesian optimization are used to jointly solve for the raw material ratio, enzyme type, and enzyme addition amount. This invention can automatically explore the complex feasible region composed of multiple decision variables and output the optimal formula combination. Therefore, while meeting the animal's growth needs, it maximizes the potential of background trace elements, significantly reduces the amount of exogenous premix added, and reduces heavy metal emissions in feces from the source, achieving green farming. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0019] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1, please refer to Figure 1 As shown in this embodiment, the intelligent feed formulation optimization method based on the availability of background trace elements includes: S1. Obtain baseline data for various feed ingredients, including the content of at least one trace element in each feed ingredient and the physicochemical characteristic parameters of each feed ingredient. S2. For each of the multiple candidate enzymes, in an in vitro digestion system simulating the animal digestive tract environment, the release rate of at least one trace element of each feed ingredient after treatment with the candidate enzyme is measured, and a feed ingredient-enzyme-release rate dataset is constructed. S3. Extract the physicochemical characteristic parameters and the enzyme feature vectors of the candidate enzymes as multidimensional input variables, and use the release rate as the output variable to train a machine learning model to obtain a trace element release rate prediction model; wherein, the enzyme feature vectors contain at least the identification information of the candidate enzymes or the functional category parameters of the enzymes. S4. Input the target raw material combination to be optimized and its corresponding candidate enzyme into the release rate prediction model, and output the expected trace element release rate of the target raw material combination under the treatment of the candidate enzyme. S5. Based on the expected trace element release rate and the preset animal nutritional requirements standard, the optimal raw material ratio, enzyme type and enzyme addition amount are determined by a multi-objective optimization algorithm to minimize the amount of exogenous trace elements added.
[0022] This invention deeply integrates machine learning, multi-objective optimization, and animal nutrition to construct a complete intelligent closed-loop system from "baseline data collection - release rate prediction - automatic formula generation". It effectively solves the core technical problems of existing technologies, such as the inability to predict across raw materials and enzymes, the inability to handle commercial enzymes with unknown components, and the inability to achieve release rate-driven optimization. It has important technological progress significance and industrial application value for promoting the green and sustainable development of animal husbandry.
[0023] In one embodiment, step S1 may include: acquiring background data for multiple feed ingredients. The background data may include the content of at least one trace element in each feed ingredient, for example, the background concentration of elements such as iron, copper, zinc, and manganese in the ingredient may be determined using inductively coupled plasma mass spectrometry (ICP-MS) or atomic absorption spectrometry. Furthermore, the background data may also include physicochemical characteristic parameters of each feed ingredient, such as phytic acid content, crude fiber content, crude protein content, or near-infrared spectral characteristics.
[0024] By obtaining the aforementioned baseline data, we can accurately determine the original reserves of trace elements in the raw materials and their binding status with anti-nutritional factors such as phytic acid and fiber, thereby providing basic data support for subsequent evaluation of the release potential after enzyme treatment.
[0025] This step links the physicochemical characteristics of raw materials with the content of trace elements, enabling subsequent prediction models to learn the mapping law of "raw material intrinsic properties → release efficiency," thus avoiding the shortcomings of traditional methods that rely solely on database averages and ignore batch differences.
[0026] In one embodiment, step S2 may include: for each of a plurality of candidate enzymes, in an in vitro digestion system simulating an animal digestive tract environment, determining the release rate of at least one trace element of each feed ingredient after treatment with that candidate enzyme, thereby constructing a feed ingredient-enzyme-release rate dataset. Specifically, the simulated animal digestive tract environment may simulate two consecutive digestion processes of gastric juice (e.g., pH 2.0–4.0, containing pepsin) and / or small intestinal juice (e.g., pH 6.0–7.5, containing pancreatic enzymes). For cross-combinations of each feed ingredient and each candidate enzyme, the supernatant can be collected by centrifugation after digestion, and the concentrations of elements such as iron, copper, zinc, and manganese can be determined. Calculate the release rate: ; All content detection data were obtained by using the average value of three parallel repeated measurements to avoid random experimental errors, and were uniformly included in the raw material-enzyme-release rate standard dataset for archiving and modeling. This step yields a series of discrete experimental data points on "raw material characteristics - enzyme characteristics - release rate". It quantifies the actual release effect of commonly used industrial enzymes under in vitro physiological conditions, providing real training samples for subsequent machine learning models and solving the problem of existing technologies that rely solely on experience to infer enzyme effects without sufficient data support.
[0027] In one embodiment, step S3 may include: extracting the physicochemical characteristic parameters and the enzyme feature vector of the candidate enzyme as multidimensional input variables, using the release rate as the output variable, training a machine learning model, and obtaining a trace element release rate prediction model. The enzyme feature vector may at least contain the identification information of the candidate enzyme (e.g., the trade name and batch number of the enzyme) or the functional category parameter of the enzyme (e.g., phytase, xylanase, or complex enzyme type).
[0028] In one specific implementation, the machine learning model can be a random forest model. This random forest model can be configured with the following parameters: the number of decision trees is 100–500, the maximum depth of each tree is 10–20 layers, the minimum number of samples for node splitting is 5–10, and the splitting criterion uses mean squared error (MSE) or mean absolute error (MAE). Input features include the phytic acid content, crude fiber content, crude protein content of each raw material, and the identification information of candidate enzymes. The output is the predicted release rate of at least one trace element among iron, copper, zinc, and manganese.
[0029] During training, 80% of the dataset obtained in steps S1 and S2 can be used as the training set, and 20% as the validation set. Out-of-bag error is used to evaluate model performance. The advantage of this random forest model lies in its ability to handle high-dimensional features and nonlinear relationships, and it can provide feature importance ranking, which helps identify which physicochemical parameters have the greatest impact on trace element release.
[0030] In another specific implementation, the machine learning model can employ a feedforward neural network.
[0031] This neural network can include: an input layer with the number of nodes equal to the dimension of the input features (e.g., 20-50 dimensions); 2-3 fully connected hidden layers, each with 64-256 neurons, using ReLU as the activation function; and an output layer with the number of nodes equal to the number of micro-elements to be predicted (e.g., 2-4). The output layer has no activation function or uses a linear activation function to output continuous values. During training, mean squared error can be used as the loss function, and the Adam optimizer can be used for parameter updates. The batch size can be set to 16-64, and the number of training epochs can be set to 100-500. Early stopping is used to prevent overfitting. The advantage of this neural network is its ability to learn highly nonlinear and complex mapping relationships. When sufficient training data is available, its prediction accuracy is usually better than that of traditional regression models.
[0032] By using the datasets obtained in steps S1 and S2 as the training set, any of the above machine learning models can automatically learn the nonlinear mapping relationship from the physicochemical properties and enzyme properties of raw materials to the release rate of trace elements.
[0033] This step constructs a predictive tool that can extrapolate the release patterns of known raw material-enzyme combinations to unknown combinations. Once the model is trained, there is no need to repeat time-consuming in vitro digestion experiments for each new raw material or enzyme, which significantly reduces the cost and time of formulation optimization.
[0034] In one embodiment, step S4 may include: inputting the target raw material combination to be optimized and its corresponding candidate enzyme into the release rate prediction model, and outputting the expected trace element release rate of the target raw material combination under the treatment of the candidate enzyme. For example, when a company needs to formulate a batch of pig feed containing raw materials such as corn, soybean meal, and wheat bran, it can input the physicochemical parameters of this batch of raw materials (such as phytic acid content and near-infrared spectrum) and the information of the commercial enzyme to be selected into the model, and the model can quickly output the predicted release rates of iron, copper, zinc, and manganese under simulated digestion conditions.
[0035] This step enables "instant prediction," allowing formulators to obtain data within minutes that would otherwise require weeks of in vitro experiments, providing real-time input for multi-objective optimization.
[0036] In one embodiment, step S5 may include: determining the optimal raw material ratio, enzyme type, and enzyme addition amount based on the expected trace element release rate and a preset animal nutritional requirement standard (e.g., the minimum trace element requirement in growing pig diets recommended by the NRC), using a multi-objective optimization algorithm to minimize the amount of exogenous trace elements added. For example, the first objective may be to maximize the expected trace element release rate (i.e., to utilize as much background element as possible), and the second objective may be to minimize the amount of exogenous trace element premix added. A genetic algorithm or particle swarm optimization algorithm may then be used to search for a Pareto optimal solution within the feasible region of the raw material ratio, enzyme type, and enzyme addition amount.
[0037] This step directly links the release rate output by the preceding prediction model to feed formulation decisions, realizing a closed loop from "data prediction" to "formulation generation". This minimizes the addition of exogenous trace elements and reduces heavy metal emissions in feces while meeting the nutritional needs of animals, thus meeting the requirements of green farming.
[0038] In one embodiment, the feed ingredients may be selected from at least two of corn, wheat, barley, sorghum, rice, wheat bran, soybean meal, cottonseed meal, rapeseed meal, peanut meal, and DDGS; the trace elements may be selected from at least two of iron, copper, zinc, and manganese.
[0039] The aforementioned raw materials cover the most commonly used energy and protein feeds in my country's livestock and poultry farming. Iron, copper, zinc, and manganese are the four trace elements essential for animal growth, reproduction, and immunity, and also the ones facing the greatest emission pressure. By including these specific types within the scope of protection, the applicability of this method in mainstream feed industry scenarios can be ensured, while avoiding loss of novelty due to over-generalization.
[0040] In one embodiment, the physicochemical characteristic parameters in step S1 may include at least one of the following: phytic acid content, phytic acid phosphorus content, total phosphorus content, crude fiber content, crude protein content, crude ash content, and near-infrared spectral characteristics. Phytic acid is a major anti-nutritional factor that chelates metal ions, while crude fiber may affect the contact efficiency between enzymes and substrates. Near-infrared spectroscopy can quickly and non-destructively reflect the overall chemical composition of raw materials. By introducing these characteristics, machine learning models can more accurately capture the differences in microstructure between different raw materials, thereby improving the prediction accuracy of release rates for raw materials from different origins and batches.
[0041] In one embodiment, the simulated animal digestive tract environment in step S2 may include a two-stage continuous simulated digestion system comprising a simulated gastric juice environment, a simulated small intestinal juice environment, or both; wherein the pH of the simulated gastric juice environment may be 2.0–4.0, and the pH of the simulated small intestinal juice environment may be 6.0–7.5. Using a two-stage continuous digestion system can more realistically simulate the physiological process of feed being digested by gastric acid and pepsin in the animal's gastrointestinal tract before entering the small intestine and being digested by pancreatic enzymes. Compared to single-stage digestion, the two-stage method can more accurately reflect the activity changes of enzymes such as phytase under different pH conditions and their temporal contribution to the release of trace elements, thereby improving the physiological relevance of release rate measurement data.
[0042] In one embodiment, the candidate enzyme may include at least one of phytase, xylanase, β-glucanase, cellulase, and a complex enzyme preparation; the candidate enzyme may include enzymes with known component information and commercially available enzymes with undisclosed component information. For commercially available enzymes with undisclosed component information, the enzyme feature vector may use the release rate data of the commercially available enzyme measured on multiple predetermined raw materials in step S2 as its functional characterization, and the functional characterization reflects the distribution of the release effect of the commercially available enzyme on the multiple predetermined raw materials for the at least one trace element.
[0043] For example, 4 to 6 representative raw materials such as corn, soybean meal, wheat, and DDGS can be selected to determine the release rates of Fe, Cu, Zn, and Mn after treatment with a commercial enzyme (components confidential). These release rate data can be combined into a multi-dimensional vector (e.g., 4 raw materials × 4 elements = 16 dimensions) and directly used as the input features of the enzyme in the prediction model. This approach breaks through the dependence of traditional methods on enzyme component information and solves the common pain point in industry that "enzyme formulation confidentiality leads to the inability to optimize," enabling even commercial compound enzymes with unknown components to be included in the intelligent optimization framework of this invention.
[0044] In one embodiment, the release rate prediction model can predict the corresponding trace element release rate for new raw material combinations and / or new candidate enzyme combinations that have not been experimentally determined in step S2, without having to re-perform the entire determination process of step S2 for those new raw material combinations and / or new candidate enzyme combinations. In other words, once the model is trained based on the release rate data of the first batch of representative raw materials and representative enzymes, the model possesses knowledge transfer capabilities.
[0045] For example, when a company introduces corn from a new origin or a new commercial enzyme, it only needs to provide the physicochemical characteristics of the corn (such as near-infrared spectroscopy and phytic acid content) or a small amount of rapid characterization data of the new enzyme (such as release rates against 2-3 standard raw materials). The model can then output the expected release rates of the enzyme in combination with various existing enzymes, without the need to conduct a complete set of in vitro digestion experiments again. This feature is the core of realizing the paradigm leap of "one-time modeling, unlimited prediction" in this invention, greatly reducing the technical application threshold for feed companies.
[0046] In one embodiment, the multi-objective optimization algorithm in step S5 may include: taking the expected trace element release rate output in step S4 as the input variable of the first objective function, maximizing the expected trace element release rate as the first objective, minimizing the amount of exogenous trace element added as the second objective, and using the NRC nutritional requirements of the animal growth stage as the constraint condition, and solving the problem using a genetic algorithm, particle swarm optimization or Bayesian optimization. The decision variables of the multi-objective optimization algorithm may include: the mass percentage of each raw material, the selected enzyme type, and the enzyme dosage. Unlike traditional formulation optimization that only considers cost or nutrient content as objectives, this invention, for the first time, uses "expected release rate" as the core of the optimization driver, and places the enzyme type and dosage with the raw material ratio in the same decision space for joint search. This joint optimization can simultaneously balance feed cost, emission reduction effect, and animal production performance. For example, it can automatically identify cross-variable synergistic solutions such as "when the corn proportion increases by 5%, adding 200 FTU / kg of a certain phytase can achieve the optimal background zinc release rate," which is difficult to achieve through manual trial and error or single-variable adjustment.
[0047] In one embodiment, the method for constructing the enzyme feature vector may include: constructing a high-dimensional functional fingerprint from the release rate data of each candidate enzyme measured against multiple predetermined raw materials in step S2, and mapping the high-dimensional functional fingerprint into a low-dimensional embedding vector using a contrastive learning or metric learning neural network, which serves as the feature input of the candidate enzyme in the prediction model.
[0048] In one specific implementation, the contrastive learning neural network can be trained using triplet loss. For a central anchor enzyme, positive samples (release rate data from different batches of the same enzyme or under different experimental conditions) and negative samples (release rate data from different enzymes), the triplet loss function can be: L=max(d(anchor,positive)-d(anchor,negative)+margin,0); In the formula, L is the calculated value of the triplet loss function; d is the Euclidean spatial distance metric operator; anchor is the feature vector of release rate of the baseline anchor sample of the same candidate enzyme; positive is the feature vector of release rate of positive samples of the same enzyme in different experimental batches and under different environments; negative is the feature vector of release rate of negative samples of candidate enzymes of different categories and activity specifications; margin is the preset interval constraint hyperparameter, and the conventional compliant value range is set to 0.5 to 1.0.
[0049] Mapping neural networks can employ 2-3 layers of fully connected networks to compress high-dimensional functional fingerprints (e.g., determining M trace elements from N predetermined raw materials, with a dimension of N×M) into a low-dimensional embedding space (e.g., 8-dimensional, 16-dimensional, or 32-dimensional). After training, the low-dimensional embedding vectors replace the original functional fingerprints as input features for the prediction model, reducing feature redundancy and improving model robustness. Through contrastive learning, the neural network can automatically learn a discriminant representation that "release rate vectors of the same enzyme in different batches should be similar, while release rate vectors of different enzymes should be far apart." This approach further enhances the robustness and generalization ability of the prediction model, making it particularly suitable for industrial scenarios involving a large number of commercially available enzymes with unknown components.
[0050] In one embodiment, in step S2, the release rate under multiple trace element coexistence conditions can also be measured to construct an extended dataset containing element interaction effects; the machine learning model in step S3 can be a graph neural network or a multi-task learning model, and its hidden layer can contain edge connections representing the interdependence between trace elements.
[0051] In one specific implementation, the graph neural network can be constructed as a graph convolutional network (GCN). Using trace elements as nodes (Fe, Cu, Zn, Mn, a total of 4 nodes), an adjacency matrix is constructed using pre-determined inter-element antagonism / cooperation coefficients (edge weights can be represented as interaction strength values, ranging from -1 to 1, with negative values indicating antagonism and positive values indicating cooperation).
[0052] The propagation layer formula for GCN can be: ; in, This is the feature matrix of the output node of the (l+1)th layer of the graph convolutional network; Since it is a non-linear activation function, the ReLU activation function can be selected in accordance with regulations. To supplement the adjacency and correlation matrix of trace elements after self-closed-loop connection, used to quantify the strength of element antagonism and cooperative interaction; This is the matching degree matrix of the corresponding adjacency matrix; This is the feature matrix of the original input nodes in the l-th layer of the network; The weight parameter matrix of the l-th layer is a neural network that can be iteratively optimized autonomously; the features of the input layer nodes can be initialized as intermediate predictions of the release rates of each trace element under set raw material-enzyme conditions. After propagation through 1 to 3 layers of GCN, the final output is a corrected release rate considering interaction effects.
[0053] For example, by simultaneously adding four elements—Zn, Cu, Fe, and Mn—to an in vitro digestion system, the release rate of each element in the coexisting state can be measured and compared with the release rate when a single element is present, quantifying antagonistic (e.g., high Zn inhibits Cu absorption) or synergistic effects. This extension can correct the unreasonable assumption of "independent release of each element" in traditional methods, making the release rate output by the model closer to the real situation of competitive absorption of multiple elements in the animal intestine, thereby improving the actual effectiveness of the formulation.
[0054] In one embodiment, the machine learning model can be a time-series prediction model; in step S2, the release rate of trace elements can be measured multiple times at different time points in the simulated digestion process to construct a release kinetic curve for each raw material-enzyme combination, and the release kinetic curve can contain release rate data at at least three time points; the output of the time-series prediction model can be the predicted release kinetic curve.
[0055] In one specific implementation, the time-series prediction model can be implemented using a Long Short-Term Memory (LSTM) network. Specifically, the release kinetics curve data for each feedstock-enzyme combination can be represented as a time-series vector of length T (e.g., T=6, corresponding to 30 min, 60 min, and 90 min in the gastric digestion stage, and 30 min, 60 min, and 120 min in the small intestine digestion stage, for a total of 6 time points). The LSTM network can be configured with 2 LSTM hidden layers, each with 64 LSTM units, followed by a Dropout layer (the dropout rate can be set to 0.2–0.5). Above the LSTM layers, a fully connected output layer can be configured, with its output dimension equal to the number of trace element species to be predicted (e.g., 4 output nodes for Fe, Cu, Zn, and Mn).
[0056] During training, a mean squared error loss function can be used, employing the Adam optimizer with a learning rate of 0.001. The time-series data input to the LSTM network can be release rate data measured at each time point, or it can further include digestive environmental parameters at each time point (such as pH, residual enzyme activity, etc.). Compared to simply predicting the endpoint release rate, kinetic curves provide richer information (such as release rate and time to reach maximum release). This information is highly correlated with the actual absorption kinetics of trace elements in animals, thus further guiding feeding strategies (such as staged feeding or adjustments to pelleting processes) to achieve deeper levels of precision nutrition.
[0057] In one embodiment, this method may further incorporate gut microbiome data as an optional optimization factor.
[0058] Specifically, intestinal contents samples from animals can be collected, and metagenomic sequencing and Raman spectroscopy can be used to identify the microbial community and its functional genes related to trace element absorption. Based on this data, the competition coefficient (reflecting the microbial fixation or consumption of trace elements) and synergy coefficient (reflecting the microbial ability to promote host absorption) for each trace element can be calculated. In the multi-objective optimization of step S5, the competition coefficient and synergy coefficient can be used as correction factors to dynamically adjust the effective utilization rate of the expected trace element release rate. For example: Effective utilization rate = Expected release rate × (1 - Competition coefficient + Synergy coefficient).
[0059] Among them, the expected release rate is the baseline effective release value of trace elements output by the machine learning model; the competition coefficient is the quantitative interference coefficient of the intestinal microbial community on the adsorption, fixation and competitive consumption of trace elements; and the synergy coefficient is the quantitative gain coefficient of the beneficial intestinal flora in assisting the intestinal mucosa of livestock and poultry hosts to adsorb and absorb trace elements. By incorporating microbial coefficients, the optimization algorithm can more accurately predict the net increase in trace elements that animals can actually utilize under real-world feeding conditions, thereby further reducing the safety redundancy of exogenous additions and achieving more precise emission reduction. This extended scheme fully considers the dual impact of gut microbiota on trace element utilization and is a beneficial supplement to the core prediction model, especially suitable for environmentally sensitive farming scenarios requiring extreme emission reduction.
[0060] In another embodiment, the method may further include adaptive adjustments to the feed processing process. After step S2, the activity residual rate of each candidate enzyme can be measured under simulated feed processing conditions (e.g., pelleting temperature 80–120°C, conditioning moisture content 12%–18%, and processing time 30–90 seconds) to establish a thermostability model for the enzyme. Then, in the multi-objective optimization of step S5, the activity residual rate is used as a multiplicative factor to calculate the expected release rate, i.e., actual usable release rate = expected release rate × activity residual rate. Here, the activity residual rate is a dimensionless percentage quantification parameter representing the percentage of remaining effective activity of the candidate enzyme after processing at 80–120°C pelleting temperature, compliant conditioning moisture content, and standard conditioning time. Furthermore, decision variables can include processing parameters (such as granulation temperature and conditioning time). This allows the optimization algorithm to automatically recommend enzymes with better thermal stability or suggest adjustments to the processing technology while ensuring release efficiency, avoiding prediction bias caused by enzyme inactivation due to high-temperature processing. This extended approach extends the optimization scope from formulation design to the production and processing stage, further improving the reliability and adaptability of this method in industrial applications.
[0061] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for intelligent feed formulation optimization based on the availability of background trace elements, characterized in that, include: S1. Obtain baseline data for various feed ingredients, including the content of at least one trace element in each feed ingredient and the physicochemical characteristic parameters of each feed ingredient. S2. For each of the multiple candidate enzymes, in an in vitro digestion system simulating the animal digestive tract environment, the release rate of at least one trace element of each feed ingredient after treatment with the candidate enzyme is measured, and a feed ingredient-enzyme-release rate dataset is constructed. S3. Extract the physicochemical characteristic parameters and the enzyme feature vectors of the candidate enzymes as multidimensional input variables, and use the release rate as the output variable to train a machine learning model to obtain a trace element release rate prediction model; wherein, the enzyme feature vectors contain at least the identification information of the candidate enzymes or the functional category parameters of the enzymes. S4. Input the target raw material combination to be optimized and its corresponding candidate enzyme into the release rate prediction model, and output the expected trace element release rate of the target raw material combination under the treatment of the candidate enzyme. S5. Based on the expected trace element release rate and the preset animal nutritional requirements standard, the optimal raw material ratio, enzyme type and enzyme addition amount are determined by a multi-objective optimization algorithm to minimize the amount of exogenous trace elements added.
2. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The feed ingredients are selected from at least two of the following: corn, wheat, barley, sorghum, rice, wheat bran, soybean meal, cottonseed meal, rapeseed meal, peanut meal, and DDGS; the trace elements are selected from at least two of the following: iron, copper, zinc, and manganese.
3. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The physicochemical characteristic parameters in S1 include at least one of the following: phytic acid content, phytic acid phosphorus content, total phosphorus content, crude fiber content, crude protein content, crude ash content, and near-infrared spectral characteristics.
4. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The simulated animal digestive tract environment in S2 includes a simulated gastric juice environment, a simulated small intestinal juice environment, or two continuous simulated digestive systems of both; wherein the pH of the simulated gastric juice environment is 2.0 to 4.0, and the pH of the simulated small intestinal juice environment is 6.0 to 7.
5.
5. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The candidate enzymes include at least one of phytase, xylanase, β-glucanase, cellulase, and complex enzyme preparations; The candidate enzymes include enzymes with known component information and commercial enzymes with undisclosed component information; For commercial enzymes whose component information is not publicly available, the enzyme feature vector uses the release rate data of the commercial enzyme on a variety of predetermined raw materials in S2 as its functional characterization. The functional characterization reflects the distribution of the release effect of the commercial enzyme on the at least one trace element on the variety of predetermined raw materials.
6. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The release rate prediction model can predict the corresponding trace element release rate for new raw material combinations and / or new candidate enzyme combinations that have not been measured in the S2 experiment, without having to re-perform the entire S2 measurement process for the new raw material combination and / or new candidate enzyme combination.
7. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The multi-objective optimization algorithm in S5 includes: The expected trace element release rate output by S4 is used as the input variable of the first objective function. The first objective is to maximize the expected trace element release rate, and the second objective is to minimize the amount of exogenous trace element added. The NRC nutritional requirements of animals in the growth stage are used as constraints. The solution is obtained by using genetic algorithm, particle swarm optimization or Bayesian optimization. The decision variables of the multi-objective optimization algorithm include: the mass percentage of each raw material, the type of enzyme selected, and the amount of enzyme added.
8. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The method for constructing the enzyme feature vector includes: The release rate data of each candidate enzyme measured against multiple predetermined raw materials in S2 are used to construct a high-dimensional functional fingerprint. The high-dimensional functional fingerprint is then mapped into a low-dimensional embedding vector using a contrastive learning or metric learning neural network, which serves as the feature input of the candidate enzyme in the prediction model. In the comparative learning, the positive sample pairs are the release rate data of the same enzyme in different batches or under different experimental conditions, and the negative sample pairs are the release rate data of different enzymes.
9. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: In step S2, the release rate of multiple trace elements under coexisting conditions is also measured to construct an extended dataset containing element interaction effects. The machine learning model in S3 is a graph neural network or a multi-task learning model, whose hidden layer contains edge connections that represent the interdependencies between trace elements.
10. The intelligent feed formulation optimization method based on the availability of background trace elements according to claim 1, characterized in that: The machine learning model is a time-series prediction model; In step S2, the release rate of trace elements is measured multiple times at different time points during the simulated digestion process to construct a release kinetic curve for each raw material-enzyme combination. The release kinetic curve contains release rate data at at least three time points. The output of the time-series prediction model is the predicted release kinetic curve.