Low-protein diversified diet design method and system based on multi-source data fusion
By constructing a directed causal graph of formula-environment-effect and using counterfactual reasoning methods, the problem of insufficient causal attribution ability in existing technologies is solved, enabling accurate differentiation of breeding effects and dynamic preference adjustment, and improving the accuracy and stability of low-protein feed formula recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LIAONING WELLHOPE AGRI TECH
- Filing Date
- 2026-03-11
- Publication Date
- 2026-06-19
AI Technical Summary
Existing low-protein feed formulation recommendation systems cannot accurately distinguish the impact of environmental factors and formulation factors on breeding effects, making preference updates susceptible to interference from confounding factors and affecting the accuracy and stability of formulation recommendations.
By fusing multi-source data, a directed causal graph of formula-environment-effect is constructed. A causal discovery algorithm is used to learn the causal relationship between variables. An inverse reinforcement learning algorithm is combined to infer the target weight preference. Counterfactual reasoning is used to distinguish the causes of effect deviation. Weight updates are triggered when the causal attribution is determined to be due to the contribution of formula factors.
This improves the accuracy and stability of preference learning in the low-protein feed formulation recommendation system, avoids false update signals caused by environmental factors, and ensures the effectiveness of formulation optimization.
Smart Images

Figure CN121811974B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of livestock breeding and feed nutrition technology, and more specifically, to a method and system for designing low-protein diversified diets based on multi-source data fusion. Background Technology
[0002] In the fields of livestock farming and feed nutrition, low-protein feed formulation recommendation systems need to provide personalized formulation recommendations based on the differentiated preferences of different farms. The system infers the target weight preferences by analyzing the historical selection behavior of farms and generates personalized formulations through multi-objective optimization based on these preferences.
[0003] In existing technologies, low-protein feed formulation recommendation systems employ a preference update method that directly uses performance feedback as a trigger signal to adjust weights. When farms report that the growth effect after adopting the recommended formulation is not as expected, the system directly adjusts the target weight preference.
[0004] However, there are two possible reasons why growth results may fall short of expectations: one is a problem with the formulation itself, where improper target weighting causes the formulation composition to deviate from the actual needs of the farm; the other is environmental factors, including external interference such as disease outbreaks, abnormal temperatures, and improper management. Existing preference update methods do not distinguish the true causes of performance deviations. When the unsatisfactory growth results are actually caused by environmental factors, incorrectly adjusting the formulation weights will increase formulation costs without improving growth results, and will also cause the preference profile to deviate from the farm's true preferences. Therefore, the lack of causal attribution ability in existing technologies makes preference updates susceptible to confounding factors, affecting the accuracy and stability of formulation recommendations. Summary of the Invention
[0005] This invention provides a method and system for designing low-protein diversified diets based on multi-source data fusion, which solves the technical problems in related technologies that make it difficult to accurately distinguish the impact of environmental factors and formulation factors on breeding effects and make it difficult to reasonably update the formulation preferences of farms.
[0006] This invention provides a method for designing low-protein diversified diets based on multi-source data fusion, comprising the following steps:
[0007] Obtain historical formula usage records, breeding effect data, breeding environment data, and formula selection behavior records from each farm to generate a multi-source dataset of formula-environment-effect-selection.
[0008] Based on the recipe feature variables, environmental variables, and effect variables in the recipe-environment-effect-selection multi-source dataset, a causal discovery algorithm is used to learn the causal relationship between the variables and generate a recipe-environment-effect directed causal graph.
[0009] Effect size estimation is performed on the causal edges pointing to the effect variables in the directed causal graph of formulation-environment-effect, the causal contribution of each formulation feature and each environmental factor to the effect is quantified, and a causal contribution benchmark table is generated.
[0010] Based on the formula selection behavior records, an inverse reinforcement learning algorithm is used to infer the implicit target weight preference from the historical formula selection behavior of the farm, and generate the farm target weight preference vector.
[0011] When feedback is received that the growth effect of the farm is not as expected, extract the formula characteristic data and environmental data of that batch, and calculate the degree of abnormality score of each variable relative to the historical normal distribution.
[0012] Based on the formula-environment-effect directed causal graph and the anomaly score, the counterfactual reasoning method is used to estimate the expected effect when the environmental variables are assumed to be normal, and to calculate the effect deviation ratio contributed by environmental factors and the effect deviation ratio contributed by formula factors.
[0013] When the proportion of effect deviation contributed by the formulation factors exceeds the preset formulation attribution threshold, the farm target weight preference vector is updated based on the direction of effect deviation; when the proportion of effect deviation contributed by the environmental factors exceeds the preset environmental attribution threshold, the farm target weight preference vector remains unchanged.
[0014] Furthermore, the step of employing a causal discovery algorithm to learn the causal relationships between variables and generate a directed causal graph of formula-environment-effect includes:
[0015] The PC algorithm is used to determine the causal relationship between variables through conditional independence test. For variables A and B, given a set of variables C, if the conditional correlation between variables A and B is not significant, it is determined that there is no direct causal edge between variables A and B. If the conditional correlation is significant, the causal edge between variables A and B is retained.
[0016] The direction of the edge is determined by combining the temporal sequence relationship and the domain causal common sense constraint. The domain causal common sense constraint includes the fact that the formula feature variable precedes the effect variable in time, there is no direct causal relationship between the environmental variable and the formula feature variable, and the effect variable cannot be used as the reason for formula feature variable or environmental variable.
[0017] Furthermore, the quantification of the causal contribution of each formulation characteristic and each environmental factor to the effect includes:
[0018] The average causal effect calculation method is used to calculate the difference between the expected value of the effect variable under the intervention and the baseline value by performing intervention operations on the variable, and obtain the average causal effect of each variable on the effect variable.
[0019] The effect size is estimated unbiasedly by using a backdoor adjustment formula based on the formulation-environment-effect directed causal graph. The backdoor adjustment eliminates the influence of confounding factors by identifying a set of adjustment variables that satisfy the backdoor criterion and then performing a weighted average on the set of adjustment variables.
[0020] Furthermore, the method of inferring the implicit target weight preference from the historical formula selection behavior of a farm using an inverse reinforcement learning algorithm includes:
[0021] The formulation selection behavior is modeled as a utility maximization decision process based on implicit objective weight preferences, and the utility function is represented as a weighted sum of the performance values of each objective.
[0022] The probability of the farm choosing the formula follows a Softmax distribution. The weight vector is solved by maximizing the likelihood function, and the gradient ascent algorithm is used for optimization. The loss function is the negative log-likelihood function.
[0023] The objectives include cost control, growth performance, and environmental protection targets, with corresponding performance values of formulation cost, expected daily weight gain, and nitrogen emissions, respectively.
[0024] Furthermore, the calculation of the abnormality score of each variable relative to its historical normal distribution includes:
[0025] For each observed value of a variable, the absolute value of the difference between the observed value and the historical mean of the variable is calculated and divided by the historical standard deviation of the variable to obtain the abnormality score of the variable.
[0026] When the abnormality score exceeds a preset abnormality threshold, the variable is determined to be in an abnormal state.
[0027] The weighted average of the abnormality scores of each environmental variable is calculated as the comprehensive environmental abnormality index. The weights are obtained by normalizing the causal contribution of each environmental variable in the causal contribution benchmark table.
[0028] Furthermore, the step of using counterfactual reasoning to estimate the expected effect under the assumption that environmental variables are normal, and calculating the percentage of effect deviation contributed by environmental factors and the percentage of effect deviation contributed by formulation factors, includes:
[0029] Based on the structural equation in the formula-environment-effect directed causal graph, the values of abnormal environmental variables are replaced with the historical normal average values of the abnormal environmental variables, while keeping the actual values of the formula characteristic variables unchanged.
[0030] Based on the replaced variable values, calculate the counterfactual expected value of the effect variable along the causal path of the formula-environment-effect directed causal graph;
[0031] The effect bias of the environmental factor contribution is calculated as the difference between the counterfactual expected value and the actual observed effect;
[0032] The percentage of effect deviation contributed by environmental factors is calculated as the ratio of the absolute value of the effect deviation contributed by environmental factors to the absolute value of the total effect deviation. The percentage of effect deviation contributed by formulation factors is calculated as 1 minus the percentage of effect deviation contributed by environmental factors.
[0033] Furthermore, the structural equation refers to the functional relationship between each node variable in the formula-environment-effect directed causal graph and its parent node variable, and the functional relationship is obtained by regression analysis on historical data;
[0034] For nonlinear structural equation models, the Monte Carlo simulation method is used to calculate the counterfactual expected value. The distribution of the random error term is repeatedly sampled, and the counterfactual value of an effect variable is calculated for each sample based on the structural equation and the values of the replaced environmental variables. After repeated sampling, the average value is taken as the counterfactual expected value.
[0035] Further, updating the farm's target weight preference vector based on the effect deviation direction includes:
[0036] Using an online learning algorithm, the updated target weight is equal to the original target weight plus the product of the learning rate and the weight adjustment amount;
[0037] The weight adjustment amount is proportional to the relative magnitude of the effect deviation, and the adjustment direction coefficient of each target is determined according to the specific manifestation of the effect deviation. When the effect deviation is manifested as insufficient growth performance, the weight of the growth performance target is increased; when the effect deviation is manifested as excessive cost, the weight of the cost control target is increased.
[0038] A gradual update strategy is adopted, setting the learning rate as the product of the base learning rate and the proportion of effect deviation contributed by the formulation factors. When the proportion of effect deviation contributed by the formulation factors is low, the weight update magnitude is reduced accordingly.
[0039] The updated weight vector is normalized so that the sum of the weights of each objective is 1.
[0040] Furthermore, it also includes: when the accumulated amount of new breeding data reaches a preset data amount threshold, the existence of edges in the formula-environment-effect directed causal graph is re-examined; if the test result is inconsistent with the original structure, the corresponding edges and their directions are updated.
[0041] The causal attribution results, weight adjustment decisions, and updated farm target weight preference vectors are recorded in the farm preference profile, generating a preference update log and attribution analysis report.
[0042] This invention also proposes a low-protein diversified diet design system based on multi-source data fusion, including:
[0043] The data acquisition module is used to acquire historical formula usage records, breeding effect data, breeding environment data and formula selection behavior records of each farm, and generate a multi-source dataset of formula-environment-effect-selection.
[0044] The causal graph construction module is used to generate a directed causal graph of formula-environment-effect based on formula feature variables, environmental variables, and effect variables, using a causal discovery algorithm.
[0045] The causal contribution estimation module is used to estimate the effect size of causal edges and generate a causal contribution benchmark table.
[0046] The preference inference module is used to infer the implicit target weight preference of the farm using an inverse reinforcement learning algorithm, and generate the target weight preference vector of the farm.
[0047] The anomaly detection module is used to calculate the degree of anomaly of each variable relative to its historical normal distribution.
[0048] The attribution analysis module is used to calculate the percentage of effect deviation contributed by environmental factors and the percentage of effect deviation contributed by formulation factors using counterfactual reasoning methods.
[0049] The conditional update module is used to determine whether to update the farm target weight preference vector based on the attribution results, and to perform weight updates when necessary.
[0050] The beneficial effects of this invention are as follows:
[0051] This invention constructs a directed causal graph of recipe-environment-effect using a PC algorithm, and performs causal attribution analysis before updating preference weights. Since the directed causal graph of recipe-environment-effect clearly represents the causal relationship structure between recipe characteristic variables, environmental variables, and effect variables, it can provide a structured causal reasoning basis for subsequent attribution analysis.
[0052] The counterfactual reasoning method is used to estimate the expected effect when the environmental variables are assumed to be normal, and then compared with the actual effect to calculate the respective contribution ratio of the effect deviation of environmental factors and formulation factors. Because the counterfactual reasoning method can evaluate the causal effect of a specific variable alone while keeping other variables constant, it can distinguish whether the failure to achieve the expected growth effect is caused by formulation factors or environmental factors, thus avoiding the misattribution of effect deviations caused by environmental factors to improper formulation preference settings.
[0053] This system uses inverse reinforcement learning to infer implicit target weight preferences from historical selection behavior in farms, and dynamically adjusts these preferences using a conditional weight update method. Specifically, weight updates are triggered only when the proportion of effect deviation attribution to formulation factors exceeds a preset formulation attribution threshold. Since the triggering condition for weight updates depends on causal attribution results rather than simple effect feedback signals, it can filter out spurious update signals caused by environmental factors, thereby improving the accuracy and stability of preference learning in the low-protein feed formulation recommendation system. Attached Figure Description
[0054] Figure 1 This is a flowchart of a low-protein feed formulation recommendation method provided in an embodiment of the present invention;
[0055] Figure 2 This is a sub-flowchart of estimating the contribution ratio of the effect deviation of environmental factors and formulation factors in the low-protein feed formulation recommendation method provided in the embodiments of the present invention;
[0056] Figure 3 This is a schematic diagram illustrating the relationship between historical batch formulation costs and daily weight gain, provided by an embodiment of the present invention.
[0057] Figure 4 This is a comparison chart of the causal contribution of the formulation and environmental variables provided in the embodiments of the present invention;
[0058] Figure 5 This is a schematic diagram of the target weight preference inference result provided by an embodiment of the present invention;
[0059] Figure 6 This is a schematic diagram of the abnormality score of the 38th batch of variables provided in the embodiments of the present invention;
[0060] Figure 7 This is a schematic diagram of counterfactual reasoning and effect bias decomposition provided by an embodiment of the present invention;
[0061] Figure 8 This is a schematic diagram illustrating the proportion of causal attribution for effect deviation provided in an embodiment of the present invention;
[0062] Figure 9 This is a historical 12-month formula selection behavior trend chart provided by an embodiment of the present invention. Detailed Implementation
[0063] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0064] At least one embodiment of the present invention discloses a method for designing low-protein diversified diets based on multi-source data fusion, such as... Figures 1-9 As shown, it includes the following steps:
[0065] S100, Obtain Formula-Environment-Effect-Selection Multi-Source Dataset: Obtain historical formula usage records, corresponding breeding effect data, concurrent breeding environment data, and formula selection behavior records from each farm to generate a Formula-Environment-Effect-Selection Multi-Source Dataset;
[0066] The data on breeding performance includes daily weight gain and feed conversion ratio, while the data on the breeding environment includes temperature, humidity, disease records, and management operation records.
[0067] Furthermore, the aforementioned historical formula usage records include formula characteristic variables such as the formula's nutritional composition, raw material ratios, and cost indicators. Formula selection behavior records refer to the actual selection results of the farm among multiple candidate formulas, used to subsequently infer the farm's implicit target weight preferences.
[0068] Furthermore, to ensure the accuracy of subsequent causal graph learning and effect estimation, data preprocessing was performed on the multi-source dataset of formulation-environment-effect-selection. Specifically, formulation feature variables, environmental variables, and effect variables were standardized to eliminate differences in dimensionality and numerical scale among different variables. For categorical environmental variables such as disease records and management operation records, one-hot encoding was used to convert the categorical environmental variables into numerical variables.
[0069] S200, Constructing a Directed Causal Graph of Recipe-Environment-Effect: Based on the recipe-environment-effect selection of multi-source datasets, recipe feature variables, environmental variables, and effect variables are used to learn the causal relationship between variables using the PC algorithm, and a directed causal graph of recipe-environment-effect is generated;
[0070] Furthermore, the PC algorithm takes observed data of formula characteristic variables, environmental variables, and effect variables as input and outputs a directed causal graph representing the causal relationships between variables. The PC algorithm determines the causal relationships between variables through conditional independence tests and determines the direction of edges by combining temporal sequence relationships and domain causal common sense constraints.
[0071] Furthermore, the aforementioned conditional independence test refers to using statistical tests to determine whether two variables are independent given other variables. Specifically, for variables... and In a given set of variables Under the following conditions, if and If the conditional correlation is not significant, then it is determined that... and Conditional independence, denoted as ,at this time and There is no direct causal edge between them; if the conditional correlation is significant, then retain... and The PC algorithm starts with the complete graph and gradually removes edges that do not have causal relationships by checking conditional independence, eventually obtaining the skeleton structure of the causal graph.
[0072] Furthermore, the causal common sense constraints in the above-mentioned fields include: formulation characteristic variables precede effect variables in time sequence; there is no direct causal relationship between environmental variables and formulation characteristic variables; and effect variables cannot be the cause of formulation characteristic variables or environmental variables.
[0073] Furthermore, to improve the accuracy of the formula-environment-effect directed causal graph, incremental correction of the causal graph structure is performed using newly added aquaculture data, based on the initial causal graph generated by the PC algorithm. Specifically, when the accumulated amount of newly added aquaculture data reaches a preset data threshold, the existence of edges in the formula-environment-effect directed causal graph is re-examined. If the examination result is inconsistent with the original structure, the corresponding edges and their directions are updated.
[0074] S300, Estimating the causal contribution of each factor to the effect: Estimating the effect size of the causal edges pointing to the effect variable in the directed causal graph of formulation-environment-effect, quantifying the causal contribution of each formulation feature and each environmental factor to the effect using the average causal effect calculation method, and generating a causal contribution benchmark table.
[0075] Furthermore, the method for calculating the average causal effect mentioned above is as follows: for variables For effect variables Average causal effect The calculation is performed using the following formula:
[0076] ;
[0077] in, Represents the variable Intervention to change variables Values , and Representing variables respectively Intervention values and baseline values, This represents the expectation operation. In actual calculations, the effect size is estimated unbiasedly using a backdoor adjustment formula based on the formulation-environment-effect directed causal graph identification.
[0078] Furthermore, the above and The specific value depends on the variable The actual business meaning is determined. For continuous variables, Take the 75th percentile value from historical data. The 25th percentile of historical data is used to assess the causal impact of raising a continuous variable from a lower to a higher level on the effect variable. For categorical variables, and Different category levels were used to assess the causal impact of category changes on the effect variable.
[0079] Furthermore, the aforementioned backdoor adjustment refers to a method of eliminating spurious causal relationships by controlling confounding variables. Specifically, for variables... For effect variables Estimating causal effects requires identifying a set of adjustment variables that satisfy the backdoor criterion. Adjust the set of variables Need to block and All false paths in between, without blocking arrive The causal path. The backdoor adjustment formula is:
[0080] ;
[0081] in, This represents the set of adjustment variables that satisfy the backdoor criterion. Represents the set of adjustment variables The specific values, Represents the set of adjustment variables Values The probability distribution. This is achieved by adjusting the set of variables. By performing a weighted average to eliminate the influence of confounding factors, the variables can be obtained. For effect variables Unbiased causal effect estimation.
[0082] S400, Inferring the target weight preference of the farm: Based on the formula selection behavior records in the formula-environment-effect-selection multi-source dataset, the inverse reinforcement learning algorithm is used to infer the implicit target weight preference from the farm's historical formula selection behavior, and generate the farm's target weight preference vector;
[0083] Furthermore, the input to the aforementioned inverse reinforcement learning algorithm is the historical formula selection record of the farm and the performance values of each formula on different objectives, and the output is the weight vector of each objective. The inverse reinforcement learning algorithm models formula selection behavior as a utility-maximizing decision-making process based on implicit objective weight preferences, and its utility function... This can be expressed as a weighted sum of all objectives:
[0084] ;
[0085] in, Indicates the total number of targets. Indicates the target index and , Indicates the first The weight of each objective, Indicates the formula in the first Performance values on each target.
[0086] Furthermore, to ensure the comparability of performance values for different targets, the performance values of each target are normalized. Specifically, for the cost control target, the formula cost is mapped to a range of 0 to 1, with a smaller value indicating a lower formula cost; for the growth performance target, the expected daily output value is mapped to a range of 0 to 1, with a larger value indicating better growth performance; and for the environmental protection target, nitrogen emissions are mapped to a range of 0 to 1, with a smaller value indicating better environmental performance.
[0087] Furthermore, the inverse reinforcement learning algorithm assumes that the farm selects the formula based on a utility function. The probability follows a Softmax (normalized exponential function, a generalized model in multi-class classification problems, used to handle classification tasks with more than two class labels) distribution:
[0088] ;
[0089] in, Indicates formula The utility value, Iterate through all candidate recipes, This represents an exponential function. The inverse reinforcement learning algorithm maximizes the likelihood function. Solving for the weight vector ,in This indicates the weight of the first objective. This indicates the weight of the second objective. Indicates the first The weight of each objective, Indicates that the farm is in the first The actual formula selected in the second selection. This represents the time index of the selected batch. The gradient ascent algorithm is used for optimization, with the loss function being the negative log-likelihood function.
[0090] Furthermore, the aforementioned objectives include cost control objectives, growth performance objectives, and environmental protection objectives, with corresponding performance values such as formulation cost, expected daily output value, and nitrogen emissions.
[0091] The expected daily output includes the egg production rate of laying hens and the weight gain ratio of non-laying hens;
[0092] S500, calculate the abnormality score of each variable: When feedback is received that the growth effect of the farm is not as expected, extract the formula characteristic data and environmental data of that batch, and calculate the abnormality score of each variable relative to the historical normal distribution.
[0093] Furthermore, the above-mentioned abnormality score is calculated as follows: for the variable Observations Based on variables Calculate the degree of abnormality score based on the historical normal distribution. :
[0094] ;
[0095] in, Representing variables Historical average, Representing variables The historical standard deviation. When When the value exceeds a preset abnormal threshold, the variable is determined. It is in an abnormal state.
[0096] Furthermore, the aforementioned preset anomaly threshold is determined based on the statistical significance level, and is usually set to 2 or 3, indicating that when the observed value deviates from the mean by 2 or 3 times the standard deviation, it is judged as an anomaly.
[0097] Furthermore, in order to comprehensively assess the anomalies of multiple environmental variables, a weighted average of the anomaly scores of each environmental variable is calculated as the comprehensive environmental anomaly index, with the weights determined based on the causal contribution of each environmental variable in the causal contribution benchmark table.
[0098] Specifically, the comprehensive environmental anomaly index The calculation formula is:
[0099] ;
[0100] in, This represents the total number of environment variables. Represents environment variable index and , Indicates the first One environment variable, Indicates the first Anomaly score for each environmental variable. Indicates the first Weighting coefficients for each environmental variable. According to the The causal contribution of each environmental variable is normalized in the causal contribution benchmark table to obtain the result, which satisfies the following conditions: .
[0101] S600, Estimating the contribution ratio of environmental factors and formulation factors to the effect deviation: Based on the structure of the formulation-environment-effect directed causal graph and the degree of anomalousness score, the counterfactual reasoning method is used to estimate the expected effect when the environmental variables are assumed to be normal. The expected effect is compared with the actual effect, and the contribution ratio of environmental factors to the effect deviation and the contribution ratio of formulation factors to the effect deviation are calculated.
[0102] Furthermore, the process of the above counterfactual reasoning method includes the following sub-steps:
[0103] S601, based on the structural equation in the directed causal graph of formulation-environment-effect, replaces the values of abnormal environmental variables with the historical normal mean of the abnormal environmental variables, while keeping the actual values of the formulation characteristic variables unchanged;
[0104] Furthermore, the structural equations mentioned above refer to the relationship between each node variable in the formula-environment-effect directed causal graph and its parent node variable. For effect variables... The structural equation is in the form of:
[0105] ;
[0106] in, Represents the set of characteristic variables of the formula. Represents a set of environment variables. Represents the random error term. This represents the functional relationship defined by the directed causal graph structure of the formula-environment-effect. In counterfactual reasoning, the values of anomalous environmental variables are replaced with their historical normal mean values. This yields the counterfactual structural equation:
[0107] ;
[0108] in, This indicates the counterfactual expected effect value.
[0109] Furthermore, the above functional relationship The specific form is obtained through regression analysis of historical data. Specifically, multiple linear regression or nonlinear regression methods are used, with formula characteristic variables and environmental variables as independent variables and effect variables as dependent variables, to fit a functional relationship. For linear structural equations, the functional relationships It is represented as a linear combination of the variables of each parent node; for nonlinear structural equations, methods such as multinomial regression, neural networks or random forests are used to fit the nonlinear functional relationship.
[0110] S602, based on the replaced variable values, calculate the counterfactual expected value of the effect variable along the causal path of the formula-environment-effect directed causal graph. ;
[0111] Furthermore, the aforementioned counterfactual expectation value The calculation is performed as follows: First, all causal paths from environmental variables to effect variables are identified based on the formula-environment-effect directed causal graph. Then, the counterfactual values of intermediate nodes are calculated sequentially along each causal path. Finally, the counterfactual expected value of the effect variable is obtained by summing these values. For linear structural equation models, the counterfactual expected value It can be directly calculated from the coefficients of the structural equations and the values of the replaced variables; for nonlinear structural equation models, the Monte Carlo simulation method is used for numerical calculation.
[0112] Furthermore, the Monte Carlo simulation method described above refers to a method of calculating numerical solutions through a large number of random samples. Specifically, for nonlinear structural equation models, sampling is repeated from the distribution of the random error term. Each sampling calculates the counterfactual value of an effect variable based on the structural equation and the values of the replaced environmental variables. After repeated sampling multiple times, the average value is taken as the counterfactual expected value. .
[0113] S603, the bias in calculating the contribution of environmental factors is... ,in For actual observation results;
[0114] S604, Calculate the percentage of effect bias in the contribution of environmental factors. The proportion of effect deviation contributed by formulation factors :
[0115] ;
[0116] ;
[0117] in, Indicates the overall effect deviation. This indicates the expected results based on historical data from the farm.
[0118] Furthermore, the above-mentioned expected effects The calculation method is as follows: Based on the historical normal batch effect data of the farm, a linear regression model is used to establish the predictive relationship between the formula characteristic variables and the effect variables. The formula characteristic variables of the current batch are then substituted into the linear regression model to obtain the expected effect value. .
[0119] Furthermore, the process of establishing the above linear regression model is as follows: Collect historical data on normal batches of feed formulation characteristic variables and corresponding effect variables from the farm; use the least squares method to fit the linear regression equation to obtain the regression coefficients of each formulation characteristic variable. Expected effect value. The calculation formula is the sum of the products of each formula characteristic variable and its regression coefficient, plus the intercept term.
[0120] S700 performs conditional weight updates based on attribution results: When the proportion of effect deviation contributed by formulation factors exceeds the preset formulation attribution threshold, it is determined that the farm's target weight preference vector needs to be adjusted. The farm's target weight preference vector is updated based on the direction of effect deviation using an online learning algorithm, and the updated farm's target weight preference vector is output. When the proportion of effect deviation contributed by environmental factors exceeds the preset environmental attribution threshold, weight updates are suppressed, the farm's target weight preference vector remains unchanged, and environmental improvement suggestions are generated based on abnormal environmental variables.
[0121] Furthermore, the aforementioned preset formula attribution threshold and preset environmental attribution threshold are set according to the actual application scenario. Typically, the preset formula attribution threshold is set between 0.5 and 0.7, and the preset environmental attribution threshold is set between 0.5 and 0.7. The sum of the preset formula attribution threshold and the preset environmental attribution threshold is 1, which is used to determine the main attribution object of the effect deviation.
[0122] Furthermore, the input to the aforementioned online learning algorithm is the current farm target weight preference vector and the direction of effect deviation, and the output is the updated farm target weight preference vector. The weight update method of the online learning algorithm is as follows:
[0123] ;
[0124] in, and They represent the first The weights of each target after the update and before the update. Indicates the learning rate. This represents the weight adjustment amount calculated based on the direction of effect deviation.
[0125] Furthermore, the aforementioned weight adjustment amount The calculation method is as follows: First, identify the specific manifestation of the effect deviation. If the effect deviation is manifested as insufficient growth performance, that is, the actual daily weight gain is lower than the expected daily weight gain, then increase the weight of the growth performance target. The weight adjustment of the growth performance target is set to a positive value, and the weight adjustment of other targets is set to a negative value. If the effect deviation is manifested as excessive cost, that is, the actual formula cost is higher than the expected formula cost, then increase the weight of the cost control target. The weight adjustment of the cost control target is set to a positive value, and the weight adjustment of other targets is set to a negative value.
[0126] The specific value of the weight adjustment is directly proportional to the magnitude of the effect deviation, and the calculation formula is as follows:
[0127] ;
[0128] in, Indicates the adjustment factor. This indicates the relative magnitude of the effect deviation. Indicates the first The adjustment direction coefficient for the first target, when it is necessary to increase the first target... When each target weight When it is necessary to reduce the first When each target weight To ensure the normalization constraint of the weight vector, after calculating the weight adjustment for all objectives, the updated weight vector is normalized so that... .
[0129] Furthermore, the aforementioned adjustment coefficients Based on the sensitivity requirements of the low-protein feed formulation recommendation system, it is usually set between 0.1 and 0.3 to control the magnitude of a single weight update and avoid drastic weight fluctuations.
[0130] Furthermore, to avoid drastic fluctuations in weights caused by single attribution errors, a gradual update strategy is adopted, adjusting the learning rate... Set as the percentage of effect deviation from the contribution of formulation factors. Positive correlation:
[0131] ;
[0132] in, This represents the base learning rate. When the contribution of formulation factors to the effect bias is relatively low, the weight update magnitude is reduced accordingly, thereby improving the stability of preference updates.
[0133] Furthermore, the aforementioned basic learning rate The convergence speed requirement of the low-protein feed formulation recommendation system is set according to the requirements of the system, usually between 0.1 and 0.5, to control the basic rate of weight updates.
[0134] Furthermore, in order to continuously optimize the accuracy of causal attribution and preference estimation, the causal attribution results, weight adjustment decisions, and updated farm target weight preference vectors are recorded in the farm preference profile, generating preference update logs and attribution analysis reports. Data is continuously accumulated for incremental correction of the formulation-environment-effect directed causal graph and preference estimation.
[0135] Furthermore, the updated farm target weight preference vector can be directly used for subsequent multi-objective optimization formulation recommendation without additional decoding or transformation processing. Specifically, in the multi-objective optimization process, the updated farm target weight preference vector is used as the weight coefficient of each objective, and the comprehensive utility value of each candidate formulation is calculated using a weighted summation method. The formulation with the highest comprehensive utility value is selected as the recommendation result.
[0136] In one embodiment of the present invention, a low-protein diversified diet design system based on multi-source data fusion is also proposed, which includes the following modules:
[0137] The data acquisition module is used to acquire historical formula usage records, breeding effect data, breeding environment data and formula selection behavior records of each farm, and generate a multi-source dataset of formula-environment-effect-selection.
[0138] The causal graph construction module is used to generate a directed causal graph of formula-environment-effect based on formula feature variables, environmental variables, and effect variables, using a causal discovery algorithm.
[0139] The causal contribution estimation module is used to estimate the effect size of causal edges and generate a causal contribution benchmark table.
[0140] The preference inference module is used to infer the implicit target weight preference of the farm using an inverse reinforcement learning algorithm, and generate the target weight preference vector of the farm.
[0141] The anomaly detection module is used to calculate the degree of anomaly of each variable relative to its historical normal distribution.
[0142] The attribution analysis module is used to calculate the percentage of effect deviation contributed by environmental factors and the percentage of effect deviation contributed by formulation factors using counterfactual reasoning methods.
[0143] The conditional update module is used to determine whether to update the farm target weight preference vector based on the attribution results, and to perform weight updates when necessary.
[0144] Based on the above-mentioned low-protein diversified diet design method and system based on multi-source data fusion, it is applied to the following application scenarios:
[0145] A large-scale pig farming enterprise operates multiple fattening farms in East China, one of which focuses on raising 100kg fattening pigs. This farm began using an intelligent low-protein feed formulation recommendation system. Based on the farm's historical formulation selection behavior, the system inferred its target weight preferences as follows: cost control 0.45, growth performance 0.40, and environmental indicators 0.15. The farm reported that after using the recommended formula, the 38th batch of fattening pigs only gained 820 grams per day, lower than the expected 950 grams per day, while the feed conversion ratio reached 3.2, higher than the expected 2.8.
[0146] The system needs to determine whether the deviation in performance is caused by improper formula weight settings or by abnormal breeding environment, and decide whether to adjust the target weight preference of the farm.
[0147] The initial conditions for this batch included: using a low-protein formula with formula number F202403A, a crude protein content of 14.2%, a lysine content of 0.95%, and a formula cost of 2.85 yuan / kg; the concurrent breeding environment data showed an average temperature of 18.3°C and an average humidity of 72%, and a mild diarrhea outbreak occurred on March 8, after which the feeders made an unconventional adjustment to the feeding time on March 10.
[0148] Application example:
[0149] Implementation example of core S100: Corresponding to S100 in the specific implementation, the system obtains historical formula usage records, breeding effect data, environmental data and formula selection behavior records of Taizhou farm, and generates a multi-source dataset of formula-environment-effect-selection.
[0150] This dataset covers complete breeding cycle data for 37 batches over the past 12 months at the farm.
[0151] Table 1. Input data of historical formula usage records for S100 (partial batches)
[0152]
[0153] Table 2. S100 Formulation-Environment-Effect Data Output (Partial Batch)
[0154]
[0155] During data preprocessing, the disease occurrence variable was coded as follows: no disease = 0, mild disease = 1, moderate disease = 2; and the management anomaly variable was coded as follows: no anomaly = 0, feeding time adjustment = 1, ventilation system failure = 2, personnel change = 3. Continuous variables such as formula cost, daily weight gain, and feed conversion ratio were standardized using Z-scores.
[0156] Implementation example of core S200: Corresponding to S200 in the specific implementation, the system is based on a multi-source dataset of recipe-environment-effect-selection and uses the PC algorithm to learn the causal relationship between variables.
[0157] The algorithm input includes 8 formula feature variables (crude protein content, lysine content, methionine content, threonine content, energy level, fiber content, calcium-phosphorus ratio, and formula cost), 6 environmental variables (temperature, humidity, disease occurrence, management abnormalities, stocking density, and ventilation), and 2 effect variables (daily weight gain and feed conversion ratio).
[0158] Table 3. Results of the conditional independence test for S200 (partial variable pairs)
[0159]
[0160] Table 4. Directed causal graph structure output of S200 (core causal edges)
[0161]
[0162] Based on the causal graph structure generated by the PC algorithm, formulation feature variables such as crude protein content and lysine content affect daily weight gain and feed conversion ratio through direct causal paths; environmental variables such as temperature, humidity, disease occurrence, and management anomalies also affect effect variables through direct causal paths; there are no causal edges between formulation feature variables and environmental variables, which conforms to the domain causal common sense constraint.
[0163] Implementation example of core S300: Corresponding to S300 in the specific implementation, the system estimates the average causal effect of the causal edges pointing to the effect variables in the directed causal graph of formula-environment-effect. Taking the causal effect of crude protein content on daily weight gain as an example, the following is set... (Historical data 75th percentile) (Historical data 25th percentile).
[0164] Regarding the causal effect of the epidemic on daily weight gain, we set... (A mild outbreak of disease) (No epidemic). Identify the set of adjustment variables that satisfy the backdoor criterion. Temperature, humidity, stocking density The unbiased causal effect is calculated by adjusting the formula through a backdoor.
[0165] Table 5. S300 Causal Contribution Benchmark Table
[0166]
[0167] Specific calculation example: When the crude protein content increases from 13.5% to 14.8%, according to the backdoor adjustment formula: ;
[0168] ;
[0169] ;
[0170] The causal contribution score was obtained through normalization, and the contribution score of crude protein content was: .
[0171] Implementation example of core S400: Corresponding to S400 in the specific implementation, the system uses an inverse reinforcement learning algorithm to infer the implicit target weight preference based on 15 formula selection records of the Taizhou farm over the past 12 months. In each selection scenario, the system provides 3 to 5 candidate formulas, and the farm selects one for actual use.
[0172] Table 6. Records of S400 formulation selection behavior (partial selection batches)
[0173]
[0174] For batch Q12, calculate the utility value of each candidate formulation:
[0175] ;
[0176] ;
[0177] By maximizing the likelihood function The weight vector is obtained by solving. That is, the weight of cost control targets Growth performance target weight Environmental protection indicator target weight .
[0178] Implementation example of core S500: In the specific implementation, when feedback is received that the growth effect of the 38th batch is not as expected, the system extracts the formula characteristic data and environmental data of the batch and calculates the abnormality score of each variable.
[0179] Table 7 Calculation of S500 Variable Abnormality Score
[0180]
[0181] Example of calculating anomaly score:
[0182] ;
[0183] ;
[0184] ;
[0185] ;
[0186] Calculation of the comprehensive environmental anomaly index:
[0187] Based on the causal contribution in Table 5, temperature weighting Epidemic weight Managing abnormal weights .
[0188] ;
[0189] Implementation example of core S600: In the specific implementation of S600, the system is based on the structural equation of the directed cause-effect graph of formula-environment-effect, and replaces abnormal environmental variables with historical normal averages: temperature is replaced with 21.2°C, disease occurrence is replaced with 0, and management abnormality is replaced with 0, while keeping the actual values of formula characteristic variables unchanged.
[0190] The linear structure equation fitted based on historical data is as follows:
[0191] ;
[0192] Table 8 Counterfactual Reasoning Calculations for S600
[0193]
[0194] Bias in calculating the contribution of environmental factors:
[0195] ;
[0196] Calculate the overall effect deviation:
[0197] ;
[0198] because This indicates that improvements in environmental factors can completely compensate for the deviation in effectiveness;
[0199] Percentage of effect bias contributed by environmental factors:
[0200] ;
[0201] The percentage of effect deviation contributed by formulation factors:
[0202] ;
[0203] Implementation example of core S700: In the specific implementation, the system makes conditional weight update decisions based on the causal attribution results. A preset recipe attribution threshold of 0.6 and a preset environmental attribution threshold of 0.6 are set.
[0204] Table 9 S700 Weight Update Decision Judgment
[0205]
[0206] Due to the bias in the contribution of environmental factors to the effect proportion If the environmental attribution threshold is exceeded, the system determines that the underperformance of this batch is mainly due to abnormal breeding environment, and suppresses weight updates to keep the farm's target weight preference vector unchanged. .
[0207] The system generates environment improvement suggestions based on abnormal environment variables:
[0208] Disease management recommendations: Strengthen biosecurity control. After the diarrhea outbreak on March 8, sick pigs should be isolated immediately. Use probiotics to regulate intestinal flora and supplement with electrolytes and multivitamins for 3 to 5 days.
[0209] Temperature control recommendations: The current average temperature of 18.3°C is below the suitable temperature range (20°C to 22°C). It is recommended to turn on the heat preservation equipment to raise the indoor temperature to above 21°C.
[0210] Management guidelines recommend avoiding unconventional feeding time adjustments and maintaining feeding time stability. It is recommended to restore the original three feeding times per day (7:00, 12:00, and 18:00).
[0211] The system records the causal attribution results, weight maintenance decisions, and environmental improvement suggestions into the farm's preference profile, generating an attribution analysis report (report number R20240315_B038) for incremental correction of the subsequent formulation-environment-effect directed causal graph and continuous optimization of preference estimation.
[0212] Throughout the implementation process, the data starts from an initial multi-source dataset of recipe-environment-effect-selection. Through PC algorithm learning, a directed causal graph structure of recipe-environment-effect is generated. This structure clarifies the causal relationships between variables and provides a structured foundation for subsequent attribution analysis. The causal edges pointing to effect variables in the causal graph are quantified into a causal contribution benchmark table through average causal effect calculation. This benchmark table is used to determine the weight coefficients of environmental variables in subsequent anomaly detection.
[0213] When feedback was received that the 38th batch did not meet expectations, the system extracted the formula characteristic data and environmental data for that batch, and calculated the degree of anomalousness score of each variable relative to the historical normal distribution. By weighted summing of the anomalousness scores of the environmental variables, an environmental comprehensive anomalousness index of 2.241 was obtained, indicating the existence of significant environmental anomalies.
[0214] The system, based on the structure equation of a directed causal graph of formula-environment-effect, replaces abnormal environmental variables with historical normal averages for counterfactual inference, calculating an expected daily weight gain of 1125.3 grams / day under the assumption of normal environmental conditions. Comparing the counterfactual expected value with the actual observed value of 820 grams / day, the system calculates that the effect deviation contributed by environmental factors is 305.3 grams / day, accounting for 1.0% of the total effect deviation, which far exceeds the preset environmental attribution threshold of 0.6.
[0215] Based on the causal attribution results, the system determined that the performance deviation in this batch was mainly caused by environmental factors such as disease outbreaks, low temperatures, and abnormal management, rather than improper formulation weight settings. Therefore, the system suppressed weight updates, maintaining the farm's target weight preference vector unchanged at [0.45, 0.40, 0.15], avoiding the erroneous attribution of performance deviations caused by environmental factors to formulation preferences, thus preventing the preference profile from deviating from the farm's true preferences. Simultaneously, the system generated targeted environmental improvement suggestions, guiding the farm to improve growth performance by strengthening disease prevention and control, increasing indoor temperature, and standardizing feeding management, ensuring that subsequent batches validate the effectiveness of the current formulation weights under normal environmental conditions. The entire data flow demonstrates a complete logical chain from learning the causal structure from historical data, quantifying causal contributions, detecting abnormal states, performing counterfactual reasoning, calculating attribution proportions, to conditional decision updates.
[0216] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A method for designing low-protein diversified diets based on multi-source data fusion, characterized in that, Includes the following steps: Obtain historical formula usage records, breeding effect data, breeding environment data, and formula selection behavior records from each farm to generate a multi-source dataset of formula-environment-effect-selection. Based on the recipe feature variables, environmental variables, and effect variables in the recipe-environment-effect-selection multi-source dataset, a causal discovery algorithm is used to learn the causal relationship between the variables and generate a recipe-environment-effect directed causal graph. Effect size estimation is performed on the causal edges pointing to the effect variables in the directed causal graph of formulation-environment-effect, the causal contribution of each formulation feature and each environmental factor to the effect is quantified, and a causal contribution benchmark table is generated. Based on the formula selection behavior records, an inverse reinforcement learning algorithm is used to infer the implicit target weight preference from the historical formula selection behavior of the farm, and generate the farm target weight preference vector. When feedback is received that the growth effect of the farm is not as expected, extract the formula characteristic data and environmental data of that batch, and calculate the degree of abnormality score of each variable relative to the historical normal distribution. Based on the formula-environment-effect directed causal graph and the anomaly score, the counterfactual reasoning method is used to estimate the expected effect when the environmental variables are assumed to be normal, and to calculate the effect deviation ratio contributed by environmental factors and the effect deviation ratio contributed by formula factors. When the proportion of effect deviation contributed by the formulation factors exceeds the preset formulation attribution threshold, the farm target weight preference vector is updated based on the direction of effect deviation; when the proportion of effect deviation contributed by the environmental factors exceeds the preset environmental attribution threshold, the farm target weight preference vector remains unchanged.
2. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 1, characterized in that, The step of using a causal discovery algorithm to learn the causal relationships between variables and generating a directed causal graph of formula-environment-effect includes: The PC algorithm is used to determine the causal relationship between variables through conditional independence test. For variables A and B, given a set of variables C, if the conditional correlation between variables A and B is not significant, it is determined that there is no direct causal edge between variables A and B. If the conditional correlation is significant, the causal edge between variables A and B is retained. The direction of the edge is determined by combining the temporal sequence relationship and the domain causal common sense constraint. The domain causal common sense constraint includes the fact that the formula feature variable precedes the effect variable in time, there is no direct causal relationship between the environmental variable and the formula feature variable, and the effect variable cannot be used as the reason for formula feature variable or environmental variable.
3. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 1, characterized in that, The quantification of the causal contribution of each formulation characteristic and each environmental factor to the effect includes: The average causal effect calculation method is used to calculate the difference between the expected value of the effect variable under the intervention and the baseline value by performing intervention operations on the variable, and obtain the average causal effect of each variable on the effect variable. The effect size is estimated unbiasedly by using a backdoor adjustment formula based on the formulation-environment-effect directed causal graph. The backdoor adjustment eliminates the influence of confounding factors by identifying a set of adjustment variables that satisfy the backdoor criterion and then performing a weighted average on the set of adjustment variables.
4. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 1, characterized in that, The method of inferring implicit target weight preferences from the historical formula selection behavior of a farm using an inverse reinforcement learning algorithm includes: The formulation selection behavior is modeled as a utility maximization decision process based on implicit objective weight preferences, and the utility function is represented as a weighted sum of the performance values of each objective. The probability of the farm choosing the formula follows a Softmax distribution. The weight vector is solved by maximizing the likelihood function, and the gradient ascent algorithm is used for optimization. The loss function is the negative log-likelihood function. The objectives include cost control objectives, growth performance objectives, and environmental protection indicators, with corresponding performance values of formula cost, expected daily output value, and nitrogen emissions, respectively. The expected daily output value includes the egg production rate of laying hens and the weight gain ratio of non-laying hens.
5. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 1, characterized in that, The calculation of the abnormality score of each variable relative to its historical normal distribution includes: For each observed value of a variable, the absolute value of the difference between the observed value and the historical mean of the variable is calculated and divided by the historical standard deviation of the variable to obtain the abnormality score of the variable. When the abnormality score exceeds a preset abnormality threshold, the variable is determined to be in an abnormal state. The weighted average of the abnormality scores of each environmental variable is calculated as the comprehensive environmental abnormality index. The weights are obtained by normalizing the causal contribution of each environmental variable in the causal contribution benchmark table.
6. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 1, characterized in that, The method of using counterfactual reasoning to estimate the expected effect under the assumption that environmental variables are normal, and calculating the percentage of effect deviation contributed by environmental factors and the percentage of effect deviation contributed by formulation factors, includes: Based on the structural equation in the formula-environment-effect directed causal graph, the values of abnormal environmental variables are replaced with the historical normal average values of the abnormal environmental variables, while keeping the actual values of the formula characteristic variables unchanged. Based on the replaced variable values, calculate the counterfactual expected value of the effect variable along the causal path of the formula-environment-effect directed causal graph; The effect bias of the environmental factor contribution is calculated as the difference between the counterfactual expected value and the actual observed effect; The percentage of effect deviation contributed by environmental factors is calculated as the ratio of the absolute value of the effect deviation contributed by environmental factors to the absolute value of the total effect deviation. The percentage of effect deviation contributed by formulation factors is calculated as 1 minus the percentage of effect deviation contributed by environmental factors.
7. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 6, characterized in that, The structural equation refers to the functional relationship between each node variable in the formula-environment-effect directed causal graph and its parent node variable, which is obtained by regression analysis on historical data; For nonlinear structural equation models, the Monte Carlo simulation method is used to calculate the counterfactual expected value. The distribution of the random error term is repeatedly sampled, and the counterfactual value of an effect variable is calculated for each sample based on the structural equation and the values of the replaced environmental variables. After repeated sampling, the average value is taken as the counterfactual expected value.
8. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 1, characterized in that, The step of updating the farm's target weight preference vector based on the direction of effect deviation includes: Using an online learning algorithm, the updated target weight is equal to the original target weight plus the product of the learning rate and the weight adjustment amount; The weight adjustment amount is proportional to the relative magnitude of the effect deviation, and the adjustment direction coefficient of each target is determined according to the specific manifestation of the effect deviation. When the effect deviation is manifested as insufficient growth performance, the weight of the growth performance target is increased; when the effect deviation is manifested as excessive cost, the weight of the cost control target is increased. A gradual update strategy is adopted, setting the learning rate as the product of the base learning rate and the proportion of effect deviation contributed by the formulation factors. When the proportion of effect deviation contributed by the formulation factors is low, the weight update magnitude is reduced accordingly. The updated weight vector is normalized so that the sum of the weights of each objective is 1.
9. The method for designing low-protein diversified diets based on multi-source data fusion according to claim 1, characterized in that, Also includes: When the accumulated amount of new aquaculture data reaches the preset data threshold, the existence of edges in the formula-environment-effect directed causal graph is re-examined. If the test result is inconsistent with the original structure, the corresponding edges and their directions are updated. The causal attribution results, weight adjustment decisions, and updated farm target weight preference vectors are recorded in the farm preference profile, generating a preference update log and attribution analysis report.
10. A low-protein diversified diet design system based on multi-source data fusion, used to execute the steps in the low-protein diversified diet design method based on multi-source data fusion as described in any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire historical formula usage records, breeding effect data, breeding environment data and formula selection behavior records of each farm, and generate a multi-source dataset of formula-environment-effect-selection. The causal graph construction module is used to generate a directed causal graph of formula-environment-effect based on formula feature variables, environmental variables, and effect variables, using a causal discovery algorithm. The causal contribution estimation module is used to estimate the effect size of causal edges and generate a causal contribution benchmark table. The preference inference module is used to infer the implicit target weight preference of the farm using an inverse reinforcement learning algorithm, and generate the target weight preference vector of the farm. The anomaly detection module is used to calculate the degree of anomaly of each variable relative to its historical normal distribution. The attribution analysis module is used to calculate the percentage of effect deviation contributed by environmental factors and the percentage of effect deviation contributed by formulation factors using counterfactual reasoning methods. The conditional update module is used to determine whether to update the farm target weight preference vector based on the attribution results, and to perform weight updates when necessary.