Epidermal growth factor adding effect prediction method based on growth performance data
By processing data on growth performance, environmental variables, and epidermal growth factor (EGF) dosage and using a dual-track learning model, the problems of time-dependent dynamic changes and individual differences in traditional prediction methods were solved, achieving accurate prediction of the effects of EGF addition.
Patent Information
- Application Number
- CN202511349369.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Traditional methods for predicting the effects of epidermal growth factor (EGF) supplementation do not adequately consider dynamic changes over time and individual differences, making it difficult to meet the needs for accurate assessment and individualized regulation.
By collecting and structuring data on the growth performance, environmental variables, and epidermal growth factor of the target object, and adding dosage data, time-series data preprocessing is performed to calculate short-term perturbation residuals and long-term drift residuals. Wavelet decomposition is then used to construct the frequency band energy spectrum. Combined with a dual-track learning model, individualized causal effect estimation and dose-time-growth performance prediction are performed, and causal significance testing and correction factor compensation are carried out.
It enables accurate and comprehensive prediction of the effects of epidermal growth factor addition, meeting the precise and individualized needs of growth performance regulation for target subjects.
Smart Images

Figure CN120849869A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of time-series data analysis technology, and in particular to a method for predicting the effects of epidermal growth factor addition based on growth performance data. Background Technology
[0002] Predicting the effects of epidermal growth factor (EGF) addition is crucial for accurately regulating the growth of target organisms, directly impacting growth performance optimization and resource utilization efficiency. Existing technologies largely rely on traditional models, using basic growth data combined with fixed algorithms to predict the effects of EGF addition. While these methods are effective in scenarios with ample data and stable environments, their limitations become apparent when applied to complex real-world scenarios, given the increasing demands for prediction accuracy. Summary of the Invention
[0003] This application addresses the technical problem that traditional methods for predicting the effects of epidermal growth factor (EGF) supplementation lack consideration for dynamic changes over time and individual differences among different target subjects, making it difficult to meet the requirements for accurate assessment and individualized regulation.
[0004] To address the aforementioned technical problems, this application proposes a method for predicting the effect of epidermal growth factor (EGF) addition based on growth performance data. The method includes: collecting and structuring data of the target object to establish an original dataset, which contains time-series records of growth performance, environmental variable records, and EGF addition dose records in chronological order; performing time-series data preprocessing on the original dataset, calculating short-term perturbation residuals and long-term drift residuals within sliding windows of multiple time scales for the growth performance time-series records in the preprocessing results, and performing wavelet decomposition on the growth performance time-series records in the preprocessing results to construct a frequency band energy spectrum; combining the short-term perturbation residuals, long-term drift residuals, and frequency band energy spectrum with environmental variable records and EGF addition dose records as input features; inputting the input features into a dual-track learning model to output individualized causal effect estimates and a dose-time-growth performance prediction mapping; performing new observations of the target object, and conducting causal significance tests on the individualized causal effect estimates based on the new observations according to the corresponding sliding windows to establish a correction factor; compensating the dose-time-growth performance prediction mapping based on the correction factor and outputting the prediction result.
[0005] This application proposes one or more technical solutions, which have at least the following technical effects:
[0006] This application collects and structures data on the growth performance time series, environmental variables, and epidermal growth factor (EGF) dosage of the target object to establish a raw dataset. After time series preprocessing, multi-timescale sliding window residual calculation, and wavelet decomposition to construct a frequency band energy spectrum, a dual-track learning model is constructed by combining multi-dimensional input features. The model outputs individualized causal effect estimates and dose-time-growth performance prediction mappings. By combining new observations to perform causal significance tests, a correction factor compensation prediction mapping is established, thereby accurately predicting the effect of EGF addition. This meets the precise and individualized requirements for the regulation of growth performance of the target object and achieves the technical effect of accurate and comprehensive prediction of the effect of EGF addition. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a flowchart illustrating the method for predicting the effect of epidermal growth factor addition based on growth performance data provided in this application embodiment.
[0009] Figure 2 This is a schematic diagram of the process for establishing a dose-time-growth performance prediction mapping in the method for predicting the effect of epidermal growth factor addition based on growth performance data provided in the embodiments of this application. Detailed Implementation
[0010] This application provides a method for predicting the effect of epidermal growth factor (EGF) addition based on growth performance data. It solves the technical problem that traditional EGF addition effect prediction methods fail to consider the dynamic changes over time and individual differences among different target subjects, making it difficult to meet the requirements for accurate assessment and individualized regulation.
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0012] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0013] like Figure 1 As shown, a method for predicting the effect of epidermal growth factor addition based on growth performance data is provided, wherein the method includes:
[0014] Step A100: Collect and structure the target object data to establish an original dataset. The original dataset includes time-series records of growth performance, records of environmental variables, and records of epidermal growth factor dosage in chronological order.
[0015] Specifically, when collecting data on target objects, three types of core information are needed to build the foundation of the original dataset. First, growth performance data is collected, recording the growth status indicators of the target object at fixed time intervals to form a time-series record of growth performance. For example, for livestock and poultry in a breeding setting, their weight and body length are recorded at fixed times each day; or for cells in a cell culture setting, their density and activity rate are recorded hourly, ensuring that the data fully reflects the dynamic changes in growth over time. Simultaneously, external environmental parameters affecting the growth of the target object are collected, forming environmental variable records. These include the temperature and humidity of the breeding environment, or the pH and dissolved oxygen levels of the cell culture environment. Each environmental data point must correspond to a specific time node to maintain consistency with the growth performance data in the time dimension. Furthermore, detailed information on each addition of epidermal growth factor (EGF) is required, forming an EGF dosage record, including the specific dosage and corresponding time for each addition, ensuring that the intervention is clearly traceable.
[0016] After data collection, the three types of data were structured in chronological order, integrating scattered information into a standardized raw dataset. Specifically, a time-series correlation approach was adopted, using time nodes as the core axis to match and integrate the growth performance time series records, environmental variable records, and epidermal growth factor (EGF) dosage records corresponding to the same time node one by one. Taking 20 days of monitoring data of a batch of experimental cells as an example, a structured data in tabular form can be constructed: the table rows represent the daily time nodes, and the columns correspond to the growth performance time series records of cell density and cell viability; environmental variable records of incubator temperature and CO2 concentration; and epidermal growth factor dosage records corresponding to the amount of EGF additive. This ensures that each row of data can fully present the growth status, environment, and EGF intervention of the target object at the corresponding time node, ultimately forming a raw dataset arranged in chronological order.
[0017] By collecting data on the growth performance, environmental variables, and epidermal growth factor dosage of the target object in a multidimensional time series, and integrating the three types of data in chronological order, a raw dataset containing time series records of growth performance, environmental variable records, and epidermal growth factor dosage records was established. This provided temporally coherent and informationally complete data support for subsequent time-series-based data preprocessing, feature extraction, and dual-track learning model input, ensuring the orderly progress of the epidermal growth factor addition effect prediction process.
[0018] Step A200: After performing time-series data preprocessing on the original dataset, the growth performance time series records in the preprocessing results are recorded in sliding windows of multiple time scales to calculate short-term perturbation residuals and long-term drift residuals, and wavelet decomposition of the growth performance time series records in the preprocessing results is performed to construct the frequency band energy spectrum.
[0019] Optionally, for the growth performance time series records in the preprocessing results, multiple sliding windows at different time scales are first determined to distinguish between short-term and long-term features. Typically, the window size is set according to the growth cycle of the target object and the data collection frequency. For example, for livestock and poultry weight data collected daily, the short-term sliding window is set to 3 time nodes (3 days), and the long-term sliding window is set to 15 time nodes (15 days). Subsequently, within each sliding window, a baseline trend curve for growth performance is fitted using linear regression or time-series smoothing algorithms. Specifically, the preprocessed growth performance time-series records are used as input. A suitable smoothing window is determined based on the data collection frequency and time span. For example, considering the multi-timescale requirements for calculating residuals later, a window of 3-15 time nodes is selected. Within each sliding window, growth performance data, including weight and daily weight gain, are calculated using methods such as moving average and exponentially weighted average. The moving average uses the arithmetic mean of the data within the window, while the exponentially weighted average assigns higher weight to recent data to highlight temporal correlation. The process then slides through the windows sequentially, connecting the smoothing results of each window. Simultaneously, the consistency between the smoothing result and the overall trend of the original time-series data is verified. If the result of a certain window deviates from the original trend by more than a preset threshold, the window parameters are adjusted and recalculated. Finally, a baseline trend curve reflecting the long-term changes in growth performance is formed, providing a reference for the subsequent calculation of short-term disturbance residuals and long-term drift residuals.
[0020] Next, the actual growth performance data at each time point within the sliding window is subtracted from the corresponding fitted value to obtain the residuals for each point. The residuals calculated within the short-term sliding window are the short-term perturbation residuals, reflecting the fluctuation of growth performance in a short period of time. For example, if the actual weights within a 3-day window are 2.0 kg, 2.1 kg, and 1.9 kg, and the fitted trend is 2.0 kg / day, the calculated short-term perturbation residuals are 0.0 kg, 0.1 kg, and -0.1 kg, respectively. The residuals calculated within the long-term sliding window are the long-term drift residuals, reflecting the degree to which growth performance deviates from the long-term trend. For example, if the fitted long-term trend within a 15-day long-term sliding window is an average daily weight gain of 0.05 kg, and the actual weight gain on the 10th day is 0.03 kg, the corresponding long-term drift residual is -0.02 kg.
[0021] Next, after calculating the short-term perturbation residual and the long-term drift residual, wavelet decomposition is performed on the preprocessed growth performance time series records. First, a suitable db4 wavelet basis function for time series signal analysis is selected. The number of decomposition levels is determined based on the time length and fluctuation characteristics of the growth performance data, generally 3-5 levels, to fully extract features from different frequency ranges. During the decomposition process, the growth performance time series is progressively decomposed into high-frequency detail signals and low-frequency approximation signals. The high-frequency detail signals correspond to the short-term rapid fluctuations in growth performance, while the low-frequency approximation signals correspond to the long-term slow changes in growth performance. Subsequently, the energy of the high-frequency detail signals and low-frequency approximation signals at each decomposition level is calculated. The energy is calculated by summing the squares of the values at each sampling point of the signal. Then, the energy of each level is arranged in descending order of frequency to construct a frequency band energy spectrum. This frequency band energy spectrum clearly shows the contribution ratio of different frequency components to the growth performance time series.
[0022] By analyzing the preprocessed growth performance time series records, short-term perturbation residuals and long-term drift residuals are calculated within sliding windows of multiple time scales. Wavelet decomposition is then performed to construct the frequency band energy spectrum. This achieves the goal of extracting short-term fluctuations, long-term trend shifts, and multi-frequency features from the growth performance time series, providing deep temporal feature support for subsequent input feature combinations and dual-track learning model modeling.
[0023] Step A300: Combine the short-term perturbation residual, long-term drift residual, frequency band energy spectrum and environmental variable records, and epidermal growth factor dosage records into input features.
[0024] In one embodiment of this application, before combining short-term perturbation residuals, long-term drift residuals, frequency band energy spectra, environmental variable records, and epidermal growth factor dosage records into input features, it is necessary to standardize various features first. Since the units of different types of features differ—for example, the unit of short-term perturbation residuals may be kilograms, reflecting fluctuations in growth performance; the unit of temperature in environmental variable records is degrees Celsius, humidity is a percentage, and the unit of epidermal growth factor dosage is micrograms—inconsistent units can cause the model to be overly sensitive to features with large numerical ranges, affecting modeling accuracy. Therefore, the Min-Max standardization method is used. The original value of a feature is subtracted from its minimum value, and then divided by the difference between its maximum and minimum values. This maps the values of various features to the [0,1] interval. After Min-Max standardization, they are unified to the same numerical interval, ensuring that various features have equal weight and influence in model training.
[0025] Following standardization, time-series alignment is performed on various features based on their temporal order. Since short-term perturbation residuals and long-term drift residuals are calculated within sliding windows across multiple time scales, and the band energy spectrum is generated through wavelet decomposition of the growth performance time series, while environmental variable records and epidermal growth factor (EGF) dosage records are raw data collected at specific time points, it is necessary to clarify the correspondence between various features and time points. For example, for growth performance data collected daily, the short-term perturbation residuals calculated using a 3-day short-term sliding window correspond to the end of the window (e.g., the residuals of a 5-7 day window correspond to day 7), and the long-term drift residuals calculated using a 15-day long-term sliding window correspond to the end of the window (e.g., the residuals of a 1-15 day window correspond to day 15). If the band energy spectrum is obtained based on a 10-day growth performance subsequence decomposition, it corresponds to each time point within those 10 days. Environmental variable records and EGF dosage records are recorded daily, directly corresponding to the corresponding daily time points. In this way, it is ensured that each time point can be matched with a complete set of features, namely the short-term perturbation residual, long-term drift residual, frequency band energy spectrum fragment, environmental variable data of the day, and epidermal growth factor dosage data of the day corresponding to that time point.
[0026] Subsequently, feature dimension integration is performed, transforming the aligned features into vector forms recognizable by the model. For a single time node, the short-term perturbation residual and the long-term drift residual each constitute one numerical feature; if the frequency band energy spectrum is decomposed into three frequency bands, these constitute three numerical features; if the environmental variable record includes temperature and humidity, it constitutes two numerical features; and the epidermal growth factor dosage record constitutes one numerical feature. Following a fixed order—short-term perturbation residual → long-term drift residual → frequency band energy spectrum sorted from high to low frequency → environmental variable records in a preset order (temperature, humidity) → epidermal growth factor dosage record—these numerical values are concatenated into a single feature vector. Similarly, the feature vectors from multiple time nodes are arranged chronologically to form the feature matrix used as input to the dual-track learning model.
[0027] By performing standardization, temporal alignment, and dimensional integration on short-term perturbation residuals, long-term drift residuals, frequency band energy spectra, environmental variable records, and epidermal growth factor dosage records, multiple types of features are combined into structured input features. This achieves the goal of providing temporally coherent, dimensionally unified, and informationally complete input data for the dual-track learning model, ensuring that the model accurately learns the correlation between features.
[0028] Step A400: Input the input features into the dual-track learning model and output individualized causal effect estimation and dose-time-growth performance prediction mapping.
[0029] Specifically, the input features are fed into a dual-track learning model containing an individualized causal effect estimation track, a dose-time-growth performance prediction track, and a dual-track interactive track. After receiving the input features, the individualized causal effect estimation track calculates and outputs the individualized causal effect function and confidence interval through a representation learning network and historical dose-growth performance data. The dose-time-growth performance prediction track receives the input features and establishes an initial prediction result through a sequence convolutional network. Then, the initial prediction result, the individualized causal effect function, and the confidence interval are sent to the dual-track interactive track. The initial prediction result is compensated by causal consistency loss constraint. Finally, a dose-time-growth performance prediction mapping is established, and the individualized causal effect estimate and the prediction mapping are output. The specific steps are explained in detail in A410-A450.
[0030] Step A500: Perform new observations on the target object, and based on the new observations, perform a causal significance test for individualized causal effect estimation according to the corresponding sliding window, and establish a correction factor.
[0031] Optionally, after performing new observations on the target object, a local growth performance characteristic representation is established based on the new observations according to the corresponding sliding window. Then, the causal significance verification analysis of the local growth performance characteristic representation is performed using individualized causal effect estimation. The significant differences are weighted and fused according to the inverse variance to establish a correction factor. The specific steps are explained in detail in A510-A520.
[0032] Step A600: Compensate the dose-time-growth performance prediction mapping based on the correction factor and output the prediction results.
[0033] Optionally, the dimensional correspondence between the correction factor and the prediction mapping should be clarified first. The correction factor is obtained by adding new observations through causal significance testing and inverse variance weighting. It corresponds to the growth performance deviation of a specific epidermal growth factor dose and time window. Therefore, it should be matched to the same dose and time dimension in the prediction mapping. For example, if the correction factor is for the growth deviation of a 5 microgram dose on days 8-12, it needs to be located in the prediction mapping to compensate for the predicted value of the corresponding time range under that dose.
[0034] Subsequently, compensation calculations are performed. Based on the physical meaning of the correction factor, which reflects the deviation between growth performance and expectations, the predicted growth performance values for the corresponding dose and time points in the prediction mapping are adjusted. If the correction factor is a comprehensive difference value after inverse variance weighting, it can be directly added to the predicted value. For example, if the predicted weight on day 12 at a certain dose is 3.2 kg, and the correction factor is 0.08 kg, indicating that the actual growth is better than expected, then the predicted value for that time point is adjusted to 3.2 + 0.08 = 3.28 kg after compensation. This process is performed sequentially for multiple time points and dose combinations to complete the initial compensation of the prediction mapping.
[0035] After compensation, the temporal consistency of the prediction mapping needs to be verified to ensure that the prediction results at different time points and under different doses conform to the growth pattern. This is done by checking whether the changes in the predicted values at adjacent time points are stable, such as whether the daily weight gain fluctuations are within a reasonable range, and whether the overall trend of the predicted values under the same dose matches the growth stage of the target object, such as linear weight gain in infancy and slower growth rate in the growth period. If any abnormalities are found, such as a sudden change in the predicted value at a certain time point, it is necessary to backtrack the correction factor matching and compensation process, investigate and correct the problems, and ensure that the temporal logic of the compensated prediction mapping is self-consistent.
[0036] Finally, after completing the above verification, the final prediction results are output based on the compensated dose-time-growth performance prediction mapping. These predictions cover growth performance indicators, including body weight and daily weight gain, for the target subject at various future time points under different epidermal growth factor (EGF) doses, and may include prediction uncertainty information such as confidence intervals, providing data support for optimizing EGF supplementation regimens. For example, if the output prediction results show that the predicted daily weight gain for the target subject over the next 5 days is 0.035 kg, 0.048 kg, and 0.042 kg at doses of 3 μg, 5 μg, and 7 μg, respectively, this prediction can help technicians select a dosage that better suits growth needs.
[0037] By matching the correction factor with the predicted mapping dimension, performing compensation calculations, verifying temporal consistency, and outputting results, the prediction mapping was corrected using the actual deviation of newly added observations, thereby improving the accuracy and dynamic adaptability of the prediction of epidermal growth factor addition effect.
[0038] Furthermore, step A100 in the method provided in this application embodiment includes:
[0039] A110: Evaluate the data volume of the original dataset and establish evaluation results.
[0040] A120: If the evaluation result fails to meet the preset threshold, a similar acquisition instruction will be generated.
[0041] A130: After performing feature extraction on the target object according to the similarity acquisition instruction, perform similarity matching of the feature extraction results, establish an additional dataset, and compensate the original dataset according to the additional dataset.
[0042] Specifically, after establishing the original dataset, the first step is to evaluate its data volume to determine whether it can support subsequent preprocessing and modeling work. The evaluation focuses on verifying the time series length and record completeness of the data. For example, the preset threshold is set to include 25 consecutive days of growth performance time series records for the target object, corresponding 25 days of environmental variable records, and records of each dose of epidermal growth factor added during this period. During the evaluation, the number of valid days and missing records for each of the three types of records in the original dataset will be counted one by one. If the original dataset only contains 12 days of growth performance time series records and 6 days of blank environmental variable records, the evaluation result is determined to be that the preset threshold is not met.
[0043] Next, when the evaluation result fails to reach the preset threshold, a similarity acquisition instruction is generated. This instruction specifies the target object feature types to be extracted to ensure matching accuracy. Based on the similarity acquisition instruction, feature extraction is performed on the target object, covering characteristics such as variety, initial growth indicators (initial weight, body height), and environmental preferences (including suitable humidity range and light requirements). Then, the extracted features are matched with historical data. By calculating feature similarity, historical object data that meets the matching requirements is selected. Specifically:
[0044] The extracted target object features, such as variety, initial weight, initial height, suitable humidity range, and suitable light duration, are standardized to correspond with historical objects. Feature values of different dimensions are uniformly mapped to the [0,1] interval to eliminate the interference of dimensional differences on similarity calculation. This process is the same as in step A300 and will not be repeated here. Subsequently, feature vectors of the target object and historical objects are constructed. Using the target object's feature vector as a benchmark, methods such as cosine similarity or Euclidean distance are used to calculate its similarity to the feature vector of each historical object. If cosine similarity is used… Directional similarity is measured by calculating the cosine of the angle between two vectors; the closer the value is to 1, the more similar the features. If Euclidean distance is used, the numerical difference is measured by calculating the straight-line distance between vectors; the smaller the value, the more similar the features. Then, a preset similarity threshold of 0.8 is set, and historical object data with similarity values that meet the threshold requirements are selected, i.e., cosine similarity ≥ 0.8 or Euclidean distance ≤ 0.2. At the same time, abnormal historical data with similarity far below the threshold are removed. Finally, historical object data that matches the target object's features are obtained, laying the foundation for the subsequent construction of additional datasets.
[0045] For example, if the target is a certain breed of broiler chicken, after extracting its initial weight of 0.5kg and suitable humidity of 50%-60%, two sets of historical broiler chicken data of the same breed with an initial weight of 0.48-0.52kg and suitable humidity of 48%-62% are matched. These historical data are then organized into an additional dataset that includes time series records of growth performance, records of environmental variables, and records of epidermal growth factor dosage.
[0046] After constructing the supplementary dataset, it is integrated and compensated with the original dataset in chronological order. For example, if the original dataset in the above example is missing 13 days of growth performance and environmental variable records, the corresponding records for the 13 days in the supplementary dataset are added to the original dataset according to the time nodes. At the same time, the epidermal growth factor dosage records during the period are checked and added, so that the original dataset finally has 25 consecutive days of complete records of the three categories, which meets the data volume requirements for subsequent processing.
[0047] By evaluating the amount of data in the original dataset, generating similar collection instructions and extracting target object features to create an additional dataset when the threshold is not reached, and using the additional dataset to compensate for the original dataset, the effect of ensuring that the amount of data in the original dataset is sufficient and meets the requirements of subsequent data preprocessing and model input is achieved.
[0048] Furthermore, step A200 in the method provided in this application embodiment includes:
[0049] A210: The data preprocessing includes performing missing value imputation, noise reduction, and batch identification in chronological order.
[0050] Optionally, when performing time-series-based data preprocessing on the original dataset, missing value imputation is performed first. Since the original dataset contains time-series records of growth performance, environmental variables, and epidermal growth factor dosage in chronological order, missing values may appear at specific time points in any of these record types. Therefore, it is necessary to first examine the data along the timeline to locate the missing time points and their corresponding record types. During imputation, priority is given to constructing temporal relationships based on valid data adjacent to the missing time points. For example, if the weight data for day 5 is missing in the growth performance time-series records, and the weight on day 4 is 2.1 kg and on day 6 is 2.3 kg, linear interpolation will be used to calculate the weight for day 5 as 2.2 kg. If the humidity data for day 8 is missing in the environmental variable records, and the humidity for the three days before and after is consistently around 55%, the nearest neighbor mean method will be used to imput the humidity for day 8 to 55%, ensuring that the imputed data conforms to the growth and environmental change patterns over time and avoiding disruption of temporal continuity.
[0051] After imputation of missing values, denoising is performed sequentially over time. The original data may contain outliers caused by equipment errors or temporary environmental interference. These outliers can interfere with the accuracy of subsequent feature extraction; therefore, a time-series smoothing algorithm is needed to filter noise. The specific steps are as follows: First, a reasonable outlier threshold is set. For example, data in the growth performance time series that deviates from the mean of three adjacent time points by ±15% are considered outliers, and data in environmental variable records that exceed the normal fluctuation range, such as temperature fluctuations of ±2℃, are considered outliers. Then, a sliding window averaging method is used to correct the outliers. A sliding window is defined as five time points, and the mean of the valid data within that window replaces the outliers. Taking the epidermal growth factor (EGF) dosage record as an example, if the dosage on day 12 is mistakenly recorded as 10 micrograms, and the dosage at adjacent time points is around 5 micrograms, and there is no dosage adjustment record, the average value of 5 micrograms can be calculated by using a sliding window of 5 micrograms on day 11, 5 micrograms on day 13, and 5 micrograms on day 14. This value can be used to replace the abnormal data of 10 micrograms, ensuring that the noise-reduced data can truly reflect the actual situation of EGF addition and the growth trend of the target object.
[0052] Finally, batch identification is performed in chronological order, dividing the data into batches and adding labels based on the time range of data collection. Specifically, the time span of data collection is analyzed, and continuous time periods with consistent experimental conditions are divided into batches. For example, data collected from days 1 to 20 for the same batch of young animals are divided into the first batch, and data collected from days 21 to 40 for another batch of young animals are divided into the second batch. A batch label field is then added to the original dataset, labeling each data point with its corresponding batch number in chronological order. This allows for clear differentiation of data from different batches during subsequent processing, laying the foundation for accurate adaptation of individualized causal effect estimation and prediction models.
[0053] By performing preprocessing steps such as missing value imputation, noise reduction, and batch identification on the original dataset in chronological order, the system achieves the effects of restoring data integrity, filtering interference noise, and distinguishing data batches, thus providing high-quality time-series data for subsequent short-term perturbation residual calculation, wavelet decomposition, and dual-track learning model input.
[0054] Furthermore, step A400 in the method provided in this application embodiment includes:
[0055] A410: The dual-track learning model includes an individualized causal effect estimation track and a dose-time-growth performance prediction track, as well as a dual-track interactive track.
[0056] A420: Wherein, the individualized causal effect estimation trajectory, after receiving the input features, learns the target object feature embedding through a representation learning network.
[0057] A430: Based on historical dose-growth performance data, calculate the individual causal effect and uncertainty of the target object at each dose, and output the individualized causal effect function and confidence interval.
[0058] A440: The dose-time-growth performance prediction track is used to receive the input features and then perform growth performance prediction through a sequence convolutional network to establish initial prediction results.
[0059] A450: The initial prediction results, the individualized causal effect function, and the confidence interval are sent to the dual-track interactive track. The initial prediction results are compensated through the causal consistency loss constraint to establish a dose-time-growth performance prediction mapping.
[0060] Specifically, firstly, the construction of the individualized causal effect estimation track requires determining short-term perturbation residuals, long-term drift residuals, frequency band energy spectra, environmental variable records, and epidermal growth factor dosage records as input features. A representation learning network is then built to transform the input features into target object features. Historical dose-growth performance data is then integrated, and individual causal effects and uncertainties at each dose are calculated using causal inference methods. Finally, the individualized causal effect function and confidence interval are output. Secondly, the dose-time-growth performance prediction track is constructed based on the aforementioned input features. A sequence convolutional network with multi-scale convolutional kernels is built, and temporal correlation features are captured through convolution and pooling operations. Initial prediction results for growth performance at each time point and dose are output. Thirdly, the construction of the dual-track interactive track requires designing a data receiving module to synchronously acquire the output data of the first two tracks. A causal consistency loss constraint module is built to calculate the difference between the effect size and the causal effect function and determine the compensation strength. Finally, a compensation calculation unit is integrated to adjust the initial prediction results, ultimately establishing a dose-time-growth performance prediction mapping.
[0061] Next, when inputting the input features into the dual-track learning model, the process first enters the individualized causal effect estimation track. After receiving the integrated input features, this track first uses a representation learning network to deeply extract and embed the features of the target object. The input features include short-term perturbation residuals, long-term drift residuals, environmental variable features, and epidermal growth factor dosage features from the time-series features of the target object's growth performance. The representation learning network uses a multi-layer neural network structure to filter and strengthen implicit information related to the individual attributes of the target object. For example, it extracts the growth stability features of the target object from the growth performance residuals and associates the target object's adaptation features to specific environments from environmental variable records. Finally, it transforms the high-dimensional, multi-type input features into a low-dimensional, highly representative target object feature embedding vector, enabling the model to accurately capture the individual differences of the target object.
[0062] After feature embedding is completed, the individualized causal effect estimation trajectory calls upon historical dose-growth performance data to calculate the individual causal effect and uncertainty of the target object at each epidermal growth factor dose. The historical dose-growth performance data needs to be similarly matched with the feature embedding vector of the current target object to filter out historical target object data of the same type and with similar growth baseline conditions. Then, by comparing the differences in growth performance of these historical target objects at different epidermal growth factor doses—for example, the difference in average daily weight gain among different microgram dose groups within the same feeding cycle—and combining propensity score matching in statistical analysis to eliminate confounding factors such as environmental variables, the individual causal effect of the current target object at each dose is determined.
[0063] Next, after generating feature embedding vectors that accurately characterize the individual attributes of the target object, such as growth baseline and environmental adaptability, in the individualized causal effect estimation trajectory, a database corresponding to historical doses and growth performance is accessed. Historical samples with high similarity to the current target object's feature embeddings and small differences in environmental variables are selected to eliminate interference data due to excessive differences in feeding environment and initial growth state. Then, causal inference methods such as propensity score matching are used to eliminate the influence of confounding factors such as environmental fluctuations and initial weight on growth performance. The change in growth performance indicators of the target object relative to the absence of added dose at each different epidermal growth factor dose is calculated, i.e., the individual causal effect value. Subsequently, the causal effect value at each dose is repeatedly calculated multiple times using the Bootstrap sampling method to obtain the standard error of the effect value. Combined with a preset confidence level of 95%, the confidence interval of each individual causal effect value is calculated using the standard error and the critical value of the corresponding statistical distribution (such as t-distribution), thereby quantifying the uncertainty of the effect estimation. Finally, the correspondence between the target object feature vector, epidermal growth factor dose, individual causal effect value, and confidence interval is constructed into a mathematical function with epidermal growth factor dose as the independent variable and individual causal effect value as the dependent variable through piecewise linear regression or nonlinear fitting methods. This function is the individualized causal effect function that reflects the relationship between the three, and the confidence interval corresponding to each dose is output as an uncertainty reference.
[0064] Subsequently, the input features are simultaneously processed in the dose-time-growth performance prediction track. This track models the temporal attributes of the input features, such as the residual changes in the growth performance time series, the fluctuations of environmental variables over time, and the temporal distribution of epidermal growth factor (EGF) dosage, using a sequence convolutional network. By setting convolutional kernels of different sizes, the sequence convolutional network can capture the local correlation features of the input features in the time dimension. For example, a convolutional kernel with three time nodes is used to extract the impact of short-term environmental variable fluctuations on growth performance, while a convolutional kernel with seven time nodes is used to capture the medium- and long-term trends of growth performance after EGF addition. Through multi-layer convolution and pooling operations, the model gradually learns the temporal mapping relationship between the input features and growth performance, ultimately outputting initial prediction results of the target object's growth performance for different EGF dosages and different time nodes.
[0065] Finally, the dose-time-growth performance prediction track will send the generated initial prediction results, i.e., the numerical sequence of growth performance under each time-dose combination, together with the individualized causal effect function and confidence interval output by the individualized causal effect estimation track, to the dual-track interactive track. This prepares the data for subsequent compensation of the initial prediction results through causal consistency loss constraints and the establishment of an accurate dose-time-growth performance prediction mapping. This process is described in detail in steps A451-A454.
[0066] By extracting individual characteristics of the target object and calculating causal effects through the individualized causal effect estimation track, capturing temporal correlations and outputting initial predictions through the dose-time-growth performance prediction track, and receiving dual-track output data through the dual-track interactive track, the core data support is provided for subsequent optimization of prediction results and the establishment of an epidermal growth factor addition effect prediction mapping that combines individualization and temporal accuracy.
[0067] Furthermore, step A410 in the method provided in this application embodiment includes:
[0068] A411: The individualized causal effect estimation track and the dose-time-growth performance prediction track in the dual-track learning model are established through a dynamic attention gating unit. The dynamic attention gating unit adaptively adjusts the constraint weight of the individualized causal effect estimation track on the dose-time-growth performance prediction track according to the time series stability of the target object.
[0069] Specifically, during the operation of the dual-track learning model, the dynamic attention gating unit first needs to obtain the time-series stability index of the target object. The calculation basis of this index comes from the preprocessed growth performance time-series records. Specifically, it extracts the statistical characteristics of the short-term perturbation residuals in the growth performance time series, and quantifies the time series stability by calculating the variance or coefficient of variation of the short-term perturbation residuals. The smaller the variance, the smaller the fluctuation of the target object's growth performance in the short term, and the more stable the time series; the larger the variance, the more frequent the fluctuation of growth performance, and the weaker the time series stability.
[0070] Next, the dynamic attention gating unit adaptively adjusts the constraint weights of the individualized causal effect estimation track on the dose-time-growth performance prediction track according to the preset weight mapping rules. When the time series stability is high, it means that the sequence convolutional network on which the dose-time-growth performance prediction track depends can accurately capture changes in growth performance through stable temporal patterns. In this case, the constraint weights of the individualized causal effect estimation track need to be reduced to decrease its interference with the prediction track, allowing the prediction track to output results more autonomously based on temporal features. When the time series stability is low, the sequence convolutional network can hardly achieve accurate prediction based solely on time series data with large fluctuations. The constraint weights of the individualized causal effect estimation track need to be increased to allow the individualized causal effect function output by this track to participate more in the prediction process, in order to correct the bias caused by temporal fluctuations.
[0071] Subsequently, the dynamic attention gating unit, based on adjusted constraint weights, establishes an interactive bridge between the individualized causal effect estimation track and the dose-time-growth performance prediction track. On one hand, the gating unit receives the individualized causal effect function and confidence interval output from the individualized causal effect estimation track, extracting the causal effect information corresponding to the current prediction time point and epidermal growth factor dose. On the other hand, it receives the initial prediction results generated by the dose-time-growth performance prediction track and analyzes the fit between these results and the causal effect information. Then, according to the constraint weights, the causal effect information is integrated into the modeling process of the prediction track according to the weight ratio. If the weight is high, the causal effect function will directly participate in the intermediate correction of the initial prediction results, such as correcting the convolution kernel parameters of the sequence convolutional network to make the prediction more closely match the individual causal patterns of the target object. If the weight is low, the causal effect information is only used as an auxiliary reference to ensure the temporal modeling of the prediction track is dominant. Taking cell culture as an example, when the cell growth environment is stable, the gating unit allows the predicted trajectory to mainly capture the temporal changes in cell density through a sequence convolutional network; when the environmental pH value fluctuates, causing the cell growth timeline to become unstable, the gating unit increases the weight of the causal trajectory and uses the inherent proliferation effect of epidermal growth factor on this type of cell to correct any possible deviations in the predicted trajectory.
[0072] By first calculating the time series stability of the target object using a dynamic attention gating unit, then adaptively adjusting the constraint weights of the individualized causal effect estimation orbit on the dose-time-growth performance prediction orbit, and finally achieving dual-track interaction, the dual-track learning model is adapted to different growth time series states of the target object, thereby improving the accuracy of dose-time-growth performance prediction mapping.
[0073] Furthermore, such as Figure 2 As shown, step A450 in the method provided in this application embodiment includes:
[0074] A451: Aggregate the initial prediction results into the effect size at the corresponding dose.
[0075] A452: Use the effect size and the individualized causal effect function to conduct a difference analysis and establish the difference analysis results.
[0076] A453: Based on the difference analysis results and confidence intervals, perform weighted fusion to establish the compensation strength.
[0077] A454: Perform initial prediction result compensation based on the stated compensation intensity.
[0078] Specifically, after receiving the initial prediction results from the dose-time-growth performance prediction track, the dual-track interactive track first groups the initial prediction results according to the epidermal growth factor (EGF) dosage. The initial prediction results are growth performance prediction values corresponding to different time points and dosages. Multiple time point prediction values under the same dosage collectively reflect the overall impact of that dosage on growth performance and need to be converted into the corresponding dose effect size through aggregation. During aggregation, an appropriate statistical method is selected based on the type of growth performance indicator. If the growth performance indicator is daily weight gain, the arithmetic mean of the initial predicted daily weight gain at all time points under the same dosage is calculated; if it is cumulative weight, the cumulative value of the initial predicted weight within a specified period under the same dosage is calculated. For example, for the 5 microgram dose group of epidermal growth factor, the initial prediction results included the average daily weight gain data for 7 consecutive days, which were 0.03 kg, 0.04 kg, 0.03 kg, 0.05 kg, 0.04 kg, 0.03 kg, and 0.04 kg, respectively. Calculated by the arithmetic mean, the effect size corresponding to this dose was (0.03+0.04+0.03+0.05+0.04+0.03+0.04) / 7≈0.037 kg, thus completing the conversion from the initial prediction results to the effect size.
[0079] Next, after obtaining the effect size for each dose, a difference analysis was performed using the individualized causal effect function output from the individualized causal effect estimation trajectory. The individualized causal effect function clearly defines the individual causal effect value for each epidermal growth factor dose. The difference analysis calculates the difference between the effect size and the individual causal effect value at the same dose, and then uses the confidence interval corresponding to the individualized causal effect function to determine the significance of this difference. For example, at a 5 μg dose of epidermal growth factor, the individual causal effect value output by the individualized causal effect function is 0.04 kg, with a corresponding 95% confidence interval of 0.035 kg to 0.045 kg. The effect size for this dose is 0.037 kg, resulting in a difference of 0.003 kg. This difference falls within the confidence interval, indicating a small difference between the initial prediction and the individual causal law. If the effect size is 0.032 kg, the difference is 0.008 kg, close to the lower limit of the confidence interval, indicating a relatively significant difference that requires close attention to subsequent compensation.
[0080] Then, based on the difference analysis results and confidence intervals, the compensation strength is determined through weighted fusion. The width of the confidence interval directly reflects the uncertainty of the individualized causal effect estimate; the narrower the interval, the higher the reliability of the individual causal effect value, and it is assigned a higher weight in the weighting. The magnitude of the difference in the difference analysis results reflects the degree of deviation between the initial prediction and the individual causal law; the larger the difference, the more the initial prediction needs to be corrected, and it is also assigned a higher weight in the weighting. In practice, the confidence interval width is first converted into a weight coefficient (if the interval width and weight are negatively correlated), then the difference magnitude is converted into a deviation coefficient (if the difference and deviation coefficient are positively correlated), and finally, the weight coefficient and the deviation coefficient are multiplied to obtain the compensation strength. For example, if the difference in dosage is 0.008 kg (deviation coefficient of 0.8) and the confidence interval width is 0.01 kg (weighting coefficient of 1.0), then the compensation intensity is 0.8 × 1.0 = 0.8. If the difference in dosage is 0.003 kg (deviation coefficient of 0.3) and the confidence interval width is 0.02 kg (weighting coefficient of 0.5), then the compensation intensity is 0.3 × 0.5 = 0.15, thus achieving differentiated setting of compensation intensity.
[0081] Finally, compensation is performed on the initial prediction results output from the dose-time-growth performance prediction trajectory according to this intensity. During the compensation calculation, the initial predicted value at each time point is added to the result of the compensation intensity multiplied by the difference, causing the initial predicted value to converge towards the individualized causal effect value, with the degree of convergence determined by the compensation intensity. For example, at a dose of 5 micrograms of epidermal growth factor, if the initial predicted daily weight gain at a certain time point is 0.03 kg, the compensation intensity is 0.8, and the difference is 0.008 kg, then the compensated predicted value is 0.03 + (0.8 × 0.008) = 0.0364 kg; if the initial predicted value at another time point is 0.05 kg, the compensation intensity is 0.15, and the difference is 0.003 kg, then the compensated predicted value is 0.05 + (0.15 × 0.003) = 0.05045 kg, ensuring that the prediction results at each time point are reasonably corrected according to individual causal laws.
[0082] By aggregating the initial prediction results into the corresponding dose effect size, performing difference analysis using the effect size and individualized causal effect function, establishing the compensation strength by combining the difference analysis results with the confidence interval in a weighted fusion, and performing compensation based on the initial prediction results according to the compensation strength, the effect of making the dose-time-growth performance prediction mapping fit the individualized causal law of the target object is achieved, thereby improving the prediction accuracy of the epidermal growth factor addition effect.
[0083] Furthermore, step A450 in the method provided in this application embodiment includes:
[0084] A455: .
[0085] A456: Among them, The compensated dose-time-growth performance predictions, For the initial prediction results, x represents the feature vector of the target object, d is the dosage of epidermal growth factor, and t is the time node. Characterizing individualized causal effect estimation, For effect size, The uncertainty in the estimation of individualized causal effects is determined by confidence intervals. The uncertainty of the initial prediction result, This is the time weighting function.
[0086] In one embodiment, the compensated dose-time-growth performance prediction is calculated. First, the initial prediction results, individualized causal effect estimates, effect sizes, and the uncertainties of individualized causal effect estimates and initial prediction results are clearly defined. Simultaneously, a time weighting function is determined to reflect the weight differences of different time points in compensation. Specifically: the initial prediction results are established by receiving input features from the dose-time-growth performance prediction track and performing growth performance prediction through a sequence convolutional network; the individualized causal effect estimates are obtained by receiving input features from the individualized causal effect estimation track, learning the target object feature embedding through a representation learning network, and then calculating the individual causal effect of the target object at each dose by combining historical dose-growth performance data; the effect size is formed by aggregating the initial prediction results according to the corresponding dose; the uncertainty of the individualized causal effect estimates is derived by calculating the standard error of the individual causal effect value and combining it with a preset confidence level, and can be determined by the confidence interval; the uncertainty of the initial prediction results is calculated using statistical methods by analyzing the fluctuations of the initial prediction results at different time points and at different doses; the time weighting function is set according to the importance of different time points in compensation, generally assigning higher weights to time points closer to the current time point to reflect their different impacts on compensation.
[0087] Next, calculate the weighting coefficients. This step follows the logic that the lower the uncertainty of the result, the higher the reliability, and therefore the greater the weight should be allocated. This allows the individualized causal effect estimate and the initial prediction result to be distributed in the compensation process according to their respective reliability levels. If the uncertainty of the individualized causal effect estimate is lower, the coefficient will be closer to 1, meaning that the individualized causal effect estimate has a larger share in the compensation; conversely, if the uncertainty of the initial prediction result is lower, the coefficient will be more biased towards the value determined by the uncertainty of the initial prediction result.
[0088] Then, the difference between the individualized causal effect estimate and the effect size is calculated. This difference reflects the degree of deviation between the initial predictions, aggregated into effect sizes, and the individualized causal patterns, and is the core source of bias in subsequent compensation. Then, the weighting coefficients, the difference term, and the time weighting function are multiplied to obtain the compensation term: The time weighting function assigns different weights based on the characteristics of each time point, such as the closer the time is to the present, the stronger its predictive value, so that the compensation strength of different time points matches their temporal importance.
[0089] Finally, the initial prediction result is added to the compensation term to obtain the compensated dose-time-growth performance prediction. For example, when a target object is at a specific dose and time point, the initial prediction result may be biased due to fluctuations in the growth time series. However, if the individualized causal effect estimation is more reliable (i.e., the uncertainty is smaller), the compensation term will push the initial prediction value to adjust towards the individualized causal effect, so that the final prediction result not only incorporates the dynamic characteristics of time series modeling but also conforms to the inherent causal response law of the target object.
[0090] By sequentially determining parameters, calculating weight allocation coefficients and difference terms, generating compensation terms and adding them to the initial prediction results, the compensation of the initial prediction results is achieved by utilizing the causal consistency loss constraint. This results in the dose-time-growth performance prediction values simultaneously conforming to temporal dynamics and individualized causal laws, thereby improving the accuracy of epidermal growth factor addition effect prediction.
[0091] Furthermore, step A500 in the method provided in this application embodiment includes:
[0092] A510: Establish a representation of local growth performance characteristics based on the newly added observations using the sliding window.
[0093] A520: Using the individualized causal effect estimation to perform causal significance verification analysis on the local growth performance characteristics, significant differences are weighted and fused according to inverse variance to establish a correction factor.
[0094] Optionally, after performing new observations of the target object, the growth performance data in the new observations is first segmented and features extracted based on the sliding window previously used to process the growth performance time series records. This sliding window, consistent with the calculation of short-term disturbance residuals and long-term drift residuals (three short-term time nodes and fifteen long-term time nodes), is used to establish a local growth performance feature representation. The new observations contain the latest growth performance time series records of the target object. According to the time scale of the sliding window, the new data is truncated into several continuous local segments, each segment corresponding to a sliding window. Subsequently, for the growth performance data within each window, statistical features reflecting the local growth state are extracted, such as the mean, maximum value, and fluctuation amplitude of growth performance indicators within the window. These features are combined to form the local growth performance feature representation corresponding to each window. For example, when adding 10 days of livestock and poultry weight observation data, using a 3-day short-term sliding window, the data can be divided into three local segments: days 1-3, days 4-6, and days 7-10. If the last segment of the window is less than 3 days, it is merged into one window. The average daily weight and weight fluctuation range of each segment are calculated separately to form three sets of local growth performance characteristics, ensuring that the local characteristics are consistent with the previous time-series processing logic.
[0095] Next, after establishing the representation of local growth performance characteristics, the individualized causal effect function output from the aforementioned individualized causal effect estimation trajectory is used to perform causal significance verification analysis on each local feature representation. First, based on the time node and epidermal growth factor dosage corresponding to each local growth performance feature representation, the expected value of the growth performance effect that the target object should have at that dosage is extracted from the individualized causal effect function, i.e., the growth performance reference range based on individual causal laws. Then, the actual growth performance indicators in the local growth performance feature representation, such as the local average daily weight, are compared with the expected value. Statistical tests (such as t-tests) are used to determine whether the difference between the two is significant. If the actual value falls outside the confidence interval corresponding to the expected value, a significant difference is determined, and the specific value of the difference is recorded. If the actual value is within the confidence interval, the difference is determined to be insignificant and is not included in subsequent calculations. At the same time, during the verification process, the variance corresponding to each significant difference is recorded. This variance is calculated from the fluctuation of the local growth performance data and reflects the reliability of the difference results. The smaller the fluctuation, the smaller the variance, and the more reliable the difference. For example, a local feature corresponds to a 5 μg dose of epidermal growth factor. The individualized causal effect function gives an expected daily weight gain of 0.04 kg with a confidence interval of 0.035 kg to 0.045 kg. However, the actual daily weight gain for this local feature is 0.032 kg, which exceeds the lower limit of the confidence interval and is therefore considered a significant difference with a value of -0.008 kg. Furthermore, the local data has small fluctuations, and the calculated variance is 0.000064.
[0096] Finally, a correction factor is established using the inverse variance weighted fusion method. The core logic of inverse variance weighting is that the smaller the variance of a significant difference, the higher the reliability of the result, and it should be given a higher weight in the fusion process. Specifically, the weight value of each significant difference is first calculated; the weight is the reciprocal of the variance of that difference, i.e., weight = 1 / variance. Then, each significant difference is multiplied by its corresponding weight to obtain the weighted difference value. Subsequently, all weighted difference values are summed and divided by the sum of all weight values to obtain the fused comprehensive difference value. This comprehensive difference value is the correction factor used to compensate for the dose-time-growth performance prediction mapping. For example, after validation analysis, two significant differences are obtained: the first difference is -0.008 kg, variance 0.000064, weight = 1 / 0.000064 = 15625; the second difference is -0.006 kg, variance 0.0001, weight = 1 / 0.0001 = 10000. After weighted fusion, the overall difference value = (-0.008×15625+(-0.006)×10000)÷(15625+10000) = (-125-60)÷25625≈-0.00722 kg, which is the final established correction factor.
[0097] By establishing local growth performance characteristics from newly added observations according to the corresponding sliding window, performing causal significance verification analysis on local characteristics using individualized causal effect estimation, and merging significant differences by inverse variance weighting, a correction factor that can reflect the deviation between newly added observations and historical causal patterns is established. This achieves the effect of providing a precise adjustment basis for subsequent compensation dose-time-growth performance prediction mapping and improving the dynamic adaptability of epidermal growth factor addition effect prediction.
[0098] Furthermore, step A520 in the method provided in this application embodiment includes:
[0099] A521: If the causal significance test analysis shows a significant result, then a compensation instruction is generated.
[0100] A522: According to the compensation instruction, significant differences are weighted and fused by inverse variance to establish a correction factor.
[0101] A523: If the causal significance verification analysis fails, an anomaly warning is generated, and anomaly reporting management is performed based on the anomaly warning.
[0102] Optionally, the criteria for passing and failing the validation should first be clearly defined. These criteria are based on the confidence intervals of the individualized causal effect estimation output. Specifically, the expected growth performance value and confidence interval for the corresponding epidermal growth factor dose are extracted from the individualized causal effect function to represent local growth performance characteristics. For example, if the expected daily weight gain at a certain dose is 0.04 kg, the 95% confidence interval is 0.035 kg to 0.045 kg. The actual growth performance indicators (such as local daily weight gain) in the local growth performance characteristic representation are then compared with the expected value and confidence interval. If the actual indicator falls within the confidence interval, and statistical methods such as t-tests verify that the difference is not statistically significant (i.e., P-value > 0.05), the causal significance validation is considered passed. If the actual indicator exceeds the confidence interval, or statistical tests show that the difference is statistically significant (P-value ≤ 0.05), the causal significance validation is considered failed.
[0103] Next, when the causal significance verification analysis results in a significant value, the system will automatically generate a compensation instruction. The compensation instruction must include key information: first, the time window and epidermal growth factor dosage corresponding to the local growth performance feature, ensuring that subsequent processing accurately matches the corresponding dimension of the prediction mapping; second, the significant difference between each verified local feature and the expected value (e.g., the difference between the actual daily weight gain of 0.042 kg and the expected 0.04 kg is 0.002 kg); and third, the variance corresponding to each significant difference (e.g., the variance calculated from the local data fluctuation is 0.00004), providing data support for inverse variance weighted fusion. After generating the compensation instruction, the system will calculate the weights according to the difference values and variances in the instruction, following the inverse variance weighting rule (weight = 1 / variance, the smaller the variance, the higher the weight). Then, the difference values are multiplied by their corresponding weights, summed, and divided by the sum of all weights to obtain the fused comprehensive difference value, which is the correction factor. For example, if two locally verified differences are 0.002 kg (variance 0.00004, weight 25000) and 0.003 kg (variance 0.00009, weight 11111.11), the weighted fusion correction factor is approximately 0.00231 kg.
[0104] If the causal significance test fails, an anomaly warning is immediately generated. This warning must clearly indicate key anomaly information, including the time point corresponding to the failed local growth performance characteristic, the corresponding epidermal growth factor dosage, the specific difference between the actual growth performance index and the expected value, and the p-value of the statistical test result to ensure accurate anomaly localization. After generating the anomaly warning, anomaly reporting management is implemented based on the warning information: on the one hand, anomaly details are recorded in the corresponding system log, including the time of anomaly occurrence, involved parameters, and difference data, to facilitate subsequent traceability analysis; on the other hand, relevant operators are notified through preset notification mechanisms, including system pop-ups and email reminders, while triggering preliminary investigation suggestions, such as checking whether the data acquisition equipment is malfunctioning and whether the actual dosage of epidermal growth factor added is consistent with the record.
[0105] By first clarifying the criteria for determining the significance of causality, then generating compensation instructions and establishing correction factors for the results that pass the verification, and generating abnormal warnings and implementing abnormal reporting management for the results that fail the verification, the system achieves the effect of ensuring the reliability of the correction factor construction, timely identifying and handling abnormal situations that affect the accuracy of prediction, and maintaining the stability of the epidermal growth factor addition effect prediction process.
[0106] Furthermore, step A600 in the method provided in this application embodiment includes:
[0107] A610: Establish a calibration effect curve, and use the calibration effect curve as a comparison standard to perform a deviation comparison of the prediction results, and establish a deviation comparison result, which includes deviation nodes and deviation magnitudes.
[0108] A620: Manage deviation anomaly reporting based on the deviation comparison results.
[0109] Optionally, after outputting the prediction results, a calibration curve can be established. This curve is based on actual growth performance data of similar target objects under different epidermal growth factor doses in history. Those skilled in the art can generate this curve using nonlinear regression statistical fitting, representing the growth response pattern of the target object under ideal conditions. For example, by collecting daily average weight gain data from multiple batches of healthy piglets at 3 micrograms, 5 micrograms, and 7 micrograms of epidermal growth factor doses, a smooth curve of daily average weight gain over time corresponding to each dose can be fitted and used as a standard for subsequent comparison.
[0110] Next, a deviation comparison is performed between the predicted results and the calibration curve. For each epidermal growth factor dose and each time point in the predicted results, the predicted growth performance value is compared with the calibration value of the calibration curve at the same dose and time point. The difference is calculated to obtain the deviation amplitude, and the corresponding deviation point is recorded. For example, at a dose of 5 micrograms, the predicted average daily weight gain on day 7 is 0.045 kg, while the calibration curve at the same dose on day 7 has a calibration value of 0.05 kg. In this case, the deviation amplitude is -0.005 kg, and the deviation point is day 7. After performing this calculation for all predicted time-dose combinations, the results are summarized to form a deviation comparison result including the deviation point and the deviation amplitude.
[0111] Subsequently, deviation anomaly reporting management is implemented based on the deviation comparison results. Those skilled in the art can pre-set deviation thresholds according to actual experimental or project needs; for example, the daily average weight gain deviation threshold is ±0.01 kg. If the deviation exceeds this threshold, it is determined to be an abnormal deviation and an anomaly report is triggered. When an anomaly is reported, information such as the deviation node, deviation magnitude, and corresponding dosage is recorded and transmitted to relevant personnel via system notification or report. Simultaneously, the processing procedure is initiated based on the deviation severity, the magnitude of the deviation, and the number of nodes involved. Minor deviations prompt further observation, while severe deviations investigate influencing factors in the prediction model or the actual growth process.
[0112] By establishing a calibration curve as a comparison standard, performing deviation comparison of prediction results, and managing anomaly reporting based on the results, the system effectively identifies deviations between prediction results and historical ideal patterns, ensuring the reliability of epidermal growth factor addition effect prediction.
[0113] In summary, the method for predicting the effect of epidermal growth factor addition based on growth performance data provided in this application has the following technical advantages:
[0114] This application establishes a raw dataset by collecting and structuring time-series records of growth performance, environmental variables, and epidermal growth factor (EGF) dosage data of the target object. Through time-series data preprocessing, calculation of short-term perturbation residuals and long-term drift residuals within a multi-timescale sliding window, and wavelet decomposition of the growth performance time-series records, short-term perturbation residuals, long-term drift residuals, and frequency band energy spectra are obtained. These are combined with environmental variable records and EGF dosage records to form a dual-track learning model with input features. Individualized causal effect estimation and dose-time-growth performance prediction mapping are calculated. A correction factor is established by combining newly added observations of the target object with a causal significance test of the individualized causal effect estimation according to the corresponding sliding window. Based on the correction factor, the dose-time-growth performance prediction mapping is compensated, thereby accurately predicting the impact of different EGF dosages on the growth performance of the target object at different time points. This makes the prediction results of EGF addition effects more accurate and reliable, meeting the needs of precise assessment and individualized regulation of target object growth performance, and achieving the technical effect of accurate and comprehensive prediction of EGF addition effects.
[0115] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0116] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for predicting the effect of epidermal growth factor supplementation based on growth performance data, characterized in that, The method includes: Data on the target object is collected and structured to establish an original dataset, which includes time-series records of growth performance, records of environmental variables, and records of epidermal growth factor dosage in chronological order. After performing time-series data preprocessing on the original dataset, the growth performance time series records in the preprocessing results are recorded in sliding windows of multiple time scales to calculate short-term perturbation residuals and long-term drift residuals, and wavelet decomposition of the growth performance time series records in the preprocessing results is performed to construct the frequency band energy spectrum. The short-term perturbation residual, long-term drift residual, frequency band energy spectrum and environmental variable records, and epidermal growth factor dosage records are combined into input features; The input features are fed into a dual-track learning model, which outputs individualized causal effect estimates and dose-time-growth performance prediction mappings. Perform new observations on the target object, and based on the new observations, conduct a causal significance test for individualized causal effect estimation according to the corresponding sliding window, and establish a correction factor; The prediction results are output based on the correction factor compensation of the dose-time-growth performance prediction mapping.
2. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 1, characterized in that, The step of inputting the input features into a dual-track learning model and outputting individualized causal effect estimates and dose-time-growth performance prediction mappings includes: The dual-track learning model includes an individualized causal effect estimation track and a dose-time-growth performance prediction track, as well as a dual-track interactive track; The individualized causal effect estimation trajectory learns the target object feature embedding through a representation learning network after receiving the input features; Based on historical dose-growth performance data, the individual causal effect and uncertainty of the target object at each dose are calculated, and the individualized causal effect function and confidence interval are output. The dose-time-growth performance prediction track is used to receive the input features and then perform growth performance prediction through a sequence convolutional network to establish an initial prediction result. The initial prediction results, the individualized causal effect function, and the confidence interval are sent to the dual-track interactive track. The initial prediction results are compensated by causal consistency loss constraints, and a dose-time-growth performance prediction mapping is established.
3. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 2, characterized in that, The individualized causal effect estimation track and the dose-time-growth performance prediction track in the dual-track learning model interact through a dynamic attention gating unit. The dynamic attention gating unit adaptively adjusts the constraint weight of the individualized causal effect estimation track on the dose-time-growth performance prediction track according to the time series stability of the target object.
4. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 2, characterized in that, The process of compensating for initial prediction results through causal consistency loss constraints and establishing a dose-time-growth performance prediction mapping includes: The initial prediction results are aggregated into the effect size at the corresponding dose; The effect size and the individualized causal effect function were used to conduct a difference analysis, and the difference analysis results were established. Based on the difference analysis results and confidence intervals, a weighted fusion is performed to establish the compensation strength. Compensation is performed based on the initial prediction results according to the compensation intensity.
5. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 4, characterized in that, The compensation for the initial prediction result is performed using the causal consistency loss constraint, calculated as follows: ; in, The compensated dose-time-growth performance predictions, For the initial prediction results, x represents the feature vector of the target object, d is the dosage of epidermal growth factor, and t is the time node. Characterizing individualized causal effect estimation, For effect size, The uncertainty in the estimation of individualized causal effects is determined by confidence intervals. The uncertainty of the initial prediction result, This is the time weighting function.
6. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 1, characterized in that, The newly added observations of the target object are used to perform a causal significance test on the individualized causal effect estimation based on the newly added observations according to the corresponding sliding window, and a correction factor is established, including: A local growth performance characteristic representation is established based on the newly added observations using the sliding window described above; The causal significance of the local growth performance characteristics represented by the individualized causal effect estimation is verified by analyzing the causal significance of the individualized causal effect estimation. The significant differences are then weighted by inverse variance to establish a correction factor.
7. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 6, characterized in that, The causal significance verification analysis of the local growth performance feature representation using the individualized causal effect estimation further includes: If the causal significance test results in a significant pass, a compensation instruction is generated. Based on the compensation instructions, significant differences are weighted and fused according to inverse variance to establish a correction factor; If the causal significance test analysis fails, an anomaly warning is generated, and anomaly reporting management is performed based on the anomaly warning.
8. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 1, characterized in that, The data preprocessing includes performing missing value imputation, noise reduction, and batch identification in chronological order.
9. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 1, characterized in that, The output prediction result includes: Establish a calibration effect curve, and use the calibration effect curve as a comparison standard to perform a deviation comparison of the prediction results, and establish a deviation comparison result, which includes deviation nodes and deviation magnitude; Based on the deviation comparison results, deviation anomaly reporting management is performed.
10. The method for predicting the effect of epidermal growth factor addition based on growth performance data as described in claim 1, characterized in that, After establishing the original dataset, the following steps are included: The original dataset is evaluated for its data volume, and evaluation results are established. If the evaluation result fails to meet the preset threshold, a similar collection instruction will be generated; After performing feature extraction on the target object according to the similarity acquisition instruction, similarity matching of the feature extraction results is performed to establish an additional dataset, and the original dataset is compensated according to the additional dataset.
Citation Information
Patent Citations
Deep learning-based medicine intelligent management and prediction analysis method
CN120636851A
Myopia ocular predictive technology and integrated characterization system
US12274503B1