An artificial intelligence-based fermentation strain yield prediction method

By aligning process observations with material ledger data, a conversion efficiency characterization sequence is generated and combined with strain characteristics. This trains a yield prediction model and calculates a reliability coefficient, solving the problem of inaccurate strain yield prediction in biofermentation and achieving more reliable production management decision support.

CN121812014BActive Publication Date: 2026-06-02汉中天然谷生物科技股份有限公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
汉中天然谷生物科技股份有限公司
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately predict strain yield during bio-fermentation, which makes it difficult to form reliable strain screening and process intervention decisions in production management. This can easily lead to situations where monitoring indicators appear normal, but the final yield deviates from expectations.

Method used

By aligning the process observation sequence of historical batches with the material ledger sequence, a conversion efficiency characterization sequence is generated and bound to the strain feature vector. The yield prediction model is trained, and the consistency parameters of the ledger-observation volume trajectory and the steady-state proportion parameter of the conversion contribution are calculated to generate a decision credibility coefficient, thereby realizing online prediction and recording of new batches.

Benefits of technology

It improves the reliability and comparability of strain yield prediction, and can proactively identify whether the prediction basis is reliable when there are batch differences and fluctuations in operating conditions, avoid misjudgment, and enhance the certainty and feasibility of strain screening and process intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121812014B_ABST
    Figure CN121812014B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on artificial intelligence's fermentation bacterial yield prediction method, specifically relates to biological fermentation production field, for solving the concentration prediction in prior art is difficult to directly map yield, resulting in monitoring index normal but final yield deviates from expectation, cannot support bacterial reliable comparison and production management decision problem;By aligning the historical batch process observation sequence with the material account sequence to generate a conversion efficiency characterization sequence to form a yield learning sample to train a yield prediction model, generate a conversion efficiency characterization sequence on a batch to input the model to obtain a yield prediction value, simultaneously calculate the account observation volume trajectory consistency parameter and the conversion contribution steady-state proportion parameter input coefficient analysis model to obtain the decision confidence coefficient, based on the decision confidence coefficient, the yield prediction value generates the yield risk judgment result, realizes new batch online prediction record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bio-fermentation production, and more specifically, to a method for predicting the yield of fermentation strains based on artificial intelligence. Background Technology

[0002] In the research and development and production practice of bio-fermentation, the selection of microbial strains and the setting of process parameters often determine whether the input raw materials can be efficiently transformed into the target product, thereby affecting production planning, cost control, and delivery stability. Due to the dynamic nature of the fermentation process, batch variations, and operational disturbances, companies hope to form reliable judgments on the yield trend of microbial strains while there is still room for production adjustments. This information can support a series of management decisions, such as strain selection, process scale-up, process intervention, and resource allocation, rather than passively confirming the merits or demerits only through backtesting results after fermentation is completed.

[0003] Among existing published patents, for example, the published invention patent application "A Method for Predicting Bio-fermentation Data Based on Deep Neural Networks" (application number 202110528150.X) proposes to predict the solution composition during fermentation by combining spectral data with a neural network model, thereby avoiding the monitoring lag caused by the time-consuming direct measurement. It also outputs the concentration results of specified components through feature extraction and regression prediction for process monitoring and anomaly detection. This type of technical approach can alleviate the problem of slow measurement, but it is often difficult to implement directly in production management focused on strain yield. This is because yield is not a simple mapping of the concentration of a single component at a certain moment, but a comprehensive result of the long-term coupling of substrate utilization, metabolic distribution, inhibition effects, and process dynamics. The same concentration change may be caused by completely different reasons, which may indicate that the cell growth and metabolism are in good condition, or it may be an apparent phenomenon caused by feeding strategies, mass transfer limitations, or by-product accumulation. Therefore, in practical applications, it is easy to encounter situations where monitoring indicators appear normal, but the final yield deviates from expectations. This makes it difficult for the prediction results to support reliable comparisons between different strains and for actionable process decisions. On-site, it is still necessary to rely on experience and trial and error to determine whether to adjust the strategy or change the strain, making it difficult to form a reusable strain yield prediction and management capability.

[0004] To address the aforementioned problems, a technical solution is provided. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an artificial intelligence-based method for predicting the yield of fermentation strains. This method generates a conversion efficiency characterization sequence by aligning historical batch process observation sequences with material ledger sequences, forming a yield learning sample to train the yield prediction model. A conversion efficiency characterization sequence is generated on each batch and input into the model to obtain the predicted yield value. Simultaneously, the consistency parameter of the ledger-observation volume trajectory and the steady-state proportion parameter of the conversion contribution are calculated and input into the coefficient analysis model to obtain the decision reliability coefficient. Based on the decision reliability coefficient, a yield risk assessment result is generated for the predicted yield value, enabling online prediction and recording of new batches, thereby solving the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] S1: Align and unify the historical batch process observation sequence and material ledger sequence according to batch identifier, unify strain identifier and product caliber, and generate batch alignment data package and corresponding yield label.

[0008] S2: Based on batch-aligned data packages, the process observation sequence and the material ledger sequence are combined and converted to generate a conversion efficiency characterization sequence. The conversion efficiency characterization sequence is then bound to the strain feature vector to form a yield learning sample.

[0009] S3: Train the yield prediction model using yield learning samples, and generate a transformation efficiency characterization sequence on the batch to be evaluated according to step S2. Input the sequence and the strain feature vector into the yield prediction model to obtain the predicted yield value. At the same time, calculate the apparent volume trajectory consistency parameter and the transformation contribution steady-state ratio parameter. Input the apparent volume trajectory consistency parameter and the transformation contribution steady-state ratio parameter into the pre-trained coefficient analysis model to obtain the decision credibility coefficient. Based on the decision credibility coefficient, generate the yield risk judgment result for the predicted yield value.

[0010] S4: For the new batch of online access process observation sequence and material account sequence, generate conversion efficiency characterization sequence according to step S2 and input it into the yield prediction model together with the strain feature vector. At the same time, calculate the decision confidence coefficient according to step S3 and output the yield risk judgment result and prediction record associated with the batch identifier.

[0011] Furthermore, in step S1, the process observation sequence and material ledger sequence are extracted and aligned to a unified timeline using batch identifiers. Incomplete batches are removed, and units are standardized, data is cleaned, strain identifiers and product calibers are standardized to generate batch alignment data packages and yield labels.

[0012] Furthermore, in step S2, the target product concentration trajectory, observed liquid volume trajectory, cumulative substrate input trajectory, and net liquid volume cumulative trajectory are extracted from the batch aligned data package. The harmonic volume is calculated by harmonic averaging of observed volume and book volume, and the instantaneous product generation rate and substrate consumption rate are calculated by point-by-point difference.

[0013] Furthermore, in step S2, a three-dimensional conversion efficiency characterization sequence is generated, including instantaneous yield, cumulative conversion rate and efficiency deviation rate. Instantaneous yield and efficiency deviation rate are divided by the sum of substrate consumption rate and small positive numbers. The efficiency deviation rate is introduced with a theoretical stoichiometric coefficient. The conversion efficiency characterization sequence is flattened and concatenated with the feature vector of the uniquely thermocoded strain, and paired with the yield label to form a yield learning sample.

[0014] Furthermore, in step S3, the yield prediction model is trained using yield learning samples. For the batch to be evaluated, a transformation efficiency characterization sequence is generated in the manner of step S2 and concatenated with the strain feature vector before being input into the yield prediction model, and the predicted yield value is output.

[0015] Furthermore, in step S3, the liquid phase volume trajectory is obtained from the material ledger sequence through cumulative operation. The observed liquid phase volume trajectory is obtained by converting and smoothing from the process observation sequence. The relative absolute deviation is calculated point by point and averaged, and then subtracted by 1 to obtain the consistency parameter of the ledger volume trajectory. The instantaneous yield of the conversion efficiency characterization sequence is taken to calculate the batch median baseline. The segment is divided by change point detection and the segment above the baseline is selected as the yield contribution period. The steady-state segment of the process range is judged in combination with the process observation sequence, and the steady-state segment duration is calculated to obtain the conversion contribution steady-state proportion parameter.

[0016] Furthermore, in step S3, the consistency parameter of the apparent volume trajectory and the steady-state proportion parameter of the conversion contribution are input into the coefficient analysis model, and the decision credibility coefficient is output. Based on the comparison between the yield prediction value and the risk threshold and the level of the decision credibility coefficient, the yield risk judgment result is generated in a graded manner and associated with the batch identifier.

[0017] Furthermore, in step S4, for each update of the data in the online access process observation sequence and material ledger sequence of the new batch, a batch alignment data package is dynamically constructed. The conversion efficiency characterization sequence is generated according to step S2 and concatenated with the strain feature vector before being input into the yield prediction model, and the yield prediction value is output.

[0018] Furthermore, in step S4, the decision credibility coefficient is calculated simultaneously as in step S3. Based on the comparison between the yield prediction value and the risk threshold and the level of the decision credibility coefficient, the yield risk judgment result is generated. The yield prediction value, the decision credibility coefficient and the yield risk judgment result are associated with the batch identifier and output as the prediction record.

[0019] The technical effects and advantages of the artificial intelligence-based fermentation strain yield prediction method of this invention are as follows:

[0020] This invention aligns process observation information and material ledger information of fermentation batches under the same batch caliber. It first transforms scattered monitoring signals into a conversion efficiency characterization that reflects the relationship between input and output. Then, it completes the learning prediction of strain yield based on this. Furthermore, it utilizes the consistency between ledger and on-site observations and the stability of key conversion stages to generate a decision credibility coefficient. It applies usability constraints and risk assessments to the yield prediction results, thereby elevating the prediction results from a monitoring reference that can only reflect changes in certain components to a yield judgment that can support strain comparison and production management decisions.

[0021] Furthermore, when batch differences, execution deviations, or fluctuations in operating conditions exist, it can proactively identify whether the prediction basis is reliable and adjust the judgment intensity accordingly, avoiding misjudging seemingly normal monitoring phenomena as reliable yields, improving the ability to identify yield deviations in advance and the comparability of different strains under different batch conditions, thereby enhancing the certainty and feasibility of strain screening, production scheduling assessment, and process intervention. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a method for predicting the yield of fermentation strains based on artificial intelligence, according to the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Example 1: Figure 1 This invention presents a method for predicting the yield of fermentation strains based on artificial intelligence, comprising:

[0025] S1: Align and unify the historical batch process observation sequence and material ledger sequence according to batch identifier, unify strain identifier and product caliber, and generate batch alignment data package and corresponding yield label.

[0026] S2: Based on batch-aligned data packages, the process observation sequence and the material ledger sequence are combined and converted to generate a conversion efficiency characterization sequence. The conversion efficiency characterization sequence is then bound to the strain feature vector to form a yield learning sample.

[0027] S3: Train the yield prediction model using yield learning samples, and generate a transformation efficiency characterization sequence on the batch to be evaluated according to step S2. Input the sequence and the strain feature vector into the yield prediction model to obtain the predicted yield value. At the same time, calculate the apparent volume trajectory consistency parameter and the transformation contribution steady-state ratio parameter. Input the apparent volume trajectory consistency parameter and the transformation contribution steady-state ratio parameter into the pre-trained coefficient analysis model to obtain the decision credibility coefficient. Based on the decision credibility coefficient, generate the yield risk judgment result for the predicted yield value.

[0028] S4: For the new batch of online access process observation sequence and material account sequence, generate conversion efficiency characterization sequence according to step S2 and input it into the yield prediction model together with the strain feature vector. At the same time, calculate the decision confidence coefficient according to step S3 and output the yield risk judgment result and prediction record associated with the batch identifier.

[0029] In biofermentation production, process observation sequences (such as online signals recorded over time for dissolved oxygen, pH, temperature, and liquid level) and material ledger sequences (such as material inflow and outflow records for initial filling, feeding, sampling, and discharge) from different batches often suffer from inconsistencies in time references, units of measurement, or caliber due to differences in recording systems, sampling frequencies, execution biases, or data gaps. Directly using this heterogeneous data for yield prediction model training can easily lead to feature distortion and batch-to-batch incomparability, ultimately making it difficult for the model output to support strain selection and production management decisions. In step S1 of this invention, the historical batch data undergoes rigorous alignment and standardization to establish a unified timeline and data caliber. This ensures that the conversion efficiency characterization sequence generated in subsequent step S2 accurately reflects the input-output relationship. Simultaneously, a reliable yield label is generated for each batch as a monitoring signal, providing a high-quality, reusable learning foundation for the yield prediction model. This step is the data cornerstone of the entire technical solution; only with accurate alignment and standardization can the subsequent conversion efficiency characterization and decision reliability coefficient calculations have practical production significance.

[0030] Detailed implementation of step S1:

[0031] S101 collects and initially pairs process observation sequences with material ledger sequences according to batch identification.

[0032] At the outset of historical batch data processing, it is crucial to clearly define the scope of each batch to prevent data from different batches from becoming mixed. In the production facility, each fermentation batch is typically assigned a unique batch identifier, such as a combination of production batch number or tank number and date. Using this batch identifier as an index, the entire process observation sequence for that batch is extracted from the process control system database. This includes the timestamps and corresponding values ​​recorded by all online instruments from the start to the end of fermentation, including signals such as dissolved oxygen concentration, pH value, temperature, liquid level, and tank weight. The material ledger sequence under the same batch identifier is then extracted from the material management system database. This includes all records related to material receipts and payments. Each record contains the timestamp of the operation, the operation type, and the corresponding quantity. The operation types are specifically categorized into four types: initial filling, replenishment, sampling, and discharge.

[0033] After extraction, immediately check whether a complete process observation sequence and material ledger sequence exist simultaneously under the same batch identifier. If either sequence is missing or the number of records is less than the preset minimum requirement, the entire batch data is directly discarded and not proceeded to subsequent processing. This process retains only batches with complete dual sequences, ensuring that all subsequent calculations are based on complete input-output information.

[0034] In one embodiment, a fermenter batch is identified as "2025-0112-A". The process control system exports all instrument records from 08:00 on January 12th to 20:00 on January 18th, forming a process observation sequence. The material management system exports the initial 1000 liters of liquid filling, multiple glucose replenishment records, three sampling records, and the final discharge record under the same identifier, forming a material ledger sequence. With both sequences complete, the batch proceeds to the next processing step.

[0035] S102 establishes a unified batch timeline and completes time alignment.

[0036] The recording frequency of process observation sequences is usually high and uneven, while material ledger sequences only have records when operations occur. The inconsistency in their time bases directly affects the accuracy of subsequent input-output conversion. Therefore, it is necessary to construct a unified timeline for each batch, from the start to the end of fermentation. The timeline is divided at fixed time intervals, such as every 5 minutes as a time node, forming a complete time point sequence arranged sequentially from the start time.

[0037] For the process observation sequence, the original records are mapped onto a unified time axis: for each time node, if an original record exists, the value is used directly; if there is no exact match, the two closest original records are taken, the time interval between them is calculated, and the values ​​before and after are weighted and averaged using this ratio to obtain the observation value of that time node. This process is repeated to fill in all time nodes in turn to form a time-aligned process observation sequence.

[0038] For the material ledger sequence, all operation quantities are converted into cumulative form and mapped onto a unified time axis: starting from the start of fermentation, the initial cumulative balance is zero; each operation record is traversed, and the cumulative balance is adjusted according to the operation type, with initial liquid filling and replenishment increasing the balance, and sampling and discharge decreasing the balance; for each time node on the unified time axis, the cumulative balance after the last operation before that node is taken as the material ledger value of that node, and if there are no operations before that node, the value is kept at zero, and all time nodes are filled in sequence to form a time-aligned material ledger sequence.

[0039] In one embodiment, the fermentation of batch "2025-0112-A" started at 08:00 on January 12th and ended at 20:00 on January 18th, constructing a unified timeline with 5-minute intervals. Dissolved oxygen records in the process observation sequence show values ​​at 08:03 and 08:09. For the time node 08:05, the replenishment value is obtained by weighting the values ​​from 08:03 and 08:09 according to their time distance. If the material ledger sequence has a 500-liter replenishment operation at 10:30, then from the time node corresponding to 10:30, the cumulative balance of all subsequent nodes increases by 500 liters from the previous level until the next operation occurs.

[0040] S103 ensures data consistency and unification of strain identification and product specifications.

[0041] After time alignment, the two sequences have the same time reference, but there may still be issues such as inconsistent units, missing values, or outliers. These problems can distort the calculation of subsequent transformation efficiency. Therefore, it is necessary to perform unit conversion and data cleaning on the aligned sequences.

[0042] First, all values ​​are standardized to the same unit system: when converting signals such as liquid level and tank weight in the process observation sequence to volume, liters are used as the unit; all quantity values ​​in the material ledger sequence are also standardized to liters. Concentration values ​​are standardized to grams per liter.

[0043] Secondly, missing and outlier values ​​are handled: The time-aligned process observation sequence is traversed. If a time node with no valid value is found for more than 30 consecutive minutes, the batch of data is deemed unreliable and the entire batch is discarded. For isolated missing nodes, forward filling is performed using the same value as the previous valid node. For outliers that significantly exceed the instrument's range or differ greatly from the preceding and following values, the average of the adjacent valid values ​​is used for replacement.

[0044] Meanwhile, to ensure fairness in comparing strains between batches, a standardized strain numbering system is used to assign a unique strain identifier to each batch. If different batches actually use the same strain, they must be assigned the exact same strain identifier code. To ensure the comparability of yield labels, the target product is clearly defined as a specific metabolite, and the final product concentration records for all batches are uniformly converted to concentration values ​​under the same chemical definition.

[0045] In one embodiment, the tank weight signal in batch "2025-0112-A" was originally in kilograms, but after conversion to volume, it was uniformly converted to liters; the initial liquid filling record in the material ledger was originally in cubic meters, but after conversion, it was uniformly converted to 1000 liters. If a 15-minute isolated missing segment appeared in the process observation sequence, the dissolved oxygen value of the previous valid node was directly used to fill it; if an abnormal pH value jump to 15 was found, it was directly replaced with the average of the normal values ​​before and after.

[0046] S104 generates batch-aligned data packets and yield tags.

[0047] After the aforementioned processing, each batch has a clean and consistent time-aligned sequence. Using the batch identifier as an index, the unified timeline, the time-aligned process observation sequence, the time-aligned material ledger sequence, the standardized strain identification code, and the unified target product caliber are packaged together to form a structured batch-aligned data package.

[0048] Simultaneously calculate the yield label: take the target product concentration in the process observation sequence at the end of the batch, multiply it by the cumulative volume in the material ledger sequence at the end of the batch to obtain the final cumulative product amount; take the cumulative amount of all substrate-related operations in the material ledger sequence as the total input substrate amount; divide the two to obtain the yield label value, and associate it with the batch alignment data package.

[0049] In one embodiment, after batch "2025-0112-A" is processed, the batch-aligned data package contains all observations and cumulative material balances for every 5-minute node from start to finish, with the strain identifier uniformly designated as "Strain-X-2024". At the end, the target product concentration is 120 g / L, the cumulative volume is 1500 L, and the final cumulative product amount is 180 kg; the total glucose input is converted to substrate amount of 300 kg, and the yield label is calculated as 0.6, all stored in association with the data package.

[0050] At this point, step S1 is complete, and all historical batches have been transformed into high-quality batch-aligned data packages and corresponding yield labels that can be directly used for subsequent generation of transformation efficiency characterization sequences. Step S1 transforms the original heterogeneous historical data into structured, comparable batch-aligned data packages and corresponding yield labels through strict pairing by batch identifier, unified timeline resampling, consistency of dimensions and caliber, and unification of strain and product identifiers. This processing ensures that the subsequent step S2 can accurately convert transformation efficiency characterization sequences under a unified caliber, thereby supporting reliable training of the yield prediction model and fair inter-batch comparisons.

[0051] After step S1, all historical batches have been converted into structured batch-aligned data packages, where the process observation sequence and the material ledger sequence are synchronized and consistent on a unified timeline, and are accompanied by yield labels. However, directly using raw observation values ​​or cumulative material values ​​to train the yield prediction model still makes it difficult to capture the core mechanism of the fermentation process—the dynamic conversion efficiency of substrate to target product. This is because yield is essentially a comprehensive reflection of the effective utilization and stable conversion of input substrate into product, rather than isolated changes in concentration or volume. In step S2, based on the batch-aligned data package, this invention generates a sequence specifically for characterizing the change in conversion efficiency over time through targeted combination and conversion of process observation sequences and material ledger sequences. This sequence highlights the coupling relationship between substrate consumption and product generation, avoiding the misleading impression of seemingly normal concentrations but inefficient conversion. Subsequently, this sequence is bound to the strain feature vector to form high-quality yield learning samples, thereby providing a supervised learning foundation that directly reflects the strain's yield potential for model training in step S3. This step is a key transformation that elevates fragmented process data into core features for yield prediction, ensuring that the patterns learned by the model originate from real metabolic conversion efficiency, rather than superficial monitoring indicators.

[0052] Detailed implementation of step S2:

[0053] S201 extracts the synchronization trajectory component from the batch alignment data packet.

[0054] The batch-aligned data package provides process observation sequences and material ledger sequences on a unified timeline, with complete time point correspondence between the two. To accurately calculate the conversion efficiency of the substrate to the target product, it is necessary to first isolate the key trajectories directly involved in the input-output calculation. The process observation sequence isolates the trajectory of the target product concentration changing over time, as well as the trajectory of the observed liquid volume changing over time, calculated from liquid level or tank weight signals. The material ledger sequence isolates the trajectory of the cumulative substrate input up to each time point changing over time, as well as the trajectory of the cumulative net liquid volume up to each time point changing over time. These trajectory components correspond point-by-point on the unified timeline, providing a synchronized data foundation for subsequent rate and efficiency conversions.

[0055] S202 calculates the instantaneous product formation rate and substrate consumption rate.

[0056] The essence of fermentation yield lies in the temporal matching between substrate consumption and product generation; directly using concentration or cumulative amount is insufficient to reflect the dynamic transformation process. Therefore, it is necessary to first calculate the instantaneous rate to capture the changing trend. Liquid volume is calculated as the fused average of observed volume and recorded volume. The observed and recorded volumes are weighted and summed according to preset weights to obtain the fused volume, reducing the impact of single-source volume deviations on the conversion of cumulative product amount, thus obtaining the harmonic volume value at each time point. The cumulative product amount up to each time point is calculated point by point, i.e., the target product concentration multiplied by the harmonic volume. From the second time point onwards, the product generation rate is calculated as the current cumulative product amount minus the previous time point's cumulative product amount, divided by the time interval between the two. The substrate input rate is calculated as the current cumulative substrate input amount minus the previous time point's cumulative substrate input amount, divided by the time interval; when a write-back correction results in a negative difference, the substrate input rate at that time point is recorded as zero to avoid abnormal accounting interfering with subsequent characterization calculations. This processing yields an instantaneous product generation rate sequence and substrate consumption rate sequence consistent with a unified time axis length.

[0057] In one embodiment: at two adjacent time points on a unified time axis, the observed volume of a batch is 1500 liters, the book volume is 1520 liters, and the harmonized volume is calculated to be 1510 liters; the target product concentrations are 80 g / L and 85 g / L, respectively, and the cumulative product amounts are 120800 kg and 128350 kg, respectively. The product generation rate is calculated by subtracting the previous value from the current cumulative product amount and dividing by a 5-minute interval. Simultaneously, the cumulative substrate input increases from 200 kg to 250 kg, and the substrate consumption rate is 50 kg divided by the time interval.

[0058] S203 generates a conversion efficiency characterization sequence.

[0059] While the instantaneous rate sequence already reflects dynamic changes, individual rate ratios can still be affected by noise or transient perturbations. To comprehensively characterize conversion efficiency, a multidimensional sequence needs to be constructed to integrate instantaneous, cumulative, and mechanistic bias information. The conversion efficiency characterization sequence is a three-dimensional vector sequence. The first dimension is the instantaneous yield, which is the sum of the instantaneous product generation rate divided by the substrate consumption rate and a small positive number. This small positive number is a fixed value, such as 0.001, to prevent division by zero errors when the substrate consumption rate approaches zero. The second dimension is the cumulative conversion rate, which is the sum of the cumulative product amount up to the current time point divided by the cumulative substrate input and the same small positive number. The third dimension is the efficiency bias rate, which is the difference between the instantaneous product generation rate and the theoretical stoichiometric coefficient multiplied by the substrate consumption rate, divided by the sum of the substrate consumption rate and the small positive number. The theoretical stoichiometric coefficient is determined based on the chemical reaction equation between the target product and the substrate, such as the number of grams of product theoretically generated per gram of substrate. After point-by-point calculation, a three-dimensional conversion efficiency characterization sequence with a uniform time axis length is formed.

[0060] In one embodiment: Continuing the previous example, the instantaneous product generation rate corresponds to a positive value at the time point, the substrate consumption rate is 50 kg / time, and the instantaneous yield is calculated as the generation rate divided by the consumption rate plus 0.001; the cumulative conversion rate is 128350 kg of product divided by 250 kg of substrate plus 0.001; the theoretical stoichiometric coefficient is 0.5, then the efficiency deviation rate is calculated as the difference between the generation rate and 0.5 times the consumption rate, divided by the consumption rate plus 0.001, and the three-dimensional numerical combination is the representation vector at that time point.

[0061] S204 binds the bacterial strain feature vector to form a yield learning sample.

[0062] The transformation efficiency characterization sequence has captured the dynamic transformation characteristics at the process level, but the differences in yield potential among different strains need to be reflected through the inherent attributes of the strains. Therefore, a strain feature vector is constructed. Using one-hot encoding combined with numerical attributes, the set of strain identification codes involved in all historical batches is one-hot encoded, meaning each strain has one position set to 1 and the rest to 0. Simultaneously, the strain generation number or known metabolic characteristic value is appended to form a fixed-length strain feature vector. This strain feature vector is placed at the beginning and concatenated with the long vector of the transformation efficiency characterization sequence, flattened point-by-point in chronological order, to form the complete input feature. This is then paired with the yield label generated in step S1 to form yield learning samples.

[0063] At this point, step S2 is complete. All historical batches are transformed into yield learning samples with conversion efficiency characterization sequences as the core and strain characteristics bound together, which can be directly used for subsequent yield prediction model training. Step S2 extracts synchronization components from batch-aligned data packets, calculates instantaneous rates, generates conversion efficiency characterization sequences through multi-dimensional combinations, and binds them to strain feature vectors, ultimately producing yield learning samples. By dynamically transforming the process into time-series features that directly reflect substrate conversion efficiency, it ensures that the yield prediction model trained in step S3 can learn the strain's yield performance under real metabolic coupling, thereby supporting reliable inter-batch comparisons and decisions.

[0064] After step S2 is completed, all historical batches have been transformed into yield learning samples with conversion efficiency characterization sequences as the core and strain feature vectors as the binding. These samples directly reflect the dynamic conversion efficiency of the substrate to the target product and the inherent properties of the strain. However, in actual production, even if the yield prediction model is trained based on these samples to obtain the predicted yield value, it may still lose the reliability of decision-making due to execution deviations, data recording errors, or fluctuations in operating conditions of the current batch. For example, inconsistencies between the recorded materials and on-site observations can lead to distortion of the volume background, or the conversion process may have high instantaneous efficiency but lack a sustained steady state, resulting in a low final yield. In step S3 of this invention, the yield prediction model is first trained offline using the yield learning samples. Then, for the batch to be evaluated, while generating the conversion efficiency characterization sequence and inputting it into the model to obtain the predicted yield value, the consistency parameter of the recorded volume trajectory (to assess the credibility of data recordings) and the steady-state proportion parameter of the conversion contribution (to assess the stability of the conversion process) are calculated in parallel. Then, the two are combined into a decision credibility coefficient through a coefficient analysis model, and finally, this coefficient is used to constrain the risk assessment strength of the yield prediction value. The core of this step lies in proactively introducing a two-parameter reliable assessment mechanism, which elevates pure numerical prediction to a management signal with reliability constraints. This ensures that when batch differences exist, misjudging normal surface phenomena is avoided, thereby supporting deterministic decision-making for strain comparison and process intervention.

[0065] Detailed implementation of step S3:

[0066] S301 uses yield learning samples to train a yield prediction model.

[0067] The yield training samples already contain the input formed by concatenating the flattened temporal features of the conversion efficiency representation sequence with the strain feature vector, along with corresponding yield labels. To enable the model to directly map the final yield from the process dynamics, the yield prediction model needs to be trained offline using this sample set. The yield prediction model employs a sequence regression structure, using the conversion efficiency representation sequence as a multi-dimensional sequence input arranged in chronological order, and the strain feature vector as a static conditional feature input. During sequence modeling, the strain feature vector is mapped and concatenated with the representation vectors at each time point to form a conditional sequence, from which the sequence model outputs the predicted yield value for the corresponding batch. The training process iteratively adjusts the model parameters until the validation set loss function reaches the preset convergence condition.

[0068] In one embodiment, the yield prediction model is constructed using a sequence regression architecture based on a Long Short-Term Memory (LSTM) network. Specifically, the model's input layer first receives a fixed-length concatenated feature vector, where the bacterial species feature vector is 50-dimensional (assuming historical batches involve 40 bacterial species with one-hot encoding plus 10-dimensional numerical attributes), and the transformation efficiency representation sequence, after flattening, is 3000-dimensional (assuming a unified timeline contains 1000 time points, each represented in three dimensions), for a total input dimension of 3050. A fully connected layer is then connected to the input layer to reduce the dimension to 512, compressing the features and extracting an initial abstract representation.

[0069] The network then proceeds to the Long Short-Term Memory (LSTM) backbone, which consists of three bidirectional LSM units. Each unit has a hidden state dimension of 256, with a dropout rate of 0.2 to prevent overfitting and improve generalization. The LSM outputs the hidden state at the final time step (512-dimensional, bidirectionally concatenated), followed by two fully connected layers. The first layer has a dimension reduced to 128 and uses ReLU as the activation function. The second layer has a 1-dimensional output with no activation function, directly generating the productivity prediction.

[0070] The training process uses all yield training samples, with a total sample size of 500 historical batches. These samples are randomly divided into a training set (400 samples), a validation set (50 samples), and a test set (50 samples) in an 8:1:1 ratio. The loss function is mean squared error, calculated as the average of the squared differences between the predicted yield and the yield label. The optimizer is Adam, with an initial learning rate of 0.001 and a weight decay coefficient of 1e-5. The first-moment estimate exponential decay rate beta1 is 0.9, and the second-moment estimate exponential decay rate beta2 is 0.999.

[0071] Training was conducted using mini-batch gradient descent with 32 samples per batch and a maximum of 200 training epochs. An early stopping mechanism was employed to monitor the validation set loss; training was stopped early if the validation loss did not decrease for 20 consecutive epochs. Simultaneously, the learning rate decayed to 0.5 times its original value if the validation loss did not improve for 10 consecutive epochs, with a minimum learning rate of 1e-6. After training, model performance was evaluated on the test set. The final mean squared error was controlled within 0.002, indicating that the model's regression accuracy for the yield label met production requirements.

[0072] In this embodiment, all hyperparameters were optimized and determined on the validation set through grid search. For example, the number of Long Short-Term Memory (LSTM) layers was tried between 1 and 4, and the hidden dimension was selected between 128 and 512. The above configuration was ultimately chosen to balance computational efficiency and prediction accuracy. Model parameters were initialized using Xavier uniform initialization, and the bias of fully connected layers was initialized to zero. The training environment was based on the PyTorch framework, using a single GPU for acceleration, and a single complete training session took approximately 30 minutes. After training, the model parameters were saved and used for calculating online yield predictions for subsequent batches to be evaluated.

[0073] S302 generates a conversion efficiency characterization sequence for the batch to be evaluated and inputs it into the yield prediction model to obtain the predicted yield value.

[0074] After the process observation sequence and material ledger sequence of the batch to be evaluated are connected online, the feature generation method must be consistent with historical samples to ensure consistent prediction standards. Trajectory components, harmonic volume, instantaneous rate, and three-dimensional conversion efficiency characterization sequences are extracted in the same manner as in step S2, and concatenated with the current batch's microbial feature vector to form input features. This input feature is fed into the yield prediction model, and the model outputs a single numerical value as the predicted yield, representing the expected final yield for that batch. Whenever the batch data is updated, the conversion efficiency characterization sequence up to the current moment is recalculated and input into the model to obtain the updated yield prediction value.

[0075] Whether the forecast results can be used for production management depends first on whether two types of premises hold true:

[0076] First, whether the batch data truly reflects the input and dilution background, to avoid systematic deviations between the input side and the field side on the same time axis due to missing records, execution deviations, or sensor drift;

[0077] Secondly, the yield is mainly determined by the critical transformation stage, and whether this stage has a repeatable and comparable stable operating state directly affects whether the process signals on which the prediction is based have generalizability.

[0078] The consistency parameter of the material ledger and the process observation is aimed at the degree of consistency between the input record and the field observation in terms of the basic dimension of volume. It can directly reveal whether there is a deviation between the material ledger and the process observation that would distort the overall relationship without relying on specific components or specific instruments.

[0079] The steady-state contribution ratio parameter of conversion targets the operational quality of key conversion stages. It makes explicit which time periods contribute to the yield and whether these time periods are in a controllable and stable state, thereby distinguishing between effective process information that can be used to infer the yield and fluctuation information that, although there are monitoring signals, is difficult to support yield judgment.

[0080] Compared to indicators constructed around single components, local fluctuations, or equipment-specific quantities, the above two parameters cover the two shortest and most critical paths of data credibility and process usability, respectively. They also have clear physical meaning and interpretable management orientation, which can provide a stable and transferable input basis for subsequent coefficient analysis models, so that the decision credibility coefficient in step S3 has a consistent caliber and a practical judgment basis.

[0081] S303 reconstructs the book liquid volume trajectory and the observed liquid volume trajectory and calculates the consistency parameters of the book and observed volume trajectories.

[0082] Yield predictions rely on accurate volumetric background values; significant discrepancies between book records and on-site observations can distort conversion efficiency representations. Therefore, the initial liquid loading, positive replenishment, sampling, and negative discharge phases are directly accumulated from the material ledger sequence according to operation type to obtain point-by-point liquid phase volume trajectories. The observed liquid phase volume trajectory is then calculated by multiplying the liquid level or tank weight signals from the process observation sequence by a calibration coefficient, and a moving average of three adjacent points is applied to the trajectory for smoothing and noise reduction.

[0083] After aligning the two trajectories, calculate the consistency parameter of the account volume trajectory: calculate the absolute value of the difference between the account volume and the observed volume at each point and divide it by the observed volume. Calculate the average value of this relative absolute deviation at all time points, and then subtract this average value from 1 to obtain the consistency parameter. The parameter value range is between 0 and 1. The closer the value is to 1, the higher the consistency between the two trajectories. It is only used in the calculation at time points where the observed volume is non-zero.

[0084] In one embodiment: the book volume of a certain batch of materials to be evaluated shows an increase of 500 liters after replenishment, but the observed volume converted from the tank weight signal in the process observation sequence only increases by 480 liters. The consistency parameter calculated by averaging the relative deviations at each point is 0.92, indicating that the data record is basically reliable; if the sampling operation is not recorded in time, resulting in an overestimation of the book volume, the consistency parameter drops to 0.75.

[0085] S304 locates the yield contribution period based on the conversion efficiency characterization sequence and calculates the steady-state proportion parameter of conversion contribution.

[0086] The yield prediction considers the dynamics of the entire process, but the final yield is mainly determined by the period during which the conversion efficiency is consistently higher than the batch's own level. Therefore, the instantaneous yield, the first dimension of the conversion efficiency characterization sequence, is taken as the analysis object. The median of all time points in the sequence is first calculated as the batch's own baseline. The instantaneous yield sequence is divided into continuous segments, and segment boundaries are formed by marking points where the cumulative deviation exceeds the preset control limit. Segments where the instantaneous yield is continuously higher than 1.2 times the baseline and the total duration of the segment exceeds 10% of the total batch duration are selected as yield contribution periods. Within the yield contribution period, the dissolved oxygen, pH, and temperature in the process observation sequence are checked point by point to see if they simultaneously fall within 5% of the process setting range and the change amplitude of adjacent points is less than the preset threshold. If these conditions are met, they are marked as steady-state segments. The total duration of all steady-state segments is calculated and divided by the total duration of the yield contribution period to obtain the steady-state contribution ratio parameter. The parameter value ranges from 0 to 1; a higher value indicates a more stable conversion process.

[0087] In one embodiment, a batch of conversion efficiency characterization sequence showed that the instantaneous yield in the middle and late stages was consistently higher than the median baseline. Change point detection divided the corresponding segment as the yield contribution period. During this segment, dissolved oxygen and pH were stable within the set range, and temperature fluctuations were small. The steady-state segment ratio was calculated to be 0.85. In another batch, although there was a brief period of high instantaneous yield, it was accompanied by drastic pH fluctuations, and the steady-state segment ratio was only 0.40.

[0088] S305 inputs the consistency parameter of the account volume trajectory and the steady-state proportion parameter of the transformation contribution into the coefficient analysis model to obtain the decision credibility coefficient.

[0089] The coefficient analysis model takes the consistency parameter of the apparent volume trajectory and the steady-state proportion parameter of the conversion contribution obtained from historical batches as input. For each historical batch, cross-validation or hold-out methods are used to generate corresponding yield predictions, and the prediction deviation is calculated by comparing these predictions with the final yield of that batch. These prediction deviations are then mapped to graded labels according to a preset interval as training labels to learn the correspondence between the two parameters and the prediction reliability. The training labels are the reciprocal of the absolute value of the deviation between the yield prediction and the final yield, categorized into five levels (the smaller the deviation, the higher the level). After training, for the batch to be evaluated, the currently calculated two parameters are input into the coefficient analysis model, and the model outputs a decision confidence coefficient. The coefficient ranges from 0 to 1; a higher value indicates that the current batch's yield prediction is more suitable for decision-making.

[0090] In one embodiment, the coefficient analysis model is constructed using a regression architecture based on gradient boosting decision trees. Specifically, the model is implemented using the LightGBM framework, and the input layer directly receives two-dimensional feature vectors, namely the account volume trajectory consistency parameter and the transformation contribution steady-state proportion parameter. The total input dimension is 2-dimensional, requiring no additional embedding or fully connected preprocessing layers.

[0091] The model's backbone is configured with gradient boosting decision tree parameters as follows: 200 trees, a maximum depth of 6 per tree, and a minimum number of samples per leaf node (5) to control model complexity and prevent overfitting. The learning rate is 0.05, the subsample ratio is 0.8, the feature sampling ratio is 0.8, and regularization parameters include L1 regularization of 0.1 and L2 regularization of 0.1. The objective function uses mean squared error, calculating the average of the squared differences between the model output and the training labels. The output layer is normalized to the 0-1 range using the sigmoid function, directly producing the decision confidence coefficient.

[0092] The training process uses historical batch data, with a total sample size of 500. These samples are randomly divided into a training set (400 samples), a validation set (50 samples), and a test set (50 samples) in an 8:1:1 ratio. Training labels are defined as the reciprocal of the absolute value of the deviation between the predicted yield and the final yield, quantized in five levels: 1.0 for a deviation less than 0.01, 0.8 for a deviation between 0.01 and 0.03, 0.6 for a deviation between 0.03 and 0.05, 0.4 for a deviation between 0.05 and 0.1, and 0.2 for a deviation greater than 0.1.

[0093] The optimizer is built into LightGBM and uses coordinate descent to iteratively build the tree. Training is performed in mini-batch mode, with each batch containing the entire dataset. An early stopping mechanism monitors the mean squared error on the validation set; if there is no decrease in the validation error for 30 consecutive rounds, training is stopped early. Simultaneously, the learning rate decays to 0.7 times its original value when the validation error shows no improvement for 15 consecutive rounds. After training, the model performance is evaluated on the test set. The final mean squared error is controlled within 0.005, and the correlation coefficient between the decision confidence coefficient and the inverse of the historical bias exceeds 0.85, indicating that the model's mapping accuracy to confidence levels meets production requirements.

[0094] In this embodiment, all hyperparameters were optimized and determined on the validation set through grid search. For example, the number of trees was tested between 100 and 500, the maximum depth was selected between 4 and 8, and the learning rate was adjusted between 0.01 and 0.1. The above configuration was ultimately chosen to balance model interpretability and prediction stability. Model parameters were initialized using the framework's default random seed, and the full-tree split gain threshold was 0.0. The training environment was based on Python's LightGBM library, using CPU multi-threading acceleration, with a single complete training session taking approximately 5 minutes. After training, the model parameters were saved and used for calculating the online decision confidence coefficient in subsequent batches to be evaluated.

[0095] S306 generates a yield risk assessment result based on the decision confidence coefficient and the predicted yield value.

[0096] Yield forecasts must be considered in conjunction with their reliability before being translated into management actions. Therefore, a preset risk threshold can be set at 0.9 times the historical average batch yield. If the predicted yield is below the risk threshold and the decision reliability coefficient is above 0.7, it is classified as high risk and an intervention record is triggered; if the predicted yield is below the risk threshold but the decision reliability coefficient is not higher than 0.7, it is classified as requiring manual review; if the predicted yield is not lower than the risk threshold, it is classified as low risk. The judgment results are stored in association with the batch identifier to form the yield risk judgment result.

[0097] Example: The predicted yield of a certain batch is 0.85 times the historical average, with a consistency parameter of 0.95 and a steady-state percentage of 0.88. After inputting the two parameters, the decision confidence coefficient is 0.82, which is higher than 0.7. Therefore, it is directly judged as high risk and intervention is recorded. Another batch has a similar predicted value but a steady-state percentage of only 0.35. The decision confidence coefficient drops to 0.60. Therefore, it is only marked as needing manual review and no automatic action is triggered.

[0098] Thus, step S3, through offline training of the yield prediction model, parallel calculation of dual parameters and decision credibility coefficients in the batch to be evaluated, and final constraint generation of risk judgment results, elevates the yield prediction from a pure numerical output to a management decision signal with credibility assessment. This ensures that the judgment strength is proactively reduced when there is data deviation or process instability, thereby improving the reliability of strain yield judgment and the feasibility of production intervention.

[0099] After step S3, the yield prediction model and coefficient analysis model have been trained offline based on historical yield learning samples, enabling them to output predicted yield values ​​from the conversion efficiency characterization sequence and to map decision confidence coefficients from two parameters. However, in actual production, the fermentation process of new batches continues, and process observation sequences and material ledger sequences are generated in real time. If evaluation is only conducted offline after the batch ends, the opportunity for early intervention will be lost, failing to support timely decisions regarding strain selection, process adjustment, or resource allocation. This invention implements online application for new batches in step S4: as data is gradually integrated, the current batch data is dynamically constructed. The conversion efficiency characterization sequence generated in step S2 is input into the yield prediction model to obtain real-time predicted yield values. Simultaneously, the decision confidence coefficient and risk assessment results are calculated in step S3, ultimately outputting a prediction record associated with the batch identifier. This step transforms the offline training results into an online management tool for the production site, ensuring that yield risk signals with confidence constraints are obtained during fermentation, supporting proactive intervention rather than passive confirmation by enterprises.

[0100] Detailed implementation of step S4:

[0101] S401 integrates the process observation sequence and material ledger sequence of a new batch online and dynamically constructs a batch alignment data package.

[0102] After a new batch of fermentation starts, process observation sequences and material ledger sequences are continuously generated. To support rolling updates of predicted values, these data must be collected and synchronized in real time. Using the new batch identifier as an index, a data access buffer is established. At fixed intervals or when a new material operation record is detected, the newly added process observation sequence and material ledger sequence are extracted. The unified time axis is updated in the same manner as in step S1. The newly added observation values ​​are resampled and aligned, the newly added operation quantities are cumulatively mapped and aligned, and unit conversion and missing data filling are performed, forming a continuously appending batch alignment data package. The batch alignment data package thus gradually grows from empty until the fermentation end marker is reached.

[0103] In one embodiment, after a new batch is started, the process control system pushes dissolved oxygen and liquid level records every minute, the material management system immediately reports the records after the replenishment operation is completed, the buffer immediately adds alignment after detecting an update, and the batch alignment data packet is filled with new observations and cumulative volumes point by point from the start time on a unified time axis.

[0104] S402 generates a conversion efficiency characterization sequence according to step S2 and inputs it into the yield prediction model to obtain the yield prediction value.

[0105] After the batch alignment data package is updated, it needs to be immediately converted into model-recognizable features to reflect the current yield expectation. Following the same method as step S2, the target product concentration trajectory, observed liquid volume trajectory, cumulative substrate input trajectory, and cumulative net liquid volume trajectory up to the current moment are extracted from the latest batch alignment data package. The harmonic volume is calculated, and the instantaneous product generation rate and substrate consumption rate are obtained through point-by-point differencing. A three-dimensional conversion efficiency characterization sequence is constructed and concatenated with the new batch's bacterial strain feature vector to form the input feature. This input feature is fed into the yield prediction model, and the model outputs the current yield prediction value. Each time the batch alignment data package is updated, the above process is repeated to generate the updated yield prediction value.

[0106] S403 simultaneously calculates the decision credibility coefficient according to step S3.

[0107] To ensure the decision-making usability of the yield predictions, the reliability of the data and processes must be evaluated in parallel with the generation of the predicted yield values. Based on the same latest batch of aligned data packages and conversion efficiency characterization sequences, the book liquid phase volume trajectory and the observed liquid phase volume trajectory are reconstructed in the same manner as in step S3. The consistency parameter of the book and observed volume trajectories is calculated, the yield contribution period is located, and the steady-state proportion parameter of the conversion contribution is calculated. The two parameters are input into the coefficient analysis model, and the model outputs the decision reliability coefficient at the current moment. This process is triggered synchronously with the yield prediction calculation, ensuring that each prediction corresponds to the latest reliability assessment.

[0108] S404 yield risk assessment results.

[0109] After both the yield forecast and the decision credibility coefficient have been updated, a comprehensive risk level assessment is required to translate them into specific management instructions. The current yield forecast is compared to a preset risk threshold. If the yield forecast is lower than the risk threshold and the decision credibility coefficient is higher than the warning line, it is classified as high risk; if the yield forecast is lower than the risk threshold but the decision credibility coefficient is not higher than the warning line, it is classified as requiring manual review; if the yield forecast is not lower than the risk threshold, it is classified as low risk. The assessment result includes a text description and is linked to the current timestamp.

[0110] In one embodiment, when the batch is in the middle or late stage, the predicted yield drops below the risk threshold, and the decision confidence coefficient is higher than the warning line. The judgment result is directly marked as high risk, and the production personnel check the replenishment strategy immediately after receiving the record. At another time, when the decision confidence coefficient is low, the same predicted value is only marked as requiring manual review, and the record is stored for subsequent confirmation.

[0111] S405 outputs yield risk assessment results and prediction records associated with batch identifiers.

[0112] Once all calculation results are generated, this information needs to be structured, stored, and pushed to support production traceability and on-site response. The current timestamp, predicted yield value, decision confidence coefficient, and yield risk assessment result are associated with the new batch identifier and written to a dedicated prediction record database. Simultaneously, the latest assessment results and predicted values ​​are pushed to the production monitoring interface. The pushed content includes numerical values ​​and text descriptions for easy viewing by management personnel.

[0113] Step S4 dynamically constructs batch-aligned data packets through online data access, synchronously executes the core calculations of steps S2 and S3, and finally outputs yield risk assessment results and prediction records with timestamps, applying the entire technical solution to the new batch production process. This achieves an online processing flow from data access, feature generation, prediction output to risk assessment result output, supporting enterprises to promptly perform strain comparison, process intervention, or resource adjustment based on reliable predictions during batch production.

[0114] Specifically, the above are merely preferred embodiments of this application and are not intended to limit this application.

[0115] The preset threshold or risk threshold can be pre-calibrated through offline simulation testing, or set as a fixed value according to the on-site operating procedures.

[0116] In the description of this specification, references to terms such as "an embodiment," "example," and "specific example" indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0117] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for predicting the yield of fermentation strains based on artificial intelligence, characterized in that, Including the following steps: S1: Align and unify the historical batch process observation sequence and material ledger sequence according to batch identifier, unify strain identifier and product caliber, and generate batch alignment data package and corresponding yield label. S2: Based on batch-aligned data packages, the process observation sequence and the material ledger sequence are combined and converted to generate a conversion efficiency characterization sequence. The conversion efficiency characterization sequence is then bound to the strain feature vector to form a yield learning sample. S3: Train the yield prediction model using yield learning samples, and generate a conversion efficiency characterization sequence on the batch to be evaluated according to step S2. This sequence, along with the strain feature vector, is input into the yield prediction model to obtain the predicted yield value. Simultaneously, calculate the apparent volume trajectory consistency parameter and the steady-state contribution ratio parameter. Specifically, the apparent liquid phase volume trajectory is obtained from the material account sequence through cumulative operations, and the observed liquid phase volume trajectory is obtained by converting and smoothing the process observation sequence. The relative absolute deviation is calculated point-by-point and averaged, then subtracted by 1 to obtain the apparent volume trajectory consistency parameter. Take the median baseline of the instantaneous yield calculation batch from the conversion efficiency characterization sequence. Divide the batch into segments using change point detection and select segments above the baseline as yield contribution periods. Combine this with the process observation sequence to determine the steady-state segment within the process range, and calculate the steady-state segment duration ratio to obtain the steady-state contribution ratio parameter. Input the apparent volume trajectory consistency parameter and the steady-state contribution ratio parameter into the pre-trained coefficient analysis model to obtain the decision confidence coefficient. Based on the decision confidence coefficient, generate a yield risk assessment result for the predicted yield value. S4: For the new batch of online access process observation sequence and material account sequence, generate conversion efficiency characterization sequence according to step S2 and input it into the yield prediction model together with the strain feature vector. At the same time, calculate the decision confidence coefficient according to step S3 and output the yield risk judgment result and prediction record associated with the batch identifier.

2. The method for predicting the yield of fermentation strains based on artificial intelligence according to claim 1, characterized in that: Step S1 extracts and aligns the process observation sequence and material ledger sequence to a unified timeline using batch identifiers, removes incomplete batches, unifies units, cleans data, standardizes strain identifiers and product calibers, and generates batch alignment data packages and yield labels.

3. The method for predicting the yield of fermentation strains based on artificial intelligence according to claim 2, characterized in that: In step S2, the target product concentration trajectory, observed liquid volume trajectory, cumulative substrate input trajectory, and net liquid volume cumulative trajectory are extracted from the batch alignment data package. The observed volume and the book volume are weighted and summed according to preset weights to obtain the harmonic volume at each time point. The instantaneous product generation rate and substrate consumption rate are calculated point by point using differential calculation.

4. The method for predicting the yield of fermentation strains based on artificial intelligence according to claim 3, characterized in that: In step S2, a three-dimensional conversion efficiency characterization sequence is generated, including instantaneous yield, cumulative conversion rate and efficiency deviation rate. Instantaneous yield and efficiency deviation rate are divided by the sum of substrate consumption rate and small positive number. The efficiency deviation rate is introduced with a theoretical stoichiometric coefficient. The conversion efficiency characterization sequence is flattened and concatenated with the feature vector of the uniquely thermocoded strain, and paired with the yield label to form a yield learning sample.

5. The method for predicting the yield of fermentation strains based on artificial intelligence according to claim 4, characterized in that: In step S3, the yield prediction model is trained using yield learning samples. For the batch to be evaluated, the transformation efficiency characterization sequence is generated in the manner of step S2 and concatenated with the strain feature vector before being input into the yield prediction model, and the predicted yield value is output.

6. The method for predicting the yield of fermentation strains based on artificial intelligence according to claim 5, characterized in that: In step S3, the consistency parameters of the account volume trajectory and the steady-state proportion of the conversion contribution are input into the coefficient analysis model, and the decision credibility coefficient is output. Based on the comparison between the yield prediction value and the risk threshold and the level of the decision credibility coefficient, the yield risk judgment result is generated in a graded manner and associated with the batch identifier.

7. The method for predicting the yield of fermentation strains based on artificial intelligence according to claim 6, characterized in that: In step S4, the online access process observation sequence and material ledger sequence of the new batch are updated. Each time the data is updated, a batch alignment data package is dynamically constructed. The conversion efficiency characterization sequence is generated according to step S2 and concatenated with the strain feature vector. The result is then input into the yield prediction model, and the yield prediction value is output.

8. The method for predicting the yield of fermentation strains based on artificial intelligence according to claim 7, characterized in that: In step S4, the decision credibility coefficient is calculated synchronously as in step S3. Based on the comparison between the yield prediction value and the risk threshold and the level of the decision credibility coefficient, the yield risk judgment result is generated. The yield prediction value, the decision credibility coefficient and the yield risk judgment result are associated with the batch identifier and output as the prediction record.