Flavor directional regulation fermentation method and device of phyllium vinegar based on reinforcement learning

By constructing a fermentation state space model and flavor target encoding based on reinforcement learning, and using a physical information deep Q-network for action decision-making, the flavor of yellow peel fruit vinegar was directionally regulated, solving the problem of unstable flavor quality in traditional fermentation processes and improving product consistency and control precision.

CN122290720APending Publication Date: 2026-06-26GUANGDONG XINGYAO BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG XINGYAO BIOTECHNOLOGY CO LTD
Filing Date
2026-05-09
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Traditional fermentation processes for yellow peel fruit vinegar suffer from several problems, including large batch-to-batch fluctuations in flavor quality, unclear relationships between process parameters and flavor formation, uneven physical fields within the fermentation tank, lack of online sensing and closed-loop control methods for flavor, and neglect of the correlation between microbial physical damage and flavor deterioration.

Method used

By employing a reinforcement learning-based approach, a fermentation state space model is constructed to encode flavor targets and set flux-oriented reward functions. A deep Q-network based on physical information is used for continuous action decision-making, and multi-level intervention is implemented. Combined with self-evolutionary learning and flux feedback optimization, the flavor of yellow peel fruit vinegar is directionally regulated.

Benefits of technology

This method improves batch consistency and control safety of the flavor of yellow peel fruit vinegar, solves the problems of poor flavor reproducibility and control in traditional methods, and ensures precise control of the fermentation process and stable product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290720A_ABST
    Figure CN122290720A_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for flavor-oriented fermentation regulation of wampee vinegar based on reinforcement learning, relating to the field of artificial intelligence learning. The method includes: constructing a fermentation state-space model containing terpene concentration; setting flavor target encoding and flux-oriented reward function; using a physical information deep Q-network for continuous action decision-making; implementing multi-level intervention execution; and optimizing strategies through self-evolutionary learning and flux feedback. This invention achieves specialized monitoring and regulation of the characteristic aroma of wampee vinegar, solves the problem of non-monotonic coupling control of two microbial communities, reduces training samples by embedding prior knowledge of strains, ensures action safety by embedding physical and biological constraints, and improves batch-to-batch flavor consistency through delay compensation and self-evolutionary mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence learning, and in particular to a method and apparatus for directional flavor regulation fermentation of yellow peel fruit vinegar based on reinforcement learning. Background Technology

[0002] Clausena lansium vinegar is made from the Lingnan specialty fruit, and is produced through two stages: alcoholic fermentation and acetic acid fermentation. Its flavor profile is a synergistic effect of over a hundred compounds, including esters (ethyl acetate, ethyl butyrate, ethyl phenylacetate, etc.), alcohols (isoamyl alcohol, phenylethyl alcohol, etc.), acids (acetic acid, citric acid, characteristic organic acids of Clausena lansium), and terpenes (limonene, β-caryophyllene, etc.). Traditional fermentation processes face the following core technological bottlenecks:

[0003] (1) Flavor quality fluctuates greatly from batch to batch. Static temperature control and timed aeration cannot respond to the dynamic changes in the fermentation process, resulting in poor reproducibility of flavor compound profiles between different batches, and the rate of high-quality products is usually less than 70%.

[0004] (2) The mapping relationship between process parameters and flavor formation is unclear. The regulation of flavor metabolic network by operating parameters such as temperature, dissolved oxygen, and stirring has highly nonlinear, time-varying and multi-objective coupling characteristics, making it difficult to establish a precise regulation model using traditional orthogonal optimization methods.

[0005] (3) The uneven physical field in the fermenter leads to spatial heterogeneity of flavor. Industrial-scale fermenters generally have physical problems such as temperature stratification, dissolved oxygen gradient, uneven distribution of shear stress, and dead zone retention, which result in significant differences in the flavor quality of fermentation liquid in different spatial locations within the same batch.

[0006] (4) Lack of online sensing and closed-loop control methods for flavor. Flavor compound detection relies on offline chromatographic analysis, which has a long detection cycle and slow feedback, and cannot support real-time optimization decisions.

[0007] (5) The relationship between physical damage to bacteria and flavor deterioration has been overlooked. Mechanical stress such as stirring, shearing, and bubble rupture can cause sublethal damage to acetic acid bacteria, which can induce stress metabolism and produce undesirable flavor precursors such as acetaldehyde and diacetyl. Existing processes lack targeted inhibition strategies. Summary of the Invention

[0008] To address the technical problems in the prior art, this invention provides a method and apparatus for directional flavor regulation fermentation of yellow peel fruit vinegar based on reinforcement learning.

[0009] This invention is achieved through the following technical solution: A method for flavor-directed fermentation regulation of yellow-skinned fruit vinegar based on reinforcement learning, comprising: Fermentation state-space modeling includes constructing a joint state vector containing physical sensing variables, biological metabolic variables, and biological response delay variables, and constructing state dynamic equations; Flavor target encoding and flux-oriented reward function setting include quantifying flavor into target flavor fingerprint vectors and target metabolic flux distributions. The reward function includes flavor matching reward, metabolic efficiency reward, microbial health reward and regulatory cost. Continuous action decision-making includes action decision-making based on a physical information deep Q-network. The actions include temperature change rate setpoint, dissolved oxygen change rate setpoint, nutrient salt flow valve opening, and functional bacteria supplementation switch. The network includes an input layer, a dual-colony coupled feature encoding layer, a shared feature extraction layer, a strain parameter embedding layer, an Actor output layer, a Critic output layer, and a physical-biological constraint barrier layer. Multi-level intervention execution includes translating the output of continuous actions into actual fermenter execution instructions, while compensating for biological response delays and continuous regulation guided by metabolic flux ratio, and implementing three-level regulation based on environmental parameters, microbial community and metabolic pathways. Self-evolutionary learning and flux feedback optimization, including continuously updating the physical information deep Q network parameters and flux model based on the deviation between actual fermentation results and predictions.

[0010] Furthermore, the biological response delay variables are dynamically identified through online impulse response experiments, including the characteristic time of enzyme activity regulation under temperature stimulation, the time of respiratory chain rearrangement after dissolved oxygen changes, osmotic pressure response time, and gene expression delay.

[0011] Furthermore, the flavor matching reward employs modified cosine similarity, taking into account both the direction and absolute concentration of the flavor profile: ; in S is the characteristic ester concentration vector, and S is the state vector. It is the target flavor fingerprint vector. It's a dot product operation. It is the sum of squares of the differences between the actual flavor vector and the target flavor vector, and α is the modulus penalty weighting coefficient.

[0012] Furthermore, the dual-colony coupling feature encoding layer in the physical information-based deep Q-network nonlinearly combines the yeast and acetic acid bacteria related features in the original state vector to generate a coupling representation vector, resulting in five coupling features: dual-colony synergistic activity index, synergistic activity change rate, unit yeast ethanol yield, unit acetic acid bacteria acid production, and dual-colony abundance ratio.

[0013] Furthermore, the strain parameter embedding layer in the physical information-based deep Q-network includes defining a strain-specific parameter vector and calculating the optimal action benchmark for the strain and a strain adaptability penalty term based on the parameter vector.

[0014] Furthermore, the physical-biological constraint barrier layer includes five constraint functions, namely, ethanol toxicity boundary, pH inhibition boundary, temperature shock boundary, oxygen transfer rate limit, and total acid accumulation boundary.

[0015] Furthermore, the compensation for biological response delay includes introducing a delay pre-compensator to calculate the actual physical action that should be applied. : ; The action settings before delay pre-compensation. This is the biological response delay vector. Let be the Jacobian matrix of metabolic flux versus action, representing how changes in each control action affect the metabolic flux of the microorganism. It is the integral variable.

[0016] Furthermore, the continuous regulation guided by the metabolic flux ratio employs model predictive control to drive the metabolic flux distribution to evolve along the optimal trajectory, with the accumulation of flux deviation and the accumulation of control costs as the control objective.

[0017] This invention also provides a fermentation device for flavor-directed regulation of wampee vinegar based on reinforcement learning, which, based on the above-described reinforcement learning-based fermentation method for flavor-directed regulation of wampee vinegar, includes: The fermentation state-space modeling module is used to construct a joint state vector that includes physical sensing variables, biological metabolic variables, and biological response delay variables. The flavor target encoding and flux-oriented reward function setting module is used to quantify flavor into a target flavor fingerprint vector and a target metabolic flux distribution. The reward function includes flavor matching, metabolic efficiency, microbial health and regulatory cost. A continuous action decision module is used to make action decisions based on a physical information deep Q-network. The actions include temperature change rate setpoint, dissolved oxygen change rate setpoint, nutrient salt flow valve opening, and functional bacteria supplementation switch. The network includes an input layer, a dual-colony coupled feature encoding layer, a shared feature extraction layer, a strain parameter embedding layer, an Actor output layer, a Critic output layer, and a physical-biological constraint barrier layer. The multi-level intervention execution module is used to convert the output continuous actions into execution instructions for the actual fermenter, while compensating for biological response delays and implementing three-level regulation based on environmental parameters, microbial communities, and metabolic pathways. The self-evolutionary learning and flux feedback optimization module is used to continuously update the PI-DQN network parameters and flux model based on the deviation between the actual fermentation results and the predictions.

[0018] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing program instructions for a reinforcement learning-based fermentation method for flavor-oriented regulation of yellow peel vinegar. The program instructions for the reinforcement learning-based fermentation method for flavor-oriented regulation of yellow peel vinegar can be executed by one or more processors to implement the steps of the reinforcement learning-based fermentation method for flavor-oriented regulation of yellow peel vinegar as described above.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention achieves specialized monitoring and regulation of the characteristic aroma components of yellow peel fruit vinegar by using state-space modeling that includes terpene concentration, thus solving the problem of insufficient characteristic flavor of the product caused by neglecting terpene retention in traditional methods.

[0020] By employing a dual-colony coupling feature encoding layer and a collaborative reward function, an effective modeling of the coupling dynamics between yeast and acetic acid bacteria in simultaneous fermentation was achieved, solving the control challenge of non-monotonic coupling between the two colonies. A strain parameter embedding layer incorporated the physiological prior knowledge of strain A-5 into the network, significantly reducing the number of samples required for training and avoiding the risk of colony collapse due to trial-and-error learning. A physical-biological constraint barrier layer embedded the physical limits and biological tolerance boundaries of the fermenter into the Q-value calculation, ensuring that the output action remains within the feasible region and improving control safety.

[0021] By pre-compensating for microbial response lag using a delay compensation module, more precise metabolic flux-oriented regulation was achieved. Through priority experience replay and posterior reward mechanisms, continuous self-evolution of the strategy was realized, improving batch-to-batch flavor consistency. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of a fermentation method for flavor-oriented regulation of yellow peel fruit vinegar based on reinforcement learning, according to an embodiment of this application. Detailed Implementation

[0023] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0024] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0025] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. The illustrations only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the shape, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0026] See Figure 1 A method for flavor-directed fermentation regulation of yellow-skinned fruit vinegar based on reinforcement learning includes the following steps: S1: Fermentation state space modeling; Construct a joint state vector that includes physical sensing variables, biological metabolic variables, and biological response delay variables. It is used to fully describe the dynamics of the fermentation system.

[0027] ; in, For physical sensing variables, As a biological metabolic variable, This is a biological response delay variable.

[0028] S11: Physical state variables; The physical state variables Obtained through direct measurement or soft measurement using online sensors: ; The fermentation temperature (°C) is obtained by thermocouple measurement. Dissolved oxygen concentration (mg / L) is measured using a polarographic electrode; pH is acidity / alkalinity measured using a glass electrode; SG is specific gravity measured using a tuning fork densitometer, indirectly reflecting total solids and ethanol content. The stirring power (W) is calculated using the motor current. The ventilation rate is measured using a mass flow meter.

[0029] S12: Biological metabolic variables; The biological metabolic variables Online near-infrared spectroscopy acquisition, combined with real-time estimation using a partial least squares (PLS) model: ; in, The viable cell density (CFU / mL) was described using a Logistic growth model. ; in For the maximum specific growth rate, For maximum carrying capacity, This represents the initial seeding cell density.

[0030] The concentration is ethanol, calculated using a specific gravity and temperature compensation model.

[0031] Total acidity (in acetic acid) is estimated by combining pH and conductivity.

[0032] It represents the concentration of various characteristic esters, including ethyl acetate, isoamyl acetate, etc.

[0033] The concentration refers to terpenoids, including terpinene-4-ol, juniper, and limonene. Terpinene-4-ol is a characteristic aroma component in the peel of wampee, giving wampee vinegar its unique wampee flavor base. Juniper is one of the characteristic aroma substances of wampee, and limonene is an aroma substance common to citrus fruits.

[0034] The metabolic flux allocation coefficient is estimated from measurable concentrations using an extended Kalman filter.

[0035] Traditional fruit vinegar state modeling focuses only on esters and acids, but the unique flavor of yellow-skinned fruit vinegar largely comes from the retention of these terpenes. During fermentation, terpenes are prone to oxidation, cyclization, and other transformation reactions, requiring specialized monitoring and control.

[0036] S13: Delayed variables in biological response; The biological response delay variable Define the characteristic response times of microorganisms to different physical stimuli and dynamically identify them through online impulse response experiments: ; in, This is a characteristic time for enzyme activity regulation under temperature stimulation, and is related to membrane lipid phase transition and enzyme conformational rearrangement. The time required for respiratory chain rearrangement following changes in dissolved oxygen reflects the lag in the expression of electron transport chain complexes. The osmotic pressure response time involves the synthesis and transport of compatible solutes (trehalose, proline); Delayed gene expression is determined by transcription, translation, and protein folding processes.

[0037] These delayed variables were identified and updated through an online system: small-amplitude step perturbations were actively applied during fermentation, metabolite concentration response curves were recorded, and the results were fitted using a first-order inertial plus pure hysteresis model. .

[0038] S14: Constructing the state dynamics equations The rate of change of the state of the fermentation system over time is determined by two parts: one part is the deterministic evolutionary trend under the combined effects of microbial metabolism and environmental regulation, described by the fermentation kinetic function F(S,a); the other part is the random fluctuations caused by various unmeasurable perturbations and model errors, described by ξ(t). This can be represented by a system of nonlinear differential equations: ; Where S is the control action, a is the control action vector, and F(·,·) is the nonlinear fermentation kinetic function; It is a zero-mean Gaussian process noise.

[0039] All states are min-max normalized before being input into the decision network: ; S2: Flavor Target Encoding and Flux-Oriented Reward Function The optimal flavor preferred by consumers is quantified as a target flavor fingerprint vector and a target metabolic flux distribution. The reward function is designed as a four-term weighted sum, taking into account flavor matching, metabolic efficiency, microbial health, and regulatory costs.

[0040] S21: Flavor Matching Bonus Modified cosine similarity is used, taking into account both the direction and absolute concentration of the flavor profile: ; The first term is cosine similarity, which measures the directional consistency of flavor profiles; the second term is modulus penalty. The weighting coefficients are used to avoid pursuing only contour similarity while ignoring absolute concentration.

[0041] It is a characteristic ester concentration vector, represented as , where p is the number of characteristic esters. Each component This indicates the concentration of the j-th characteristic ester in the fermentation broth, in milligrams per liter. Typical characteristic esters in yellow peel fruit vinegar include ethyl acetate (contributing to fruit aroma), isoamyl acetate (contributing to banana aroma), and ethyl butyrate (contributing to pineapple aroma), etc. It is estimated in real time by combining online Raman spectroscopy with a partial least squares model.

[0042] It is the target flavor fingerprint vector, represented as Each component This represents the target concentration of the j-th characteristic ester, with units equal to... same. It was determined based on consumer sensory evaluation experiments and represents the ester concentration combination corresponding to the optimal flavor.

[0043] It is the dot product operation, which represents the sum of the products of corresponding components of two vectors. The larger the dot product value, the better the consistency between the actual ester concentration vector and the target vector in terms of direction.

[0044] Cosine similarity measures the directional consistency between two vectors, regardless of their magnitudes. When the actual flavor profile perfectly matches the target flavor profile, the cosine similarity is 1; when their directions are completely opposite, the cosine similarity is 0. Using only cosine similarity leads to a problem: even with extremely low total ester concentrations, high scores can still be achieved as long as the proportions of each ester are correct. Therefore, a second term is needed for correction.

[0045] This is the sum of squares of the differences between the actual flavor vector and the target flavor vector. This term penalizes the absolute deviation in concentration. The larger the difference between the actual and target concentrations, the larger this term becomes.

[0046] α is the modulus penalty weighting coefficient, with a value of 0.05. α controls the degree to which concentration deviation affects the reward. The larger α is, the heavier the penalty for concentration deviation; the smaller α is, the greater the weight of directional consistency. A value of 0.05 means that when the sum of squares of concentration deviations reaches 20, the second term (1-0.05×20)=0, and the reward drops to zero; when it exceeds 20, the reward becomes negative. This ensures that the strategy not only pursues the correct flavor profile but also ensures that the concentrations of each ester reach the target level.

[0047] S22: Metabolic Flux Matching Reward Drive metabolic flux toward the optimal esterification pathway, and define a flux deviation penalty: ; in, Sensitivity coefficient The vector of coefficients for actual metabolic flux allocation is represented as follows: The meanings of these four components are as follows: It is the proportion of esterification flux to total carbon flux. It is the proportion of acetic acid production flux to total carbon flux. It represents the proportion of ethanol consumption flux to total carbon flux. It represents the proportion of terpene retention flux to total carbon flux. The sum of the four components is 1.

[0048] It is the target metabolic flux allocation coefficient vector, denoted as .

[0049] It is the squared error between the actual flux ratio and the target flux ratio. The smaller the squared error, the closer the distribution of metabolic flux is to the target.

[0050] S23: Microbial Health Reward Integrating cell activity and environmental stress factors: ; in, This is the current live bacterial cell density. During the simultaneous fermentation of wampee fruit vinegar, It is the sum of yeast density and acetic acid bacteria density, estimated in real time using online capacitance method or near-infrared spectral model.

[0051] Maximum carrying capacity, that is, the maximum cell density that microorganisms can achieve in the fermenter, when the cell density approaches... When the ratio approaches 1, it indicates that the bacterial community has reached a plateau and further growth is limited.

[0052] This is the current ethanol concentration. Ethanol is the main metabolic product of yeast and also a substrate for acetic acid bacteria. However, excessively high ethanol concentrations can be toxic to both bacteria.

[0053] This is the ethanol tolerance concentration, which represents the maximum ethanol concentration that microorganisms can tolerate. Exceeding this value will severely inhibit their growth and metabolism.

[0054] This item means: when hour, Since max(0, negative) = 0, this term does not incur a penalty; conversely, it incurs a positive penalty. This reflects the threshold effect, meaning that ethanol is harmless below the tolerable concentration; however, its toxicity increases rapidly beyond the tolerable concentration.

[0055] γ1 is the ethanol toxicity sensitivity coefficient. The role of γ1 is to control the severity of the penalty after ethanol exceeds the limit.

[0056] |dT / dt| is the absolute value of the rate of temperature change, expressed in degrees Celsius per hour. The rate of temperature change reflects the drastic nature of temperature adjustment. Microorganisms are sensitive to abrupt temperature changes; drastic temperature fluctuations trigger the expression of heat shock proteins, leading to metabolic remodeling and decreased cell activity.

[0057] It is the maximum rate of temperature change that microorganisms can tolerate.

[0058] The meaning is similar to the ethanol penalty: when the rate of temperature change is within the tolerance range, it is 0, with no penalty; when it exceeds the tolerance range, a positive penalty is generated.

[0059] γ2 is the heat shock sensitivity coefficient.

[0060] The product of the three factors constitutes When cell density is high, ethanol levels are within limits, and temperature changes are gradual, Approaching 1; when any condition worsens, The corresponding decrease.

[0061] S24: Regulatory Costs and Penalties Avoid energy waste and excessive actuator movement: ; 'a' is the current control action vector, represented as... .in Set the temperature change rate in degrees Celsius per hour. Set the dissolved oxygen change rate, in milligrams per liter per hour; The valve opening is set to the nutrient salt flow, with a value ranging from 0 to 1; Add a switch to the functional bacteria, with a value of 0 or 1.

[0062] It is the reference action vector, representing the baseline action when there is no optimization requirement.

[0063] It is the sum of squares of the differences between the current action and the reference action. This term penalizes the deviation of the action from the reference.

[0064] ||da / dt|| is the modulus of the rate of change of the action, representing the degree of drastic change of the action over time.

[0065] λ² is the penalty coefficient for action variation. Its function is to suppress drastic fluctuations in action. In fermentation control, smooth operations (such as slow temperature adjustments) are easier to achieve and more friendly to equipment and microorganisms than static deviations (such as maintaining a non-zero rate of temperature change).

[0066] Since it's a negative value, it's actually a penalty term. The more the movement deviates from the baseline, and the more drastic the change in movement, the more likely it is to be penalized. The larger the negative value (the larger the absolute value), the lower the total reward.

[0067] S25: Total Reward Function ; Weights are determined using Bayesian optimization. .

[0068] S3: Continuous Action Decision Based on Physical Information Deep Q-Network (PI-DQN) This step outputs continuous control actions. And ensure that the actions are within the physically feasible domain and the biosafety domain.

[0069] S31: Action Space Definition ; in Set the temperature change rate in degrees Celsius per hour. Set the dissolved oxygen change rate, in milligrams per liter per hour; The valve opening is set to the nutrient salt flow, with a value ranging from 0 to 1; Add a switch to the functional bacteria, with a value of 0 or 1.

[0070] S32: Establish a Physical Information-Based DQN Network The goal of the DQN network is to learn a mapping strategy from fermentation states to optimal control actions. This strategy must be able to handle multiple complex factors in the simultaneous fermentation of yellow peel fruit vinegar, including the coupled dynamics of two microbial communities, strain-specific physiological parameters, non-stationarity of fermentation stages, and physical and biological constraints. The network input is the coupled state vector. The vector contains the original state and the dual-colony coupling features, and the output is a continuous control action a(t).

[0071] The Physical Information Deep Q-Network (PQN) employs an Actor-Critic architecture, where the Actor network outputs actions and the Critic network evaluates state-action value. The network's unique feature lies in embedding dual-colony coupling representations, strain-specific parameters, and physical constraints specific to wampee vinegar into the network structure in a differentiable hierarchical manner. The PQN comprises the following layers from input to output: input layer, dual-colony coupling feature encoding layer, shared feature extraction layer, strain parameter embedding layer, Actor output layer, Critic output layer, and physical-biological constraint barrier layer. These layers are connected sequentially according to the data flow direction, forming an end-to-end trainable network.

[0072] (a) Establishment of the dual-species coupled feature coding layer The dual-colony coupled feature encoding layer is the first processing layer of the network. Its function is to nonlinearly combine the yeast and acetic acid bacteria related features in the original state vector to generate a coupled representation vector. This layer does not contain trainable parameters, but instead designs a fixed feature transformation function based on fermentation engineering knowledge.

[0073] The input to this layer is the community-related component in the original state vector, including yeast density. Acetic acid bacteria density ethanol concentration Total acidity .

[0074] This layer computes five coupling features.

[0075] The first coupling feature is the dual-community synergistic activity index. The calculation formula is: ; in and This represents the initial inoculation density. This indicator comprehensively reflects the relative activity levels of the two bacterial communities.

[0076] The second coupling characteristic is the rate of change of synergistic activity. The calculation formula is: ; Where Δt = 1 hour. This feature reflects the evolutionary trend of the cooperative state.

[0077] The third coupling feature is the yield of ethanol per unit of yeast. The calculation formula is: ; This characteristic reflects the efficiency of yeast alcoholic fermentation.

[0078] The fourth coupling characteristic is the acid production per unit of acetic acid bacteria. The calculation formula is: ; This characteristic reflects the acetic acid fermentation efficiency of acetic acid bacteria.

[0079] The fifth coupling feature is the abundance ratio of the two bacterial communities. The calculation formula is: ; This feature is used to determine which type of bacteria is dominant.

[0080] The output of this layer is a coupled feature vector. This vector, when concatenated with the original state vector, forms a complete coupled state representation. .

[0081] (b) Establishment of a shared feature extraction layer The shared feature extraction layer consists of three fully connected hidden layers. Its function is to extract high-level abstract features from the coupled state, which will be used for both Actor decision-making and Critic evaluation.

[0082] The calculation for the first hidden layer is as follows: ; Where W1 is the weight matrix, with dimensions 128×dim( b1 is the bias vector with a dimension of 128; ReLU is the activation function. The calculation for the second hidden layer is as follows: ; Where W2 is a 128×128 weight matrix and b2 is a 128-dimensional bias vector.

[0083] The calculation for the third hidden layer is as follows: ; Where W3 is a 128×128 weight matrix and b3 is a 128-dimensional bias vector. The 128-dimensional shared features will serve as the common input for all subsequent task heads.

[0084] (c) Establishment of strain parameter embedding layer The role of the strain parameter embedding layer is to embed the physiological parameters of Acetic Acid Bacteria A-5 strain into the network in the form of trainable basis vectors, so that the strategy can obtain prior knowledge of strain characteristics in the early stage of training.

[0085] This layer first defines a vector of strain-specific parameters. : ; Based on experimental data from strain A-5, these parameters can optionally be initialized as follows: , , , , , , , These parameters can be adjusted slightly during training (limited to ±20% of the initial values) to adapt to the characteristics of different batches of raw materials.

[0086] This layer calculates the optimal action benchmark for the strain. : ; in , , , This benchmark represents the recommended control actions under optimal strain conditions.

[0087] This layer also calculates a strain fitness penalty term, which is used to adjust the Q-value in the Critic network: ; in is the penalty coefficient. This penalty term results in lower Q values ​​for actions that deviate from the strain's optimal conditions, thus guiding the Actor network to tend to output actions that conform to the strain's characteristics.

[0088] This layer also contains a strain switching decision subnetwork, which is structured as a fully connected layer with a sigmoid activation function: ; in For shared features, The weight matrix is ​​4×128. For scalar bias. Output This indicates the probability that functional bacteria need to be added.

[0089] (d) Establishment of the Actor output layer The Actor output layer maps shared features to specific control actions. This layer contains two sub-layers: an action generation layer and an action scaling layer.

[0090] The action generation layer is a fully connected layer: ; in The weight matrix is ​​4×128. Given a 4-dimensional bias vector, the tanh activation function restricts the output to the interval (-1, 1). This is a 4-dimensional original action vector.

[0091] The motion scaling layer maps the original motion to the actual control range. The mapping formula for the temperature change rate setpoint is: ; The output range is -0.5 to +0.5 degrees Celsius per hour.

[0092] The mapping formula for the dissolved oxygen change rate setpoint is: ; The output range is -0.3 to +0.3 mg / L / hour.

[0093] The mapping formula for the nutrient salt flow valve opening is: ; The output range is 0 to 1.

[0094] The mapping formula for the functional bacteria supplementation switch is: like >0.7 and [3]>0, then ,otherwise .

[0095] That is, the strain switching subnetwork and the Actor network jointly determine whether to add functional bacteria.

[0096] (e) Establishment of the Critic output layer The Critic output layer evaluates the state-action value Q. This layer receives shared features. And action 'a' as input.

[0097] First, the shared features and actions are concatenated: ;

[0098] The concatenated vector has a dimension of 132 (128+4).

[0099] Then it is processed through two fully connected layers: ; in The weight matrix is ​​128×132. It is a 128-dimensional bias vector.

[0100] ; in The weight matrix is ​​64×128. It is a 64-dimensional bias vector.

[0101] Finally, output the original Q value: ; in The weight matrix is ​​1×64. This is a scalar bias.

[0102] (f) Establishment of a physical-biological constraint barrier layer The physical-biological constraint barrier layer serves to embed the physical limits and biological tolerance boundaries of wampee vinegar fermentation into the Q-value calculation. This layer receives the raw Q-value output from the Critic output layer. Given the current state S as input, output the constrained Q value. .

[0103] This layer first defines five constraint functions.

[0104] The first constraint is the ethanol toxicity boundary: ; in This is the ethanol tolerance concentration.

[0105] The second constraint is the pH inhibition boundary: ; The third constraint is the temperature shock boundary: ; in This represents the maximum rate of temperature change that microorganisms can tolerate. Set the temperature change rate.

[0106] The fourth constraint is the oxygen transfer rate limit: ; in, Determined by the maximum stirring power and the aeration flow rate, This represents the saturated dissolved oxygen concentration. This represents the current dissolved oxygen concentration in the fermentation broth. Set the dissolved oxygen change rate as a value.

[0107] The fifth constraint is the total acid accumulation boundary: ;

[0108] This layer calculates the barrier penalty term for each constraint. For the i-th constraint, the formula for calculating the barrier penalty term is: ; in To prevent small quantities with negative infinity logarithms. The adaptive penalty coefficient is calculated using the following formula: ; in The warning threshold is set to 0.8 times the upper limit of each constraint. When the constraint function value... Less than the warning threshold When, the penalty coefficient is 10; when As the penalty coefficient is further reduced, it increases linearly.

[0109] This layer calculates the total barrier penalty: ; Finally, output the constrained Q value: ; When the action approaches the constraint boundary This forces the strategy to avoid choosing such actions.

[0110] S4: Multi-level intervention implementation integrating bio-physical approaches The continuous output actions are translated into actual fermentation tank execution instructions, while compensating for biological response delays and implementing three-level regulation from environment to metabolism.

[0111] S41: Bio-physical fusion actuator based on delay compensation module Because microorganisms exhibit a lag in their response to stimuli, directly applying an action will cause the effective stimulus to lag behind the set value. A delay pre-compensator is introduced to calculate the actual physical action that should be applied. : ; Set the action value for the output layer of step S3Actor. This is the biological response delay vector. Let be the Jacobian matrix of metabolic flux versus action, representing how changes in each control action affect the metabolic flux of the microorganism. It is the integral variable.

[0112] Compensated actions It is sent to the underlying execution mechanism.

[0113] S42: Continuous regulation guided by metabolic flux ratio Traditional methods only maintain the alcohol-acid ratio within a certain range, while this method uses model predictive control (MPC) to drive the metabolic flux distribution to evolve along the optimal trajectory.

[0114] Define flux distribution coefficient: ; For esterification reaction flux, For acid production flux, This represents the ethanol consumption flux. ,satisfy .

[0115] Flux kinetics model: ; Among them, parameters The catalytic constant of the esterification reaction, The Michaelis constant for esterification reaction, The degradation rate constant for esters is determined through offline experiments.

[0116] Similar equations are applicable and .

[0117] Real-time flux estimation: using an extended Kalman filter (EKF): ; in, This represents the metabolic flux allocation coefficient at time t, using all measurement information up to time t. The optimal estimate made; This represents the predicted value of the metabolic flux allocation coefficient at time t, using measurement information up to time t-1. The Kalman gain matrix; Let be the vector of measured values ​​at time t; C is the observation matrix.

[0118] Optimal flux trajectory planning: Solving the optimal control problem to minimize the flavor deviation throughout the fermentation cycle. ; in, This is the point at which fermentation ends; This is the actual flux allocation coefficient. The target metabolic flux allocation coefficient; The weighted squared error represents the sum of squares of the deviations between the actual flux distribution and the target flux distribution, weighted by the weight matrix Q.

[0119] The physical meaning of this formula is: to find an optimal sequence of control actions a(t) that minimizes two objectives simultaneously. The first objective is the accumulation of flux deviation, i.e., the desire for carbon flow distribution to remain close to the ideal state throughout the fermentation process; the second objective is the accumulation of control costs, i.e., the desire to achieve the objective with the minimum control intensity.

[0120] Constraints include flux dynamics and physical-biological constraints. A rolling time-domain optimization method is employed, solving the problem every 15 minutes to obtain the current optimal action MPC output. Then, it is weighted and fused with the action output of PI-DQN: ; Early stage As training progressed, the value gradually decreased to 0.2.

[0121] S43: Three-level intervention implementation Level 1: Environmental parameter regulation (based on) Temperature, dissolved oxygen, nutrient salt flow) Temperature control: The maximum rate of change of circulating water through the jacket is constrained by the temperature shock boundary. Dissolved oxygen control: cascade control, with the inner loop controlling the stirring speed and the outer loop controlling the ventilation volume;

[0122] Nutrient salt flow valve opening: according to Add a compound nitrogen source in a specific ratio.

[0123] Level 2: Microbial Community Regulation When flux estimation shows And when this continues for more than 2 hours, it triggers the replenishment of functional bacteria ( ), and supplement with domesticated high-esterase-producing acetic acid bacteria and aroma-producing yeast.

[0124] Level 3: Regulation of metabolic pathways Pre-treatment feeding: When acetic acid concentration >30g / L and ester concentration <1mg / L, follow the instructions. Ethanol was added to maintain the alcohol-acid molar ratio between 1:1 and 2:1, and the real-time alcohol-acid molar ratio was monitored by NIR.

[0125] Enzyme activity regulation: The ratio of alcohol dehydrogenase (ADH) to aldehyde dehydrogenase (ALDH) activity is indirectly controlled by adjusting dissolved oxygen levels: low dissolved oxygen (DO < 0.5 mg / L) promotes ADH activity, which is beneficial for esterification; high dissolved oxygen promotes ALDH activity, which is beneficial for acid production.

[0126] S5: Self-evolutionary learning and flux feedback optimization By analyzing the discrepancies between actual fermentation results and predictions, the PI-DQN network parameters and flux model are continuously updated to achieve the evolution of the control strategy.

[0127] S51: Experience Replay and Priority Sampling The experience tuples from each decision are stored in the replay buffer. Experience-first replay is used, and sampling weights are assigned based on the absolute value of the temporal difference error. ; in, Let be the sampling probability of the i-th sample. For timing difference error, This is the priority index.

[0128] S52: Posterior Rewards and Batch Stability Constraints After fermentation, the actual flavor concentration and final flux were detected offline by GC-MS, and the posterior reward was calculated. ; in, This represents the actual ester concentration vector after fermentation. For the target flavor fingerprint vector; This represents the batch-to-batch flavor variance. This is the allowable variance threshold; This is the penalty coefficient.

[0129] like If the predicted reward is less than the set target, the batch of data is marked as a difficult sample, its sampling priority is increased, and model fine-tuning is triggered.

[0130] S53: Transfer Learning and Mass Production Adaptation When switching raw material batches, such as wampee fruit from different origins, or expanding the scale of fermentation tanks, parameters can be migrated to quickly adapt to the new environment: This includes freezing the feature extraction layer of PI-DQN, i.e., the first two fully connected layers; By only fine-tuning the coefficients of the output layer and constraint barrier layer, the required training samples are reduced from approximately 80 batches to less than 20 batches; The isoparameters in the flux kinetic model were recalibrated through 3-5 batches of short-term experiments.

[0131] S54: Continuous Evolutionary Cycle Automatic execution after each fermentation cycle: (1) Calculate the posterior reward Update flavor targets The Bayesian posterior distribution is used for consumer preference drift; (2) Compress the complete batch trajectory into keyframes (state-action-reward sequence) and store them in long-term memory; (3) Replay historical best batches in the long-term memory bank offline every week to perform strategy distillation and prevent catastrophic forgetting.

[0132] In this embodiment, specialized monitoring and regulation of the characteristic aroma components of yellow peel fruit vinegar were achieved, solving the problem of insufficient product characteristic flavor caused by neglecting terpene retention in traditional methods. Through a dual-microbial community coupling feature encoding layer and a cooperative reward function, effective modeling of the coupling dynamics between yeast and acetic acid bacteria in simultaneous fermentation was achieved, solving the control problem of non-monotonic coupling between the two microbial communities. By incorporating the physiological prior knowledge of strain A-5 into the network through a strain parameter embedding layer, the amount of samples required for training was significantly reduced, avoiding the risk of microbial community collapse caused by trial-and-error learning. By embedding the physical limits and biological tolerance boundaries of the fermenter into the Q-value calculation through a physical-biological constraint barrier layer, the output action was ensured to be within the feasible region, improving control safety. More precise metabolic flux-oriented regulation was achieved through a delay compensation module to pre-compensate for microbial response lag. Continuous self-evolution of the strategy was achieved through priority experience replay and posterior reward mechanisms, improving batch-to-batch flavor consistency.

[0133] This invention also proposes a reinforcement learning-based fermentation device for flavor-directed regulation of wampee vinegar, based on the reinforcement learning-based fermentation method for flavor-directed regulation of wampee vinegar described above, comprising: The fermentation state-space modeling module is used to construct a joint state vector that includes physical sensing variables, biological metabolic variables, and biological response delay variables. The flavor target encoding and flux-oriented reward function setting module is used to quantify flavor into a target flavor fingerprint vector and a target metabolic flux distribution. The reward function includes flavor matching, metabolic efficiency, microbial health and regulatory cost. A continuous action decision module is used to make action decisions based on a physical information deep Q-network. The actions include temperature change rate setpoint, dissolved oxygen change rate setpoint, nutrient salt flow valve opening, and functional bacteria supplementation switch. The network includes an input layer, a dual-colony coupled feature encoding layer, a shared feature extraction layer, a strain parameter embedding layer, an Actor output layer, a Critic output layer, and a physical-biological constraint barrier layer. The multi-level intervention execution module is used to convert the output continuous actions into execution instructions for the actual fermenter, while compensating for biological response delays and implementing three-level regulation based on environmental parameters, microbial communities, and metabolic pathways. The self-evolutionary learning and flux feedback optimization module is used to continuously update the PI-DQN network parameters and flux model based on the deviation between the actual fermentation results and the predictions.

[0134] Furthermore, this embodiment of the invention also proposes a computer-readable storage medium storing program instructions for a reinforcement learning-based fermentation method for flavor-directed regulation of yellow peel vinegar. These program instructions can be executed by one or more processors to implement the steps of the reinforcement learning-based fermentation method for flavor-directed regulation of yellow peel vinegar as described above.

[0135] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for directional flavor regulation fermentation of yellow-skinned fruit vinegar based on reinforcement learning, characterized in that, include: Fermentation state-space modeling includes constructing a joint state vector containing physical sensing variables, biological metabolic variables, and biological response delay variables, and constructing state dynamic equations; Flavor target encoding and flux-oriented reward function setting include quantifying flavor into target flavor fingerprint vectors and target metabolic flux distributions. The reward function includes flavor matching reward, metabolic efficiency reward, microbial health reward and regulatory cost. Continuous action decision-making includes action decision-making based on a physical information deep Q-network. The actions include temperature change rate setpoint, dissolved oxygen change rate setpoint, nutrient salt flow valve opening, and functional bacteria supplementation switch. The network includes an input layer, a dual-colony coupled feature encoding layer, a shared feature extraction layer, a strain parameter embedding layer, an Actor output layer, a Critic output layer, and a physical-biological constraint barrier layer. Multi-level intervention execution includes translating the output of continuous actions into actual fermenter execution instructions, while compensating for biological response delays and continuous regulation guided by metabolic flux ratio, and implementing three-level regulation based on environmental parameters, microbial community and metabolic pathways. Self-evolutionary learning and flux feedback optimization, including continuously updating the physical information deep Q network parameters and flux model based on the deviation between actual fermentation results and predictions.

2. The method for flavor-directed fermentation regulation of yellow peel fruit vinegar based on reinforcement learning according to claim 1, characterized in that, The biological response delay variables are dynamically identified through online impulse response experiments, including the characteristic time of enzyme activity regulation under temperature stimulation, the time of respiratory chain rearrangement after dissolved oxygen changes, osmotic pressure response time, and gene expression delay.

3. The method for flavor-directed fermentation regulation of yellow peel fruit vinegar based on reinforcement learning according to claim 1, characterized in that, The flavor matching reward uses modified cosine similarity, taking into account both the direction and absolute concentration of the flavor profile: ; in S is the characteristic ester concentration vector, and S is the state vector. It is the target flavor fingerprint vector. It's a dot product operation. It is the sum of squares of the differences between the actual flavor vector and the target flavor vector, and α is the modulus penalty weighting coefficient.

4. The method for flavor-directed fermentation regulation of yellow peel fruit vinegar based on reinforcement learning according to claim 1, characterized in that, The dual-colony coupling feature encoding layer in the physical information-based deep Q-network nonlinearly combines the yeast and acetic acid bacteria related features in the original state vector to generate a coupling representation vector, resulting in five coupling features: dual-colony synergistic activity index, synergistic activity change rate, unit yeast ethanol yield, unit acetic acid bacteria acid production, and dual-colony abundance ratio.

5. The method for flavor-directed fermentation regulation of yellow peel fruit vinegar based on reinforcement learning according to claim 4, characterized in that, The strain parameter embedding layer in the physical information-based deep Q-network includes defining a strain-specific parameter vector and calculating the optimal action benchmark for the strain and a strain adaptability penalty term based on the parameter vector.

6. The method for flavor-directed fermentation regulation of yellow peel fruit vinegar based on reinforcement learning according to claim 5, characterized in that, The physical-biological constraint barrier layer includes five constraint functions: ethanol toxicity boundary, pH inhibition boundary, temperature shock boundary, oxygen transfer rate limit, and total acid accumulation boundary.

7. The method for flavor-directed fermentation regulation of yellow peel fruit vinegar based on reinforcement learning according to claim 1, characterized in that, The compensation for biological response delay includes introducing a delay pre-compensator to calculate the actual physical action that should be applied. : ; The action settings before delay pre-compensation. This is the biological response delay vector. Let be the Jacobian matrix of metabolic flux versus action, representing how changes in each control action affect the metabolic flux of the microorganism. It is the integral variable.

8. The method for flavor-directed fermentation regulation of yellow peel fruit vinegar based on reinforcement learning according to claim 7, characterized in that, The metabolic flux ratio-guided continuous regulation employs model predictive control to drive the metabolic flux distribution to evolve along the optimal trajectory, with the accumulation of flux deviation and control costs serving as the control objective.

9. A fermentation device for flavor-directed regulation of wampee vinegar based on reinforcement learning, based on the flavor-directed regulation fermentation method for wampee vinegar based on reinforcement learning as described in any one of claims 1 to 8, characterized in that, include: The fermentation state-space modeling module is used to construct a joint state vector that includes physical sensing variables, biological metabolic variables, and biological response delay variables. The flavor target encoding and flux-oriented reward function setting module is used to quantify flavor into a target flavor fingerprint vector and a target metabolic flux distribution. The reward function includes flavor matching, metabolic efficiency, microbial health and regulatory cost. A continuous action decision module is used to make action decisions based on a physical information deep Q-network. The actions include temperature change rate setpoint, dissolved oxygen change rate setpoint, nutrient salt flow valve opening, and functional bacteria supplementation switch. The network includes an input layer, a dual-colony coupled feature encoding layer, a shared feature extraction layer, a strain parameter embedding layer, an Actor output layer, a Critic output layer, and a physical-biological constraint barrier layer. The multi-level intervention execution module is used to convert the output continuous actions into execution instructions for the actual fermenter, while compensating for biological response delays and implementing three-level regulation based on environmental parameters, microbial communities, and metabolic pathways. The self-evolutionary learning and flux feedback optimization module is used to continuously update the PI-DQN network parameters and flux model based on the deviation between the actual fermentation results and the predictions.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions for a reinforcement learning-based fermentation method for flavor-directed regulation of yellow peel vinegar. These program instructions can be executed by one or more processors to implement the steps of the reinforcement learning-based fermentation method for flavor-directed regulation of yellow peel vinegar as described in any one of claims 1 to 8.