Seasoning low-temperature fermentation dynamic control method based on reinforcement learning
Through flavor semantic analysis and metabolic path map construction based on reinforcement learning, combined with Riemann manifold optimization, the problem of flavor inconsistency in the low-temperature fermentation of condiments is solved, and intelligent metabolic path structure-driven control is realized, which improves the flavor consistency and response sensitivity of condiments.
Patent Information
- Application Number
- CN202510780794.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The prior art is difficult to achieve dynamic control of the microbial metabolism process during the low-temperature fermentation of condiments, especially when facing different raw material batches, environmental conditions or microbial community variations, the ability to respond to internal structural changes in the metabolic process is lacking, resulting in inconsistent flavor performance.
Using reinforcement learning-based methods, the metabolic path map construction and Riemann manifold reinforcement learning are achieved through flavor semantic analysis, the metabolic path structure-driven dynamic regulation. The method includes receiving flavor target input, constructing a metabolic intent vector, mapping to a non-Euclidean embedding space, calculating structural bias, and generating a control strategy output to adjust fermentation environment parameters.
It realizes intelligent and refined control of the low-temperature fermentation process of condiments, improves flavor consistency and response sensitivity, and can achieve precise metabolic path regulation in nonlinear and strongly coupled industrial biological processes.
Smart Images

Figure CN120353135A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control of fermentation processes, and particularly to a dynamic control method for low-temperature fermentation of condiments based on reinforcement learning. Background Art
[0002] Condiments are an essential part of the food industry, and their flavor quality is affected by multiple factors such as raw materials, processes, and fermentation process control. Especially in low-temperature fermentation condiments based on natural microbial fermentation, such as soy sauce, sauces, and fermented vinegar, the microbial metabolic activities are complex, the time span is long, and the process is significantly affected by environmental disturbances. Therefore, the ability to dynamically adjust fermentation environment parameters becomes an important factor affecting the flavor performance of end products.
[0003] Traditional control methods for condiment fermentation processes are mostly based on empirical rules or static control logics, mainly relying on the setting experience of process personnel for process parameters such as temperature, humidity, ventilation, and stirring, supplemented by regular sampling and analysis means for fermentation process monitoring. When facing different raw material batches, environmental conditions, or microbial community variations, such methods lack the ability to respond to the internal structural changes of the metabolic process and are difficult to achieve precise dynamic control. In addition, the more advanced methods in the prior art mainly focus on establishing physicochemical sensor networks or applying data-driven shallow optimization models, such as fuzzy control and genetic algorithm parameter tuning, and still have not broken through the complete intelligent regulation goal of structure perception - strategy generation - environmental intervention.
[0004] In recent years, with the wide application of artificial intelligence technology, some studies have attempted to introduce machine learning methods into fermentation process modeling and prediction. For example, neural networks are used to fit the mapping relationship between environmental parameters and metabolite concentrations to assist process regulation. However, such methods generally have problems such as strong dependence of the model on specific samples, insufficient generalization ability, inability to explain the structural evolution of the metabolic process, and most models only focus on result prediction and lack the ability to generate control strategies based on metabolic pathway mechanisms. In addition, traditional deep reinforcement learning algorithms have defects such as low convergence efficiency and large search direction deviation in dealing with complex graph structures, path topologies, and dynamic policy optimization in non-Euclidean spaces, and are difficult to be directly applied to industrial biological processes such as condiment fermentation that are nonlinear, strongly coupled, and have strong structural dynamics.
[0005] In view of the above problems, the existing technologies have not provided a complete intelligent fermentation control method that can simultaneously combine semantic input of flavor objectives, representation of metabolic pathway structures, quantification of differences in the embedding space, and output of structural perturbation control. In particular, the path-oriented control idea of establishing an embedding representation in the metabolic structure space and performing policy optimization and intervention output within this space has not been disclosed in the published literature and known patent technologies. The existing technologies generally ignore that the microbial metabolic activities during the fermentation process are essentially a process of structural dynamic evolution, and the core problem lies not only in parameter control itself, but in whether it is possible to perform policy intervention based on the structural expression of the metabolic mechanism.
[0006] Therefore, how to provide a dynamic control method for low-temperature fermentation of condiments based on reinforcement learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0007] An object of the present invention is to propose a dynamic control method for low-temperature fermentation of condiments based on reinforcement learning. The present invention integrates flavor semantic parsing, construction of metabolic pathway maps, and Riemannian manifold reinforcement learning methods, and realizes a dynamic regulation mechanism driven by the metabolic pathway structure during the low-temperature fermentation process of condiments. By perceiving the difference between the current and target metabolic structures in the non-Euclidean embedding space, outputting the metabolic perturbation direction and mapping it into an environmental control instruction, a control system of intelligent perception - policy optimization - environmental intervention is constructed, which has the advantages of strong intelligence, fine control, sensitive response, and high flavor consistency.
[0008] The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to an embodiment of the present invention includes the following steps: S1. Receive the flavor objective input, where the flavor objective is the text description information of the condiment, and convert the flavor objective into a metabolic intention vector; S2. Construct a target metabolic pathway map according to the metabolic intention vector. The target metabolic pathway map is composed of metabolic pathway nodes and connection relationships, and is used to represent the metabolic structure path corresponding to flavor formation; S3. Periodically collect the current fermentation state data at a set time interval during the fermentation process; S4. Construct a current metabolic pathway map based on the current fermentation state data, and map the current metabolic pathway map to a non-Euclidean embedding space to obtain a current metabolic pathway embedding representation; S5. Map the target metabolic pathway map to the same non-Euclidean embedding space as the current metabolic pathway map to obtain a target metabolic pathway embedding representation, and compare the target metabolic pathway embedding representation with the current metabolic pathway embedding representation to calculate the structural deviation; S6. Use the Riemannian manifold strategy optimization method for strategy training and update based on the structural deviation to generate a control strategy output in the non-Euclidean embedding space, where the control strategy output is represented by the metabolic pathway perturbation direction; S7. Convert the control strategy output into fermentation process control instructions, where the fermentation process control instructions include temperature set value, humidity set value, ventilation rate, and stirring frequency; S8. Execute the control instructions to adjust the fermentation environment parameters and complete the regulation of the metabolic pathway evolution direction during the fermentation process.
[0009] Optionally, the S1 specifically includes: S11. Receive the condiment text description information input by the user, where the condiment text description information includes natural language expressions of flavor, aroma, taste, process type, and regional style; S12. Perform word segmentation on the condiment text description information to extract keywords with sensory feature meanings; S13. Match the keywords with the preset flavor-metabolite mapping dictionary, where the flavor-metabolite mapping dictionary records the association relationships between flavor terms and specific metabolites, and obtain one or more metabolite names associated with the keywords; S14. By querying the metabolomic pathway information table established locally, retrieve the metabolomic pathway numbers of each metabolite name and generate a metabolomic pathway index list. The metabolomic pathway information table is a pre-organized structured data file containing metabolites and corresponding metabolomic pathway numbers, starting substrates, intermediate metabolites, end products, and reaction sequences; S15. Based on the metabolomic pathway index list, obtain the association frequencies between the metabolomic pathways and the flavor targets, use the proportion of the association frequencies as the importance weight values of each metabolomic pathway, and construct a metabolic intention vector. The metabolic intention vector is a multi-dimensional numerical sequence used to represent the importance of different metabolomic pathways, where each dimension of the numerical value is obtained by normalizing the association frequency of the corresponding metabolomic pathway.
[0010] Optionally, the S2 specifically includes: S21. Sort the metabolomic pathways according to the importance weight values in the metabolic intention vector, and select the metabolomic pathways with importance weight values higher than the set weight threshold as the target metabolomic pathway set; S22. Search for the structural information corresponding to the target metabolomic pathway set, where the structural information includes the starting substrates, intermediate metabolites, end products, and reaction sequences involved in the target metabolomic pathway; S23. Extract the metabolite entries in the structural information to establish a target metabolomic pathway node set, where each node corresponds to a metabolite and is attached with a position identifier in the metabolomic pathway; S24. Construct a directed connection relationship between the target metabolic pathway nodes according to the sequential reaction relationship of each metabolite in the metabolic pathway, form a structured metabolic pathway dependency chain, and constitute a target metabolic pathway map with the target metabolic pathway nodes and the connection relationship.
[0011] Optionally, the current fermentation state data specifically includes temperature, humidity, pH value, oxygen concentration, carbon dioxide concentration, and metabolite concentration.
[0012] Optionally, the S4 specifically includes: S41. Perform a difference operation on the metabolite concentrations at adjacent sampling time points, calculate the change rate of each metabolite concentration, and mark the metabolites with a change rate higher than the preset rate threshold as active metabolites; S42. Based on the corresponding relationship in the metabolic pathway information table established locally by the active metabolites, construct a set of current metabolic pathway nodes; S43. According to the reaction sequence and conversion direction of the active metabolites in the metabolic pathway, establish a directed connection relationship between the current metabolic pathway nodes, and construct a current metabolic pathway map; S44. Based on the current metabolic pathway map, calculate the path distance between the current metabolic pathway nodes, generate a distance matrix, and the elements of the distance matrix are the shortest reaction step lengths between two nodes in the metabolic pathway; S45. Use Laplacian eigenmaps to process the distance matrix, and map the current metabolic pathway map to a non-Euclidean embedding space by preserving the relative position relationship and topological structure characteristics between the current metabolic pathway nodes. The non-Euclidean embedding space is a structured representation space that does not satisfy Euclidean symmetry and the Pythagorean theorem.
[0013] Optionally, the calculation process of the structural deviation specifically includes: One-to-one correspondence pairing is performed on the nodes with the same metabolite identifier in the target metabolic pathway embedding representation and the current metabolic pathway embedding representation to form a set of node pairs; For each pair of paired nodes, the embedding vector coordinates in the non-Euclidean embedding space are respectively extracted, and the embedding distance offset value of the paired nodes is calculated. The embedding distance offset value adopts the Euclidean distance calculation method and is defined as the square root of the sum of the squares of the coordinate differences of each dimension between the embedding vectors of the two nodes; The node weighting factors of the two nodes in each pair of paired nodes are respectively calculated. The node weighting factors are obtained through a linear weighting method based on the importance weight value of the node in the metabolic pathway and the local topological connectivity in the corresponding metabolic pathway map; Arithmetically average the node weighting factors of the two nodes in the paired node pair as the final weighting factor of the paired node, and multiply the distance offset value of the paired node by the corresponding final weighting factor to obtain a weighted offset value; Sum the weighted offset values of each pair of paired nodes and divide by the sum of the weighting factors of all node pairs to obtain a structural deviation value, which is used to quantify the overall structural difference between the current metabolic pathway map and the target metabolic pathway map in the non-Euclidean embedding space.
[0014] Optionally, the S6 specifically includes: S61. Use the structural deviation value as the basis of the state input, and combine it with the current metabolic pathway embedding representation to construct a state vector in the non-Euclidean embedding space ; S62. Based on the state vector , construct a policy function , where represents the conditional probability distribution of outputting an action under the given state vector , represents the policy function parameter, represents the metabolic pathway perturbation direction vector output at time step ; S63. Define an immediate reward function to represent the structural optimization effect: ; where represents the reward value at time step , represents the structural deviation value at time step , represents the structural deviation value at time step ; S64. Construct a Fisher information matrix based on the current policy function: ; where represents the expectation operator, represents the gradient of the policy logarithmic function with respect to the parameter, represents the transpose operation; S65. Update the policy parameters using the natural gradient optimization method on the Riemannian manifold: ; where represents the policy parameter at time step , represents the policy parameter at time step ; denotes the learning rate, denotes the time step the inverse of the Fisher information matrix at time denotes the discount factor, denotes the time step the reward value at time denotes the gradient of the cumulative expected reward of the policy with respect to the policy parameters; S66. Apply the updated policy parameters to the current state vector to generate a metabolic pathway perturbation direction vector, which is defined in a non-Euclidean embedding space and is used to represent the adjustment direction of the metabolic pathway map in the structural space, as the control policy output.
[0015] Optionally, the S7 specifically includes: S71. According to the metabolic pathway perturbation direction vector output by the control policy, extract the perturbation direction components on each principal component axis in the non-Euclidean embedding space, and map them to the fermentation process parameter adjustment direction and adjustment amplitude according to the positive and negative of the change of the perturbation direction components and the magnitude; S72. Set the mapping rules between the perturbation direction components and each control parameter, where: the temperature set value is positively correlated with the component of the first principal component axis in the perturbation direction component, the humidity set value is positively correlated with the component of the second axis in the perturbation direction component, the ventilation rate is positively correlated with the component of the third axis in the perturbation direction component, and the stirring frequency is positively correlated with the component of the fourth axis in the perturbation direction component; S73. According to the value range of the perturbation direction components of the metabolic pathway perturbation direction vector in the non-Euclidean embedding space, construct a hierarchical adjustment rule for the fermentation process parameters, where: when the absolute value of the perturbation direction component is less than 0.1, it is a weak adjustment, and the adjustment amplitude is set to ±2% of the original set value; when the absolute value of the perturbation direction component is between 0.1 and 0.3, it is a medium adjustment, and the adjustment amplitude is ±5%; when the absolute value of the perturbation direction component is greater than 0.3, it is a strong adjustment, and the adjustment amplitude is ±10%; S74. When the perturbation direction component is positive, increase the set amplitude of the corresponding fermentation process parameter; when the perturbation component is negative, decrease the set amplitude of the corresponding fermentation process parameter, and finally generate a fermentation process control instruction, which includes the temperature set value, humidity set value, ventilation rate, and stirring frequency; S75. Package the fermentation process control instruction into a fermentation control device-level parameter configuration instruction and send it to the fermentation control device control execution unit to complete the dynamic adjustment of the fermentation environment.
[0016] The beneficial effects of the present invention are: First, starting from the flavor target text information, the present invention constructs a semantic association between flavor terms and metabolites, and combines the local metabolic pathway knowledge table to establish a metabolic intention vector, realizing the target mapping from sensory requirements to the metabolic pathway space. This transformation from natural language description to metabolic structure abstraction breaks through the limitation that traditional control systems can only be set based on physical or chemical indicators, providing a basis for personalized flavor control.
[0017] Secondly, by collecting the changes in metabolite concentrations in real time during the fermentation process, constructing a path map reflecting the current microbial metabolic state, and then embedding it into a non-Euclidean space to represent the structural state, the present invention realizes the dynamic structural perception of the metabolic evolution process. This representation method can retain the topological relationship and reaction sequence information of the metabolic pathway map, significantly improving the ability to capture the microscopic changes in complex metabolic processes, and providing a unified and stable representation basis for subsequent quantification of structural deviations.
[0018] In addition, the present invention introduces the Riemannian manifold natural gradient optimization method to train and update the control strategy driven by structural deviations in the non-Euclidean embedding space. Compared with the traditional gradient strategy optimization, this method can perform more natural updates on the geometric structure of the parameter space, improving the efficiency and convergence stability of strategy learning. The resulting strategy output is the metabolic pathway perturbation direction vector, which directly corresponds to the evolution direction in the structural space, and has the advantages of strong interpretability and controllable behavior.
[0019] Finally, the present invention constructs a mapping mechanism between the perturbation vector and the fermentation process parameters, realizes the transformation from structural control to physical control, and establishes a closed-loop fermentation environment regulation driven by the metabolic mechanism. By converting the components of the perturbation direction on different principal component axes into adjustment instructions for parameters such as temperature, humidity, ventilation rate, and stirring frequency, the effective connection between the guidance of the metabolic pathway structure evolution and the production process execution level is realized. Description of the Drawings
[0020] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings: Figure 1 is the overall flowchart of the dynamic control method for low-temperature fermentation of seasonings based on reinforcement learning proposed by the present invention; Figure 2 is the flowchart of the structural deviation calculation of the dynamic control method for low-temperature fermentation of seasonings based on reinforcement learning proposed by the present invention; Figure 3 is the flowchart of the control strategy training process for Riemannian manifold strategy optimization based on structural deviation of the dynamic control method for low-temperature fermentation of seasonings based on reinforcement learning proposed by the present invention. Detailed implementation mode
[0021] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0022] Reference Figures 1-3 , a dynamic control method for low-temperature fermentation of condiments based on reinforcement learning, includes the following steps: S1. Receive a flavor target input, where the flavor target is condiment text description information, and convert the flavor target into a metabolic intention vector; S2. Construct a target metabolic pathway map according to the metabolic intention vector. The target metabolic pathway map is composed of metabolic pathway nodes and connection relationships, and is used to represent the metabolic structure pathway corresponding to flavor formation; S3. Periodically collect current fermentation state data at a set time interval during the fermentation process; S4. Construct a current metabolic pathway map based on the current fermentation state data, and map the current metabolic pathway map to a non-Euclidean embedding space to obtain a current metabolic pathway embedding representation; S5. Map the target metabolic pathway map to the same non-Euclidean embedding space as the current metabolic pathway map to obtain a target metabolic pathway embedding representation, and compare the target metabolic pathway embedding representation with the current metabolic pathway embedding representation to calculate the structural deviation; S6. Based on the structural deviation, use the Riemannian manifold strategy optimization method for policy training and update to generate a control policy output located in the non-Euclidean embedding space. The control policy output is represented by the metabolic pathway perturbation direction; S7. Convert the control policy output into a fermentation process control instruction. The fermentation process control instruction includes a temperature set value, a humidity set value, a ventilation rate, and a stirring frequency; S8. Execute the control instruction to adjust the fermentation environment parameters to complete the regulation of the evolution direction of the metabolic pathway during the fermentation process.
[0023] The present invention proposes a dynamic control mechanism for low-temperature fermentation of condiments based on reinforcement learning, and establishes a complete technical process around multiple key steps including flavor target drive, metabolic pathway structure modeling, non-Euclidean space representation, structural deviation calculation and strategy optimization. By introducing the linkage between flavor semantic analysis and metabolic map, the fuzzy sensory target is converted into a quantifiable metabolic intention expression, and a structured metabolic pathway target map is constructed. The evolution of the metabolic structure is dynamically represented using a non-Euclidean embedding space, and the optimal control of the disturbance direction is achieved in the space, thereby opening up the path between the flavor target, microstructure and fermentation process. Finally, the control strategy output is converted into executable process parameter adjustment instructions, which significantly improves the intelligence level and control accuracy of the fermentation process.
[0024] In this implementation, S1 specifically includes: S11, receiving condiment text description information input by a user, wherein the condiment text description information includes natural language expressions of flavor, aroma, taste, process type and regional style; S12, performing word segmentation processing on the condiment text description information to extract keywords with sensory characteristic meanings; S13, matching the keyword with a preset flavor-metabolite mapping dictionary, wherein the flavor-metabolite mapping dictionary records the association between flavor terms and specific metabolites, to obtain one or more metabolite names associated with the keyword; S14, by querying the locally established metabolic pathway information table, retrieving the metabolic pathway number of each metabolite name, and generating a metabolic pathway index list, wherein the metabolic pathway information table is a pre-organized structured data file containing metabolites and corresponding metabolic pathway numbers, starting substrates, intermediate metabolites, final products, and reaction sequences; S15. Based on the metabolic pathway index list, the association frequency between the metabolic pathway and the flavor target is obtained, the proportion of the association frequency is used as the importance weight value of each metabolic pathway, and a metabolic intention vector is constructed. The metabolic intention vector is a multidimensional numerical sequence used to represent the importance of different metabolic pathways, wherein each dimensional value is obtained by normalizing the association frequency of the corresponding metabolic pathway.
[0025] By parsing the flavor target information of condiments from the text description, combining keyword tokenization, flavor-metabolite mapping dictionary matching, and metabolic pathway information query, the conversion from sensory language to the metabolic structure level is effectively achieved. The words such as aroma and taste in natural language are mapped to specific metabolites, and further associated with the known metabolic pathway numbers to form a metabolic pathway index list, providing a structured input for subsequent map construction and control strategy output. The construction of a metabolic intention vector using the frequency normalization between flavor targets and pathways not only enhances the interpretability of target guidance but also provides a quantifiable semantic basis for measuring structural deviations in non-Euclidean spaces, improving the system's ability to respond to diverse flavor targets.
[0026] In this embodiment, the specific steps of S2 include: S21. Sort the metabolic pathways according to the importance weight values in the metabolic intention vector, and select the metabolic pathways with importance weight values higher than the set weight threshold as the target metabolic pathway set; S22. Search for the corresponding structural information of the target metabolic pathway set, where the structural information includes the starting substrates, intermediate metabolites, end products, and reaction order involved in the target metabolic pathway; S23. Extract the metabolite entries in the structural information to establish a target metabolic pathway node set, where each node corresponds to a metabolite and is attached with a position identifier in the metabolic pathway; S24. According to the sequential reaction relationship of each metabolite in the metabolic pathway, construct a directed connection relationship between the target metabolic pathway nodes to form a structured metabolic pathway dependence chain, and constitute a target metabolic pathway map with the target metabolic pathway nodes and connection relationships.
[0027] This part realizes the specific construction from the metabolic intention vector to the structure map. Based on the path weights, highly relevant metabolic pathways are screened, and their reaction order and substance composition information are extracted to construct a metabolic pathway map structure for the target flavor. By establishing a node set and a directed connection relationship, not only the microscopic reaction process in the pathway is restored, but also a foundation for subsequent graph embedding representation is laid. This map expression method has good topological structure integrity and biological rationality, can truly reflect the metabolic evolution path in the flavor formation process, and provides a unified format for comparative analysis with the real-time state map, making the target path control more targeted and scientifically based.
[0028] In this embodiment, the current fermentation state data specifically includes temperature, humidity, pH value, oxygen concentration, carbon dioxide concentration, and metabolite concentration.
[0029] In this embodiment, the specific steps of S4 include: S41. Perform a difference operation on the metabolite concentrations at adjacent sampling time points, calculate the change rate of each metabolite concentration, and mark the metabolites with a change rate higher than the preset rate threshold as active metabolites; S42. Based on the corresponding relationships in the metabolic pathway information table established locally for the active metabolites, construct the current metabolic pathway node set; S43. According to the reaction order and conversion direction of the active metabolites in the metabolic pathway, establish a directed connection relationship between the current metabolic pathway nodes, and construct the current metabolic pathway map; S44. Based on the current metabolic pathway map, calculate the path distances between the current metabolic pathway nodes, generate a distance matrix, and the elements of the distance matrix are the shortest reaction step lengths between two nodes in the metabolic pathway; S45. Use Laplacian eigenmaps to process the distance matrix, and by retaining the relative position relationship and topological structure features between the current metabolic pathway nodes, map the current metabolic pathway map to a non-Euclidean embedding space, and the non-Euclidean embedding space is a structured representation space that does not satisfy Euclidean symmetry and the Pythagorean theorem.
[0030] Combined with the metabolite change rate, identify active metabolites and construct the current metabolic pathway map, and then map it to a non-Euclidean embedding space to complete the dynamic representation of the structural state. This process retains the topological features of the graph and the reaction logic between nodes, improves the accuracy of structural comparison and the model generalization ability. Using Laplacian eigenmaps further improves the consistency and robustness of graph embedding, provides a spatial basis for subsequent structural deviation quantification and strategy training, and realizes an effective transition from data perception to structural expression.
[0031] In this embodiment, the calculation process of the structural deviation specifically includes: Pair up the nodes with the same metabolite identification in the target metabolic pathway embedding representation and the current metabolic pathway embedding representation one by one to form a set of node pairs; For each pair of paired nodes, extract the embedding vector coordinates in the non-Euclidean embedding space respectively, calculate the embedding distance offset value of the paired nodes, and the embedding distance offset value adopts the Euclidean distance calculation method, which is defined as the square root of the sum of the squares of the coordinate differences of each dimension between the two node embedding vectors; Calculate the node weighting factors of the two nodes in each pair of paired nodes respectively, and the node weighting factors are obtained through a linear weighting method according to the importance weight value of the node in the metabolic pathway and the local topological connectivity in the corresponding metabolic pathway map; Arithmetically average the node weighting factors of the two nodes in the paired node pair as the final weighting factor of the paired nodes, and multiply the distance offset value of the paired nodes by the corresponding final weighting factor to obtain the weighted offset value; Sum the weighted offset values of each pair of paired nodes and divide by the sum of the weighted factors of all node pairs to obtain a structural deviation value, which is used to quantify the overall structural difference between the current metabolic pathway map and the target metabolic pathway map in the non-Euclidean embedding space.
[0032] A structural deviation quantification mechanism is proposed, which can accurately measure the structural difference between the target metabolic pathway embedding representation and the current metabolic pathway embedding representation. By constructing a set of node pairs, calculating the embedding distance, and combining a weighting mechanism based on path weight and connectivity, a normalized structural difference value is finally formed. This method not only improves the sensitivity to the metabolic deviation trend, but also provides a state input with directionality and strong interpretability for policy training, enabling the control policy to be continuously iteratively optimized and converge to the target structure.
[0033] In this embodiment, step S6 specifically includes: S61. Use the structural deviation value as the basis of the state input, and combine it with the current metabolic pathway embedding representation to construct a state vector in the non-Euclidean embedding space ; S62. Based on the state vector , construct a policy function , where represents the conditional probability distribution of outputting action under the given state vector , represents the policy function parameters, represents the metabolic pathway perturbation direction vector output at time step ; S63. Define an immediate reward function to represent the structural optimization effect: ; where, represents the reward value at time step , represents the structural deviation value at time step , represents the structural deviation value at time step ; S64. Construct a Fisher information matrix based on the current policy function: ; where, represents the expectation operator, represents the gradient of the policy logarithmic function with respect to the parameters, represents the transpose operation; S65. Update the policy parameters using the natural gradient optimization method on the Riemannian manifold: ; where represents the policy parameters at time step , represents the policy parameters at time step , represents the learning rate, represents the inverse of the Fisher information matrix at time step , represents the discount factor, represents the reward value at time step , represents the gradient of the cumulative expected reward of the policy with respect to the policy parameters; S66. Apply the updated policy parameters to the current state vector to generate a metabolic pathway perturbation direction vector, which is defined in a non-Euclidean embedding space and represents the adjustment direction of the metabolic pathway map in the structural space, as the control policy output.
[0034] Introduce the Riemannian manifold policy optimization method to implement the training and update of the control policy in the non-Euclidean structural space. By constructing the state vector, defining the reward function, constructing the Fisher information matrix, and using the natural gradient for parameter iteration, the efficient optimization of the policy function in the structural space is achieved. The output form of the control policy is the metabolic pathway perturbation direction vector, which can accurately guide the metabolic pathway map to evolve along the optimal structural direction, thereby realizing behavior intervention at the map level, greatly improving the interpretability, response speed, and biological mechanism adaptation ability of the control.
[0035] The Fisher information matrix is a symmetric positive definite matrix used to describe the sensitivity of the policy output to parameter changes. In the present invention, this matrix is used to capture the change trend of the policy function in the non-Euclidean space and is an important basis for measuring the change intensity of the policy in different directions. Through this matrix, the most natural and effective update direction can be selected in the parameter space, making the control policy more in line with the actual structural characteristics of the metabolic pathway evolution, and improving the training convergence speed and control accuracy.
[0036] In this embodiment, the specific steps of S7 are as follows: S71. Extract the perturbation direction components on each principal component axis of the non-Euclidean embedding space according to the metabolic pathway perturbation direction vector output by the control policy, and map them to the fermentation process parameter adjustment direction and adjustment amplitude according to the positive and negative and magnitude of the change of the perturbation direction components. S72. Set the mapping rules between the disturbance direction components and each control parameter, where: the set temperature value is positively correlated with the component of the first principal component axis in the disturbance direction component, the set humidity value is positively correlated with the component of the second axis in the disturbance direction component, the ventilation rate is positively correlated with the component of the third axis in the disturbance direction component, and the stirring frequency is positively correlated with the component of the fourth axis in the disturbance direction component; S73. According to the value range of the disturbance direction component of the metabolic pathway disturbance direction vector in the non-Euclidean embedding space, construct the hierarchical adjustment rules for the fermentation process parameters, where: when the absolute value of the disturbance direction component is less than 0.1, it is a weak adjustment, and the adjustment range is set to ±2% of the original set value; when the absolute value of the disturbance direction component is between 0.1 and 0.3, it is a medium adjustment, and the adjustment range is ±5%; when the absolute value of the disturbance direction component is greater than 0.3, it is a strong adjustment, and the adjustment range is ±10%; S74. When the disturbance direction component is positive, the corresponding fermentation process parameter increases by the set range; when the disturbance component is negative, the corresponding fermentation process parameter decreases by the set range, and finally generate the fermentation process control instruction, where the fermentation process control instruction includes the set temperature value, the set humidity value, the ventilation rate, and the stirring frequency; S75. Package the fermentation process control instruction into a fermentation control device-level parameter configuration instruction and send it to the control execution unit of the fermentation control device to complete the dynamic adjustment of the fermentation environment.
[0037] Mapping the metabolic pathway disturbance direction vector in the non-Euclidean embedding space to an actually executable fermentation control instruction realizes the conversion closed-loop from the structural strategy to the environmental intervention. By establishing the association rules between the disturbance components and temperature, humidity, ventilation, and stirring, and setting the multi-level adjustment intensity standard, it is ensured that the strategy output can be accurately converted into a control signal recognizable by the device, realizing the automatic and gradual adjustment of the fermentation environment, and improving the process stability and flavor consistency.
[0038] Example 1
[0039] To verify the feasibility of the present invention in implementation, the present invention is applied to the low-temperature fermentation link of a traditional condiment production enterprise with an annual production scale of more than 20,000 tons. This enterprise has been producing sauce-flavored and compound-flavored condiments for a long time, and there is a prominent quality control problem - due to the large differences in the microbial metabolic states of different batches, the fermentation flavors show obvious fluctuations. Even if the raw materials are the same and the control parameters are stable, the sensory properties of the flavors still deviate greatly from the expectations, posing a challenge to product consistency.
[0040] Especially under low-temperature fermentation conditions, the enterprise conducts fermentation operations within a temperature control range of 15°C to 30°C, aiming to extend the metabolic cycle and enhance the fineness of aroma. However, in a low-temperature environment, the metabolic rate of microorganisms generally decreases, and the enzymatic activity of some pathways is inhibited, which easily leads to insufficient production of key flavor substances. In addition, the traditional constant control strategy centered on temperature and humidity is sluggish in response in a low-temperature environment and is difficult to compensate for metabolic shifts in a timely manner, further exacerbating the instability of flavor.
[0041] To solve the above problems, the dynamic control method for low-temperature fermentation of seasonings based on reinforcement learning proposed in the present invention is applied to the low-temperature fermentation test of the soy sauce-five-flavor fusion type seasonings of the enterprise. First, 5 senior flavorists describe the seasonings in text, including strong soy sauce aroma, obvious aftertaste, slightly sour fruity aroma, and low pungency. The system extracts keywords through the flavor analysis module, constructs a metabolic intention vector, and generates 18 target metabolic pathways, mainly including the synthesis pathways of phenylacetate, butyl acetate, glutamic acid, and medium-chain fatty acids.
[0042] In this embodiment, the experimental group uses two parallel 1500L low-temperature fermentation tanks, which are started synchronously at a set initial temperature of 20°C. One group is the traditional static control control group, and the other group is connected to the intelligent control system of the present invention. The control group uses fixed parameters: the temperature is constantly 28°C, the humidity is 65%, the ventilation rate is 3 air changes per hour, and the stirring period is once every 6 hours. The system of the experimental group updates the state map once an hour according to the state information such as the temperature and humidity, pH value, oxygen / carbon dioxide concentration, and metabolite concentration collected in real time, and compares it with the target path map through non-Euclidean embedding representation, outputs the structural deviation value, and drives the optimization process of the Riemannian manifold strategy.
[0043] During the entire 72-hour low-temperature fermentation cycle, the system completed a total of 114 control strategy updates, automatically adjusted the temperature, ventilation, or stirring frequency about 1.6 times per hour on average, and achieved fully autonomous dynamic control. The results of flavor sample analysis show that the concentration of ethyl phenylacetate in the experimental group reached 24.6 mg / L (20.1 mg / L in the control group), butyl acetate was 16.3 mg / L (13.7 mg / L in the control group), and glutamic acid increased to 89.2 mg / L (72.5 mg / L in the control group). In the sensory evaluation, the soy sauce aroma intensity increased from 7.6 points to 9.1 points, the persistence of aftertaste increased by 1.5 points, and the standard deviation of the overall flavor consistency score decreased from 0.89 to 0.35.
[0044] Under low-temperature conditions, the system specifically addresses the problem of insufficient activation of the fatty acid pathway caused by temperature inhibition by automatically increasing the stirring frequency and finely adjusting the ventilation volume to enhance the diffusion of metabolic substrates and enzyme activity conditions, significantly increasing the content of medium-chain fatty acids. In the traditional method, the operator needs to adjust the control parameters about 3 times a day, while the intervention frequency of the intelligent system is only 0.5 times per shift, significantly reducing the reliance on manpower.
[0045] It can be seen that the present invention has extremely strong adaptability and dynamic adjustment ability in the low-temperature fermentation environment. It not only solves the problem of slow reaction of the metabolic pathway caused by low temperature, but also realizes fine intervention in the flavor evolution direction through the structural difference perception mechanism, truly controlling the flavor by structure, and providing an efficient and popularizable solution path for the low-temperature intelligent manufacturing of seasonings.
[0046] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A dynamic control method for low-temperature fermentation of condiments based on reinforcement learning, characterized in that It includes the following steps: S1. Receive a flavor target input, where the flavor target is condiment text description information, and convert the flavor target into a metabolic intention vector; S2. Construct a target metabolic pathway map according to the metabolic intention vector. The target metabolic pathway map consists of metabolic pathway nodes and connection relationships, and is used to represent the metabolic structure path corresponding to flavor formation; S3. Periodically collect current fermentation state data at set time intervals during the fermentation process; S4. Construct a current metabolic pathway map based on the current fermentation state data, and map the current metabolic pathway map to a non-Euclidean embedding space to obtain a current metabolic pathway embedding representation; S5. Map the target metabolic pathway map to the same non-Euclidean embedding space as the current metabolic pathway map to obtain a target metabolic pathway embedding representation, and compare the target metabolic pathway embedding representation with the current metabolic pathway embedding representation to calculate the structural deviation; S6. Based on the structural deviation, adopt a Riemannian manifold strategy optimization method for strategy training and update to generate a control strategy output located in the non-Euclidean embedding space. The control strategy output is represented by the metabolic pathway perturbation direction; S7. Convert the control strategy output into a fermentation process control instruction. The fermentation process control instruction includes a temperature set value, a humidity set value, a ventilation rate, and a stirring frequency; S8. Execute the control instruction to adjust the fermentation environment parameters and complete the regulation of the evolution direction of the metabolic pathway during the fermentation process.
2. The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to claim 1, characterized in that The specific content of S1 includes: S11. Receive the condiment text description information input by the user. The condiment text description information includes natural language expressions of flavor, aroma, taste, process type, and regional style; S12. Perform word segmentation processing on the condiment text description information to extract keywords with sensory feature meanings; S13. Match the keywords with a preset flavor-metabolite mapping dictionary. The flavor-metabolite mapping dictionary records the association relationship between flavor terms and specific metabolites, and obtain one or more metabolite names associated with the keywords; S14. By querying the locally established metabolic pathway information table, retrieve the metabolic pathway numbers of each metabolite name and generate a metabolic pathway index list. The metabolic pathway information table is a pre-organized structured data file containing metabolites and corresponding metabolic pathway numbers, starting substrates, intermediate metabolites, end products, and reaction sequences; S15. Based on the metabolic pathway index list, obtain the association frequency between the metabolic pathway and the flavor target, take the proportion of the association frequency as the importance weight value of each metabolic pathway, and construct a metabolic intention vector. The metabolic intention vector is a multi-dimensional numerical sequence used to represent the importance of different metabolic pathways, where each dimension of the numerical value is obtained by normalizing the association frequency of the corresponding metabolic pathway.
3. The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to claim 1, wherein The specific content of S2 includes: S21. Sort the metabolic pathways according to the importance weight values in the metabolic intention vector, and select the metabolic pathways with importance weight values higher than the set weight threshold as the target metabolic pathway set; S22. Search for the structural information corresponding to the set of target metabolic pathways, where the structural information includes the starting substrates, intermediate metabolites, end products, and reaction sequence involved in the target metabolic pathways; S23. Extract the metabolite entries in the structural information, establish a set of target metabolic pathway nodes, where each node corresponds to a metabolite and is attached with a position identifier in the metabolic pathway; S24. According to the sequential reaction relationship of each metabolite in the metabolic pathway, construct a directed connection relationship between the target metabolic pathway nodes, form a structured metabolic pathway dependency chain, and constitute a target metabolic pathway map with the target metabolic pathway nodes and connection relationships.
4. The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to claim 1, characterized in that The current fermentation state data specifically includes temperature, humidity, pH value, oxygen concentration, carbon dioxide concentration, and metabolite concentration.
5. The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to claim 1, wherein The S4 specifically includes: S41. Perform a difference operation on the metabolite concentrations at adjacent sampling time points, calculate the change rate of each metabolite concentration, and mark the metabolites with a change rate higher than the preset rate threshold as active metabolites; S42. Based on the corresponding relationship of the active metabolites in the locally established metabolic pathway information table, construct a set of current metabolic pathway nodes; S43. According to the reaction sequence and conversion direction of the active metabolites in the metabolic pathway, establish a directed connection relationship between the current metabolic pathway nodes, and construct a current metabolic pathway map; S44. Based on the current metabolic pathway map, calculate the path distance between the current metabolic pathway nodes, generate a distance matrix, where the elements of the distance matrix are the shortest reaction step lengths between two nodes in the metabolic pathway; S45. Use Laplacian eigenmaps to process the distance matrix, and by preserving the relative position relationship and topological structure features between the current metabolic pathway nodes, map the current metabolic pathway map to a non-Euclidean embedding space, where the non-Euclidean embedding space is a structured representation space that does not satisfy Euclidean symmetry and the Pythagorean theorem.
6. The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to claim 1, wherein, The calculation process of the structural deviation specifically includes: Pair the nodes with the same metabolite identifier in the target metabolic pathway embedding representation and the current metabolic pathway embedding representation one by one to form a set of node pairs; For each pair of paired nodes, respectively extract the embedding vector coordinates in the non-Euclidean embedding space, calculate the embedding distance offset value of the paired nodes, where the embedding distance offset value uses the Euclidean distance calculation method and is defined as the square root of the sum of the squares of the coordinate differences of each dimension between the two node embedding vectors; Calculate the node weighting factors of the two nodes in each pair of paired nodes respectively, where the node weighting factors are obtained through a linear weighting method based on the importance weight value of the node in the metabolic pathway and the local topological connectivity in the corresponding metabolic pathway map; Arithmetically average the node weighting factors of the two nodes in the paired node pair as the final weighting factor of the paired node, and multiply the distance offset value of the paired node by the corresponding final weighting factor to obtain the weighted offset value. Sum the weighted offset values of each pair of paired nodes and divide by the total sum of the weighted factors of all node pairs to obtain a structural deviation value, which is used to quantify the overall structural difference between the current metabolic pathway map and the target metabolic pathway map in the non-Euclidean embedding space.
7. The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to claim 1, characterized in that The specific steps of S6 include: S61. Use the structural deviation value as the basis for state input, and combine the current metabolic pathway embedding representation to construct a state vector in the non-Euclidean embedding space ; S62. Based on the state vector , construct a policy function , where represents the conditional probability distribution of outputting an action under the given state vector , represents the policy function parameter represents the metabolic pathway perturbation direction vector output at time step ; S63. Define an immediate reward function to represent the structural optimization effect: ; Among them, represents the reward value at the time step, represents the structural deviation value at the time step, represents the structural deviation value at the time step; S64. Construct a Fisher information matrix based on the current policy function : ; Among them, represents the expectation operator, represents the gradient of the policy logarithmic function with respect to the parameter, represents the transpose operation; S65. Update the policy parameters using the natural gradient optimization method on the Riemannian manifold: ; Among them, represents the policy parameter at time step ; represents the policy parameter at time step ; represents the learning rate; represents the inverse of the Fisher information matrix at time step ; represents the discount factor; represents the reward value at time step ; represents the gradient of the cumulative expected reward of the policy with respect to the policy parameter; S66. Apply the updated policy parameters to the current state vector to generate a metabolic pathway perturbation direction vector, which is defined in the non-Euclidean embedding space and is used to represent the adjustment direction of the metabolic pathway map in the structural space, serving as the output of the control strategy.
8. The dynamic control method for low-temperature fermentation of condiments based on reinforcement learning according to claim 1, characterized in that The specific steps of S7 include: S71. According to the metabolic pathway perturbation direction vector output by the control strategy, extract the perturbation direction components on each principal component axis in the non-Euclidean embedding space, and map them to the adjustment direction and adjustment amplitude of the fermentation process parameters according to the positive or negative and magnitude of the change of the perturbation direction components. S72. Set the mapping rules between the perturbation direction components and each control parameter. Specifically, the temperature set value is positively correlated with the component of the first principal component axis in the perturbation direction components, the humidity set value is positively correlated with the component of the second axis in the perturbation direction components, the ventilation rate is positively correlated with the component of the third principal component axis in the perturbation direction components, and the stirring frequency is positively correlated with the component of the fourth principal component axis in the perturbation direction components. S73. According to the range of the perturbation direction component values of the metabolic pathway perturbation direction vector in the non-Euclidean embedding space, construct a hierarchical adjustment rule for the fermentation process parameters. Specifically, when the absolute value of the perturbation direction component is less than 0.1, it is a weak adjustment, and the adjustment amplitude is set to ±2% of the original set value; when the absolute value of the perturbation direction component is between 0.1 and 0.3, it is a medium adjustment, and the adjustment amplitude is ±5%; when the absolute value of the perturbation direction component is greater than 0.3, it is a strong adjustment, and the adjustment amplitude is ±10%. S74. When the perturbation direction component is positive, increase the set amplitude of the corresponding fermentation process parameter; when the perturbation component is negative, decrease the set amplitude of the corresponding fermentation process parameter, and finally generate a fermentation process control instruction, which includes the temperature set value, humidity set value, ventilation rate, and stirring frequency. S75. Package the fermentation process control instruction into a fermentation control device-level parameter configuration instruction and send it to the control execution unit of the fermentation control device to complete the dynamic adjustment of the fermentation environment.
Citation Information
Patent Citations
Biological process control method based on knowledge map directed graph
CN109243528A
Intelligent routing system and method based on knowledge definition network and graph reinforcement learning
CN118282918A
Method for dynamically adjusting tea fermentation environment by utilizing fuzzy logic control
CN119148799A
Flavor analysis method, device, equipment, storage medium and program product
CN119940074A
Optimization solving method and system for metabolic flux in biomanufacturing process
WO2025015986A1