Intelligent decision-making method and system for ultra-high pressure processing parameter of multi-objective optimization
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN SHANG PIN FOOD CO LTD
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-07
AI Technical Summary
这种由于缺乏对负载热力学特性与包装力学响应的实时协同控制,导致了在保障食品安全、维持高品质外观与降低能耗三者之间形成了难以调和的问题,影响了HPP技术在高端敏感食品领域的应用效能
[0034] This invention constructs a multi-dimensional state space containing a dynamic spectrum of fluid compressibility modulus and an instantaneous energy efficiency ratio fingerprint, transforming the ultra-high pressure processing process from traditional "black box" control to transparent control capable of sensing load rheological characteristics and equipment energy efficiency status. This effectively solves the technical challenge of being unable to identify differences in food matrix (such as different adiabatic temperature rises caused by high moisture and high fat) in a single pressure dimension. Based on the path mapping and dynamic correction mechanism of the Pareto optimal frontier library, it overcomes the problem of traditional static conservative parameter settings struggling to balance safety, quality maintenance, and energy consumption control. It can automatically find the optimal process balance point to reduce energy consumption and minimize the loss of heat-sensitive nutrients while meeting the safety threshold for microbial inactivation. In particular, for modified atmosphere packaging products, by generating an adaptive variable rate depressurization command sequence, it achieves refined variable rate control of the depressurization process, matching the depressurization rate with the expansion-equilibrium dynamics of the gas inside the packaging. This effectively avoids bulging, cracking, or delamination caused by excessive pressure difference between the inside and outside of the packaging while ensuring production efficiency, thus improving the packaging integrity and appearance quality of the product.
Smart Images

Figure CN122526129A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of process parameter control technology, and more specifically, to a method and system for intelligent decision-making of ultra-high pressure machining process parameters based on multi-objective optimization. Background Technology
[0002] Ultra-high pressure processing (HPP), a non-thermal sterilization technology, utilizes hydrostatic pressure of 100MPa to 600MPa at room temperature or low temperature to kill pathogenic and spoilage bacteria in food. It has been widely applied in fruit and vegetable juices, meat products, and seafood shelling. Compared to traditional heat processing, HPP can better preserve the sensory quality and nutritional components of food. Chinese patent application CN120509519A discloses a decision-making system based on machine learning to predict the flavor formation mechanism and optimize flavor during food processing. It receives flavor perception probability distribution data, combines it with equipment physical constraints, and uses a multi-objective particle swarm optimization algorithm to solve for the Pareto optimal solution set to generate processing parameter schemes. The aim is to achieve closed-loop precise control of processing parameters through dynamic threshold modeling and spatiotemporal preference mapping correction, thereby improving the flavor quality of food. Chinese patent application CN121235039A discloses a method and system for joint optimization of multi-objective hyperparameters for deep learning models. Although it mainly involves deep learning models, it solves the efficiency and local optimum problems in multi-objective tasks by constructing a Pareto optimal front and using a non-dominated sorting update strategy.
[0003] However, despite the optimizations made to the aforementioned technologies in specific dimensions, setting process parameters still presents challenges in industrial HPP applications. In actual production, to ensure compliance with specified microbial reduction levels, control strategies typically employ static and conservative "worst-case" settings (such as uniformly setting a pressure of 600 MPa for 6 minutes). This crude control approach ignores the complex dynamic response characteristics of food matrices and packaging materials. Specifically, different food matrices (such as high-moisture chicken breast and high-fat bacon) exhibit drastically different adiabatic compression heat-generating characteristics under pressure. Static pressure settings cannot detect these differences in "physical fingerprints," leading to accelerated lipid oxidation and nutrient degradation in high-heat-generating materials due to excessively high actual temperatures, or prolonged ineffective pressure holding time for low-heat-generating materials due to insufficient temperatures, resulting in energy waste. More critically, for products containing modified atmosphere packaging (MAP), existing technologies generally employ linear depressurization logic, completely ignoring the physical fact that the expansion dynamics of high-density residual gases inside the packaging lag behind the rate of external pressure reduction. When external pressure drops abruptly, the gas inside the packaging expands explosively, creating a pressure difference exceeding the yield strength of the packaging material. This directly leads to packaging rupture, bulging, or delamination of the seal. The lack of real-time coordinated control over the thermodynamic characteristics of the load and the mechanical response of the packaging creates an irreconcilable problem between ensuring food safety, maintaining a high-quality appearance, and reducing energy consumption, thus affecting the effectiveness of HPP technology in the high-end sensitive food sector. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of existing technologies, this invention provides a multi-objective optimized intelligent decision-making method and system for ultra-high pressure processing parameters, capable of real-time sensing of the rheological characteristics and energy efficiency status of the load within the chamber. By constructing a Pareto optimal frontier library and an adaptive execution control mechanism, this invention can automatically find the optimal balance between quality preservation and energy consumption control while ensuring microbial inactivation safety standards. Furthermore, it utilizes an adaptive variable rate depressurization strategy to effectively prevent the modified atmosphere packaging from rupture and deformation during depressurization, achieving efficient, low-loss, and energy-saving intelligent production.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A multi-objective optimization-based intelligent decision-making method for ultra-high pressure machining process parameters includes:
[0007] Construct a product feature dataset, a dynamic spectrum of fluid compressibility modulus, and an instantaneous energy efficiency ratio fingerprint; generate an initial state vector and a real-time state vector.
[0008] Construct a Pareto optimal front library, map the initial state vector to the Pareto optimal front library to generate the initial target process path, determine whether the dynamic adjustment condition is triggered based on the real-time state vector, and if it is triggered, correct the initial target process path and output the corrected target process path.
[0009] The final target process path is determined based on the initial target process path and the revised target process path. During the pressurization and pressure holding stages, the booster pump is controlled to execute the final target process path. The final target process path includes at least a target pressure relief mode. If the target pressure relief mode is an adaptive pressure relief mode, an adaptive variable rate pressure relief command sequence is generated after the pressure holding is terminated, and the pressure relief valve is controlled to perform the pressure relief operation according to the adaptive variable rate pressure relief command sequence.
[0010] The method for constructing the product feature dataset includes:
[0011] Read the product metadata of the batch to be processed and collect the initial environmental signal. Integrate the product metadata and the initial environmental signal to form a product feature dataset. The product metadata includes product ID, packaging type and packaging material yield strength threshold.
[0012] The method for constructing the dynamic spectrum of fluid compressibility modulus includes:
[0013] During the pressurization and pressure holding phases, basic electromechanical signals of the booster pump and the ultra-high pressure chamber are collected. These basic electromechanical signals include the real-time current value of the booster pump servo motor, the plunger displacement speed of the booster pump, the real-time pressure in the ultra-high pressure chamber, the real-time fluid temperature, and the real-time fluid flow rate.
[0014] Based on the plunger displacement velocity and the real-time pressure inside the ultra-high pressure chamber, the pressure rise rate and volume compressibility are calculated, and a dynamic spectrum of fluid compressibility modulus is generated based on the pressure rise rate and volume compressibility.
[0015] The method for constructing the instantaneous energy efficiency ratio fingerprint includes:
[0016] The motor input power is calculated based on the real-time current value; the hydraulic power is calculated based on the real-time pressure and the real-time fluid flow rate; the ratio of the hydraulic power to the motor input power is used as the instantaneous energy efficiency ratio; and the instantaneous energy efficiency ratio fingerprint is obtained based on the instantaneous energy efficiency ratio.
[0017] The method for generating the initial state vector includes:
[0018] Before the boost begins, the historical rheological feature template corresponding to the product ID in the product feature dataset is retrieved from the historical database, and the product feature dataset and the historical rheological feature template are fused to generate the initial state vector.
[0019] The method for generating the real-time state vector includes:
[0020] During the pressurization and pressure holding processes, the real-time temperature rise slope is calculated based on the real-time fluid temperature and real-time pressure. The fluid compressibility modulus dynamic spectrum, instantaneous energy efficiency ratio fingerprint, real-time temperature rise slope, and initial state vector are fused in real time to continuously update and generate the real-time state vector.
[0021] The method for constructing the Pareto optimal front surface library includes:
[0022] Establish an energy consumption cost function, and combine it with a preset safety function and quality loss function to form a set of objective functions for a multi-objective optimization problem;
[0023] Using the objective function set as the optimization objective, a non-dominated sorting genetic algorithm is executed for different product categories to generate Pareto optimal fronts for each product category, and these are then compiled into a Pareto optimal front library.
[0024] The method for determining whether the dynamic adjustment condition has been triggered is as follows:
[0025] The real-time temperature rise slope is extracted from the real-time state vector, the historical reference temperature rise slope is retrieved from the historical database, the deviation rate between the real-time temperature rise slope and the historical reference temperature rise slope is calculated, and when the deviation rate exceeds the preset deviation threshold, the dynamic adjustment condition is triggered.
[0026] The pressure holding process terminates upon receiving a pressure holding termination command.
[0027] The conditions for generating the pressure holding termination command are:
[0028] During the pressure holding phase, the cumulative microbial lethality rate and marginal sterilization benefit are continuously calculated. When the pressure holding termination condition is met, a pressure holding termination command is output. The pressure holding termination condition is that the cumulative microbial lethality rate meets the safety redundancy condition and the marginal sterilization benefit meets the benefit threshold condition.
[0029] A multi-objective optimization intelligent decision-making system for ultra-high pressure machining process parameters, used to implement the aforementioned multi-objective optimization intelligent decision-making method for ultra-high pressure machining process parameters, the system comprising:
[0030] State vector generation module: used to construct product feature dataset, fluid compressibility modulus dynamic spectrum and instantaneous energy efficiency ratio fingerprint, and generate initial state vector and real-time state vector;
[0031] Process path decision module: It is used to construct the Pareto optimal front library, map the initial state vector to the Pareto optimal front library to generate the initial target process path, and determine whether the dynamic adjustment condition is triggered based on the real-time state vector. If it is triggered, the initial target process path is corrected and the corrected target process path is output.
[0032] Adaptive execution control module: Determines the final target process path based on the initial target process path and the corrected target process path. During the pressurization and pressure holding stages, it controls the booster pump to execute the final target process path. The final target process path includes at least a target depressurization mode. If the target depressurization mode is an adaptive depressurization mode, an adaptive variable rate depressurization command sequence is generated after the pressure holding ends, and the depressurization valve is controlled to perform depressurization operation according to the adaptive variable rate depressurization command sequence.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] This invention constructs a multi-dimensional state space containing a dynamic spectrum of fluid compressibility modulus and an instantaneous energy efficiency ratio fingerprint, transforming the ultra-high pressure processing process from traditional "black box" control to transparent control capable of sensing load rheological characteristics and equipment energy efficiency status. This effectively solves the technical challenge of being unable to identify differences in food matrix (such as different adiabatic temperature rises caused by high moisture and high fat) in a single pressure dimension. Based on the path mapping and dynamic correction mechanism of the Pareto optimal frontier library, it overcomes the problem of traditional static conservative parameter settings struggling to balance safety, quality maintenance, and energy consumption control. It can automatically find the optimal process balance point to reduce energy consumption and minimize the loss of heat-sensitive nutrients while meeting the safety threshold for microbial inactivation. In particular, for modified atmosphere packaging products, by generating an adaptive variable rate depressurization command sequence, it achieves refined variable rate control of the depressurization process, matching the depressurization rate with the expansion-equilibrium dynamics of the gas inside the packaging. This effectively avoids bulging, cracking, or delamination caused by excessive pressure difference between the inside and outside of the packaging while ensuring production efficiency, thus improving the packaging integrity and appearance quality of the product. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart of the intelligent decision-making method for multi-objective optimization of ultra-high pressure processing parameters provided in this embodiment of the invention;
[0037] Figure 2 A schematic diagram of the Pareto optimal frontier provided in an embodiment of the present invention;
[0038] Figure 3 This is a flowchart of the initial target process path correction and judgment provided in an embodiment of the present invention;
[0039] Figure 4This is a schematic diagram of gas expansion during the depressurization process of modified atmosphere packaging provided in an embodiment of the present invention;
[0040] Figure 5 This is a schematic diagram of the adaptive variable rate depressurization curve provided in an embodiment of the present invention;
[0041] Figure 6 A functional block diagram of the intelligent decision-making system for multi-objective optimized ultra-high pressure processing parameters provided in this embodiment of the invention. Detailed Implementation
[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] Example 1
[0044] Please see Figure 1 As shown, this embodiment provides a multi-objective optimized intelligent decision-making method for ultra-high pressure machining process parameters, including:
[0045] Step S10: Read the product metadata of the batch to be processed and collect the initial environmental signals, and integrate them to form a product feature dataset; during the pressurization stage and the pressure holding stage, collect the basic electromechanical signals of the booster pump and the ultra-high pressure chamber simultaneously, construct the dynamic spectrum of fluid compression modulus and instantaneous energy efficiency ratio fingerprint, and generate the initial state vector and real-time state vector.
[0046] Step S10 constructs a multi-dimensional state space based on second-order rheological characteristics and energy efficiency fingerprints. Its core lies in utilizing the inherent mechanical characteristics of the hydraulic system of the ultra-high pressure processing equipment as a soft sensor to derive the rheological characteristics of the load within the ultra-high pressure chamber. A soft sensor refers to a virtual sensing method that does not rely on additional dedicated physical sensors, but rather indirectly obtains physical quantities or state information that are difficult to measure directly by mathematically processing and modeling existing electrical and mechanical signals during equipment operation. Second-order rheological characteristics refer to the second derivative characteristics of the volume compressibility of food materials under pressure, reflecting the phase change behavior occurring within the food matrix. For example, when the pressure reaches a specific range, the fat inside high-fat bacon undergoes a crystalline phase change, leading to a rapid volume contraction. At this point, the compressibility curve shows a clear inflection point, which is a typical representation of second-order rheological characteristics. Traditional ultra-high pressure processing control systems rely solely on pressure sensors installed within the ultra-high pressure chamber to obtain pressure readings and on timers to record holding time. These two dimensions of information are insufficient to reflect the true physical state of the food materials within the chamber, and the system cannot distinguish between different food matrices, forming a so-called black-box control mode. Step S10 constructs a multi-dimensional state space, integrating product metadata, initial environmental signals, fluid compressibility modulus dynamic spectrum, instantaneous energy efficiency ratio fingerprint, and real-time temperature rise slope into unified decision input variables. This enables the deep reinforcement learning agent in subsequent step S20 to obtain sufficient environmental state information, thereby enabling precise process path selection and dynamic adjustment on the Pareto optimal frontier.
[0047] Further, step S10 includes:
[0048] Step S11: Read the product metadata of the batch to be processed and collect the initial environmental signal, and integrate the product metadata with the initial environmental signal to form a product feature dataset; the product metadata includes product ID, packaging type and packaging material yield strength threshold.
[0049] Product metadata refers to a set of structured information related to the batch of food to be processed, pre-stored in a production management system or a pre-stored database. Product ID is used to uniquely identify the product category. For example, the product ID can be set as a label such as "high-fat thick-cut bacon" or "low-fat chicken breast." This label directly corresponds to the product category partition in the Pareto optimal frontier library in subsequent step S22, enabling step S23 to quickly match the corresponding Pareto optimal frontier based on the product ID. Packaging type refers to the gas isolation and preservation method used in food packaging. Common types include vacuum packaging, skin packaging, and modified atmosphere packaging. Vacuum packaging refers to a packaging method where the food is placed in a bag, the air inside is removed, and the bag is sealed, creating a negative pressure state inside the packaging. Skin packaging refers to a packaging form where a thermoplastic film is heated and softened, then tightly adhered to the surface of the food and sealed with a bottom tray. Modified atmosphere packaging refers to a packaging method where the air inside the packaging is replaced with a protective gas mixture in a specific proportion before sealing. For example, modified atmosphere packaging can be filled with a mixture of 20% nitrogen and 80% carbon dioxide by volume. The packaging type information directly determines the selection logic of the depressurization mode in step S31. When the packaging type is modified atmosphere packaging, the system needs to call the adaptive depressurization mode to prevent the rapid expansion of residual gas inside the packaging during the depressurization process from causing the packaging to rupture. The yield strength threshold of the packaging material refers to the upper limit of the pressure difference that the packaging material can withstand without permanent deformation or damage when subjected to internal and external pressure differences. This threshold is provided by the packaging material supplier or obtained through standard mechanical testing. For example, the yield strength threshold range of commonly used ethylene-vinyl alcohol copolymer composite films is usually between tens of MPa and hundreds of MPa.
[0050] The initial environmental signal refers to sensor readings collected from the production site before the start of the processing batch, reflecting the physical state of the raw materials and environmental conditions. This includes the raw material entry temperature, ambient temperature, and loading density. The raw material entry temperature refers to the measured temperature of the food material before it is loaded into the ultra-high pressure chamber. This temperature is affected by the food's previous storage conditions; for example, food from a cold storage facility typically has a lower entry temperature than similar food from a room-temperature transit area. The ambient temperature refers to the real-time air temperature within the ultra-high pressure processing workshop, which affects the initial thermal state of the ultra-high pressure chamber and the pressure-transmitting medium. The loading density refers to the volume ratio of the food package to the pressure-transmitting medium within the ultra-high pressure chamber, which affects the overall heat capacity and pressure transmission efficiency during the pressurization process. Integrating product metadata with the initial environmental signal forms a product feature dataset, enabling the system to obtain complete prior information about the batch to be processed before the processing cycle begins. This product feature dataset serves as one of the basic data sources for generating the initial state vector in step S15, ensuring that the initial state vector comprehensively characterizes the material properties and environmental conditions of the batch to be processed.
[0051] Step S12: During the pressurization and pressure holding stages, the basic electromechanical signals of the booster pump and the ultra-high pressure chamber are collected simultaneously. The basic electromechanical signals include the real-time current value of the booster pump servo motor, the plunger displacement speed of the booster pump, the real-time pressure in the ultra-high pressure chamber, the real-time fluid temperature, and the real-time fluid flow rate.
[0052] The pressurization phase refers to the process by which the ultra-high pressure processing equipment starts the booster pump to inject the pressure-transmitting medium into the ultra-high pressure chamber, gradually increasing the pressure from atmospheric pressure to the target pressure. The pressure holding phase refers to the process of maintaining the pressure level within the allowable fluctuation range near the target pressure after the pressure inside the ultra-high pressure chamber reaches the target pressure, by controlling the intermittent operation of the booster pump or adjusting the pressure compensation valve. Basic electromechanical signals refer to the electrical and mechanical motion parameters generated by the ultra-high pressure processing equipment during operation that can be directly acquired by sensors. The real-time current value of the booster pump servo motor refers to the current amplitude consumed by the servo motor driving the booster pump plunger at any given moment; this current value is acquired in real-time by the Hall effect current sensor built into the servo driver. The plunger displacement speed of the booster pump refers to the instantaneous speed of the booster pump plunger moving axially within the cylinder; this speed is obtained by a linear encoder installed on the booster pump or by converting the rotation angle of the servo motor. The real-time pressure inside the ultra-high pressure chamber refers to the instantaneous value of the hydrostatic pressure borne by the pressure-transmitting medium inside the ultra-high pressure chamber; this pressure is acquired by a high-precision pressure sensor installed on the wall of the ultra-high pressure chamber. Real-time fluid temperature refers to the instantaneous temperature of the pressure-transmitting medium within the ultra-high pressure chamber. This temperature is acquired by temperature sensors installed inside the ultra-high pressure chamber or along the flow path of the pressure-transmitting medium. Real-time fluid flow rate refers to the volume of pressure-transmitting medium injected into the ultra-high pressure chamber per unit time via a booster pump. This flow rate can be directly measured by a flow meter installed on the high-pressure pipeline, or indirectly calculated by multiplying the plunger cross-sectional area by the plunger displacement velocity. Millisecond-level sampling frequencies are used to synchronously acquire the aforementioned basic electromechanical signals. Synchronous acquisition means simultaneously recording the values of multiple signals at the same timestamp, ensuring time-domain alignment between different signals. Millisecond-level sampling frequencies can capture transient changes during pressurization and pressure holding processes. For example, when the fat component in the food inside the chamber undergoes a crystalline phase transition, the change in compressibility occurs within a timescale of tens to hundreds of milliseconds. If the sampling frequency is too low, the inflection point of this phase transition cannot be identified.
[0053] Step S13: Calculate the pressure rise rate and volume compressibility based on the plunger displacement velocity and the real-time pressure inside the ultra-high pressure chamber, and generate a dynamic spectrum of fluid compressibility modulus based on the pressure rise rate and volume compressibility.
[0054] The dynamic spectrum of fluid compressibility modulus refers to a two-dimensional curve constructed with real-time pressure inside the ultra-high pressure chamber as the abscissa and equivalent compressibility modulus as the ordinate. This curve reflects the volume change characteristics of the load inside the chamber in different pressure ranges. The physical meaning of equivalent compressibility modulus is the pressure increment required to compress a fluid per unit relative volume; the larger the value, the more difficult the fluid is to compress. The pressure rise rate is calculated by performing time-domain differentiation on the real-time pressure sequence composed of real-time pressures. Specifically, it can be obtained by dividing the pressure difference between adjacent sampling points by the sampling time interval; that is, the pressure rise rate equals the pressure difference between two adjacent sampling times divided by the time interval between the two sampling times. The volume compressibility is calculated based on the geometric characteristics of the booster pump. When the plunger moves forward in the cylinder, it forces the pressure-transmitting medium into the ultra-high pressure chamber. The product of the plunger cross-sectional area and the plunger displacement velocity is the volume increment of the medium injected into the ultra-high pressure chamber per unit time. Since the ultra-high pressure chamber is a rigid sealed container, the injected medium volume cannot increase the internal space of the chamber. Instead, it further compresses the original fluid and food material mixture system inside the chamber. Therefore, the volume increment of the injected medium is numerically equal to the volume reduction of the fluid inside the chamber due to compression. The volume compressibility is defined as the volume reduction of the fluid inside the chamber due to compression per unit time, and its value is equal to the plunger cross-sectional area multiplied by the plunger displacement velocity. The formula for calculating the equivalent compressive modulus K(P) is K(P) = -V × (dP / dV). This formula originates from the standard definition of bulk modulus in fluid mechanics and thermodynamics. Bulk modulus is a physical quantity that measures a substance's ability to resist uniform compression. Its physical meaning is the amount of pressure change required to cause a unit relative volume change in a substance. In ultra-high pressure processing scenarios, this definition can accurately characterize the compressibility of the fluid-food mixture system in the chamber at different pressure levels. Here, V is the volume of the fluid in the ultra-high pressure chamber, P is the real-time pressure, dP is the pressure change, and dV is the volume change. The negative sign is introduced because the volume decreases when the pressure increases, and the two changes in opposite directions. The negative sign keeps the equivalent compressive modulus positive. In practical calculations, according to the chain rule of calculus, dP / dV can be obtained by dividing dP / dt by dV / dt, where dP / dt is the rate of pressure rise. Since the fluid in the chamber is compressed during the pressurization process, resulting in a volume reduction, the rate of change of volume over time, dV / dt, is negative. The volumetric compressibility is defined as the rate of change of the absolute value of the volume reduction over time, i.e., the volumetric compressibility is equal to negative dV / dt. Therefore, dP / dV is equal to the rate of pressure rise divided by the negative volumetric compressibility. Substituting this result into the formula K(P)=-V×(dP / dV), the two negative signs cancel each other out. The final equivalent compression modulus can be calculated using the simplified form K(P)=V×(rate of pressure rise ÷ volumetric compressibility). This calculation method transforms two directly measurable or indirectly acquired signals into physical parameters reflecting load characteristics, achieving an effective conversion from measurable signals to target parameters, enabling the system to sense the compression behavior of the material in the chamber.The initial value of the fluid volume V inside the ultra-high pressure chamber is obtained by subtracting the volume of the food package from the rated volume of the ultra-high pressure chamber. The volume of the food package is calculated based on the number of packages loaded and the nominal volume of each package, or by weighing before loading. During pressurization, the real-time value of V is updated by integrating the volumetric compressibility over time. Specifically, V(t) is equal to the initial value of V minus the definite integral of the volumetric compressibility from time zero to the current time t. This integral value represents the cumulative volume of medium injected into the chamber from the start of pressurization to the current time, and is numerically equal to the cumulative volume reduction of the fluid inside the chamber due to compression.
[0055] The method for generating the dynamic spectrum of fluid compressibility modulus is as follows: As the pressurization process progresses, the system calculates the equivalent compressibility modulus at the current moment in each sampling period, and records this compressibility modulus value and the corresponding real-time pressure value as a data point. All data points are arranged in chronological order to form a dynamic curve with pressure as the horizontal axis and compressibility modulus as the vertical axis. This curve can characterize the hardness change characteristics of the load inside the chamber. When only pure water is used as the pressure transmission medium inside the chamber, the compressibility modulus curve shows an approximately smooth monotonically increasing trend. When the chamber is loaded with food materials, different components in the food exhibit differentiated compression responses during the pressurization process. For example, the compressibility modulus curve of food with high moisture content is close to that of pure water, while the compressibility modulus curve of food with high fat content will show an inflection point with a sudden change in slope due to the solidification phase transition of fat in a specific pressure range. This inflection point is the specific manifestation of the second-order rheological characteristics. By identifying the inflection point positions and amplitudes in the dynamic spectrum of fluid compressibility modulus, the compositional differences of the materials loaded in the chamber can be distinguished. For example, when a sharp increase in the slope of the compressibility modulus curve is detected in a certain pressure range, it can be determined that the materials in the chamber have a high fat content or contain other components that are prone to phase change. This determination result is passed to step S24 as a component of the real-time state vector, enabling the deep reinforcement learning agent to adjust the target pressure or plan the holding time accordingly. If step S13 is missing, the system will not be able to obtain information reflecting the compressibility characteristics of the materials in the chamber, the real-time state vector generated in step S15 will lack the second-order rheological feature dimension, and the deep reinforcement learning agent in step S24 will not be able to distinguish the differences in physical properties of different food matrices when making dynamic path adjustments, and can only use the same control strategy for all products, resulting in the loss of the control system's adaptability to changes in materials.
[0056] Step S14: Generate an instantaneous energy efficiency ratio fingerprint based on the real-time current value of the booster pump servo motor, the real-time pressure in the ultra-high pressure chamber, and the real-time fluid flow rate.
[0057] The method for generating an instantaneous energy efficiency ratio fingerprint includes: calculating the motor input power based on the real-time current value; calculating the hydraulic power output based on the real-time pressure and the real-time fluid flow rate; using the ratio of the hydraulic power output to the motor input power as the instantaneous energy efficiency ratio; and obtaining the instantaneous energy efficiency ratio fingerprint based on the instantaneous energy efficiency ratio.
[0058] Instantaneous energy efficiency ratio fingerprint refers to a time-domain curve constructed with time as the horizontal axis and instantaneous energy efficiency ratio as the vertical axis. This curve reflects the efficiency level of ultra-high pressure processing equipment in converting electrical energy into effective hydraulic work at different times. The instantaneous energy efficiency ratio is a dimensionless ratio, physically representing the ratio between the effective power output of the hydraulic system and the electrical power input to the motor. A higher ratio indicates higher energy conversion efficiency, and vice versa. The motor input power equals the motor's rated operating voltage multiplied by the real-time current value. The motor's rated operating voltage is a design parameter of the equipment, read from the equipment parameter database during system initialization. Hydraulic power output equals the real-time pressure within the ultra-high pressure chamber multiplied by the real-time fluid flow rate. This formula originates from the fundamental definition of hydraulic power in fluid mechanics. Its derivation is as follows: work equals force multiplied by displacement, and power equals force multiplied by velocity. In a hydraulic system, force equals pressure multiplied by the area of action, and velocity equals the fluid velocity. Therefore, power equals pressure multiplied by the area multiplied by the velocity. The product of area and velocity is the volumetric flow rate. Thus, hydraulic power output equals the product of pressure and volumetric flow rate. This formula is the standard method for calculating the output power of hydraulic pumps in the field of hydraulic engineering. The instantaneous energy efficiency ratio (EER) is calculated by dividing the hydraulic power output by the motor input power, quantifying the energy conversion efficiency of the equipment at the current moment. This calculation method conforms to the general definition of energy conversion efficiency, which is the ratio of effective output power to total input power. In ultra-high pressure processing equipment, the motor input power represents the rate of consumption of total electrical energy obtained from the grid, while the hydraulic power output represents the rate of energy output converted into effective hydraulic energy and used to compress the fluid within the chamber. The ratio of these two values comprehensively reflects the energy loss in various aspects, including motor efficiency, mechanical transmission efficiency, and hydraulic system efficiency. It is a direct indicator for evaluating the energy utilization level of the equipment. The instantaneous energy efficiency ratio fingerprint is generated as follows: As the pressurization and holding processes progress, the system calculates the instantaneous energy efficiency ratio at the current moment in each sampling period, and records this energy efficiency ratio value and the corresponding timestamp as a data point. All data points are arranged in chronological order to form a time-domain curve with time as the horizontal axis and energy efficiency ratio as the vertical axis. This curve can reveal the energy consumption characteristics of the equipment in different pressure ranges. For example, in the initial stage of pressurization when the pressure is low, the fluid is more easily compressed, and the motor only needs to output a small torque to push the plunger forward. At this time, the current value is low and the pressure rises rapidly, and the instantaneous energy efficiency ratio is at a high level. As the pressure gradually increases, the fluid becomes more difficult to compress, and the motor needs to output a larger torque to overcome the fluid resistance. The current value increases accordingly, but the rate of pressure rise gradually slows down, and the instantaneous energy efficiency ratio shows a downward trend. At the end of the pressurization process, when the pressure is close to the equipment limit, the fluid compressibility approaches zero, the motor current increases sharply, but the pressure hardly rises anymore, and the instantaneous energy efficiency ratio drops sharply to an extremely low value.The instantaneous energy efficiency ratio fingerprint provides a quantitative basis for establishing the energy consumption cost function in step S21. The total energy consumption value of the entire processing cycle can be obtained by integrating the instantaneous energy efficiency ratio fingerprint in the time domain. This total energy consumption value serves as the output variable of the energy consumption cost function in the multi-objective optimization problem. The instantaneous energy efficiency ratio fingerprint also provides real-time feedback on energy efficiency for dynamic path adjustment in step S24. For example, when the system observes that the instantaneous energy efficiency ratio has dropped to an extremely low level, it indicates that continuing to increase the pressure or extend the holding pressure will consume a large amount of electrical energy, but the resulting sterilization effect gain is negligible. The deep reinforcement learning agent can trigger the decision to terminate the holding pressure in advance based on this. If step S14 is missing, the system will not be able to obtain information reflecting the energy conversion efficiency of the equipment. Step S21 will not be able to establish an energy consumption cost function based on actual operating data, and can only use theoretical energy consumption estimation based on equipment nameplate parameters, resulting in a disconnect between the energy consumption optimization target and the actual situation. At the same time, step S24 will not be able to obtain real-time feedback on energy consumption efficiency. When the deep reinforcement learning agent makes dynamic adjustments, it will lack decision-making basis in the energy consumption dimension and will not be able to achieve multi-objective collaborative optimization of safety, quality and energy consumption.
[0059] Step S15: Before pressurization begins, retrieve the historical rheological feature template corresponding to the product ID in the product feature dataset from the historical database, and fuse the product feature dataset with the historical rheological feature template to generate an initial state vector; during pressurization and pressure holding, calculate the real-time temperature rise slope based on the real-time fluid temperature and real-time pressure; fuse the fluid compressibility modulus dynamic spectrum, instantaneous energy efficiency ratio fingerprint, real-time temperature rise slope with the initial state vector in real time, and continuously update to generate a real-time state vector.
[0060] The initial state vector refers to the multi-dimensional numerical vector formed by fusing the product feature dataset with the historical rheological feature template before the pressurization begins. This vector serves as the environmental state input for the deep reinforcement learning agent to generate the initial target process path in step S23. The historical rheological feature template refers to the set of typical fluid compressibility modulus dynamic spectrum feature parameters obtained based on statistical analysis of historical processing batch data for a specific product category. For example, the historical rheological feature template for high-fat bacon products may include typical pressure ranges where the compressibility modulus curve of this type of product shows phase transition inflection points during historical processing, typical slope change amplitudes at the inflection points, and other feature parameters. The method for retrieving the historical rheological feature template from the historical database is as follows: using the product ID in the product feature dataset as the search keyword, matching the historical processing records corresponding to the product ID in the historical database, and extracting the statistical mean or typical curve template of the rheological feature parameters from the matched historical records. The product feature dataset and historical rheological feature template are fused using vector concatenation technology. Specifically, the product ID from the product metadata is encoded as a uniquely heated vector or embedded vector; the packaging type is encoded as a category label vector; numerical parameters such as the yield strength threshold of the packaging material, the raw material entry temperature, the ambient temperature, and the loading density are directly used as vector elements; and the feature parameter sequence from the historical rheological feature template is used as an additional dimension of the vector. These parts are concatenated in a preset order to form the initial state vector. The dimensional design of the initial state vector must be consistent with the state space definition of the deep reinforcement learning agent in step S23 to ensure that the agent can correctly interpret the physical meaning of each dimension in the input vector. The real-time state vector refers to the time-varying multidimensional numerical vector formed by fusing the fluid compressibility modulus dynamic spectrum, instantaneous energy efficiency ratio fingerprint, real-time temperature rise slope, and the initial state vector during the pressurization and pressure holding processes. This vector serves as the environmental state input for the deep reinforcement learning agent to perform dynamic path adjustment in step S24. The real-time temperature rise slope refers to the ratio of the rate of change of real-time fluid temperature to the rate of change of real-time pressure. Physically, it represents the response rate of the fluid temperature within the chamber as pressure increases. This indicator reflects the adiabatic compression heat generation characteristics of the load within the chamber. The calculation method for the real-time temperature rise slope is as follows: The rate of temperature change is calculated based on the real-time fluid temperature sequence collected in step S12, and the rate of pressure change is calculated based on the real-time pressure sequence collected in step S12. The real-time temperature rise slope is obtained by dividing the rate of temperature change by the rate of pressure change. This calculation method is based on the chain rule of calculus. According to the chain rule, the derivative of temperature with respect to pressure, dT / dP, is equal to the derivative of temperature with respect to time, dT / dt, divided by the derivative of pressure with respect to time, dP / dt. Through this transformation, the response characteristics of temperature with respect to pressure can be obtained from two directly calculable time-domain rates of change. Physically, this indicator characterizes the temperature rise caused by a unit pressure increment during adiabatic compression of the material within the chamber, reflecting intrinsic physical properties such as the material's heat capacity and compression heat generation coefficient.Different food matrices have different heat capacities and compressibility heat generation coefficients, resulting in different temperature rise rates under the same pressure increment. For example, foods with high moisture content usually have a lower temperature rise slope than foods with high fat content because water has a larger specific heat capacity.
[0061] The significance of incorporating the real-time temperature rise slope into the real-time state vector lies in the fact that the deep reinforcement learning agent can identify the differences between the actual loaded materials and the preset product categories by comparing the deviation between the real-time temperature rise slope and the historical baseline temperature rise slope, thereby triggering corresponding dynamic adjustment actions. The method for real-time fusion of the fluid compressibility modulus dynamic spectrum, instantaneous energy efficiency ratio fingerprint, real-time temperature rise slope, and initial state vector is as follows: Extract the current fluid compressibility modulus value, instantaneous energy efficiency ratio value, and real-time temperature rise slope value. Use these three values, along with derivative features such as the difference between the previous and current compressibility modulus values, as dynamic dimensions, and concatenate them with the initial state vector to form the real-time state vector. The real-time state vector is continuously updated as the pressurization and pressure holding processes progress. Within each sampling period, the system recalculates the dynamic dimension values and updates the real-time state vector based on the latest collected basic electromechanical signals, ensuring that the deep reinforcement learning agent always obtains the latest information reflecting the current state of the cabin. The initial state vector and the real-time state vector together constitute the information foundation of the closed-loop control system. The initial state vector provides input for the initial decision before the start of the processing cycle, enabling the system to generate the initial target process path based on product characteristics and historical experience before the boost start-up. The real-time state vector provides feedback for dynamic adjustments during the processing, enabling the system to correct the process path based on the actual physical response of the load in the cabin, and achieve adaptive response to unexpected situations.
[0062] Step S10 constructs a multi-dimensional state space based on second-order rheological characteristics and energy efficiency fingerprints, transforming the traditional black-box control mode, which relies solely on pressure and time, into a transparent control mode capable of sensing the physical state of the load within the chamber. By collecting the real-time current value of the booster pump servo motor and the plunger displacement velocity, and correlating them with the real-time pressure within the ultra-high-pressure chamber to construct a dynamic spectrum of fluid compressibility modulus, the system gains the ability to perceive the second-order rheological characteristics of the load within the chamber. This perception capability enables the subsequent deep reinforcement learning agent to distinguish the differences in compression response of different food matrices, thereby adopting differentiated control strategies for different product types. By calculating the ratio of motor input power to hydraulic power, an instantaneous energy efficiency ratio fingerprint is constructed, giving the system the ability to quantify the energy conversion efficiency of the equipment in real time. This quantification capability provides actual operating data support for the subsequent establishment of the energy consumption cost function, ensuring that the energy consumption optimization target is no longer divorced from the actual energy efficiency characteristics of the equipment. By integrating product metadata, initial environmental signals, fluid compressibility modulus dynamic spectrum, instantaneous energy efficiency ratio fingerprint, and real-time temperature rise slope into a unified multi-dimensional state vector, multi-source heterogeneous data is transformed into structured decision input variables. This structured processing enables deep reinforcement learning agents to perform policy reasoning in a unified state space, avoiding information parsing obstacles caused by inconsistent data formats. By distinguishing between the initial state vector and the real-time state vector and establishing the inheritance and update relationship between them, the closed-loop control system simultaneously possesses initial decision-making capabilities based on prior information and dynamic adjustment capabilities based on real-time feedback. This dual capability allows the system to quickly generate initial process paths using historical experience and correct deviations based on actual operating conditions to cope with unexpected situations. Step S10 converts the inherent mechanical characteristics of the equipment's hydraulic system into soft sensor signals, enabling the system to obtain rheological characteristics of the load inside the chamber without the need for additional dedicated physical property detection equipment. This method reduces hardware modification costs while expanding the state perception dimension of the control system. The multi-dimensional state space constructed in step S10 provides a complete information foundation for the multi-objective optimization in step S20 and the adaptive execution control in step S30. This enables the quantitative evaluation of the three optimization objectives of safety, quality and energy consumption, as well as the real-time monitoring of the pressure difference risk inside the packaging, to have calculable data support, thereby enabling the entire intelligent decision-making method to have the ability of closed-loop feedback and adaptive adjustment.
[0063] Step S20: Establish an objective function set, construct a Pareto optimal front library based on the objective function set, map the initial state vector to the Pareto optimal front library to generate an initial target process path, determine whether dynamic adjustment conditions are triggered based on the real-time state vector, if triggered, correct the initial target process path and output the corrected target process path; the objective function set includes energy consumption cost function, safety function and quality loss function.
[0064] Step S20, based on Pareto front-based multi-objective optimization of the process path, employs a hybrid architecture combining offline optimization and online decision-making. Before the batch processing begins, an evolutionary algorithm pre-calculates the Pareto optimal solution set for the multi-objective optimization problem. During processing, a pre-trained deep reinforcement learning agent performs rapid decision-making and dynamic adjustments. The Pareto optimal solution set refers to the set of solutions in a multi-objective optimization problem where no other feasible solution can improve the values of other objective functions without reducing the value of any single objective function. Each solution in the Pareto optimal solution set is called a non-dominated solution. There is no inherent superiority or inferiority among non-dominated solutions; only trade-offs exist between objective function values. Setting ultra-high pressure processing parameters involves three objectives: safety, quality maintenance, and energy consumption control. These three objectives are inherently conflicting: increasing safety margin requires higher pressure or longer holding time, but higher pressure accelerates the degradation of heat-sensitive nutrients, thus damaging quality; longer holding time increases equipment energy consumption. Reducing energy consumption requires shortening the processing cycle, but an excessively short processing cycle may not meet the safety standards for microbial inactivation. The conflict between these multiple objectives cannot be resolved by a single combination of process parameters to simultaneously achieve the extreme values of all objectives. Therefore, it is necessary to find a Pareto optimal solution set and select the process parameter combination that conforms to the current production strategy from this set. Traditional ultra-high pressure machining control systems use static and conservative parameter settings. For example, to ensure that the inactivation effect on specific pathogens meets the logarithmic reduction value required by regulations, a high target pressure and a long holding time are uniformly set regardless of the specific characteristics of the product. Although this extensive control mode can meet safety requirements, it leads to energy waste and excessive degradation of heat-sensitive nutrients. Step S20 establishes a set of objective functions including an energy cost function, a safety function, and a quality loss function to mathematically represent the conflict relationship between multiple objectives. The Pareto optimal frontier is calculated offline using a non-dominated sorting genetic algorithm, enabling the system to quickly select the required process parameter combination on the Pareto optimal frontier according to the current production strategy weights during the online decision-making stage. At the same time, a deep reinforcement learning agent dynamically monitors and adjusts the machining process, allowing the process path to adapt to real-time changes in the physical state of the load inside the chamber.
[0065] Further, step S20 includes:
[0066] Step S21: Establish an energy consumption cost function, and combine it with a preset safety function and quality loss function to form a set of objective functions for a multi-objective optimization problem;
[0067] In step S21, the energy consumption cost function is an objective function with the total energy consumption of the equipment within a single batch processing cycle as the output variable. The calculation of the energy consumption cost function is based on the motor input power obtained in step S14. The method for obtaining the motor input power is the same as described in step S14. Integrating the motor input power over time yields the total energy consumption value E for the entire processing cycle. total The mathematical expression for the energy cost function is E. total The energy cost function is equal to the definite integral of the motor input power over time from zero to the end of the processing cycle, where the motor input power is obtained by multiplying the motor's rated operating voltage and real-time current. The energy cost function uses combinations of process parameters as independent variables, including target pressure, planned holding time, and pressure ramp rate. Different combinations of process parameters correspond to different processing cycle lengths and different power consumption curves, resulting in different total energy consumption values. The energy cost function is established based on instantaneous energy efficiency ratio fingerprints rather than theoretical energy consumption estimates based on equipment nameplate parameters. This is because instantaneous energy efficiency ratio fingerprints reflect the actual energy conversion efficiency of the equipment during operation. This efficiency is affected by various factors such as the load characteristics inside the chamber, the temperature of the pressure transmission medium, and the wear state of the hydraulic system, leading to deviations from theoretical values. The energy cost function established based on actual operating data can more accurately predict the energy consumption levels corresponding to different combinations of process parameters.
[0068] The safety function is an objective function with the cumulative lethality of microorganisms as the output variable. The cumulative lethality is typically represented by a logarithmic reduction value, which is the base-10 logarithmic ratio of the initial microbial quantity to the residual quantity after treatment. The safety function is established based on the extended Arrhenius equation, which expands the temperature-dependent chemical reaction rate constant model into a microbial inactivation kinetic model that simultaneously considers the synergistic effects of pressure and temperature. The mathematical expression of the safety function is that the cumulative lethality of microorganisms equals the definite integral of the pressure-temperature synergistic lethality coefficient over time from zero to the end of the processing cycle. The pressure-temperature synergistic lethality coefficient is a function with real-time pressure and real-time fluid temperature as independent variables. The specific form of this function is determined by microbial inactivation kinetic data published by food safety regulatory agencies. For example, the European Food Safety Authority (EFSA) has published experimental data on the inactivation rate of Listeria monocytogenes under different pressure and temperature conditions. Fitting these experimental data into a pressure-temperature synergistic lethality coefficient function can then be used to calculate the safety function. The safety function also serves as a hard constraint. Hard constraints are boundary conditions that must be met during the optimization process; any solution that violates a hard constraint is deemed infeasible and excluded from the feasible region. The hard constraint corresponding to the safety function is that the logarithmic reduction of the cumulative microbial mortality rate is greater than or equal to a safety threshold. The safety threshold is set based on the mandatory requirements of food safety regulations regarding the inactivation effect of specific pathogens. For example, for the inactivation of Listeria and Salmonella, the safety threshold is typically set to a logarithmic reduction greater than or equal to 5 (i.e., meeting the 5-log kill standard), indicating that the number of residual microorganisms after treatment is reduced to less than one ten-thousandth of the initial number. The safety function is set as a hard constraint rather than a soft constraint because the microbial inactivation effect is directly related to food safety. Any combination of process parameters that does not reach the safety threshold should not be output by the system, even if that combination performs well in terms of quality maintenance and energy consumption control.
[0069] The quality loss function is an objective function with the degradation rate of thermosensitive nutrients or the degree of lipid oxidation as the output variable. Thermosensitive nutrients are those that are prone to chemical decomposition or structural denaturation at higher temperatures; for example, vitamin B1 is a thermosensitive nutrient. The quality loss function is established based on a thermal degradation kinetic model, a mathematical model describing the degradation rate of a specific substance under given temperature and time conditions, a well-known technique in the field of chemical kinetics. Adiabatic compression heat generation occurs during ultra-high pressure processing. Adiabatic compression heat generation refers to the physical process where a fluid's internal energy increases due to volume reduction during compression, resulting in a temperature rise. The real-time fluid temperature collected in step S12 reflects the temperature change within the chamber caused by this heat generation process. Higher temperatures and longer holding times accelerate the degradation of thermosensitive nutrients; therefore, the quality loss function needs to include temperature and time as input variables. The quality loss function uses the combination of process parameters and the temperature change process within the chamber as independent variables, outputting the predicted nutrient retention rate or lipid oxidation index value. The process parameter combination includes the target pressure and the planned holding time, which together determine the intensity and duration of the processing. The temperature change within the chamber is characterized by the real-time fluid temperature sequence acquired in step S12, which records the temperature change trajectory over time from the start of pressurization to the end of the holding period. The specific form of the quality loss function can be achieved using a neural network model trained on historical detection data. The input to the neural network model consists of two parts: the first part is a time-domain sequence composed of real-time pressure and real-time fluid temperature acquired in step S12, reflecting the changes in pressure and temperature over time during processing; the second part consists of parameters related to material properties from the product metadata acquired in step S11. For example, the product ID can indirectly characterize material properties such as fat content and moisture content, which affect the degradation rate of nutrients. The output of the neural network model is the predicted nutrient degradation rate, defined as the ratio of the difference between the nutrient content before and after treatment to the nutrient content before treatment. The nutrient degradation rate serves as the output value of the quality loss function. The training data for the neural network model comes from laboratory test data of historical processing batches. Each training data point includes the real-time pressure sequence, real-time fluid temperature sequence, product metadata, and nutrient content measured in the laboratory for that batch. The model parameters are trained through supervised learning, and the training objective is to minimize the mean square error between the nutrient retention rate predicted by the model and the laboratory measured value.The reason for using a neural network model to establish the quality loss function instead of a simplified analytical expression is that nutrient degradation is affected by the complex composition of food matrix. Different product categories have different fat content, moisture content, and protein structure. These differences have different degrees of regulatory effect on the nutrient degradation rate. These complex interactions are difficult to describe accurately with simple analytical expressions, while neural network models have the ability to approximate arbitrarily complex functions and can learn these implicit interaction patterns from historical data.
[0070] The objective function set consists of an energy consumption cost function, a safety function, and a quality loss function. The constraint set includes hard constraints corresponding to the safety function and boundary constraints corresponding to the equipment hardware limits. Equipment hardware limit constraints include the target pressure not exceeding the upper limit of the equipment's rated pressure, the pressurization rate not exceeding the upper limit of the rate corresponding to the booster pump's rated flow rate, and the pressure holding time not being less than the lower limit corresponding to the equipment control system's response time. Integrating the objective function set and the constraint set forms a complete mathematical description of the multi-objective optimization problem. This mathematical description provides the optimization objective and feasible region boundary for the subsequent step S22, which executes a non-dominated sorting genetic algorithm.
[0071] Step S22: Using the objective function set as the optimization objective, execute the non-dominated sorting genetic algorithm for different product categories to generate Pareto optimal front surfaces corresponding to each product category, and summarize them to form a Pareto optimal front surface library.
[0072] In step S22, the non-dominated sorting genetic algorithm is a type of evolutionary algorithm specifically designed for multi-objective optimization problems. Its core mechanisms include non-dominated sorting, crowding distance calculation, and an elite retention strategy. Non-dominated sorting refers to hierarchically sorting individuals in the population according to Pareto dominance. The first layer contains all individuals not dominated by any other individual; the second layer contains individuals not dominated by any remaining individuals after removing individuals from the first layer, and so on until all individuals are assigned to a certain layer. Crowding distance is an indicator of the sparseness of the distribution of individuals in the same non-dominated layer within the objective space. Individuals with larger crowding distances are located on the edge or in sparse regions of the Pareto front; retaining these individuals helps maintain the diversity and integrity of the Pareto front. The elite retention strategy involves merging the current population with the offspring population in each generation of evolution and then selecting individuals according to the non-dominated sorting layer and crowding distance to ensure that superior individuals are not lost during the evolutionary process. The non-dominated sorting genetic algorithm used in step S22 is an improved third-generation non-dominated sorting genetic algorithm. This algorithm introduces a reference point mechanism on the basis of the traditional non-dominated sorting genetic algorithm. The reference point mechanism guides the population to evolve towards these reference points by pre-setting uniformly distributed reference points in the target space, thereby obtaining a more uniformly distributed Pareto front in the high-dimensional target space.
[0073] The execution of a non-dominated sorting genetic algorithm requires defining a mapping relationship between the decision variable space and the target space. Decision variables include target pressure, planned holding time, and pressurization rate. The target depressurization mode is predetermined by the packaging type in the product metadata and is not used as an optimization decision variable in the evolutionary algorithm. The range of values for the decision variables is determined by equipment hardware limits and process feasibility constraints. For example, the target pressure range can be set to a continuous interval from the lower limit of the equipment's effective operating pressure to the upper limit of the equipment's rated pressure, and the planned holding time range can be set to a continuous interval from the lower limit of the control system's response time to the upper limit of process economics. The target space consists of the output values of the energy cost function, the safety function, and the quality loss function. Each set of decision variable values corresponds to a point in the target space. The goal of the non-dominated sorting genetic algorithm is to find the set of decision variable values that makes the points in the target space as close as possible to the Pareto optimal frontier.
[0074] The operating parameters of a non-dominated sorting genetic algorithm include population size, number of generations, crossover probability, and mutation probability. Population size determines the number of candidate solutions participating in optimization in each generation. A larger population size helps improve the comprehensiveness of the search but increases computation time. The population size needs to be set in a balance between search effectiveness and computational efficiency; for example, the population size can be set to several hundred individuals. The number of generations determines the duration of the evolutionary process. More generations help the population converge more fully to the Pareto optimal front. The number of generations needs to be set in a balance between convergence quality and computational time; for example, the number of generations can be set to several hundred. Crossover probability and mutation probability control the frequency of application of the crossover and mutation operators in genetic operations, respectively. The crossover probability is usually set to a higher value to promote the recombination of superior genes in the population, while the mutation probability is usually set to a lower value to introduce new genetic diversity while maintaining population stability.
[0075] The non-dominated sorting genetic algorithm is executed separately for each product category. The product category division is based on differences in food matrix characteristics. For example, high-fat meat products and high-moisture poultry products should be classified as different product categories due to significant differences in their compressive heat generation characteristics and nutrient degradation sensitivity. After executing the non-dominated sorting genetic algorithm for each product category, the Pareto optimal front for that product category is obtained. Please refer to [link to relevant documentation]. Figure 2 The figure shown is a schematic diagram of the Pareto optimal frontier provided in an embodiment of this application. Figure 2 This presents a three-dimensional target space comprised of three coordinate axes: safety, energy consumption, and quality retention rate. Within this three-dimensional target space, a large number of filtered non-dominated solutions fit together to form a spatial surface, namely the Pareto optimal frontier. Several discrete data points are distributed on this surface, each representing a specific combination of process parameters. For example... Figure 2The point circled and marked as the "optimal solution" is located on the leading edge surface and visually reflects the optimal parameter point selected by the system after weighing multiple objectives. Distributed in... Figure 2 Each point on the surface represents a set of process parameter combinations. These combinations achieve a Pareto optimal balance among the three objectives of safety, quality maintenance, and energy consumption control, and it is impossible to improve the other objectives without sacrificing one of them. Pareto optimal frontiers for all product categories are compiled into a Pareto optimal frontier library. The data structure of the Pareto optimal frontier library uses key-value pairs with the product ID as the index key and the Pareto optimal frontier data as the index value, facilitating the subsequent step S23 to quickly retrieve and match the corresponding Pareto optimal frontier based on the product ID.
[0076] Step S22 is executed during the system deployment and periodic update phases, rather than in real-time at the start of each processing batch. This is because the computational complexity of the non-dominated sorting genetic algorithm is high, and completing a full Pareto optimal front calculation takes a considerable amount of time, failing to meet the rapid response requirements for processing batch startup in industrial production. By moving the computationally intensive Pareto optimal front generation process to the offline phase, the online phase only requires a lightweight operation of querying and matching from the Pareto optimal front library. This allows the system to complete process path generation decisions within milliseconds, meeting the real-time requirements of industrial settings. The periodic update mechanism of the Pareto optimal front library enables the system to continuously absorb operational data from new processing batches. When a sufficient amount of new batch data has accumulated in the historical database, the system triggers a recalculation of the Pareto optimal front library, incorporating equipment state changes and process optimization experience reflected in the new data into the Pareto optimal front, achieving continuous improvement of the control strategy.
[0077] Step S23: Match the corresponding Pareto optimal frontier from the Pareto optimal frontier library based on the product ID in the initial state vector, and select the optimal solution on the matched Pareto optimal frontier in combination with the preset strategy weight vector. Decode the optimal solution to generate the initial target process path. The initial target process path includes target pressure, planned pressure holding time, target pressure increase curve and target pressure relief mode. The target pressure relief mode includes linear pressure relief mode and adaptive pressure relief mode.
[0078] In step S23, the deep reinforcement learning agent is a reinforcement learning agent that uses a deep neural network as a policy function approximator. Its task is to select actions from a predefined action space that maximize long-term cumulative rewards based on the current environmental state. The state space of the deep reinforcement learning agent is defined as the dimensional space of the initial state vector. The action space of the deep reinforcement learning agent is defined as the discrete point index or continuous interpolation coefficients on the Pareto optimal front surface. When the action space uses discrete point indexes, the agent's output action is the index number of a non-dominated solution in the Pareto optimal front surface library. When the action space uses continuous interpolation coefficients, the agent's output action is the weight coefficients used for linear interpolation between adjacent non-dominated solutions. The reward function design of the deep reinforcement learning agent needs to comprehensively consider three objectives: safety, quality preservation, and energy consumption control. The mathematical expression of the reward function is as follows: ,in, For security weights; Weighted by quality; Energy consumption weighting; To normalize the safety margin, S is calculated as follows: when the logarithmic reduction of the cumulative microbial mortality rate is less than the safety threshold, S equals zero; when the logarithmic reduction of the cumulative microbial mortality rate is not less than the safety threshold, S equals the difference between the logarithmic reduction and the safety threshold, divided by the safety margin normalization benchmark value, and the calculation result is limited to the range of zero to one. The safety margin normalization benchmark value is preset according to the equipment process capability and product characteristics. For example, the safety margin normalization benchmark value can be set to two units of the logarithmic reduction value. For quality retention rate, , This is the output value of the quality loss function; The normalized energy saving rate is the ratio of one minus the actual energy consumption to the reference energy consumption limit. The safety weight, quality weight, and energy consumption weight in the reward function correspond to the three components of the strategy weight vector. The strategy weight vector is preset by production managers according to the current production strategy. For example, in the strategy weight vector corresponding to the safety priority mode, the safety weight is higher than the quality weight and energy consumption weight; in the strategy weight vector corresponding to the quality priority mode, the quality weight is higher than the safety weight and energy consumption weight; and in the strategy weight vector corresponding to the energy priority mode, the energy consumption weight is higher than the safety weight and quality weight.
[0079] The neural network architecture of the deep reinforcement learning agent adopts a fully connected network structure. The number of neurons in the input layer is equal to the dimension of the initial state vector. The hidden layers use a multi-layer fully connected structure, with each hidden layer followed by an activation function to achieve a nonlinear transformation. The number of neurons in the output layer is equal to the dimension of the action space. The number of hidden layers and the number of neurons per layer are determined based on the complexity of the state space and the complexity of the action space. For example, a three-layer hidden layer structure can be used, with each layer containing hundreds of neurons. The activation function uses the modified linear unit function or its variants. The modified linear unit function outputs the input value when the input is greater than zero, and outputs zero when the input is less than or equal to zero. This function has the advantages of high computational efficiency and the ability to alleviate the gradient vanishing problem. The training of the deep reinforcement learning agent uses an ultra-high pressure processing simulator built based on historical production data as the training environment. The simulator simulates the pressure, temperature, and flow rate changes during the pressurization, holding, and depressurization processes according to the input process parameter combinations and the load characteristics inside the chamber, and calculates the corresponding cumulative microbial mortality rate, quality loss rate, and total energy consumption. The training process employs either the proximal policy optimization algorithm or the soft actor critic algorithm. Both of these algorithms belong to the policy gradient class and have the advantages of good training stability and high sample efficiency. The hyperparameters of the training process include the learning rate, discount factor, experience replay buffer capacity, and batch size. The learning rate controls the step size of each parameter update, the discount factor controls the weight of future rewards relative to immediate rewards, the experience replay buffer is used to store historical interaction experiences to achieve off-policy learning, and the batch size controls the number of experiences sampled for each parameter update.
[0080] The deep reinforcement learning agent performs decision-making inference at the start of each processing batch. This inference process includes four stages: product matching, policy weight loading, forward propagation, and action decoding. In the product matching stage, a matching Pareto optimal front is retrieved from the Pareto optimal front library based on the product ID in the initial state vector. In the policy weight loading stage, the currently active policy weight vector is read from the production management system. The sum of the three components of the policy weight vector equals one, ensuring mathematical consistency in weight allocation between different objectives. In the forward propagation stage, the initial state vector and policy weight vector are concatenated and input into the deep reinforcement learning agent's neural network. The neural network performs calculations layer by layer and generates action outputs at the output layer. In the action decoding stage, the neural network output is converted into specific combinations of process parameters. When the action space is a discrete point index, the index corresponding to the component with the largest value in the neural network output is taken as the selected Pareto optimal solution index, and the process parameter combination corresponding to that index is extracted from the Pareto optimal front. When the action space is a continuous interpolation coefficient, a weighted average is calculated between adjacent Pareto optimal solutions based on the interpolation coefficients output by the neural network to obtain the interpolated process parameter combination.
[0081] The decoded process parameters form the initial target process path, which includes four components: target pressure, planned holding time, target pressure rise curve, and target depressurization mode. The target pressure refers to the pressure setpoint that needs to be reached in the ultra-high pressure chamber at the end of the pressure rise phase. The planned holding time refers to the preset duration for maintaining the pressure level after reaching the target pressure. The target pressure rise curve is the preset trajectory of pressure change over time during the process of rising from atmospheric pressure to the target pressure. Target pressure rise curves can take the form of linear pressure rise, stepped pressure rise, or exponential pressure rise, with different forms corresponding to different compression heat generation characteristics and different energy consumption distributions. The target depressurization mode refers to the strategy type for performing depressurization operations after the holding phase. Target depressurization modes include linear depressurization mode and adaptive depressurization mode. Linear depressurization mode is suitable for vacuum-packed or skin-packaged products, while adaptive depressurization mode is suitable for modified atmosphere packaging products. The deep reinforcement learning agent automatically selects the corresponding target depressurization mode based on the packaging type information in the initial state vector. When the packaging type is modified atmosphere packaging, it outputs the adaptive depressurization mode; when the packaging type is vacuum packaging or skin-packaged products, it outputs the linear depressurization mode. The reason for using packaging type as the basis for selecting the depressurization mode, rather than a uniform depressurization procedure, is that the residual internal gas content differs significantly between different packaging types. Vacuum packaging and skin packaging have near-zero residual internal gas content, eliminating the pressure difference accumulation issue caused by gas expansion during depressurization. Rapid linear depressurization can shorten the processing cycle per batch, thereby increasing equipment capacity. Modified atmosphere packaging, on the other hand, has significant residual protective gas. Rapid depressurization would cause these gases to expand rapidly and generate destructive stress on the packaging, requiring an adaptive depressurization strategy to balance depressurization efficiency and packaging protection. Without step S23, although the system possesses a Pareto optimal front library, it cannot select appropriate process parameter combinations based on the specific characteristics and production strategy of the current batch. The Pareto optimal front library becomes unusable static data, and the system can only be controlled using fixed process parameters or manual selection, losing its intelligent decision-making capability.
[0082] Step S24: During the pressurization and holding stages, receive the real-time state vector, determine whether the dynamic adjustment condition is triggered based on the real-time state vector, and if the dynamic adjustment condition is triggered, perform a correction on the initial target process path and output the corrected target process path; during the holding stage, continuously calculate the cumulative microbial lethality rate and marginal sterilization benefit, and output the holding termination command when the holding termination condition is met.
[0083] Further, step S24 includes:
[0084] Step S241, see Figure 3The system extracts the real-time temperature rise slope from the real-time state vector, retrieves the historical reference temperature rise slope from the historical database, calculates the deviation rate between the real-time temperature rise slope and the historical reference temperature rise slope, and corrects the target pressure and planned holding time in the initial target process path when the deviation rate exceeds the preset deviation threshold, and outputs the corrected target process path; otherwise, it continues to monitor.
[0085] Step S242: Calculate the cumulative lethality rate of microorganisms and the marginal sterilization benefit. When the pressure holding termination condition is met, output the pressure holding termination command. The pressure holding termination condition is that the cumulative lethality rate of microorganisms meets the safety redundancy condition and the marginal sterilization benefit meets the benefit threshold condition.
[0086] In step S24, the dynamic adjustment condition refers to a situation where, during the pressurization and pressurization process, there is a significant deviation between the physical state of the cabin load reflected by the real-time state vector and the expected state at the time of initial decision-making. The judgment of the dynamic adjustment condition is based on the deviation rate between the real-time temperature rise slope and the historical baseline temperature rise slope. The deviation rate is calculated as follows: the deviation rate equals the difference between the real-time temperature rise slope and the historical baseline temperature rise slope, divided by the absolute value of the historical baseline temperature rise slope. This formula uses a relative deviation form rather than an absolute deviation form because the historical baseline temperature rise slope values vary significantly across different product categories. Using a relative deviation form allows the setting of the deviation threshold to be independent of the product category, facilitating unified threshold management. The historical baseline temperature rise slope refers to the statistical mean of the temperature rise slope in historical processing batches for a specific product category. This mean is retrieved from the historical database by searching the historical processing records of the product category based on the product ID, extracting the temperature rise slope data for each batch, and calculating the statistical mean and standard deviation. The deviation threshold is set based on the statistical distribution characteristics of historical temperature rise slope data. The deviation threshold is usually set as a multiple of the historical standard deviation. For example, the deviation threshold can be set as two to three times the historical standard deviation, so that temperature rise slope changes within the normal process fluctuation range do not trigger dynamic adjustment, while abnormal temperature rise slope deviations can be identified and trigger adjustment actions.
[0087] In step S241, when the deviation rate exceeds the deviation threshold, the deep reinforcement learning agent performs decision-making reasoning for dynamic adjustment actions. The type of dynamic adjustment action is determined based on the deviation direction. When the real-time temperature rise slope is higher than the historical baseline temperature rise slope, it indicates that the compression heat generation rate of the load in the chamber is higher than expected. Possible reasons include the fat content of the actual loaded material being higher than the value indicated on the product label, the raw material entering the chamber being higher than the recorded value, or the loading density deviating from the preset value. In this case, if the pressurization to the target pressure is continued according to the initial target process path, the temperature inside the chamber will exceed the expected level, which may lead to excessive degradation of heat-sensitive nutrients or deterioration of food taste. The deep reinforcement learning agent outputs a corrective action to reduce the target pressure and extend the planned pressure holding time, thereby reducing the target pressure. Force can slow down the rate of heat generation during the subsequent pressurization phase, and extending the planned holding time can compensate for the cumulative lethality of microorganisms at lower pressure levels to ensure the achievement of safety objectives. When the real-time temperature rise slope is lower than the historical baseline temperature rise slope, it indicates that the rate of heat generation during compression of the load in the chamber is lower than expected. Possible reasons include that the moisture content of the actual loaded material is higher than expected or the loading density is lower than the preset value. In this case, the deep reinforcement learning agent outputs corrective actions to maintain or moderately increase the target pressure and shorten the planned holding time. Shortening the planned holding time can reduce equipment operating time and thus reduce energy consumption while meeting safety objectives.
[0088] The magnitude of the correction action is determined by the output quantization of the deep reinforcement learning agent. The agent outputs the target pressure adjustment amount and the planned holding time adjustment amount corresponding to the deviation rate. The mapping relationship between the target pressure adjustment amount and the deviation rate is learned through the agent's training process. During training, the agent experiences numerous dynamic adjustment scenarios at different deviation rate levels, learning, through feedback from reward signals, the optimal adjustment strategy that achieves the best overall performance in terms of safety, quality maintenance, and energy consumption control in the adjusted process path. The corrected target process path includes the adjusted target pressure, the adjusted planned holding time, and the corresponding updated target pressure rise curve. The update method for the target pressure rise curve is as follows: the final pressure value of the pressure rise is re-determined based on the adjusted target pressure, while maintaining the original pressure rise curve form and pressure rise rate. The end time of the pressure rise stage is adjusted accordingly so that the pressure rise curve reaches the adjusted target pressure at the new end time. The target depressurization mode usually remains unchanged because the selection of the depressurization mode depends on the packaging type attribute, which does not change during processing.
[0089] In step S242, the real-time calculation of the cumulative microbial mortality rate uses the same mathematical model as the safety function in step S21. The real-time pressure and fluid temperature sequences from the start of the pressurization phase to the current moment are substituted into the pressure-temperature co-lethality coefficient function. Numerical integration of the mortality coefficient yields the cumulative microbial mortality rate at the current moment. Numerical integration is achieved using the trapezoidal rule or Simpson's rule. Within each sampling period, the mortality coefficient between the current sampling point and the previous sampling point is approximated as a linear distribution. The mortality rate increment within that sampling period is calculated and accumulated to the cumulative value. The marginal sterilization benefit refers to the increment of the cumulative microbial mortality rate corresponding to a unit increase in energy consumption. The formula for calculating the marginal sterilization benefit is: marginal sterilization benefit equals the time derivative of the cumulative microbial mortality rate divided by the motor input power. The numerator of this formula represents the sterilization rate at the current moment, and the denominator represents the energy consumption rate at the current moment. The ratio of the two reflects the cost-effectiveness of continuing the pressure-holding operation.
[0090] The pressure holding termination condition consists of two parts: a safety redundancy condition and a benefit threshold condition. Pressure holding termination is triggered only when both conditions are met simultaneously. The safety redundancy condition requires the logarithmic reduction of the cumulative microbial mortality rate to be greater than or equal to the safety threshold plus the safety redundancy value. The safety redundancy value is an additional margin added to the safety threshold, intended to provide a buffer for potential model prediction errors and fluctuations in actual sterilization effectiveness, ensuring that the actual treatment effect statistically meets safety requirements. The safety redundancy value is set based on the deviation distribution between historical batch laboratory validation data and model predictions. For example, the safety redundancy value can be set to a fraction of the logarithmic reduction value. The benefit threshold condition requires the marginal sterilization benefit to be lower than the benefit threshold. The benefit threshold is a critical value that measures whether continuing pressure holding is economically viable. When the marginal sterilization benefit falls below the benefit threshold, it indicates that the sterilization gain from continuing pressure holding is very small, while the energy cost is relatively high. From a cost-effectiveness perspective, pressure holding should be terminated. The benefit threshold is set based on the economic relationship between energy consumption cost and product value. For example, the benefit threshold can be set as the marginal sterilization benefit value when the energy consumption cost of continuing to hold pressure for a unit time is equal to the increase in product value corresponding to the sterilization effect gain during that time.
[0091] When the pressure holding termination condition is met, step S242 outputs a pressure holding termination command. This command triggers the control system to stop the pressure holding operation of the booster pump and enter the depressurization preparation state. The dynamic termination judgment mechanism, rather than a fixed pressure holding time mechanism, is used to adaptively determine the pressure holding end time based on the actual sterilization process. When the actual sterilization rate is higher than expected, the pressure holding is terminated early to save energy; when the actual sterilization rate is lower than expected, the pressure holding is extended to ensure the achievement of safety targets. This adaptive mechanism allows the pressure holding time to accurately match actual needs without redundancy or insufficiency. Without step S24, the system would only be able to perform fixed pressurization and pressure holding operations according to the initial target process path, unable to adjust based on the actual physical response of the load inside the chamber observed during pressurization. When there is a deviation between the actual material characteristics and the values indicated on the product label, the system cannot compensate, potentially leading to failure to achieve safety targets or excessive quality loss. The continuously updated real-time state vector in step S15 provides the real-time environmental state information required for dynamic adjustment in step S24. Step S24 determines whether the adjustment condition is triggered based on features such as the temperature rise slope in the real-time state vector and outputs the correction action to form a complete closed-loop control loop.
[0092] Step S20 transforms the qualitative description of the ultra-high pressure processing parameter optimization problem into a quantitative multi-objective optimization mathematical model by establishing a set of objective functions and a set of constraints. This allows the conflict between the three objectives of safety, quality maintenance, and energy consumption control to be quantitatively represented through a Pareto optimal solution set. An offline non-dominated sorting genetic algorithm is used to generate a Pareto optimal front library, decoupling the computationally intensive multi-objective optimization solution process from the real-time-critical online decision-making process. This enables the system to achieve multi-objective optimization while meeting the real-time requirements of industrial settings. By introducing a pre-trained deep reinforcement learning agent as the online decision engine, the system can automatically select appropriate combinations of process parameters on the Pareto optimal front based on the specific characteristics of the current batch and the production strategy weights, avoiding the subjectivity and inefficiency of manual parameter selection. By continuously receiving real-time state vectors and judging dynamic adjustment conditions during pressurization and pressure holding, the process path can be corrected according to the actual physical response of the load inside the chamber, compensating for potential deviations between product label values and actual material characteristics, and improving the matching accuracy between process parameters and actual needs. By employing a dynamic pressure holding termination judgment mechanism based on cumulative microbial lethality and marginal sterilization benefits, the pressure holding time can be adaptively determined according to the actual sterilization process, eliminating the pressure holding redundancy commonly found in fixed pressure holding time modes. This reduces unnecessary energy consumption and processing cycle extensions while ensuring safety targets are met. Step S20 transforms the ultra-high pressure processing control system from an open-loop control mode using static conservative parameters to a closed-loop control mode using dynamically optimized parameters. Process parameters are no longer pre-fixed values but are dynamically generated optimization results based on product characteristics, production strategies, and real-time feedback. This transformation enables the system to achieve synergistic optimization of quality maintenance and energy consumption control without sacrificing safety, fully leveraging the comprehensive efficiency of the ultra-high pressure processing equipment.
[0093] Step S30: Determine the final target process path based on the initial target process path and the corrected target process path. During the pressurization and pressure holding stages, control the booster pump to execute the final target process path. If the target pressure relief mode in the final target process path is the adaptive pressure relief mode, generate an adaptive variable rate pressure relief command sequence after the pressure holding ends, and control the pressure relief valve to perform the pressure relief operation according to the adaptive variable rate pressure relief command sequence.
[0094] Specifically, step S30, based on adaptive execution control of the packaging medium response, focuses on dynamically adjusting the opening sequence of the pressure relief valve by monitoring physical feedback signals during the pressure relief process, thereby achieving differentiated pressure relief control for different packaging types. The pressure relief stage refers to the process where, after the pressure holding phase ends during ultra-high pressure processing, the pressure inside the ultra-high pressure chamber gradually decreases from the target pressure to atmospheric pressure. This process achieves controlled release of the pressure-transmitting medium within the chamber by controlling the opening of the pressure relief valve. Traditional ultra-high pressure processing control systems typically employ a linear pressure relief program (i.e., linear pressure relief mode) during the pressure relief stage. A linear pressure relief program refers to a control method that monotonically reduces the pressure inside the chamber to atmospheric pressure at a constant pressure relief rate or a constant valve opening. Linear pressure relief programs are suitable for vacuum-packed or skin-packaged products because these types of packaging contain almost no residual gas, and the pressure difference between the inside and outside of the packaging remains balanced during the pressure relief process, preventing stress concentration caused by gas expansion. Modified atmosphere packaging products face unique physical challenges during ultra-high pressure processing; please refer to [link to relevant documentation]. Figure 4 , Figure 4 This illustrates the gas response process of modified atmosphere packaging from a high-pressure state through depressurization to rapid depressurization. For example... Figure 4 As shown in the high-pressure state, the protective gas filled inside the modified atmosphere packaging is compressed to an extremely high density under external high pressure. When the pressure is released and the external pressure drops sharply, as... Figure 4 As shown in the diagram after rapid depressurization, the compressed gases expand rapidly, causing the packaging to expand and creating a pressure difference between the inside and outside of the packaging. If the expansion rate exceeds the rate at which the gas escapes through the micropores of the packaging film or the sealing gaps, a positive pressure difference will form inside the packaging. When this positive pressure difference exceeds the yield strength of the packaging material, the packaging will suffer damage such as rupture, bulging, or delamination of the seal. Step S30 monitors the microscopic pressure rebound characteristics in the early stage of depressurization, assesses the risk of pressure difference between the inside and outside of the packaging in real time, and dynamically generates an adaptive variable rate depressurization command sequence containing rapid depressurization commands, deceleration depressurization commands, and dwell commands based on the risk value. This allows the depressurization process to adapt to the expansion-equilibrium dynamics of the gas inside the packaging, avoiding packaging damage while ensuring depressurization efficiency.
[0095] Further, step S30 includes:
[0096] Step S31: If the dynamic adjustment condition is not triggered, the final target process path is the initial target process path; if the dynamic adjustment condition is triggered, the final target process path is the corrected target process path, and the booster pump is controlled to execute the final target process path.
[0097] In step S31, the determination of the final target process path follows conditional branching logic. This logic integrates the initial decision and dynamic adjustment results from step S20 to form the final combination of process parameters. When step S24 does not detect a deviation rate exceeding the deviation threshold during the pressurization and holding processes, it is determined that the dynamic adjustment condition has not been triggered. In this case, the final target process path directly adopts the initial target process path output in step S23. The initial target process path includes four components: target pressure, planned holding time, target pressurization curve, and target depressurization mode. When step S241 detects that the deviation rate between the real-time temperature rise slope and the historical baseline temperature rise slope exceeds the deviation threshold, it is determined that the dynamic adjustment condition has been triggered. In this case, the final target process path adopts the corrected target process path output in step S241. The corrected target process path includes the adjusted target pressure, the adjusted planned holding time, and the corresponding updated target pressurization curve. The target depressurization mode usually remains unchanged because the selection of the depressurization mode depends on the packaging type attribute, and the packaging type does not change during processing.
[0098] The method for controlling the booster pump to execute the final target process path is as follows: the target pressure and target pressure rise curve in the final target process path are converted into a speed command sequence for the booster pump servo motor. After receiving the speed command sequence, the servo driver controls the servo motor to drive the plunger forward according to the preset speed curve. The advancement of the plunger forces the pressure transmission medium into the ultra-high pressure chamber, causing the pressure inside the chamber to gradually rise along the target pressure rise curve. When the real-time pressure collected in step S12 reaches the target pressure, the system enters the pressure holding stage. During the pressure holding stage, the pressure inside the chamber is maintained within the allowable fluctuation range near the target pressure by controlling the intermittent operation of the booster pump or by adjusting the pressure compensation valve. When step S242 determines that the pressure holding termination condition is met and outputs the pressure holding termination command, the system stops the pressure holding operation of the booster pump and prepares to enter the pressure relief stage. After receiving the pressure holding termination command, the system reads the target pressure relief mode in the final target process path. When the target pressure relief mode is the linear pressure relief mode, the system controls the pressure relief valve to be fully open or operate at a constant opening, so that the pressure inside the chamber decreases monotonically to atmospheric pressure at a high pressure relief rate. The entire pressure relief process does not require dynamic adjustment. When the target depressurization mode is the adaptive depressurization mode, the system enters steps S32 and S33 to execute the explosion-proof adaptive depressurization strategy.
[0099] Step S32: If the target pressure relief mode is the adaptive pressure relief mode, after receiving the pressure holding termination command, control the pressure relief valve to perform a rapid opening-short-time closing action, collect the pressure micro rebound amplitude during the valve closing period, read the packaging material yield strength threshold from the product feature dataset, and calculate the pressure difference risk value inside the packaging based on the pressure micro rebound amplitude and the packaging material yield strength threshold.
[0100] In step S32, the rapid opening-short-time closing action is an active excitation-response detection method used to detect the expansion state of gas inside the packaging. Rapid opening refers to controlling the pressure relief valve to switch from a closed state to a fully open state in a very short time, causing a rapid, step-like drop in pressure inside the chamber. Short-time closing refers to controlling the pressure relief valve to quickly return to the closed state and maintain this state for a short time window after the rapid opening causes a certain pressure drop. During valve closure, the pressure change behavior inside the chamber reflects the expansion state of the gas inside the packaging: if there is no significant residual gas inside the packaging or the gas expansion rate is slow, the pressure inside the chamber will remain relatively stable or only fluctuate slightly after the valve closes; if there is significant residual gas inside the packaging and the gas is expanding rapidly, the packaging will expand outward like an inflating balloon, compressing the surrounding pressure-transmitting medium, resulting in a measurable small rebound in the total pressure inside the chamber. The amplitude of this rebound is the pressure micro-rebound amplitude.
[0101] The micro-pressure rebound amplitude is acquired using a high-precision pressure sensor. Pressure readings inside the chamber are continuously recorded during valve closure. The local maximum value of the pressure curve after valve closure is compared with the pressure value at the instant of valve closure; the difference between the two is the micro-pressure rebound amplitude. The physical meaning of the micro-pressure rebound amplitude is the manifestation of the squeezing effect of the gas expansion inside the packaging on the surrounding pressure-transmitting medium in the total pressure inside the chamber. A larger amplitude indicates more intense gas expansion inside the packaging, and a higher risk of pressure difference between the inside and outside of the packaging. The micro-pressure rebound amplitude is used as an indirect means of detecting packaging status rather than directly measuring the internal pressure of the packaging. This is because directly measuring the internal pressure of the packaging requires installing miniature pressure sensors on each food package or using invasive measurement methods, which is neither economical nor practical in industrial mass production. The micro-pressure rebound amplitude can be obtained using existing internal pressure sensors, without additional hardware investment.
[0102] The calculation of the pressure difference risk value inside the packaging needs to comprehensively consider two factors: the pressure micro-rebound amplitude and the yield strength threshold of the packaging material. The yield strength threshold of the packaging material is read from the product metadata in step S11 and stored in the product feature dataset. This threshold represents the upper limit of the internal and external pressure difference that the packaging material can withstand. There is an approximately proportional relationship between the pressure micro-rebound amplitude and the actual internal and external pressure difference of the packaging. The physical basis is that when the gas inside the packaging expands, the increase in packaging volume is limited by the elastic modulus of the packaging material and the internal and external pressure difference. This increase in volume has a squeezing effect on the pressure transmission medium inside the chamber, resulting in a corresponding slight increase in the total pressure inside the chamber. In industrial applications where the volume of the ultra-high pressure chamber is much larger than the total volume of the food package, the pressure micro-rebound amplitude and the internal and external pressure difference of the packaging are approximately linearly proportional. Therefore, the pressure micro-rebound amplitude can be used as a substitute indicator for the internal and external pressure difference of the packaging for risk assessment. The formula for calculating the internal pressure difference risk value is Risk = ΔP_rebound / σ_yield, where Risk is the internal pressure difference risk value, ΔP_rebound is the micro-pressure rebound amplitude, and σ_yield is the yield strength threshold of the packaging material. This formula uses a ratio for normalization, making the range of risk values independent of the specific yield strength of the packaging material, facilitating a unified risk threshold setting and risk assessment. When the micro-pressure rebound amplitude equals the yield strength threshold of the packaging material, the risk value is one, indicating that the internal and external pressure difference has reached the material's bearing capacity limit, and continued depressurization will lead to packaging damage. When the micro-pressure rebound amplitude is much smaller than the yield strength threshold of the packaging material, the risk value is much smaller than one, indicating that the packaging is in a safe state and depressurization can continue at a higher rate. When the micro-pressure rebound amplitude is close to but has not yet reached the yield strength threshold of the packaging material, the risk value is close to one, indicating that the packaging is in a critical state, requiring a reduction in the depressurization rate or a pause in depressurization to allow time for the gas inside the packaging to rebalance. The risk value is calculated using a ratio instead of directly using the micro-pressure rebound amplitude as the judgment criterion because the yield strength of different packaging materials varies significantly. For example, the yield strength of some high-strength composite films may be several times that of ordinary polyethylene films. If the micro-pressure rebound amplitude is used directly as the judgment criterion, different judgment thresholds need to be set for each packaging material, increasing the complexity of system configuration. However, by using a normalized risk value, a unified risk threshold can be used for judgment, simplifying the implementation of control logic. If step S32 is missing, the system will not be able to obtain real-time information on the gas expansion state inside the packaging. The subsequent step S33 cannot dynamically adjust the depressurization strategy according to the actual packaging state and can only use a fixed depressurization curve based on preset parameters. This fixed curve cannot adapt to the differences in the residual gas content in the packaging between different batches of products, resulting in packaging damage in some batches or excessively long depressurization times in others.
[0103] Step S33: Compare the internal pressure difference risk value with the preset risk threshold, dynamically adjust the valve opening of the pressure relief valve according to the comparison result, generate an adaptive variable rate pressure relief command sequence including rapid pressure relief command, deceleration pressure relief command and dwell command, and control the pressure relief valve to perform pressure relief operation according to the adaptive variable rate pressure relief command sequence.
[0104] In step S33, the adaptive variable rate pressure relief command sequence is a set of discrete control commands dynamically generated based on the real-time changes in the differential pressure risk value within the packaging. This command sequence controls the pressure relief valve to adopt different opening degrees and operating modes in different pressure ranges. Please refer to [link / reference]. Figure 5 The figure shown is a schematic diagram of the adaptive variable rate depressurization curve provided in the embodiment of this application. Figure 5 Using time as the horizontal axis and cabin pressure as the vertical axis, the comparison between solid and dashed lines visually demonstrates the pressure change characteristics of the adaptive variable rate depressurization mode compared to the traditional linear depressurization mode (shown by dashed lines). Figure 5 As shown by the solid line, the adaptive variable rate depressurization command sequence includes three types of commands: rapid depressurization command, deceleration depressurization command, and dwell command. The rapid depressurization command instructs the depressurization valve to operate at a large opening, causing the chamber pressure to decrease at a high rate. Figure 5 The steep descending line segment marked "rapid depressurization" indicates a large curve slope; this command is generated when the pressure differential risk value inside the packaging is below the risk threshold. The deceleration depressurization command instructs the depressurization valve to operate at a medium or small opening, causing the pressure inside the compartment to decrease at a slower rate. Figure 5 The gently descent line segment marked "Deceleration and Pressure Relief" indicates a decrease in the curve's slope. This instruction is generated when the pressure differential risk within the packaging increases but has not yet exceeded the risk threshold. Its purpose is to reduce the pressure relief rate to slow the accumulation of pressure differential between the inside and outside of the packaging. The Dwell instruction instructs the pressure relief valve to fully close and maintain the current chamber pressure unchanged. Figure 5The horizontal platform segment marked "Dwell" maintains constant pressure within the chamber to facilitate "gas balance." This command is generated when the pressure difference risk value inside the packaging exceeds the risk threshold. Its purpose is to pause the depressurization process, allowing time for the gas inside the packaging to escape through the micropores of the packaging film or the seal gaps, or to redissolve in the food matrix. Depressurization resumes only after the pressure inside and outside the packaging has returned to equilibrium. The risk threshold includes two boundary values: a lower risk threshold and an upper risk threshold. The setting of the risk threshold needs to strike a balance between packaging safety and depressurization efficiency. Setting the upper risk threshold too low will cause the system to trigger deceleration or dwell commands too frequently, prolonging the depressurization time and reducing equipment capacity. Setting the upper risk threshold too high will cause the packaging to continue depressurizing rapidly even when it is close to the breakage threshold, increasing the probability of packaging damage. The method for determining the upper limit of the risk threshold is as follows: For specific types of packaging materials and specific modified atmosphere packaging formulations, gradient decompression tests are conducted under laboratory conditions. The normalized risk value at which micro-damage begins to appear in the packaging is recorded. This value is multiplied by a safety factor less than one to obtain the upper limit of the risk threshold. The value of the safety factor is determined based on product quality requirements and production experience; for example, the safety factor can be set between 0.7 and 0.9. The lower limit of the risk threshold is set based on maximizing the application ratio of rapid decompression while ensuring a smooth transition in decompression control. The lower limit of the risk threshold is usually set as a fixed proportion of the upper limit of the risk threshold; for example, the lower limit of the risk threshold can be set to 0.5 to 0.7 times the upper limit of the risk threshold.
[0105] The generation of the adaptive variable rate depressurization command sequence follows the control logic as follows: the entire depressurization process is divided into multiple pressure ranges, a rapid opening-short closing action is performed in each pressure range to obtain the current pressure micro-rebound amplitude, the pressure difference risk value inside the packaging is calculated, and the depressurization command for the next pressure range is generated based on the comparison result of the risk value and the risk threshold. When the risk value is below the lower limit of the risk threshold, the system generates a rapid pressure relief command, controlling the pressure relief valve to release pressure quickly with a large opening until entering the next pressure range. When the risk value is between the lower and upper limits of the risk threshold, the system generates a deceleration pressure relief command, controlling the pressure relief valve to operate with an opening inversely proportional to the risk value; the higher the risk value, the smaller the opening and the lower the pressure relief rate. When the risk value exceeds the upper limit of the risk threshold, the system generates a dwell command, controlling the pressure relief valve to close completely. During the dwell period, the system continuously monitors the change in the micro-pressure rebound amplitude. When the micro-pressure rebound amplitude begins to decay, it indicates that the gas inside the packaging is being discharged outward through the micropores of the packaging film or the sealing gaps or is redissolved in the food matrix, and the pressure difference between the inside and outside of the packaging is gradually balancing. The system continues to wait until the risk value drops below the risk threshold before generating a rapid pressure relief command or a deceleration pressure relief command to continue the pressure relief process.
[0106] The pressure relief valve is controlled to perform pressure relief operations according to the adaptive variable rate pressure relief command sequence as follows: The system converts each command in the adaptive variable rate pressure relief command sequence into a control signal for the pressure relief valve servo actuator. The servo actuator adjusts the valve core position according to the control signal, thereby changing the valve opening. A rapid pressure relief command corresponds to a large valve opening, a deceleration pressure relief command corresponds to a medium or small valve opening, and a lingering command corresponds to a fully closed valve position. The relationship between valve opening and pressure relief rate follows the valve flow characteristic curve in fluid mechanics, which is usually non-linear. The system converts the desired pressure relief rate into the corresponding valve opening command based on the pre-calibrated valve flow characteristic curve. The pressure relief process continues until the real-time pressure inside the chamber, collected in step S12, drops to atmospheric pressure. At this point, the entire processing cycle is complete, and the system enters batch data archiving mode.
[0107] Step S30 introduces a real-time packaging status detection mechanism based on pressure micro-rebound characteristics, transforming the pressure relief control from a traditional open-loop preset mode to a closed-loop adaptive mode. Traditional pressure relief control uses a pre-set fixed pressure relief curve, which cannot detect differences in residual gas levels between different batches of products, nor can it adapt to individual differences between different food packages within the same batch. This leads to packaging damage in some products due to excessively rapid pressure relief, or extended processing cycles in others due to excessively slow pressure relief. Step S30, by monitoring the pressure micro-rebound amplitude in real time during the pressure relief process, transforms the previously invisible physical quantity of gas expansion state inside the packaging into a measurable and calculable risk indicator, enabling the control system to dynamically adjust the pressure relief strategy based on the actual packaging status. By normalizing and comparing the pressure micro-rebound amplitude with the yield strength threshold of the packaging material to calculate the internal pressure difference risk value, risk judgment is decoupled from specific packaging material characteristics. Different types of packaging materials can use a unified risk threshold for judgment, simplifying the implementation of control logic and improving the system's versatility. By generating an adaptive variable-rate depressurization command sequence that includes rapid depressurization commands, deceleration depressurization commands, and dwell commands, the depressurization process can employ differentiated control strategies in different pressure ranges. Rapid depressurization improves efficiency when risk is low, while deceleration or dwelling protects the packaging when risk is high, achieving a dynamic balance between efficiency and safety. By continuously monitoring the decay trend of the pressure micro-rebound amplitude during the dwell period, the system can determine whether the internal and external pressures of the packaging have rebalanced. This avoids situations where the dwell time is too short, leading to continued packaging damage after depressurization, or too long, resulting in unnecessary processing delays. Step S30 ensures that the modified atmosphere packaging product maintains its intact packaging appearance and sealing performance after ultra-high pressure processing, avoiding defects such as micro-cracks in the seal, film delamination, or bulging caused by improper depressurization, resulting in a smooth and tight packaging appearance at the retail terminal. By selecting differentiated depressurization modes, vacuum-packed and skin-packaged products can employ rapid linear depressurization to shorten processing cycles, while modified atmosphere packaging products can utilize adaptive depressurization to protect packaging integrity. This optimizes overall production capacity without sacrificing the quality of any product type. The constructed adaptive variable rate depressurization control mechanism transforms the depressurization stage of the ultra-high pressure processing from coarse fixed-program control to refined real-time feedback control. Together with the multi-dimensional state space constructed in step S10 and the multi-objective optimization decision achieved in step S20, it forms a closed-loop intelligent control system covering the entire process of pressurization, pressure holding, and depressurization. This allows each stage of the ultra-high pressure processing to adaptively adjust according to the actual state, thereby achieving synergistic optimization in four dimensions: safety, quality maintenance, energy consumption control, and packaging integrity.
[0108] Example 2
[0109] This embodiment, based on Embodiment 1, provides a multi-objective optimized intelligent decision-making system for ultra-high pressure processing parameters, such as... Figure 6 As shown, it includes:
[0110] State vector generation module: used to construct product feature dataset, fluid compressibility modulus dynamic spectrum and instantaneous energy efficiency ratio fingerprint, and generate initial state vector and real-time state vector;
[0111] Process path decision module: It is used to construct the Pareto optimal front library, map the initial state vector to the Pareto optimal front library to generate the initial target process path, and determine whether the dynamic adjustment condition is triggered based on the real-time state vector. If it is triggered, the initial target process path is corrected and the corrected target process path is output.
[0112] Adaptive execution control module: Determines the final target process path based on the initial target process path and the corrected target process path. During the pressurization and pressure holding stages, it controls the booster pump to execute the final target process path. The final target process path includes at least a target depressurization mode. If the target depressurization mode is an adaptive depressurization mode, an adaptive variable rate depressurization command sequence is generated after the pressure holding ends, and the depressurization valve is controlled to perform depressurization operation according to the adaptive variable rate depressurization command sequence.
[0113] Furthermore, in the state vector generation module, the method for constructing the product feature dataset includes:
[0114] Read the product metadata of the batch to be processed and collect the initial environmental signal. Integrate the product metadata and the initial environmental signal to form a product feature dataset. The product metadata includes product ID, packaging type and packaging material yield strength threshold.
[0115] The method for constructing the dynamic spectrum of fluid compressibility modulus includes:
[0116] During the pressurization and pressure holding phases, basic electromechanical signals of the booster pump and the ultra-high pressure chamber are collected. These basic electromechanical signals include the real-time current value of the booster pump servo motor, the plunger displacement speed of the booster pump, the real-time pressure inside the ultra-high pressure chamber, the real-time fluid temperature, and the real-time fluid flow rate. Based on the plunger displacement speed and the real-time pressure inside the ultra-high pressure chamber, the pressure rise rate and volume compressibility are calculated, and a dynamic spectrum of fluid compressibility modulus is generated based on the pressure rise rate and volume compressibility.
[0117] The method for constructing the instantaneous energy efficiency ratio fingerprint includes:
[0118] The motor input power is calculated based on the real-time current value; the hydraulic power is calculated based on the real-time pressure and the real-time fluid flow rate; the ratio of the hydraulic power to the motor input power is used as the instantaneous energy efficiency ratio; and the instantaneous energy efficiency ratio fingerprint is obtained based on the instantaneous energy efficiency ratio.
[0119] Furthermore, in the process path decision module, the method for constructing the Pareto optimal frontier library includes:
[0120] An energy consumption cost function is established, which is combined with a preset safety function and a quality loss function to form a set of objective functions for a multi-objective optimization problem. Using the set of objective functions as the optimization objective, a non-dominated sorting genetic algorithm is executed for different product categories to generate Pareto optimal fronts for each product category, and these are compiled into a Pareto optimal front library.
[0121] The method for determining whether the dynamic adjustment condition has been triggered is as follows:
[0122] The real-time temperature rise slope is extracted from the real-time state vector, the historical reference temperature rise slope is retrieved from the historical database, the deviation rate between the real-time temperature rise slope and the historical reference temperature rise slope is calculated, and when the deviation rate exceeds the preset deviation threshold, the dynamic adjustment condition is triggered.
[0123] The methods and systems of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this application are not limited to the order specifically described above, unless otherwise specifically stated.
[0124] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.
[0125] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-objective optimization intelligent decision-making method for ultra-high pressure machining process parameters, characterized in that, The method includes: Construct a product feature dataset, a dynamic spectrum of fluid compressibility modulus, and an instantaneous energy efficiency ratio fingerprint; generate an initial state vector and a real-time state vector. Construct a Pareto optimal front library, map the initial state vector to the Pareto optimal front library to generate the initial target process path, determine whether the dynamic adjustment condition is triggered based on the real-time state vector, and if it is triggered, correct the initial target process path and output the corrected target process path. The final target process path is determined based on the initial target process path and the revised target process path. During the pressurization and pressure holding stages, the booster pump is controlled to execute the final target process path. The final target process path includes at least a target pressure relief mode. If the target pressure relief mode is an adaptive pressure relief mode, an adaptive variable rate pressure relief command sequence is generated after the pressure holding ends, and the pressure relief valve is controlled to perform the pressure relief operation according to the adaptive variable rate pressure relief command sequence.
2. The intelligent decision-making method for multi-objective optimization of ultra-high pressure processing parameters according to claim 1, characterized in that, The method for constructing the product feature dataset includes: Read the product metadata of the batch to be processed and collect the initial environmental signal. Integrate the product metadata and the initial environmental signal to form a product feature dataset. The product metadata includes product ID, packaging type and packaging material yield strength threshold.
3. The intelligent decision-making method for multi-objective optimization of ultra-high pressure processing parameters according to claim 2, characterized in that, The method for constructing the dynamic spectrum of fluid compressibility modulus includes: During the pressurization and pressure holding phases, basic electromechanical signals of the booster pump and the ultra-high pressure chamber are collected. These basic electromechanical signals include the real-time current value of the booster pump servo motor, the plunger displacement speed of the booster pump, the real-time pressure in the ultra-high pressure chamber, the real-time fluid temperature, and the real-time fluid flow rate. Based on the plunger displacement velocity and the real-time pressure inside the ultra-high pressure chamber, the pressure rise rate and volume compressibility are calculated, and a dynamic spectrum of fluid compressibility modulus is generated based on the pressure rise rate and volume compressibility.
4. The intelligent decision-making method for multi-objective optimization of ultra-high pressure machining process parameters according to claim 3, characterized in that, The method for constructing the instantaneous energy efficiency ratio fingerprint includes: The motor input power is calculated based on the real-time current value; the hydraulic power is calculated based on the real-time pressure and the real-time fluid flow rate; the ratio of the hydraulic power to the motor input power is used as the instantaneous energy efficiency ratio; and the instantaneous energy efficiency ratio fingerprint is obtained based on the instantaneous energy efficiency ratio.
5. The intelligent decision-making method for multi-objective optimization of ultra-high pressure machining process parameters according to claim 4, characterized in that, The method for generating the initial state vector includes: Before the boost begins, the historical rheological feature template corresponding to the product ID in the product feature dataset is retrieved from the historical database, and the product feature dataset and the historical rheological feature template are fused to generate the initial state vector.
6. The intelligent decision-making method for multi-objective optimization of ultra-high pressure machining process parameters according to claim 5, characterized in that, The method for generating the real-time state vector includes: During the pressurization and pressure holding processes, the real-time temperature rise slope is calculated based on the real-time fluid temperature and real-time pressure. The fluid compressibility modulus dynamic spectrum, instantaneous energy efficiency ratio fingerprint, real-time temperature rise slope, and initial state vector are fused in real time to continuously update and generate the real-time state vector.
7. The intelligent decision-making method for multi-objective optimization of ultra-high pressure machining process parameters according to claim 6, characterized in that, The method for constructing the Pareto optimal front surface library includes: Establish an energy consumption cost function, and combine it with a preset safety function and quality loss function to form a set of objective functions for a multi-objective optimization problem; Using the objective function set as the optimization objective, a non-dominated sorting genetic algorithm is executed for different product categories to generate Pareto optimal fronts for each product category, and these are then compiled into a Pareto optimal front library.
8. The intelligent decision-making method for ultra-high pressure machining process parameters with multi-objective optimization according to claim 7, characterized in that, The method for determining whether the dynamic adjustment condition has been triggered is as follows: The real-time temperature rise slope is extracted from the real-time state vector, the historical reference temperature rise slope is retrieved from the historical database, the deviation rate between the real-time temperature rise slope and the historical reference temperature rise slope is calculated, and when the deviation rate exceeds the preset deviation threshold, the dynamic adjustment condition is triggered.
9. The intelligent decision-making method for ultra-high pressure machining process parameters with multi-objective optimization according to claim 8, characterized in that, The pressure holding process terminates upon receiving a pressure holding termination command. The conditions for generating the pressure holding termination command are: During the pressure holding phase, the cumulative microbial lethality rate and marginal sterilization benefit are continuously calculated. When the pressure holding termination condition is met, a pressure holding termination command is output. The pressure holding termination condition is that the cumulative microbial lethality rate meets the safety redundancy condition and the marginal sterilization benefit meets the benefit threshold condition.
10. A multi-objective optimization intelligent decision-making system for ultra-high pressure machining process parameters, used to implement the multi-objective optimization intelligent decision-making method for ultra-high pressure machining process parameters as described in any one of claims 1-9, characterized in that, The system includes: State vector generation module: used to construct product feature dataset, fluid compressibility modulus dynamic spectrum and instantaneous energy efficiency ratio fingerprint, and generate initial state vector and real-time state vector; Process path decision module: It is used to construct the Pareto optimal front library, map the initial state vector to the Pareto optimal front library to generate the initial target process path, and determine whether the dynamic adjustment condition is triggered based on the real-time state vector. If it is triggered, the initial target process path is corrected and the corrected target process path is output. Adaptive execution control module: Determines the final target process path based on the initial target process path and the corrected target process path. During the pressurization and pressure holding stages, it controls the booster pump to execute the final target process path. The final target process path includes at least a target depressurization mode. If the target depressurization mode is an adaptive depressurization mode, an adaptive variable rate depressurization command sequence is generated after the pressure holding ends, and the depressurization valve is controlled to perform depressurization operation according to the adaptive variable rate depressurization command sequence.
Citation Information
Patent Citations
Decision-making system for predicting flavor formation mechanism and flavor optimization in food processing based on machine learning
CN120509519A
Deep learning model-oriented multi-target hyper-parameter joint optimization method and system
CN121235039A