A control method for biofeed conversion process based on multi-objective optimization

Through the deep integration of the adaptive multi-task and multi-objective reinforcement learning algorithm and the multi-modal ant colony collaborative optimization algorithm, real-time dynamic and precise regulation of multiple parameters in the biological feed conversion process is achieved, solving the problems of insufficient parameter coordination control accuracy and delayed response in existing technologies, improving conversion efficiency and resource utilization, and reducing environmental impact.

CN120471091BActive Publication Date: 2025-09-19SHENYANG FENGSUO ANIMAL HUSBANDRY FEED CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510969519.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-19
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively balancing the nonlinear relationship between parameters such as temperature, pH, bacterial flora ratio and nutrient ratio during the biofeed conversion process, resulting in insufficient control accuracy, low resource utilization, large environmental impact, and delayed regulatory response, which limits the improvement of conversion efficiency and industrial economic efficiency.

Method used

A deep fusion method based on adaptive multi-task and multi-objective reinforcement learning algorithm and multi-modal ant colony collaborative optimization algorithm is adopted. The process parameters are perceived in real time through deep neural network, the ant colony exploration mode is dynamically adjusted, and a real-time control strategy is generated. The industrial Internet of Things is used to feedback control instructions to achieve precise and real-time closed-loop control of multiple parameters.

Benefits of technology

It significantly improves the biofeed conversion efficiency and resource utilization, reduces environmental impact, enhances the control accuracy and stability of process parameters, solves the problems of insufficient parameter coordination control accuracy and dynamic response lag in existing technologies, and improves economic and environmental benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471091B_ABST
    Figure CN120471091B_ABST
Patent Text Reader

Abstract

The present invention discloses a biofeed conversion process control method based on multi-objective optimization, comprising the following steps: establishing a multi-objective optimization mathematical model based on real-time collected temperature values, pH values, bacterial flora ratio values, and nutrient ratio values ​​to obtain a multidimensional decision space; utilizing a deep neural network to perceive the state of the decision space in real time, determine the optimization target priorities and weights, and generate a real-time control strategy; adjusting a multimodal ant colony collaborative optimization algorithm based on the real-time control strategy to obtain a set of candidate solutions; selecting the optimal control scheme from the set of candidate solutions using a reinforcement learning strategy; converting the control scheme into process control instructions and controlling them in real time; and providing feedback on the process parameters after real-time control, dynamically updating the decision space, and continuously optimizing the conversion process. The present invention improves conversion efficiency and resource utilization, and reduces environmental impact.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent optimization and control, and in particular to a biological feed conversion process control method based on multi-objective optimization. Background Art

[0002] Biofeed conversion process control technology is one of the important supporting technologies in the field of modern biomanufacturing. It mainly achieves the efficient conversion of feed raw materials into target products by adjusting key process parameters such as temperature, pH, bacterial flora ratio, and nutrient ratio. In recent years, with the development of industrial Internet of Things and online sensing technology, relevant technicians have developed a variety of automated monitoring and control technologies. Existing technical solutions usually use industrial sensing equipment to monitor changes in process parameters in real time and adjust the parameters with the help of feedback control principles to achieve the goal of stable production. In addition, to improve the accuracy of production process control, existing technologies have introduced intelligent optimization algorithms, such as ant colony algorithms, genetic algorithms, reinforcement learning algorithms, etc., which are respectively applied to the optimization of process parameters or decision-making strategy optimization.

[0003] At present, most existing technical solutions use a single algorithm or a simple combination algorithm, and use traditional control principles to achieve parameter optimization and real-time adjustment. These technologies generally optimize the process parameters independently first, and then determine the specific instructions for parameter control based on experience. Although such technical solutions can improve the process control effect to a certain extent, in the complex biological feed conversion process, due to the highly nonlinear relationship and mutual influence between parameters such as temperature, pH, bacterial population ratio and nutrient ratio, the existing independent optimization control solutions are difficult to effectively balance the contradictions between multiple objectives in actual applications. The parameter control accuracy is limited, the resource utilization rate is not high, and it may even aggravate the emission of waste in the production process and increase the environmental burden. In addition, relying solely on empirical control methods cannot fully adapt to the dynamic change trend of process parameters. In actual industrial production processes, there are problems such as control response lag, insufficient control stability, and difficulty in continuous process optimization, which seriously limit the improvement of feed conversion efficiency and the improvement of industrial economic efficiency.

[0004] Therefore, how to provide a biological feed conversion process regulation method based on multi-objective optimization is a problem that those skilled in the art urgently need to solve. Summary of the Invention

[0005] One purpose of the present invention is to propose a method for controlling the biofeed conversion process based on multi-objective optimization. In order to solve the problem of how to scientifically control multiple parameters such as temperature, pH, bacterial flora ratio and nutrient ratio during the conversion process, a deep fusion method based on an adaptive multi-task multi-objective reinforcement learning algorithm and a multimodal ant colony collaborative optimization algorithm is proposed to achieve accurate and real-time closed-loop control of process parameters. The present invention has the advantages of significantly improving conversion efficiency, improving resource utilization and reducing environmental impact.

[0006] The biological feed conversion process control method based on multi-objective optimization according to an embodiment of the present invention includes:

[0007] Based on the real-time data collected during the biofeed conversion process, a multi-objective optimization mathematical model is established to obtain a multi-dimensional decision space;

[0008] Through deep neural networks, the system can perceive the environmental status information of the multidimensional decision space in the current bio-feed conversion process in real time, evaluate the conflict degree between the current optimization objectives and the resource constraints, and generate real-time control strategies;

[0009] Based on the real-time control strategy, the exploration mode and exploration direction of the global exploration ant colony and the local optimization ant colony in the multimodal ant colony collaborative optimization algorithm are dynamically adjusted to obtain the candidate solution set;

[0010] Evaluate the effectiveness of each solution in the candidate solution set through reinforcement learning strategy and select the optimal control solution;

[0011] Utilize the real-time collection of bio-feed conversion process parameters from the Industrial Internet of Things and online sensor networks to convert the optimal control scheme into specific process control instructions and feed them back to the bio-feed conversion process control equipment;

[0012] The biological feed conversion process parameters after real-time regulation are fed back to the multi-objective optimization mathematical model to dynamically update the multi-dimensional decision space.

[0013] Optionally, a multi-objective optimization mathematical model is established based on the real-time data collected during the bio-feed conversion process to obtain a multi-dimensional decision space, specifically:

[0014] Real-time collection of temperature values, pH values, bacterial flora ratio values, and nutrient ratio values ​​during the biological feed conversion process to form optimization decision variables;

[0015] Taking the conversion efficiency value, resource utilization efficiency value and environmental impact index value as the optimization target, the optimization target is expressed as an objective function composed of optimization decision variables;

[0016] Determine the value range of optimization decision variables and form a multi-dimensional optimization decision variable constraint space;

[0017] The nonlinear fitting method is used to determine the mathematical relationship between the optimization decision variables and the optimization objectives, and the mathematical expression of the objective function is obtained;

[0018] Construct a multi-objective optimization mathematical model based on the mathematical expression of the objective function and the constraint space of the optimization decision variables;

[0019] Calculate the objective function value using the multi-objective optimization mathematical model solution algorithm and obtain the feasible solution set of the objective function;

[0020] According to the obtained feasible solution set of the objective function, a multidimensional decision space of the biofeed conversion process is generated.

[0021] Optionally, the deep neural network is used to perceive the environmental status information of the multidimensional decision space in the current bio-feed conversion process in real time, and to evaluate the conflict degree and resource constraints between the current optimization objectives to generate a real-time control strategy, specifically:

[0022] Real-time acquisition of temperature values, pH values, bacterial flora ratio values, and nutrient ratio values ​​in the multidimensional decision space, and standardization of these values ​​to obtain normalized state data.

[0023] The normalized state data is input into the trained deep neural network, and the multi-layer feedforward connection structure inside the deep neural network is used for layer-by-layer nonlinear mapping to obtain the specific numerical values ​​of the degree of conflict between the optimization objectives;

[0024] According to the specific values ​​of the conflict degree between the optimization objectives, the linear weighted summation method is used to calculate the quantitative index of the comprehensive conflict degree;

[0025] According to the comprehensive conflict degree quantitative index, the priority ranking of each optimization goal is determined through the goal priority decision matrix;

[0026] Based on the priority sorting results, a dynamic weight allocation method is used to determine the real-time optimization weight coefficients of the conversion efficiency value, resource utilization efficiency value and environmental impact index value;

[0027] Based on the determined real-time optimization weight coefficient and the current state information of the multidimensional decision space, the optimized control target value combination of temperature value, pH value, bacterial population ratio value and nutrient ratio value is calculated to generate a real-time control strategy.

[0028] Optionally, the real-time control strategy is used to dynamically adjust the exploration mode and exploration direction of the global exploration ant colony and the local optimization ant colony in the multimodal ant colony collaborative optimization algorithm, and to update the pheromone distribution for guiding the ant colony search path in real time to obtain a candidate solution set for the multi-objective optimization of the biological feed conversion process, specifically:

[0029] Generate the initial pheromone field intensity distribution according to the real-time control strategy;

[0030] According to the initial pheromone field intensity distribution, the initial position and number of the global exploration ant colony are determined, and the initial exploration direction and step length of the global exploration ant colony are determined according to the pheromone field intensity gradient direction;

[0031] According to the exploration results of the global exploration ant colony, the pheromone concentration is dynamically adjusted, and the pheromone concentration gradient is used to determine the search area boundary of the local optimization ant colony;

[0032] According to the search area boundary, the local optimization ant colony conducts an intensive search inside the boundary and records the pheromone distribution corresponding to the search path and the objective function value of the optimization solution in real time;

[0033] According to the objective function value corresponding to the recorded search path, the pheromone enhancement coefficient and volatility coefficient of the global exploration ant colony and the local optimization ant colony are determined, and the pheromone field strength in the multidimensional decision space is dynamically updated according to the high and low objective function values ​​corresponding to the optimization solution;

[0034] According to the updated pheromone field strength and corresponding gradient changes, the exploration mode, exploration direction and step size of the global exploration ant colony and the local optimization ant colony in the next cycle are readjusted, and the pheromone sharing method between ant colonies of different modes is determined;

[0035] According to the updated pheromone field intensity distribution, the candidate solution set of multi-objective optimization of the biological feed conversion process is recovered.

[0036] Optionally, the method evaluates the effectiveness of each solution in the obtained multi-objective optimization candidate solution set through a reinforcement learning strategy, and selects the optimization solution that meets the current optimization goal priority and weight requirements as the optimal control solution, specifically:

[0037] According to the optimization target obtained by the real-time control strategy, the weight coefficient is optimized in real time, and the temperature value, pH value, bacterial population ratio value and nutrient ratio value corresponding to each candidate solution are obtained from the multi-objective optimization candidate solution set;

[0038] The temperature, pH, bacterial composition, and nutrient composition values ​​corresponding to each candidate solution are used as the optimization action to be evaluated. The comprehensive evaluation index value corresponding to each optimization action is calculated based on the reinforcement learning strategy.

[0039] According to the comprehensive evaluation index value, all candidate solutions in the multi-objective optimization candidate solution set are quantitatively ranked to obtain the ranking results of the candidate solutions;

[0040] According to the ranking results of the candidate solutions, the temperature value, pH value, bacterial population ratio value and nutrient ratio value combination corresponding to the highest-ranked candidate solution are determined as the optimal control solution that meets the current optimization target priority and weight requirements.

[0041] Optionally, the reinforcement learning strategy is specifically:

[0042] According to the temperature value, pH value, bacterial population ratio value and nutrient ratio value corresponding to each candidate solution in the multi-objective optimization candidate solution set, a state-action mapping relationship between the current state and the optimization action is constructed;

[0043] According to the state-action mapping relationship, the dynamic adjustment factor of the task priority corresponding to each optimization action is calculated based on the real-time optimization weight coefficient of each optimization target corresponding to the current state;

[0044] According to the task priority dynamic adjustment factor, the action value function value of each optimization action is updated with the goal of minimizing the degree of conflict between tasks;

[0045] Based on the updated action-value function value, an action-value function search space is generated, and the potential long-term cumulative reward function value of each optimized action in the action-value function search space is calculated;

[0046] Calculating the action value confidence interval corresponding to each optimization action based on the value of the potential long-term cumulative reward function, and obtaining the upper and lower confidence bounds of each optimization action in the action value function search space;

[0047] According to the action value function value of the optimization action in the action value function search space and the corresponding action value confidence interval, the optimal optimization action under the current state is determined with the goal of maximizing the action value and the confidence interval.

[0048] According to the temperature value, pH value, bacterial population ratio value and nutrient ratio value corresponding to the determined optimal optimization action, the corresponding comprehensive evaluation index value is calculated.

[0049] Optionally, based on the optimal control scheme, the parameters of the bio-feed conversion process collected in real time by the industrial Internet of Things and online sensor networks are used to convert the optimal control scheme into specific process control instructions, and the instructions are fed back to the bio-feed conversion process control equipment to achieve real-time control of the conversion process, specifically:

[0050] Generate multi-objective control instructions with a tolerance range based on the target value combination of temperature, pH, bacterial population ratio and nutrient ratio determined in the optimal control plan;

[0051] The industrial Internet of Things is used to collect the actual temperature, pH, bacterial flora ratio and nutrient ratio of the current biological feed conversion process in real time, and each collected value is subjected to continuous multi-point smoothing filtering to obtain the actual values ​​of each process parameter after real-time smoothing;

[0052] The difference between the actual value of each process parameter after real-time smoothing and the corresponding process parameter target value is calculated to obtain the comprehensive deviation of each process parameter;

[0053] For the comprehensive deviation, a joint assessment is conducted based on the deviation change rate and historical cumulative deviation amplitude to form a dynamic adjustment value for each process parameter in the bio-feed conversion process;

[0054] Specific process control instructions are generated based on the dynamic adjustment quantity and transmitted in real time to the execution terminal of the biological feed conversion process control equipment through the industrial Internet of Things.

[0055] Optionally, the biological feed conversion process parameters after real-time control are fed back to the multi-objective optimization mathematical model, the multi-dimensional decision space is dynamically updated, and the above steps are repeatedly performed, specifically:

[0056] The Industrial Internet of Things is used to collect the temperature, pH, bacterial composition, and nutrient composition of the biological feed conversion process in real time after real-time regulation, and to calculate the difference between the corresponding parameter values ​​before the previous real-time regulation to obtain the actual regulation response increment of each parameter;

[0057] Using the actual control response increment as the new sample data, the nonlinear fitting method in the multi-objective optimization mathematical model is used to dynamically update the functional mapping relationship between each process parameter and the conversion efficiency value, resource utilization efficiency value and environmental impact index value;

[0058] Reconstruct the objective function mathematical expression in the multi-objective optimization mathematical model using the dynamically updated function mapping relationship;

[0059] Based on the reconstructed mathematical expression of the objective function, the multidimensional decision space is updated and adjusted, and the feasible solution domain of the optimization variables is re-divided;

[0060] Based on the re-divided feasible solution domain of the optimization variables, the task priority dynamic adjustment factor in the adaptive multi-task and multi-objective reinforcement learning algorithm is dynamically adjusted to obtain a new real-time control strategy;

[0061] Based on the new real-time control strategy, the multimodal ant colony collaborative optimization algorithm is re-executed to obtain a new set of multi-objective optimization candidate solutions;

[0062] The new set of multi-objective optimization candidate solutions is used as the basis for the next real-time regulation, and the feedback closed-loop process is continuously iterated.

[0063] The beneficial effects of the present invention are:

[0064] (1) The present invention realizes real-time dynamic and precise control of multiple parameters such as temperature, pH, bacterial population ratio and nutrient ratio through a deep fusion method based on an adaptive multi-task multi-objective reinforcement learning algorithm and a multi-modal ant colony collaborative optimization algorithm, effectively improving the overall conversion efficiency and resource utilization efficiency of the biological feed conversion process, significantly reducing the environmental impact indicators generated during the production process, and enhancing the accuracy and stability of process control.

[0065] (2) The present invention realizes rapid response and stable closed-loop control of biofeed conversion process parameters through the dynamic update mechanism of the multi-objective optimization mathematical model and the adaptive generation of real-time control strategies, significantly improving the continuous stability of process parameters and the timeliness of process adjustment, and showing better adaptability in the biofeed industrial production process.

[0066] (3) In terms of parameter coordination control of multi-objective optimization, the present invention effectively solves the shortcomings of insufficient coordination control accuracy and dynamic response lag of various process parameters in the existing technology through dynamic closed-loop control with a tolerance range and a real-time multi-parameter deviation fusion mechanism, breaks through the limitations of a single optimization method, and achieves specific and significant improvements in process control accuracy and response speed, effectively improving the economic and environmental benefits of the biofeed conversion process. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0068] Figure 1 This is an overall flow chart of the biological feed conversion process control method based on multi-objective optimization proposed in the present invention. DETAILED DESCRIPTION

[0069] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0070] refer to Figure 1 , a biofeed conversion process control method based on multi-objective optimization, including:

[0071] Based on the real-time temperature values, pH values, bacterial flora ratio values, and nutrient ratio values ​​collected during the bio-feed conversion process, a multi-objective optimization mathematical model with conversion efficiency values, resource utilization efficiency values, and environmental impact index values ​​as optimization targets is established to obtain a multi-dimensional decision space;

[0072] Based on the obtained multidimensional decision space, a deep neural network is used to perceive the environmental status information of the multidimensional decision space in the current biofeed conversion process in real time, and to evaluate the conflict degree and resource constraints between the current optimization objectives, dynamically determine the priority and corresponding weight of each optimization objective, and generate a real-time control strategy corresponding to the current state;

[0073] Based on the real-time control strategy, the exploration mode and exploration direction of the global exploration ant colony and the local optimization ant colony in the multimodal ant colony collaborative optimization algorithm are dynamically adjusted, and the pheromone distribution used to guide the ant colony search path is updated in real time to obtain a candidate solution set for the multi-objective optimization of the biological feed conversion process;

[0074] Based on the obtained multi-objective optimization candidate solution set, the effectiveness of each solution in the candidate solution set is evaluated through reinforcement learning strategy, and the optimization solution that meets the current optimization goal priority and weight requirements is selected as the optimal control solution;

[0075] Based on the optimal control scheme, the bio-feed conversion process parameters collected in real time by the industrial Internet of Things and online sensor networks are used to convert the optimal control scheme into specific process control instructions, and the instructions are fed back to the bio-feed conversion process control equipment to achieve real-time control of the conversion process;

[0076] The biological feed conversion process parameters after real-time control are fed back to the multi-objective optimization mathematical model, the multidimensional decision space is dynamically updated, and the above steps are repeatedly performed to continuously keep the biological feed conversion process in an optimized state.

[0077] In this embodiment, a multi-objective optimization mathematical model is established based on the real-time data collected during the bio-feed conversion process to obtain a multi-dimensional decision space, specifically:

[0078] Real-time collection of temperature values, pH values, flora ratio values, and nutrient ratio values ​​during the bio-feed conversion process. The temperature value is collected every 5 seconds with an accuracy of 0.1°C by a temperature sensor, the pH value is collected every 10 seconds with an accuracy of 0.01 by an industrial-grade pH sensor, and the flora ratio value and nutrient ratio value are collected every 30 seconds in the form of mass percentage by an online biomass composition analyzer to form optimization decision variables.

[0079] The conversion efficiency value, resource utilization efficiency value, and environmental impact index value are used as optimization targets, wherein the conversion efficiency value is the ratio between the yield of the target product per unit mass of feed after biological conversion and the theoretical maximum yield, the resource utilization efficiency value is the actual output value of the target product under unit energy consumption or raw material consumption, and the environmental impact index value is a comprehensive evaluation index of the concentration of pollutants in waste gas and wastewater generated during the conversion process. The temperature value, pH value, bacterial population ratio value, and nutrient ratio value are used as independent variables, and the conversion efficiency value, resource utilization efficiency value, and environmental impact index value are used as dependent variables respectively. Based on historical data and real-time collected data, a nonlinear regression analysis method is used to fit the objective function that characterizes the mapping relationship between each optimization target and the optimization decision variable in the form of a polynomial function;

[0080] Based on the specific biological characteristics of the feed conversion process and actual industrial production conditions, the constraint range of temperature values ​​was determined to be 30°C to 60°C, the constraint range of pH values ​​was determined to be pH 5.0 to 8.0, the constraint range of microbial population ratio values ​​was determined to be 10% to 50% of the total microbial population mass, and the constraint range of nutrient ratio values ​​was determined to be 20% to 60% of the total mass. The value range of each optimization decision variable was clarified, and the Cartesian product method was used to construct a multidimensional optimization decision variable constraint space with temperature values, pH values, microbial population ratio values, and nutrient ratio values ​​as dimensions.

[0081] The nonlinear least squares curve fitting method is used to perform fitting analysis on the data samples between the optimization decision variables and the optimization targets. The optimization target values ​​and the optimization decision variable values ​​collected in real time are used as input and output data samples to obtain the objective function mathematical expression that can reflect the mapping relationship between the optimization decision variables and the optimization targets:

[0082] ;

[0083] in, It is the comprehensive multi-objective optimization function value. The larger the value, the better the comprehensive performance. It is the real-time temperature value during the biological feed conversion process. It is the real-time pH value during the biological feed conversion process. It is the bacterial flora ratio value collected in real time. It is the nutrient ratio value collected in real time. The actual yield of target products generated by biological feed conversion under given temperature, pH, bacterial flora ratio and nutrient ratio conditions. is the theoretical maximum yield of the biological feed conversion process, is the energy consumption value or total raw material consumption of the biological feed conversion process under given conditions, It is a comprehensive indicator of exhaust gas pollutant concentration under given conditions. is a comprehensive indicator of wastewater pollutant concentration under given conditions. 、 、 Optimize the weight coefficient for each objective, which is a non-negative number and satisfies , 、 is the comprehensive evaluation weight coefficient of waste gas and wastewater pollutants in the environmental impact indicators, ;

[0084] Based on the mathematical expression of the objective function and the constraint space of the optimization decision variables obtained above, a nonlinear multi-objective optimization mathematical model is constructed with multiple optimization goals of improving conversion efficiency, improving resource utilization efficiency, and reducing environmental impact;

[0085] The nonlinear multi-objective optimization mathematical model is used to solve the problem using a non-dominated sorting genetic algorithm, wherein the initial population size is set to 100, the crossover probability is set to 0.8, the mutation probability is set to 0.05, and the maximum evolutionary generation is 500. The genetic algorithm is used to iteratively solve the problem to obtain a feasible solution set of the objective function that meets the optimization goal.

[0086] The feasible solutions in the feasible solution set of the above objective function are converted through a multidimensional space mapping method, and are combined and classified according to the values ​​of temperature, pH, bacterial population ratio and nutrient ratio to generate a multidimensional decision space that specifically reflects the optimization and control conditions of the biological feed conversion process.

[0087] In this embodiment, the deep neural network is used to perceive the environmental status information of the multidimensional decision space in the current biological feed conversion process in real time, and to evaluate the conflict degree and resource constraints between the current optimization objectives to generate a real-time control strategy, specifically:

[0088] The temperature, pH, bacterial composition, and nutrient composition values ​​in the multidimensional decision space are acquired in real time and standardized using the minimum-maximum standardization method, converting each value into a normalized state data between 0 and 1 to eliminate the dimensional influence of each optimization decision variable.

[0089] The normalized state data after standardization is input into a deep neural network consisting of an input layer, five hidden layers, and an output layer. The number of neurons in the input layer is the number of optimization decision variables, and the number of neurons in the hidden layer is set to 128, 256, 512, 256, and 128, respectively. The number of neurons in the output layer is the number of conflict degrees between the optimization objectives. Through a multi-layer feedforward connection structure and a layer-by-layer nonlinear mapping with the activation function being the hyperbolic tangent function, the specific values ​​of the conflict degrees between the optimization objectives are obtained.

[0090] According to the specific values ​​of the degree of conflict between the optimization objectives, a weighting coefficient is assigned to each pair of optimization objective combinations, where the weighting coefficient between the conversion efficiency value and the resource utilization efficiency value is 0.4, the weighting coefficient between the conversion efficiency value and the environmental impact index value is 0.35, and the weighting coefficient between the resource utilization efficiency value and the environmental impact index value is 0.25. The product operation of each conflict degree value and the corresponding weighting coefficient is performed respectively and then summed to calculate a comprehensive conflict degree quantitative index that can clearly reflect the overall optimization objective conflict status in the current state;

[0091] Based on the comprehensive conflict degree quantitative index, combined with the target priority decision matrix constructed by the hierarchical analysis method, the rows and columns of the target priority decision matrix respectively represent the conversion efficiency value, resource utilization efficiency value, and environmental impact index value. The value of each matrix element is determined in advance based on historical operation data, and the priority ranking result of each optimization target in the current state is output in a clear numerical form:

[0092] ;

[0093] in, is the target priority decision matrix with a dimension of 3×3. Each element represents the quantitative value of the relative importance of the corresponding optimization goals. The rows and columns of the matrix correspond to the conversion efficiency value, resource utilization efficiency value, and environmental impact index value, respectively. To optimize the goal and optimization goals The initial coefficient of relative weight determined in advance based on historical process data, is the current optimization target obtained by real-time calculation of deep neural network and optimization goals The larger the value, the higher the conflict degree between the corresponding optimization objectives. is the conflict degree normalization constant, which is the sum of the conflict degree values ​​between all optimization objectives in the current process state;

[0094] According to the priority ranking results of the optimization objectives, an initial weight coefficient is assigned to each optimization objective respectively, and with the quantitative index of the degree of conflict as input, based on the dynamic weight allocation method, the initial weight coefficients of the conversion efficiency value, the resource utilization efficiency value and the environmental impact index value are adjusted in real time to obtain the real-time optimization weight coefficient that meets the requirements of the current process state. Each optimization weight coefficient is a non-negative number, and the sum of the weight coefficients is always 1;

[0095] Based on the determined real-time optimization weight coefficient and the current state information of the multidimensional decision space, the optimized control target value combination of temperature value, pH value, bacterial population ratio value and nutrient ratio value is calculated:

[0096] ;

[0097] in, Indicates the The optimized control target value combination calculated within a real-time control cycle is a vector composed of temperature values, pH values, bacterial population ratio values, and nutrient ratio values. Indicates the The current state value combination obtained from the multidimensional decision space in real time during a real-time control cycle is a state vector composed of the current temperature value, pH value, bacterial population ratio value and nutrient ratio value. Indicates the real-time optimization control step factor, which is used to control the amplitude of the control adjustment. The value range is between 0.01 and 0.2. Indicates the A diagonal matrix composed of the real-time optimization weight coefficients determined in a real-time control cycle, where the diagonal elements are the real-time optimization weight coefficients of the conversion efficiency value, the real-time optimization weight coefficients of the resource utilization efficiency value, and the real-time optimization weight coefficients of the environmental impact index value, which are used to determine the weight ratio of different optimization objectives in the real-time control strategy. Indicates the In a real-time control cycle, the first The state value combination corresponding to the feasible solution is Indicates the The first The comprehensive multi-objective optimization function value corresponding to the feasible solution, Indicates the The total number of feasible solutions obtained in the multidimensional decision space within a real-time control cycle is used to generate a real-time control strategy based on the obtained combination of optimized control target values.

[0098] In this embodiment, the real-time control strategy is based on dynamically adjusting the exploration mode and exploration direction of the global exploration ant colony and the local optimization ant colony in the multimodal ant colony collaborative optimization algorithm, and updating the pheromone distribution used to guide the ant colony search path in real time to obtain the candidate solution set for the multi-objective optimization of the biological feed conversion process, specifically:

[0099] Based on the optimized control target value combination of temperature, pH, bacterial population ratio and nutrient ratio determined in the real-time control strategy, the initial pheromone field strength distribution is generated based on the inverse of the Euclidean distance between each optimized control target value and each optimized decision variable value in the multidimensional decision space, where the initial pheromone field strength value range is limited to between 0.5 and 2.0;

[0100] According to the initial pheromone field strength distribution, the 5 to 10 areas with the highest pheromone field strength distribution values ​​are selected as the initial positions of the global exploration ant colony, and the initial number of ants in the global exploration ant colony is set to 50 to 100. The number of ants at each initial position is proportional to the corresponding pheromone field strength distribution value. The initial exploration direction of the global exploration ant colony is determined according to the gradient direction of the local pheromone field strength distribution value. The initial search step size is set to 0.05 to 0.1 times the value range of each optimization decision variable in the multidimensional decision space.

[0101] Based on the exploration paths completed by the global exploration ant colony with the initial exploration direction and initial search step size, the comprehensive multi-objective optimization function value corresponding to each path is calculated. Based on the optimization function value, the pheromone concentration adjustment direction and amplitude are determined by real-time calculation of the pheromone field intensity change rate of each path. The dynamic adjustment amplitude of the pheromone concentration is 5% to 15% of the current pheromone field intensity value. The pheromone concentration change gradient is obtained, and the search area boundary of the local optimization ant colony is delineated accordingly.

[0102] Based on the boundaries of the defined search area, a local optimization ant colony uses a grid division method to conduct an intensive search within the boundaries. The total number of ants in the local optimization ant colony is 100 to 150, and at least 3 ants are deployed in each grid cell. The search step size is set to 0.01 to 0.03 times the value range of each optimization decision variable in the multidimensional decision space. The pheromone distribution value and the comprehensive multi-objective optimization function value corresponding to each search path of the local optimization ant colony are recorded in real time.

[0103] According to the real-time recorded values ​​of the comprehensive multi-objective optimization function of each search path of the local optimization ant colony, the pheromone enhancement coefficient and pheromone volatility coefficient of the local optimization ant colony and the global exploration ant colony paths are calculated respectively. The pheromone enhancement coefficient ranges from 0.2 to 0.5, and the pheromone volatility coefficient ranges from 0.05 to 0.2. The larger the value of the comprehensive multi-objective optimization function, the larger the corresponding pheromone enhancement coefficient and the smaller the pheromone volatility coefficient. The pheromone field strength distribution in the multidimensional decision space is updated accordingly.

[0104] According to the updated pheromone field strength distribution, the exploration mode of the global exploration ant colony and the local optimization ant colony in the next cycle is re-determined based on the gradient change of the pheromone field strength distribution value. The exploration mode includes a pheromone-guided mode based on the probability transfer rule and a target tendency mode based on the gradient rule of the comprehensive optimization function. The transfer probability range of the pheromone-guided mode based on the probability transfer rule is set to 0.3 to 0.6, and the transfer probability range of the target tendency mode based on the gradient rule of the comprehensive optimization function is set to 0.4 to 0.7. In addition, the specific exploration direction and step size of the global exploration ant colony and the local optimization ant colony in the next cycle are determined in combination with the gradient change direction of the comprehensive multi-objective optimization function. The step size adjustment range is 80% to 120% of the previous search step size. At the same time, the pheromone field strength matrix is ​​used. The sharing mechanism clarifies the pheromone sharing method between the global exploration ant colony and the local optimization ant colony. The pheromone field strength matrix sharing mechanism is to establish a unified pheromone field strength matrix. The rows and columns of each element in the matrix respectively represent the value range of the optimization decision variable in the multidimensional decision space. The value of the matrix element represents the pheromone field strength value of the corresponding area. After each search cycle, the latest pheromone field strength values ​​of the global exploration ant colony and the local optimization ant colony are respectively updated to the unified pheromone field strength matrix; the shared pheromone field strength values ​​of the global exploration ant colony and the local optimization ant colony in the next search cycle are determined by normalization, and the normalized shared pheromone field strength values ​​are distributed to each ant colony in real time, so that the pheromone field strength distribution between ant colonies of different modes is kept synchronized in real time;

[0105] According to the updated pheromone field strength distribution, the threshold range of the comprehensive multi-objective optimization function is set, and the paths with comprehensive multi-objective optimization function values ​​higher than the set threshold are selected as candidate solutions, and the candidate solution set of the multi-objective optimization of the biological feed conversion process is regained.

[0106] In this embodiment, the effectiveness of each solution in the candidate solution set is evaluated by a reinforcement learning strategy based on the obtained multi-objective optimization candidate solution set, and the optimization solution that meets the current optimization goal priority and weight requirements is selected as the optimal control solution, specifically:

[0107] The real-time optimization weight coefficients of the optimization targets obtained according to the real-time control strategy, the real-time optimization weight coefficients respectively correspond to the conversion efficiency value, the resource utilization efficiency value and the environmental impact index value, and the real-time optimization weight coefficient value range is between 0 and 1, and the sum of the real-time optimization weight coefficients of all optimization targets is always equal to 1;

[0108] From the multi-objective optimization candidate solution set, the temperature value, pH value, bacterial population ratio value and nutrient ratio value corresponding to each candidate solution are obtained one by one, with the accuracy of the temperature value being 0.1°C, the accuracy of the pH value being 0.01, and the bacterial population ratio value and nutrient ratio value being expressed in mass percentage, with an accuracy of 0.1%;

[0109] Substitute the temperature, pH, bacterial population ratio, and nutrient ratio values ​​corresponding to each candidate solution into the real-time updated objective function mathematical expression to calculate the conversion efficiency, resource utilization efficiency, and environmental impact index values ​​corresponding to each candidate solution.

[0110] According to the reinforcement learning strategy, the conversion efficiency value, resource utilization efficiency value, and environmental impact index value corresponding to the candidate solution are used as the basic optimization action to be evaluated, and the state difference vector of each basic optimization action relative to the current state is calculated. The state difference vector norm corresponding to each candidate solution is calculated based on the state difference vector, and the state difference vector norm represents the state change amplitude corresponding to the candidate solution;

[0111] Taking the state difference vector norm and real-time optimization weight coefficient corresponding to each candidate solution as input, the comprehensive evaluation index value corresponding to each candidate solution is calculated. The comprehensive evaluation index value is obtained by calculating the action value function in the reinforcement learning strategy:

[0112] ;

[0113] in, is the comprehensive evaluation index value, For the The real-time optimization weight coefficient corresponding to the optimization target, The candidate solution corresponds to The optimization objective function value, is the norm of the state difference vector corresponding to the candidate solution;

[0114] According to the comprehensive evaluation index value corresponding to each candidate solution, all candidate solutions in the multi-objective optimization candidate solution set are quantitatively ranked using the descending order method to obtain the ranking result of the candidate solutions;

[0115] According to the ranking results of the candidate solutions, the temperature value, pH value, bacterial population ratio value and nutrient ratio value combination corresponding to the highest-ranked candidate solution are determined as the optimal control solution that meets the current optimization target priority and weight requirements.

[0116] In this embodiment, the reinforcement learning strategy is specifically:

[0117] Based on the temperature values, pH values, bacterial population ratio values, and nutrient ratio values ​​corresponding to each candidate solution in the multi-objective optimization candidate solution set, a state-action mapping relationship is constructed between the current state and the optimization action. The state-action mapping relationship is obtained by calculating the difference between each process parameter value in the candidate solution and the actual process parameter value collected in real time. Each state-action mapping relationship is represented in the form of a state difference vector;

[0118] Based on the state-action mapping relationship, the absolute value of the deviation of each process parameter in each state difference vector is calculated in real time, and the obtained absolute value of the deviation is multiplied by the real-time optimization weight coefficient of the corresponding process parameter to obtain the deviation impact degree of each optimization action corresponding to different optimization objectives;

[0119] Based on the degree of deviation of each optimization action corresponding to different optimization objectives, an online cluster analysis method is used to determine the quantitative value of the task conflict between each optimization action, and the quantitative value of the task conflict is specifically represented by the inter-cluster distance index value obtained by cluster analysis;

[0120] According to the quantitative value of task conflict degree, through the task priority dynamic adjustment matrix, with the goal of minimizing the degree of conflict between tasks, the task priority dynamic adjustment factor corresponding to each optimization action is calculated;

[0121] The dynamic adjustment factor is adjusted according to the obtained task priority. Based on the action value function update strategy in the reinforcement learning algorithm, a dynamic discounted cumulative reward function containing the difference vector norm is used to update the action value function value in real time:

[0122] ;

[0123] in, is the value of the action-value function after the current update, is the value of the action value function before the last update, is the action value function learning rate, ranging from 0 to 1, is the dynamic discount factor of the action value function, ranging from 0 to 1, is the value of the immediate reward function for the action, is the norm of the state difference vector corresponding to the current optimization action, is the next state after the current action is executed, All possible actions in the next state;

[0124] Based on the updated action-value function value, an action-value function search space is generated. The action-value function search space is constructed with the action-value function value as the horizontal axis and the corresponding action-state difference vector norm as the vertical axis. Based on the state difference vector norm distribution characteristics of each optimized action in the action-value function search space, the potential long-term cumulative reward function value corresponding to each optimized action is determined;

[0125] Based on the potential long-term cumulative reward function value, perform multiple random resampling calculations on the action-value function value corresponding to each optimized action to determine the upper and lower limits of the action-value confidence interval corresponding to each optimized action;

[0126] According to the action value function value of the optimization action in the action value function search space and the upper and lower limits of the corresponding action value confidence interval, with the goal of maximizing the coordinated action value function value and the action value confidence interval, the optimal optimization action in the current state is determined. The optimal optimization action corresponds to the optimization action whose action value function value is at the upper limit of the action value confidence interval.

[0127] In this embodiment, based on the optimal control scheme, the bio-feed conversion process parameters collected in real time by the industrial Internet of Things and the online sensor network are converted into specific process control instructions, and fed back to the bio-feed conversion process control equipment to achieve real-time control of the conversion process, specifically:

[0128] According to the target value combination of temperature, pH, bacterial population ratio and nutrient ratio determined by the optimal control scheme, a tolerance range of ±0.2°C is set for the temperature target value, a tolerance range of ±0.05 for the pH target value, a tolerance range of ±0.3% for the bacterial population ratio target value and a tolerance range of ±0.3% for the nutrient ratio target value are set respectively, thereby generating a multi-target control instruction with a clear tolerance range;

[0129] High-precision temperature sensors, pH sensors, and online component analyzers in the industrial Internet of Things are used to collect the actual temperature, pH, bacterial flora ratio, and nutrient ratio values ​​of the current biological feed conversion process in real time. The temperature sensor has an accuracy of 0.05°C and a sampling frequency of once every 2 seconds. The pH sensor has an accuracy of 0.01 and a sampling frequency of once every 5 seconds. The online component analyzer has a detection accuracy of 0.1% for bacterial flora ratio and nutrient ratio values ​​and a sampling frequency of once every 15 seconds. A continuous multi-point smoothing filter algorithm with a sliding window length of 10 is used on the real-time collected process parameter values ​​to obtain the actual values ​​of each process parameter after real-time smoothing.

[0130] The actual values ​​of each process parameter after real-time smoothing are calculated one by one with the corresponding process parameter target value to obtain the temperature value deviation, pH value deviation, bacterial population ratio value deviation and nutrient ratio value deviation. The real-time optimization weight coefficient of each process parameter is used as the weighting factor, and the comprehensive deviation value representing the comprehensive deviation degree is obtained through weighted summation.

[0131] For the comprehensive deviation, a joint evaluation is performed based on the real-time rate of change of the deviation and the integral value of the accumulated deviation in the continuous process. The real-time rate of change of the deviation is calculated by dividing the difference between the deviation values ​​of two adjacent sampling moments by the time difference between the adjacent sampling moments. The cumulative deviation integral value is obtained by integrating the deviation in the time dimension with the current moment as the upper limit of the integral and the initial control moment as the lower limit of the integral, thereby forming a dynamic adjustment value for each process parameter.

[0132] Based on the dynamic adjustment amount obtained, the closed-loop PID adjustment algorithm is used to calculate the specific control instructions of each process parameter in real time:

[0133] ;

[0134] in, For the The process parameters at time The real-time control value of For the The process parameters at time The comprehensive deviation of is the proportional adjustment coefficient, ranging from 0.1 to 1.0, is the integral adjustment coefficient, ranging from 0.01 to 0.5. is the differential adjustment coefficient, ranging from 0.001 to 0.1;

[0135] The specific control instructions for each process parameter obtained by calculation are encoded in real time into a control message that meets the industrial Internet of Things transmission protocol. The control message clearly includes the instruction type, control amplitude, execution priority, and execution response time limit parameters. The execution priority parameter ranges from 1 to 10, and the execution response time limit parameter is in seconds and ranges from 1 to 5 seconds. The encoded control message is transmitted in real time to the execution terminal of the biofeed conversion process control equipment through the industrial Internet of Things;

[0136] Each control device execution terminal receives the control message and parses the specific process control instructions in the message, executes the specific control action in real time according to the instruction type, control amplitude, execution priority and execution response time limit parameters, and provides real-time feedback on the actual process parameter change information after execution, forming accurate, continuous and fast-response real-time closed-loop dynamic control.

[0137] In this embodiment, the biological feed conversion process parameters after real-time control are fed back to the multi-objective optimization mathematical model, the multi-dimensional decision space is dynamically updated, and the above steps are repeatedly performed, specifically:

[0138] The temperature sensor, pH sensor and online component analyzer in the industrial Internet of Things are used to collect the temperature, pH value, bacterial flora ratio and nutrient ratio of the biological feed conversion process after real-time regulation. The temperature sensor sampling frequency is once every 2 seconds with a measurement accuracy of 0.05°C, the pH sensor sampling frequency is once every 5 seconds with a measurement accuracy of 0.01, and the online component analyzer sampling frequency is once every 15 seconds with a measurement accuracy of 0.1%. The difference between the collected process parameter values ​​and the corresponding parameter values ​​before the previous real-time regulation is calculated one by one to obtain the actual regulation response increment of each parameter;

[0139] Using the actual control response increment as the new training sample data, the nonlinear function mapping relationship between each process parameter and the conversion efficiency value, resource utilization efficiency value and environmental impact index value in the multi-objective optimization mathematical model is updated;

[0140] The dynamically updated nonlinear function mapping relationship is used to reconstruct the mathematical expressions of each objective function in the multi-objective optimization mathematical model, and the reconstructed mathematical expressions are explicitly expressed to obtain accurate objective function expressions that can be calculated in real time;

[0141] Based on the reconstructed and explicit mathematical expression of the objective function, an adaptive spatial mapping method based on space segmentation and density estimation is used to re-divide the multidimensional decision space.

[0142] Based on the re-divided feasible solution domain of the optimization variables, the number of candidate solutions and the standard deviation of the corresponding optimization objective function for each spatial unit in the solution domain are calculated in real time. The number of candidate solutions and the standard deviation are used to jointly define a quantitative index of the degree of conflict between tasks. The quantitative index is used to dynamically adjust the dynamic adjustment factor of the task priority in the adaptive multi-task and multi-objective reinforcement learning algorithm. The value range of the dynamic adjustment factor of the task priority is set to 0.1 to 1.0, and the dynamic update step size is 0.05.

[0143] The new real-time control strategy is obtained based on the dynamically updated task priority dynamic adjustment factor. The generation of the real-time control strategy is obtained by real-time evaluation of the action value function of the reinforcement learning strategy network.

[0144] Based on the new real-time control strategy, the multimodal ant colony collaborative optimization algorithm was re-executed, and the pheromone field strength of the global exploration ant colony and the local optimization ant colony was adaptively updated with gradients using the feedback parameters collected in real time. The pheromone field strength gradient update adopted a negative feedback proportional adjustment mechanism with a proportional factor ranging from 0.1 to 0.9 to obtain an updated set of candidate solutions for the multi-objective optimization.

[0145] The updated multi-objective optimization candidate solution set is used as the optimization solution input for the next real-time control. The convergence of the optimization cycle is evaluated based on the feedback difference between the candidate solution and the real-time collected process parameters. The convergence evaluation is calculated using the Euclidean distance between the current candidate solution and the previous candidate solution. The Euclidean distance convergence threshold is set to 0.001. The convergence evaluation is used to determine in real time whether the feedback closed-loop iterative process should continue.

[0146] Example 1

[0147] To verify the feasibility of the present invention, the method was applied to a real-world production scenario at a large-scale biofeed production enterprise, conducting multi-objective process parameter control for a typical batch of biofeed conversion. In this application scenario, conventional control techniques typically employ fixed parameter threshold control, independently optimizing each process parameter. This fails to comprehensively consider the interplay between temperature, pH, bacterial composition, and nutrient profiles. This results in insufficient real-time process parameter adjustment accuracy, low resource utilization efficiency, suboptimal environmental pollution control, and significant process parameter fluctuations, making sustained and stable production difficult.

[0148] During implementation, the industrial Internet of Things (IIoT) platform was first used to collect real-time data on temperature, pH, bacterial flora ratio, and nutrient ratio during the conversion process. The sampling period was 5 seconds for temperature, 10 seconds for pH, and 30 seconds for bacterial flora and nutrient ratio. This data was then transmitted to a central control server via a wireless communication network. A multi-objective optimization mathematical model was established on the server, with conversion efficiency, resource utilization efficiency, and environmental impact indicators as optimization targets. Nonlinear curve fitting analysis was performed on the optimization decision variables to determine the functional mapping relationship between the optimization decision variables and the optimization objectives, thereby constructing a multidimensional decision space.

[0149] Subsequently, the deep neural network model in the server receives real-time data on the current state of the multidimensional decision space, dynamically determines the priority and real-time weight of each optimization objective, and outputs a real-time control strategy. The system uses a multimodal ant colony collaborative optimization algorithm to obtain a set of candidate optimization solutions. It then uses a reinforcement learning strategy to evaluate the comprehensive evaluation indicators of each optimization action in the candidate solution set to determine the optimal control solution.

[0150] The system then converts the optimal control plan into specific process control instructions, transmits them in real time to the production equipment execution terminal via the Industrial Internet of Things, and dynamically updates the process parameters in real time, achieving precise closed-loop control. During the implementation of this invention, the ant colony algorithm parameters were set as follows: a global exploration ant colony of 50, a local optimization ant colony of 100, 200 iterations, a pheromone volatility coefficient of 0.5, 150 action value search iterations in the reinforcement learning strategy, and a cumulative reward function discount factor of 0.95.

[0151] After three consecutive months of actual production operation, five typical samples were selected for statistical analysis, and a comparative analysis was conducted between the predicted and measured values ​​of the process parameters before and after real-time control. The results are shown in Table 1.

[0152] Table 1 Comparison of predicted and measured parameters of biological feed conversion process

[0153]

[0154] As can be clearly seen from the comparison of predicted and measured biofeed conversion process parameters shown in Table 1, the present method demonstrates high levels of prediction accuracy for three key indicators: conversion efficiency, resource utilization, and environmental impact. Specifically, the error between the predicted and measured conversion efficiency values ​​was kept within ±1.0%. For sample SFT-002, for example, the actual conversion efficiency was 89.5%, while the predicted value reached 90.1%, with a deviation of only 0.6%. Regarding resource utilization, the error between the predicted and measured values ​​was also kept within ±0.6%. For sample SFT-005, the measured resource utilization rate was 76.3%, while the predicted value was 76.8%, with an error of 0.5%. The prediction accuracy for environmental impact indicators was also outstanding, with the error between the predicted and measured values ​​generally within 0.01. For sample SFT-003, the measured environmental index was 0.61, while the predicted value was 0.62, with an error of only 0.01, demonstrating the high accuracy of the present method in controlling environmental indicators.

[0155] Temperature control deviations remained stable within ±0.9°C for all samples, with sample SFT-004 exhibiting a control deviation of only 0.6°C, demonstrating the clear advantage of this method in precise temperature control. pH control was also strictly controlled within ±0.05, with the deviations for samples SFT-002 and SFT-004 being only 0.03. Microbial composition control deviations were kept within 1.6%, with the SFT-004 sample exhibiting an even lower deviation of 1.2%, significantly outperforming traditional empirical control methods.

[0156] Comprehensive analysis of the above embodiments shows that the method of the present invention effectively solves the problem that traditional methods cannot accurately control multi-parameter coordinated regulation through the deep integration of multi-objective dynamic closed-loop regulation and intelligent optimization algorithm, realizes precise control and efficient optimization of the process, significantly improves conversion efficiency and resource utilization efficiency, and effectively reduces environmental pollution indicators, and has obvious industrial application and promotion value.

[0157] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A bio-feed conversion process control method based on multi-objective optimization, characterized in that: include: Based on the real-time data collected during the biofeed conversion process, a multi-objective optimization mathematical model is established to obtain a multi-dimensional decision space; Through deep neural networks, the system can perceive the environmental status information of the multidimensional decision space in the current bio-feed conversion process in real time, evaluate the conflict degree between the current optimization objectives and the resource constraints, and generate real-time control strategies; Based on the real-time control strategy, the exploration mode and exploration direction of the global exploration ant colony and the local optimization ant colony in the multimodal ant colony collaborative optimization algorithm are dynamically adjusted to obtain the candidate solution set; Evaluate the effectiveness of each solution in the candidate solution set through reinforcement learning strategy and select the optimal control solution; Utilize the real-time collection of bio-feed conversion process parameters from the Industrial Internet of Things and online sensor networks to convert the optimal control scheme into specific process control instructions and feed them back to the bio-feed conversion process control equipment; Feedback the biofeed conversion process parameters after real-time control to the multi-objective optimization mathematical model to dynamically update the multi-dimensional decision space; The deep neural network is used to perceive the environmental status information of the multidimensional decision space in the current bio-feed conversion process in real time, and to evaluate the conflict degree and resource constraints between the current optimization objectives to generate a real-time control strategy, specifically: Real-time acquisition of temperature values, pH values, bacterial flora ratio values, and nutrient ratio values ​​in the multidimensional decision space, and standardization of these values ​​to obtain normalized state data. The normalized state data is input into the trained deep neural network, and the multi-layer feedforward connection structure inside the deep neural network is used for layer-by-layer nonlinear mapping to obtain the specific numerical values ​​of the degree of conflict between the optimization objectives; According to the specific values ​​of the conflict degree between the optimization objectives, the linear weighted summation method is used to calculate the quantitative index of the comprehensive conflict degree; According to the comprehensive conflict degree quantitative index, the priority ranking of each optimization goal is determined through the goal priority decision matrix; Based on the priority sorting results, a dynamic weight allocation method is used to determine the real-time optimization weight coefficients of the conversion efficiency value, resource utilization efficiency value and environmental impact index value; Based on the determined real-time optimization weight coefficient and the current state information of the multidimensional decision space, the optimized control target value combination of temperature value, pH value, bacterial population ratio value and nutrient ratio value is calculated to generate a real-time control strategy; The real-time control strategy is based on dynamically adjusting the exploration mode and exploration direction of the global exploration ant colony and the local optimization ant colony in the multimodal ant colony collaborative optimization algorithm, and updating the pheromone distribution used to guide the ant colony search path in real time to obtain the candidate solution set for the multi-objective optimization of the biological feed conversion process. Specifically, Generate the initial pheromone field intensity distribution according to the real-time control strategy; According to the initial pheromone field intensity distribution, the initial position and number of the global exploration ant colony are determined, and the initial exploration direction and step length of the global exploration ant colony are determined according to the pheromone field intensity gradient direction; According to the exploration results of the global exploration ant colony, the pheromone concentration is dynamically adjusted, and the pheromone concentration gradient is used to determine the search area boundary of the local optimization ant colony; According to the search area boundary, the local optimization ant colony conducts an intensive search inside the boundary and records the pheromone distribution corresponding to the search path and the objective function value of the optimization solution in real time; According to the objective function value corresponding to the recorded search path, the pheromone enhancement coefficient and volatility coefficient of the global exploration ant colony and the local optimization ant colony are determined, and the pheromone field strength in the multidimensional decision space is dynamically updated according to the high and low objective function values ​​corresponding to the optimization solution; According to the updated pheromone field strength and corresponding gradient changes, the exploration mode, exploration direction and step size of the global exploration ant colony and the local optimization ant colony in the next cycle are readjusted, and the pheromone sharing method between ant colonies of different modes is determined; According to the updated pheromone field intensity distribution, the candidate solution set of multi-objective optimization of the biological feed conversion process is re-obtained; The reinforcement learning strategy is specifically: According to the temperature value, pH value, bacterial population ratio value and nutrient ratio value corresponding to each candidate solution in the multi-objective optimization candidate solution set, a state-action mapping relationship between the current state and the optimization action is constructed; According to the state-action mapping relationship, the dynamic adjustment factor of the task priority corresponding to each optimization action is calculated based on the real-time optimization weight coefficient of each optimization target corresponding to the current state; According to the task priority dynamic adjustment factor, the action value function value of each optimization action is updated with the goal of minimizing the degree of conflict between tasks; Based on the updated action-value function value, an action-value function search space is generated, and the potential long-term cumulative reward function value of each optimized action in the action-value function search space is calculated; Calculating the action value confidence interval corresponding to each optimization action based on the value of the potential long-term cumulative reward function, and obtaining the upper and lower confidence bounds of each optimization action in the action value function search space; According to the action value function value of the optimization action in the action value function search space and the corresponding action value confidence interval, the optimal optimization action under the current state is determined with the goal of maximizing the action value and the confidence interval. According to the temperature value, pH value, bacterial population ratio value and nutrient ratio value corresponding to the determined optimal optimization action, the corresponding comprehensive evaluation index value is calculated.

2. The bio-feed conversion process control method based on multi-objective optimization according to claim 1, characterized in that: According to the real-time data collected during the bio-feed conversion process, a multi-objective optimization mathematical model is established to obtain a multi-dimensional decision space, specifically: Real-time collection of temperature values, pH values, bacterial flora ratio values, and nutrient ratio values ​​during the biological feed conversion process to form optimization decision variables; Taking the conversion efficiency value, resource utilization efficiency value and environmental impact index value as the optimization target, the optimization target is expressed as an objective function composed of optimization decision variables; Determine the value range of optimization decision variables and form a multi-dimensional optimization decision variable constraint space; The nonlinear fitting method is used to determine the mathematical relationship between the optimization decision variables and the optimization objectives, and the mathematical expression of the objective function is obtained; Construct a multi-objective optimization mathematical model based on the mathematical expression of the objective function and the constraint space of the optimization decision variables; Calculate the objective function value using the multi-objective optimization mathematical model solution algorithm and obtain the feasible solution set of the objective function; According to the obtained feasible solution set of the objective function, a multidimensional decision space of the biofeed conversion process is generated.

3. The bio-feed conversion process control method based on multi-objective optimization according to claim 1, characterized in that: According to the obtained multi-objective optimization candidate solution set, the effectiveness of each solution in the candidate solution set is evaluated by reinforcement learning strategy, and the optimization solution that meets the current optimization goal priority and weight requirements is selected as the optimal control solution, specifically: According to the optimization target obtained by the real-time control strategy, the weight coefficient is optimized in real time, and the temperature value, pH value, bacterial population ratio value and nutrient ratio value corresponding to each candidate solution are obtained from the multi-objective optimization candidate solution set; The temperature, pH, bacterial composition, and nutrient composition values ​​corresponding to each candidate solution are used as the optimization action to be evaluated. The comprehensive evaluation index value corresponding to each optimization action is calculated based on the reinforcement learning strategy. According to the comprehensive evaluation index value, all candidate solutions in the multi-objective optimization candidate solution set are quantitatively ranked to obtain the ranking results of the candidate solutions; According to the ranking results of the candidate solutions, the temperature value, pH value, bacterial population ratio value and nutrient ratio value combination corresponding to the highest-ranked candidate solution are determined as the optimal control solution that meets the current optimization target priority and weight requirements.

4. The bio-feed conversion process control method based on multi-objective optimization according to claim 1, characterized in that: Based on the optimal control scheme, the bio-feed conversion process parameters collected in real time by the industrial Internet of Things and online sensor networks are converted into specific process control instructions, and fed back to the bio-feed conversion process control equipment to achieve real-time control of the conversion process. Specifically, Generate multi-objective control instructions with a tolerance range based on the target value combination of temperature, pH, bacterial population ratio and nutrient ratio determined in the optimal control plan; The industrial Internet of Things is used to collect the actual temperature, pH, bacterial flora ratio and nutrient ratio of the current biological feed conversion process in real time, and each collected value is subjected to continuous multi-point smoothing filtering to obtain the actual values ​​of each process parameter after real-time smoothing; The difference between the actual value of each process parameter after real-time smoothing and the corresponding process parameter target value is calculated to obtain the comprehensive deviation of each process parameter; For the comprehensive deviation, a joint assessment is conducted based on the deviation change rate and historical cumulative deviation amplitude to form a dynamic adjustment value for each process parameter in the bio-feed conversion process; Specific process control instructions are generated based on the dynamic adjustment quantity and transmitted in real time to the execution terminal of the biological feed conversion process control equipment through the industrial Internet of Things.

5. The biological feed conversion process control method based on multi-objective optimization according to claim 1, characterized in that: The biological feed conversion process parameters after real-time control are fed back to the multi-objective optimization mathematical model, the multi-dimensional decision space is dynamically updated, and the above steps are repeatedly performed, specifically: The Industrial Internet of Things is used to collect the temperature, pH, bacterial composition, and nutrient composition of the biological feed conversion process in real time after real-time regulation, and to calculate the difference between the corresponding parameter values ​​before the previous real-time regulation to obtain the actual regulation response increment of each parameter; Using the actual control response increment as the new sample data, the nonlinear fitting method in the multi-objective optimization mathematical model is used to dynamically update the functional mapping relationship between each process parameter and the conversion efficiency value, resource utilization efficiency value and environmental impact index value; Reconstruct the objective function mathematical expression in the multi-objective optimization mathematical model using the dynamically updated function mapping relationship; Based on the reconstructed mathematical expression of the objective function, the multidimensional decision space is updated and adjusted, and the feasible solution domain of the optimization variables is re-divided; Based on the re-divided feasible solution domain of the optimization variables, the task priority dynamic adjustment factor in the adaptive multi-task and multi-objective reinforcement learning algorithm is dynamically adjusted to obtain a new real-time control strategy; Based on the new real-time control strategy, the multimodal ant colony collaborative optimization algorithm is re-executed to obtain a new set of multi-objective optimization candidate solutions; The new set of multi-objective optimization candidate solutions is used as the basis for the next real-time regulation, and the feedback closed-loop process is continuously iterated.

Citation Information

Patent Citations

  • High-precision underwater low-light target edge detection method

    CN117372462A

  • New material production process parameter optimization method and system based on artificial intelligence

    CN119397914A