Self-adaptive cooperative control method for raw material mixing and stirring process
By employing an adaptive collaborative control method and utilizing a reinforcement learning framework and agent training, the problem of poor mixing effect caused by batch differences and nonlinear variations in raw materials during the production of Sachima was solved. This resulted in efficient and stable mixing control, improving product quality and production efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-13
- Publication Date
- 2026-04-14
AI Technical Summary
Existing mixing equipment is difficult to adapt to batch differences and nonlinear changes in raw materials in the production of Sachima, resulting in poor mixing effect, affecting the taste and uniformity of the product. Existing control methods rely on precise mathematical models and have weak adaptability.
An adaptive cooperative control method is adopted, which enables the agent to learn the stirring control strategy autonomously through a reinforcement learning framework. By combining simulation and real training environments, multi-objective optimization is achieved, eliminating the dependence on precise mathematical models. Deep reinforcement learning is used to establish a mapping from state space to actions to optimize the stirring process.
It achieves adaptive control of the mixing process of Sachima raw materials, improves product qualification rate, saves energy and reduces consumption, adapts to the characteristics of different batches of raw materials, and improves the generalization ability and stability of the control process.
Smart Images

Figure CN121857342A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automated control technology for food processing, and in particular to an adaptive collaborative control method for the mixing and stirring process of raw materials. Background Technology
[0002] In the traditional production process of Sachima, the mixing and stirring of raw materials is a key step that determines the final product's texture, crispness, and uniformity. Currently, this step often uses mixing equipment with fixed speed and time control. However, batch-to-batch differences in raw materials, slight fluctuations in the amount of raw materials added, and non-linear changes in the state of raw materials during the mixing process make it difficult for fixed-parameter mixing control to always guarantee the optimal mixing effect. Over-mixing may lead to excessive formation of gluten networks, resulting in a harder product texture, while under-mixing will lead to uneven mixing of raw materials, affecting the product structure.
[0003] Existing improvement methods employ PID control or model-based adaptive control, but PID parameter tuning is difficult and its adaptability is weak. Model reference adaptive control heavily relies on the accuracy of the mathematical model of the controlled object. For complex processes such as stirring raw materials for Sachima, which have strong nonlinearity, time-varying characteristics, and involve changes in multiple physical parameters, it is extremely difficult to establish an accurate mechanism model. Model mismatch will lead to a decrease in control performance and instability. Summary of the Invention
[0004] This application provides an adaptive collaborative control method for the raw material mixing and stirring process, which overcomes the modeling difficulties and adaptive optimization limitations of existing model-based control methods in the mixing and stirring process of Sachima raw materials. This method does not rely on a precise mathematical model of the process. Through a hybrid training environment that combines simulation and reality, based on a reinforcement learning framework, the agent can autonomously learn the optimal stirring control strategy to adapt to changes in Sachima raw materials and environmental disturbances, thereby achieving multi-objective collaborative optimization.
[0005] This application provides an adaptive cooperative control method for a raw material mixing and stirring process, including: Step S10: The adaptive collaborative control of the mixing and stirring process of Sachima raw materials is transformed into a Markov decision process. Based on the stirring process data and the corresponding mixer control strategy, the state space, action space, transition probability and reward function of the decision process are defined. Step S20: Establish an agent for the Markov decision process of Sachima raw material mixing and stirring control, set up a value network, decision network, target network and experience playback buffer, and learn the control strategy of Sachima raw material mixing and stirring. Step S30: Establish a hybrid training environment to train the agent, simulate the mixing process of Sachima raw materials, calculate the target value according to the reward function to update the network parameters, optimize and adjust the control strategy of the mixer, and make corrections on a real Sachima mixer. Step S40: Based on the training results of the agent, generate the optimal adaptive control strategy for the mixing and stirring process of Sachima raw materials, and dynamically adjust the working parameters of the Sachima mixer in the actual production process.
[0006] Furthermore, the control objective of the mixing process of Sachima raw materials is set to mix the Sachima raw materials to the target uniformity with the lowest energy consumption. The constraints on the control process based on this objective include: limiting the motor current and torque of the Sachima mixer to no more than the safety limit, keeping the speed within the mechanical design range, and avoiding violent acceleration or deceleration to protect the equipment.
[0007] Within a set time window, the operating status information and sensor data of the mixing and stirring process of Sachima raw materials are collected. Based on the torque fluctuation characteristics, acoustic spectrum characteristics, and vibration dominant frequency amplitude of the mixer, the uniformity of Sachima raw materials is comprehensively quantified from multiple perspectives, including the following steps: The torque value of the main motor of the mixer is collected in real time during the mixing and stirring of Sachima raw materials, and the set time window is calculated. The average value of torque within. and standard deviation This can be expressed as a formula:
[0008]
[0009] in, This indicates the length of the time window, i.e., the number of sampling points during the data acquisition process; For time indexing, Indicates a point in time The measured torque value of the main motor of the mixer reflects the resistance of the Sachima raw materials; the ratio of the standard deviation to the average value of the torque is taken as the coefficient of variation of the torque of the main motor of the mixer. This reflects the relative fluctuation of torque. The smaller the value, the more uniform the raw materials. Sound pressure signals from the mixer during operation are collected using a sound sensor. A fast Fourier transform is performed on the sound pressure signals to obtain the power spectral density. The power at each frequency point is divided by the total power to obtain the normalized power spectral density. The spectral entropy of the sound pressure signal is then calculated, expressed by the formula:
[0010] in, The spectral entropy of a sound pressure signal is used to measure the complexity and irregularity of the signal spectrum; a higher value reflects a more uneven raw material. For frequency index, and These represent the lowest and highest frequency points of the sound pressure signal collected within the set time window, respectively. This represents the normalized power spectral density; The vibration acceleration signal of the mixer is collected, and a fast Fourier transform is performed to obtain the spectrum. The maximum value of the spectrum amplitude within a preset frequency range is obtained, expressed by the formula:
[0011] in, This indicates the amplitude of the dominant vibration frequency of the mixer. This refers to the vibration frequency range of the mixer. The spectrum representing vibration acceleration includes amplitude and phase. This indicates the absolute value operation; By combining the torque variation coefficient, sound pressure signal spectral entropy, and vibration dominant frequency amplitude of the Sachima mixer, a quantified mixing uniformity index is obtained:
[0012] in, Indicates in The estimated uniformity index of Sachima raw materials at any time. It is the coefficient of variation of the torque of the main motor of the mixer, which represents the ratio of the standard deviation of the torque to the average value; and These represent the normalized sound pressure signal spectral entropy and the normalized vibration dominant frequency amplitude, respectively. After normalization processing, , and They have the same physical dimensions; , and The weighting coefficients represent the relative importance of each feature to the evenness estimation. They are determined through regression analysis of historical data and satisfy the following conditions: The uniformity value ranges from [0,1], where 0 and 1 represent completely non-uniform and completely uniform, respectively.
[0013] Furthermore, the relevant parameter space of the Markov decision process for mixing and stirring the raw materials of Sachima is defined as follows: The state space of the intelligent agent is based on the process data of mixing and stirring of Sachima raw materials collected at each decision moment, which includes real-time sensor readings, process characteristics, process stage information and historical information. Using the control commands executed by the Sachima mixer as the workspace of the intelligent agent, the system outputs the incremental values of the set speed of the main and slave motors of the mixer, as well as the extended action commands for switching the mixing process stages, thereby realizing the process conversion from dry mixing, wet mixing to maturation. The state transition probability represents the probability that the training environment of the agent will transition to the next state based on the current state and action. The reward function for multi-objective optimization is defined as a weighted sum of quality reward, efficiency reward and energy consumption reward, which guides the training process of the agent.
[0014] Furthermore, based on Markov decision process parameters, an agent is established to achieve adaptive control of the mixing and stirring process of Sachima raw materials. The agent components consist of a policy network, a value network, a target network, and an experience replay buffer. The deep deterministic policy gradient method is used to complete the training process of the agent. The input and output of the policy network are state vector and action vector, respectively, representing the state of the Sachima raw materials and the corresponding action executed by controlling the mixer. The value network takes state and action as input and outputs a Q-value, which evaluates the long-term expected accumulated reward of the current action and guides the parameter update of the policy network. The target network is used to create replicas of the policy network and value network, and to calculate the stable Q-value target. The experience replay buffer stores the interaction data of the agent during the mixing and stirring process of Sachima raw materials, and is used for random sampling to break the temporal correlation of the data.
[0015] Furthermore, a hybrid training environment for the intelligent agent is established. Based on the physical principles of the Sachima mixer, a simulated digital twin environment is constructed to simulate the dynamics of the mixing motor and the state evolution of the Sachima raw materials. The intelligent agent is pre-trained through data interaction. In addition, the actual Sachima mixer and its sensors and actuators in production are used as the real physical training environment for strategy verification and online adjustment. In the simulation environment, the agent observes the initial state of the mixing process of Sachima raw materials, selects the control action of the mixer according to the current strategy, calculates the reward function to obtain a new state vector, samples batch data from the buffer to update the parameters of the policy network and the value network, and repeats the above process with the goal of maximizing the cumulative reward to train and obtain the first control strategy. The offline-trained first control strategy is loaded into the controller of the real equipment, and the execution command is sent and executed by the PID controller. Bounded random noise is added to the output control action, real interaction data is collected, the learning rate is reduced to update the network parameters, the first control strategy is finely adjusted, the error between simulation and reality is made up, the characteristics of the equipment and batch changes of raw materials are adapted, and the final accurate control strategy is generated to realize the online adaptive control of the Sachima mixer.
[0016] This application discloses the following technical effects: This application provides an adaptive cooperative control method for the raw material mixing and stirring process. Based on deep reinforcement learning, the method learns the optimal adaptive control strategy through direct interaction between the agent and the environment, thus eliminating the dependence on precise mathematical models and adapting to the strong nonlinearity, time-varying nature, and multi-physics coupling characteristics of the Sachima raw material mixing and stirring process, thereby achieving direct optimization control of the complex process. The method proposed in this application achieves multi-objective dynamic balance and global optimization of quality, efficiency, and energy consumption during agent training. By maximizing the quality reward function related to the uniformity of Sachima raw materials, the agent learns an adaptive control strategy to achieve the target uniformity, avoiding over-stirring that could damage the gluten network and steadily improving the product qualification rate. An energy consumption reward function is established to guide the agent to automatically adjust the speed to maintain the minimum energy consumption required for mixing, thus achieving energy saving. An efficiency reward function drives the agent to learn a faster mixing path, adaptively compressing ineffective mixing time and dynamically adjusting the mixing rhythm according to the characteristics of different batches of raw materials, thereby improving efficiency. This method utilizes the learning ability of deep reinforcement learning to establish a high-dimensional state space to action mapping, enhancing its strong adaptability to uncertainties such as batch differences in Sachima raw materials, environmental interference, and slight equipment wear, and improving the generalization ability of the control process to various working conditions. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0018] Figure 1 This is a flowchart illustrating an adaptive collaborative control method for a raw material mixing and stirring process, provided as an embodiment of this application. Detailed Implementation
[0019] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application.
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] In the following description, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only.
[0022] This application provides an adaptive cooperative control method for a raw material mixing and stirring process, such as... Figure 1 As shown, the method includes: Step S10: The adaptive collaborative control of the mixing process of Sachima raw materials is transformed into a Markov decision process. Based on the mixing process data and the corresponding mixer control strategy, the state space, action space, transition probability and reward function of the decision process are defined.
[0023] In this embodiment, the control objective of the mixing and stirring process of Sachima raw materials is set to mix the Sachima raw materials to the target uniformity with the lowest energy consumption. The constraints on the control process based on this objective include: limiting the motor current and torque of the Sachima mixer to not exceed the safety limit, keeping the speed within the mechanical design range, and avoiding violent acceleration or deceleration to protect the equipment.
[0024] Within a set time window, the operating status information and sensor data of the mixing and stirring process of Sachima raw materials are collected. Based on the torque fluctuation characteristics, acoustic spectrum characteristics, and vibration dominant frequency amplitude of the mixer, the uniformity of Sachima raw materials is comprehensively quantified from multiple perspectives, including the following steps: The torque value of the main motor of the mixer is collected in real time during the mixing and stirring of Sachima raw materials, and the set time window is calculated. The average value of torque within. and standard deviation This can be expressed as a formula:
[0025]
[0026] in, This indicates the length of the time window, i.e., the number of sampling points during the data acquisition process; For time indexing, Indicates a point in time The measured torque value of the main motor of the mixer reflects the resistance of the Sachima raw materials; the ratio of the standard deviation to the average value of the torque is taken as the coefficient of variation of the torque of the main motor of the mixer. This reflects the relative fluctuation of torque. The smaller the value, the more uniform the raw materials. Sound pressure signals from the mixer during operation are collected using a sound sensor. A fast Fourier transform is performed on the sound pressure signals to obtain the power spectral density. The power at each frequency point is divided by the total power to obtain the normalized power spectral density. The spectral entropy of the sound pressure signal is then calculated, expressed by the formula:
[0027] in, The spectral entropy of a sound pressure signal is used to measure the complexity and irregularity of the signal spectrum; a higher value reflects a more uneven raw material. For frequency index, and These represent the lowest and highest frequency points of the sound pressure signal collected within the set time window, respectively. This represents the normalized power spectral density; The vibration acceleration signal of the mixer is collected, and a fast Fourier transform is performed to obtain the spectrum. The maximum value of the spectrum amplitude within a preset frequency range is obtained, expressed by the formula:
[0028] in, This indicates the amplitude of the dominant vibration frequency of the mixer. This refers to the vibration frequency range of the mixer. The spectrum representing vibration acceleration includes amplitude and phase. This indicates the absolute value operation; By combining the torque variation coefficient, sound pressure signal spectral entropy, and vibration dominant frequency amplitude of the Sachima mixer, a quantified mixing uniformity index is obtained:
[0029] in, Indicates in The estimated uniformity index of Sachima raw materials at any time. and These represent the normalized sound pressure signal spectral entropy and the normalized vibration dominant frequency amplitude, respectively. After normalization processing, , and They have the same physical dimensions; , and The weighting coefficients represent the relative importance of each feature to the evenness estimation. They are determined through regression analysis of historical data and satisfy the following conditions: The uniformity value ranges from [0,1], where 0 and 1 represent completely non-uniform and completely uniform, respectively. The relevant parameter space of the Markov decision process for mixing and stirring the raw materials of Sachima is defined as follows: The state space of the agent is defined by the mixing and stirring process data of the Sachima raw materials collected at each decision moment, which includes real-time sensor readings, process characteristics, process stage information, and historical information. Real-time sensor data includes the main stirring motor speed, torque, and current of the Sachima mixer, the auxiliary scraper motor speed, and the temperature of key points inside the vessel; process characteristics include the torque fluctuation coefficient, vibration signal main frequency amplitude, and sound pressure signal spectral entropy calculated over the past 10 seconds; process stage information includes the current stirring time of the mixer, batch formula identification, and ambient humidity and temperature; historical information records the state vectors and action vectors of the past few time steps to help the intelligent agent perceive the dynamic trend of the mixing process; The control commands executed by the Sachima mixer serve as the workspace of the intelligent agent. The output includes the increment of the set value of the main and slave motor speeds of the mixer, as well as the extended action commands for switching the mixing process stages, to realize the process conversion from dry mixing, wet mixing to maturation. The state transition probability represents the probability that the training environment of the intelligent agent will transition to the next state based on the current state and action. Define a reward function for multi-objective optimization, expressed as a weighted sum of quality reward, efficiency reward, and energy consumption reward, to guide the training process of the agent, expressed by the formula:
[0030] in, express Total reward value at each moment , and Let represent the function values of quality reward, efficiency reward, and energy consumption reward, respectively. , and This represents the weighting coefficient, which is set empirically. The quality reward function focuses on improving the uniformity of raw materials and avoiding overmixing, and is expressed as a uniformity improvement reward. and excessive mixing of punishment The weighted sum, expressed by the formula:
[0031] in, and They represent and The uniformity of the raw materials in Sachima at any given time Indicates the threshold for excessive mixing of raw materials; and These represent the sensitivity parameter and the penalty steepness parameter, respectively, both set to 0.05; and These are weighting coefficients, set empirically. The efficiency reward function aims to quickly complete the mixing of raw materials while ensuring quality. The formula is as follows:
[0032] in, This indicates the mixing time currently used. This represents the historical average optimal time. and These represent the maximum allowable mixing time and the starting point of the time penalty, respectively. This indicates the target uniformity of the raw materials for Sachima. , and The coefficient is set empirically. The energy reward function aims to reduce energy loss during the stirring process, and its formula is as follows:
[0033]
[0034] in, and Indicates instantaneous power and the system's rated total power. and These represent the current cumulative energy consumption and energy budget, respectively. and The weighting parameters are set empirically. and This indicates the torque and speed of the main motor. and This indicates the torque and speed of the motor; The multi-objective optimization reward function guides the agent to learn an optimized control strategy that ensures both the quality of the mixture and the efficiency and energy consumption, thus meeting the actual industrial needs of Sachima production.
[0035] Step S20: Establish an agent for the Markov decision process of Sachima raw material mixing and stirring control, set up a value network, decision network, target network and experience playback buffer, and learn the control strategy of Sachima raw material mixing and stirring.
[0036] In this embodiment, an agent is established based on Markov decision process parameters to achieve adaptive control of the mixing and stirring process of Sachima raw materials. The agent components consist of a policy network, a value network, a target network, and an experience replay buffer. The deep deterministic policy gradient method is used to complete the training process of the agent. The input and output of the policy network are state vector and action vector, respectively, representing the state of the Sachima raw materials and the corresponding action executed by controlling the mixer. The value network also takes state and action as input and outputs a Q-value, which evaluates the long-term expected accumulated reward of the current state-action pair and guides the parameter update of the policy network. The target network is used to create replicas of the policy network and value network, and to calculate the stable Q-value target. The experience replay buffer stores the interaction data of the agent during the mixing and stirring process of Sachima raw materials, and is used for random sampling to break the temporal correlation of the data.
[0037] Step S30: Establish a hybrid training environment to train the agent, simulate the mixing process of Sachima raw materials, calculate the target value according to the reward function to update the network parameters, optimize and adjust the control strategy of the mixer, and make corrections on a real Sachima mixer.
[0038] In this embodiment, based on the physical principles of a Sachima mixer, a simulated digital twin environment is constructed to simulate the dynamics of the mixing motor and the evolution of the Sachima raw material state. The intelligent agent is pre-trained through data interaction. In a digital twin, a motor and transmission model is established to simulate the electromagnetic characteristics, mechanical inertia, and friction of a servo motor. Simultaneously, a hybrid modeling approach is used to establish a model of the mixing process of Sachima raw materials. Based on computational fluid dynamics, the macroscopic flow and particle collisions of Sachima raw materials are simulated. Historical production data is used to train a neural network to predict changes in output load torque and uniformity increments, thereby optimizing and correcting the simulation process. Based on the above model, a noise model that conforms to the characteristics of real sensors is added, and corresponding sensor signals are synthesized according to the physical state of the mixing process. In addition, the actual physical training environment for Sachima mixers and their sensors and actuators in production is used for strategy verification and online adjustment. In the simulation environment, the agent observes the initial state of the mixing process of Sachima raw materials, selects the control action of the mixer according to the current strategy, calculates the reward function to obtain a new state vector, samples batch data from the buffer to update the parameters of the policy network and the value network, and repeats the above process with the goal of maximizing the cumulative reward to train and obtain the first control strategy. The offline-trained first control strategy is loaded into the controller of the real equipment, and the execution command is sent and executed by the PID controller. Bounded random noise is added to the output control action, real interaction data is collected, the learning rate is reduced to update the network parameters, the first control strategy is finely adjusted, the error between simulation and reality is made up, the characteristics of the equipment and batch changes of raw materials are adapted, and the final accurate control strategy is generated to realize the online adaptive control of the Sachima mixer.
[0039] Step S40: Based on the training results of the agent, generate the optimal adaptive control strategy for the mixing and stirring process of Sachima raw materials, and dynamically adjust the working parameters of the Sachima mixer in the actual production process.
[0040] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. An adaptive cooperative control method for raw material mixing and stirring processes, characterized in that, The method includes: Step S10: The adaptive collaborative control of the mixing and stirring process of Sachima raw materials is transformed into a Markov decision process. Based on the stirring process data and the corresponding mixer control strategy, the state space, action space, transition probability and reward function of the decision process are defined. Step S20: Establish an agent for the Markov decision process of Sachima raw material mixing and stirring control, set up a value network, decision network, target network and experience playback buffer, and learn the control strategy of Sachima raw material mixing and stirring. Step S30: Establish a hybrid training environment to train the agent, simulate the mixing process of Sachima raw materials, calculate the target value according to the reward function to update the network parameters, optimize and adjust the control strategy of the mixer, and make corrections on a real Sachima mixer. Step S40: Based on the training results of the agent, generate the optimal adaptive control strategy for the mixing and stirring process of Sachima raw materials, and dynamically adjust the working parameters of the Sachima mixer in the actual production process.
2. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 1, characterized in that, In step S10, the control objective of the mixing and stirring process of Sachima raw materials is set to mix the Sachima raw materials to the target uniformity with the lowest energy consumption. The constraints on the control process based on this objective include: limiting the motor current and torque of the Sachima mixer to not exceed the safety limit and the rotation speed to within the mechanical design range.
3. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 2, characterized in that, The uniformity is calculated based on the mixer's torque fluctuation characteristics, acoustic spectrum characteristics, and vibration dominant frequency amplitude, and is expressed by the following formula: in, Indicates in The estimated uniformity index of Sachima raw materials at any time. It is the coefficient of variation of the torque of the main motor of the mixer, which represents the ratio of the standard deviation of the torque to the average value; and These represent the normalized sound pressure signal spectral entropy and the normalized vibration dominant frequency amplitude, respectively. After normalization processing, , and They have the same physical dimensions, and the sound pressure signal and vibration signal are acquired through corresponding sensors; , and The weighting coefficients represent the relative importance of each feature to the evenness estimation. They are determined through regression analysis of historical data and satisfy the following conditions: The uniformity value ranges from [0,1], where 0 and 1 represent completely non-uniform and completely uniform, respectively.
4. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 1, characterized in that, In step S10, the relevant parameter space of the Markov decision process for mixing and stirring the raw materials of Sachima is defined, including the following steps: The state space of the intelligent agent is based on the process data of mixing and stirring of Sachima raw materials collected at each decision moment, which includes real-time sensor readings, process characteristics, process stage information and historical information. Using the control commands executed by the Sachima mixer as the workspace of the intelligent agent, the system outputs the incremental values of the set speed of the main and slave motors of the mixer, as well as the extended action commands for switching the mixing process stages, thereby realizing the process conversion from dry mixing, wet mixing to maturation. The state transition probability represents the probability that the training environment of the agent will transition to the next state based on the current state and action. The reward function for multi-objective optimization is defined as a weighted sum of quality reward, efficiency reward and energy consumption reward, which guides the training process of the agent.
5. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 4, characterized in that, The reward function is expressed as a weighted sum of quality reward, efficiency reward, and energy consumption reward, and is given by the formula: in, express Total reward value at each moment , and Let represent the function values of quality reward, efficiency reward, and energy consumption reward, respectively. , and This represents the weighting coefficient, which is set empirically.
6. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 5, characterized in that, The quality reward is represented as a uniformity improvement reward. and excessive mixing of punishment The weighted sum, expressed by the formula: in, and They represent and The uniformity of the raw materials in Sachima at any given time Indicates the threshold for excessive mixing of raw materials; and These represent the sensitivity parameter and the penalty steepness parameter, respectively. and These are weighting coefficients, set empirically.
7. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 5, characterized in that, The formula for the efficiency reward is expressed as follows: in, This indicates the mixing time currently used. This represents the historical average optimal time. and These represent the maximum allowable mixing time and the starting point of the time penalty, respectively. This indicates the target uniformity of the raw materials for Sachima.
8. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 5, characterized in that, The formula for the energy consumption reward is expressed as follows: in, and Indicates instantaneous power and the system's rated total power. and These represent the current cumulative energy consumption and energy budget, respectively. and The weighting parameters are set empirically. and This indicates the torque and speed of the main motor. and This indicates the torque and speed of the motor.
9. The adaptive cooperative control method for raw material mixing and stirring process as described in claim 1, characterized in that, In step S20, based on Markov decision process parameters, an agent is established to achieve adaptive control of the mixing and stirring process of Sachima raw materials. The agent components consist of a policy network, a value network, a target network, and an experience replay buffer. The deep deterministic policy gradient method is used to complete the training process of the agent. The input and output of the policy network are state vector and action vector, respectively, representing the state of the Sachima raw materials and the corresponding action executed by controlling the mixer. The value network also takes state and action as input and outputs a Q-value, which evaluates the long-term expected accumulated reward of the current action and guides the parameter update of the policy network. The target network is used to create replicas of the policy network and value network, and to calculate the stable Q-value target. The experience replay buffer stores the interaction data of the agent during the mixing and stirring process of Sachima raw materials, and is used for random sampling to break the temporal correlation of the data.
10. The adaptive cooperative control method for a raw material mixing and stirring process as described in claim 1, characterized in that, In step S30, based on the physical principle of the Sachima mixer, a simulated digital twin environment is constructed to simulate the dynamics of the mixing motor and the state evolution of the Sachima raw materials. The intelligent agent is pre-trained through data interaction. In addition, the actual physical training environment for Sachima mixers and their sensors and actuators in production is used for strategy verification and online adjustment. In the simulation environment, the agent observes the initial state of the mixing process of Sachima raw materials, selects the control action of the mixer according to the current strategy, calculates the reward function to obtain a new state vector, samples batch data from the buffer to update the parameters of the policy network and the value network, and repeats the above process with the goal of maximizing the cumulative reward to train and obtain the first control strategy. The offline-trained first control strategy is loaded into the controller of the real device, and the execution command is sent and executed by the PID controller. Bounded random noise is added to the output control action, real interaction data is collected, the learning rate is reduced to update the network parameters, the first control strategy is finely adjusted, and the final accurate control strategy is generated to realize the online adaptive control of the Sachima mixer.