Automatic integrated processing system for coarse cereal nutrition formula powder
By integrating the signal sensing and fusion module, the nutrient retention prediction module, the processing parameter decision module, and the instruction optimization and collaborative control module, the real-time dynamic adjustment of the grain nutrient grinding process is realized, which solves the problem of unstable nutrient loss in the existing technology and improves product quality and nutrient retention rate.
Patent Information
- Application Number
- CN202512021963.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies cannot dynamically adjust process parameters during the nutrient grinding of mixed grains, resulting in unstable loss of nutrients and making it difficult to maximize the retention of active nutrients while ensuring product fineness and taste.
Employing a signal sensing and fusion module, a nutrient retention prediction module, a processing parameter decision module, and an instruction optimization and collaborative control module, the system collects signals in real time through multiple types of sensors, generates a spatiotemporally aligned multidimensional fusion feature matrix, and dynamically optimizes processing parameters to achieve precise nutrient retention by combining a deep temporal prediction model and a reinforcement learning decision model.
It enables real-time control and optimization of the processing, ensuring the stable retention of nutrients and improving product quality, and solving the instability problem caused by fixed parameters in existing technologies.
Smart Images

Figure CN121669403A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of grain grinding and processing, and more specifically, to an automated integrated processing system for mixed grain nutritional formula powder. Background Technology
[0002] With the increasing awareness of health among the public, nutritional formula powders based on whole grains are highly favored due to their rich dietary fiber, vitamins, minerals, and phytochemicals. To meet the higher demands of modern consumers for product solubility and taste, it is necessary to effectively grind the grain raw materials to obtain powders of ideal fineness. However, the heat energy converted from mechanical energy during the grinding process, as well as the strong shearing, extrusion, and other mechanochemical effects, can significantly damage the heat-sensitive and oxidation-sensitive nutrients naturally present in the raw materials, such as B vitamins, vitamin E, polyphenols, and unsaturated fatty acids. This makes it a common industry dilemma for processors when producing whole grain powders or functional formula powders with high nutritional value, where they have to make a difficult trade-off between pursuing better grinding fineness and retaining the active nutrients to the maximum extent.
[0003] Existing technical solutions have significant shortcomings in addressing the aforementioned contradictions. While equipment employing cutting principles is considered more nutrient-efficient, a single cutting method often fails to achieve the required powder fineness. Therefore, actual production lines often use a combined crushing path, such as coarse crushing via cutting followed by fine crushing via high-speed impact or grinding equipment. However, the control strategies for such multi-stage crushing systems are typically quite crude. Specifically, key process parameters, including but not limited to cutting speed, impact hammer linear velocity, grinding disc gap, and pressure, are mostly set once before batch production based on historical experience and kept constant throughout the processing. This static parameter setting cannot adapt to the natural fluctuations in initial moisture content, particle size distribution, and temperature of raw materials from different origins and batches. More importantly, existing... The technology lacks the ability to online sense and assess the dynamic changes of nutrient components during processing. Some equipment may integrate simple temperature monitoring points, but their control logic is limited to over-temperature alarms or activation of cooling devices. It has not built a quantitative model that can link the complex relationship between real-time process parameters, material state, and final nutrient retention rate. Because it is impossible to perform soft measurement or prediction of the real-time nutrient retention, the production process is in an open-loop control state. It is impossible to dynamically feedback and optimize the grinding intensity based on the actual processing effect. This operation mode, which relies on fixed parameters and experience, directly leads to the instability of processing results. It is difficult to achieve precise and maximum retention of high-value-added nutrients while ensuring product solubility and taste. This has become a key technological bottleneck restricting the upgrading of the high-quality whole grain nutritional powder industry towards intelligence and refinement. Summary of the Invention
[0004] This invention addresses the technical problems existing in the prior art by providing an automated integrated processing system for mixed grain nutritional formula powder. The system utilizes a signal sensing and fusion module, a nutrient retention prediction module, a processing parameter decision module, and an instruction optimization and collaborative control module to solve the problems mentioned in the background.
[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: specifically including: a signal sensing and fusion module, a nutrient retention prediction module, a processing parameter decision module, and an instruction optimization and collaborative control module, wherein;
[0006] Signal sensing and fusion module: In response to the processing start command of the grinding unit, it synchronously collects raw signals from multiple types of sensors deployed in the key physical action area of the grinding unit, and performs synchronization, alignment and feature extraction processing on the raw signals to generate a spatiotemporally aligned multidimensional fusion feature matrix;
[0007] Nutrient retention prediction module: receives the multidimensional fusion feature matrix from the signal sensing and fusion module, and inputs the multidimensional fusion feature matrix into the deep time series prediction model. The deep time series prediction model performs time series analysis and outputs a binary prediction result containing the predicted mean of the retention rate of key nutrients and its corresponding prediction uncertainty.
[0008] Processing parameter decision module: It is used to fuse the current raw material initial composition data, real-time equipment parameter set, target product particle size index and the binary prediction result from the nutrient retention prediction module to form the current processing state representation vector, and input the current processing state representation vector into the reinforcement learning decision model. The reinforcement learning decision model calculates a set of processing parameter adjustment amounts to be executed according to the embedded robust optimization strategy.
[0009] Instruction optimization and collaborative control module: used to receive the set of processing parameter adjustment amounts to be executed from the processing parameter decision module, and with the set as the set target, combined with the dynamic characteristic constraints of each actuator of the grinding unit and the preset global optimization target, to perform rolling time domain optimization calculation, thereby generating a sequence of collaborative control instructions to drive each actuator in the future short time domain, and distributing the collaborative control instructions corresponding to the current time to the corresponding actuator of the grinding unit;
[0010] In a preferred embodiment, the signal sensing and fusion module, in response to the processing start command, sends a synchronous acquisition command to multiple sensors deployed in the key physical action area of the grinding unit, and synchronously receives the original signals returned by each sensor according to a unified clock. The key physical action area includes the cutting zone inlet, the side wall of the impact crushing chamber, and the grinding pair gap outlet. The multiple sensors include an acoustic emission sensor for capturing transient stress waves of particle crushing, a vibration accelerometer for monitoring rotor imbalance and collision characteristics, an infrared thermal imager for sensing the apparent temperature distribution of the material, and a multispectral camera for acquiring the apparent spectral information of the material.
[0011] In a preferred embodiment, the specific operations of synchronizing, aligning, and extracting features from the original signal are as follows:
[0012] Each original signal is preprocessed, including noise reduction, detrending term and amplitude normalization, to obtain the preprocessed signal of each channel.
[0013] Based on the preset sensor physical location coordinates and material flow rate information, spatiotemporal alignment calculations are performed on each preprocessed signal to generate a spatiotemporal alignment signal corresponding to each sensor.
[0014] For the spatiotemporal aligned signals corresponding to acoustic emission sensors and vibration accelerometers, a continuous wavelet transform method is used to extract the time-frequency domain energy distribution values at multiple preset scales from the spatiotemporal aligned signals;
[0015] For the spatiotemporal aligned signals corresponding to infrared thermal imagers and multispectral cameras, statistical moments including mean, standard deviation, skewness, and kurtosis, as well as the rate of change of the spatiotemporal aligned signal within a sliding time window, are extracted from the spatiotemporal aligned signals. All the above-mentioned time-frequency domain energy distribution values, statistical moments, and rates of change extracted from each spatiotemporal aligned signal are collectively defined as multiple feature values extracted from that signal. All feature values extracted from a spatiotemporal aligned signal are arranged in a predetermined order to form a feature vector.
[0016] The feature vectors extracted from the spatiotemporally aligned signals of all sensors at the same time are concatenated to form the initial total feature vector.
[0017] In a preferred embodiment, the specific process of generating the spatiotemporally aligned multidimensional fusion feature matrix is as follows:
[0018] Calculate the covariance matrix of the initial total eigenvectors within a sliding time window;
[0019] Eigenvalue decomposition is performed on the covariance matrix, and the eigenvectors corresponding to the first set number of largest eigenvalues are selected to form a principal component projection matrix with the number of rows equal to the set number and the number of columns equal to the dimension of the initial total eigenvectors.
[0020] An adaptive diagonal weight matrix is constructed. This matrix is a square matrix with the number of rows and columns equal to the dimension of the initial total eigenvector, and only the elements on the main diagonal have non-zero values. The value of each diagonal element of the adaptive diagonal weight matrix is determined by the eigenvalue in the initial total eigenvector that is at the same index as the diagonal element. Specifically, it is dynamically calculated by a preset nonlinear mapping function based on the signal-to-noise ratio estimate of the eigenvalue in the nearest time window and the Mahalanobis distance calculated from the eigenvector containing the eigenvalue and the corresponding eigenvector in the historical baseline state.
[0021] The initial total eigenvector, principal component projection matrix, and adaptive diagonal weight matrix are input into the adaptive weighted fusion function for calculation. The calculation process is as follows: First, each element on the main diagonal of the adaptive diagonal weight matrix is multiplied by the element with the same index in the initial total eigenvector to obtain an intermediate vector. Then, the principal component projection matrix is transposed to obtain the transposed principal component projection matrix. Finally, the transposed principal component projection matrix is multiplied by the intermediate vector to obtain a fusion feature vector with a set number of rows. This fusion feature vector is a column vector of the multidimensional fusion feature matrix at the corresponding time point.
[0022] The fused feature vectors from continuous time series are stacked in chronological order to generate a spatiotemporally aligned multidimensional fused feature matrix.
[0023] In a preferred embodiment, the nutrient retention prediction module uses a deep time-series prediction model that is a time-series hybrid network based on quantile regression. Its construction and input processing specifically include:
[0024] Receive the multi-dimensional fusion feature matrix from the signal sensing and fusion module; obtain the initial key nutrient content vector of the raw material from the current processing batch information; use a preset fixed time length as a sliding window to sequentially extract the fusion feature vectors at multiple consecutive time points from the multi-dimensional fusion feature matrix to form a dynamic sequence.
[0025] The contents of the initial key nutrient content vector are copied according to the number of time points contained in the sliding window to form a static feature sequence with the same length in the time dimension as the dynamic sequence.
[0026] The dynamic sequence and the static feature sequence are merged along the data dimension to form an enhanced sequence sample;
[0027] The temporal hybrid network based on quantile regression includes an encoder and a quantile output head. The encoder is composed of a multi-scale temporal convolutional network and a multi-head self-attention mechanism connected sequentially. It is used to encode features of the enhanced sequence samples, extract multi-scale local features and global dependencies, and output a context encoding vector.
[0028] The quantile output head is connected to the output of the encoder. It is configured to receive the context encoding vector and the initial critical nutrient content vector, and output in parallel the quantile prediction values of the three critical nutrient retention rates corresponding to the preset low quantile, median quantile and high quantile.
[0029] In a preferred embodiment, the training of the deep temporal prediction model and the generation of the binary prediction results specifically include:
[0030] During the training phase, a historical dataset containing augmented sequence samples and their corresponding real key nutrient retention rate labels is used. The temporal hybrid network is trained by optimizing a quantile loss function. The value of the quantile loss function is obtained by multiplying the quantile loss values corresponding to the lower quantile, median quantile, and upper quantile respectively by preset lower, median, and upper weight coefficients, and then summing them. For any one of the preset lower, median, or upper quantile, the value of that preset quantile is denoted as the corresponding quantile value, and its corresponding quantile prediction value is the predicted value under that quantile value. The calculation rule for the quantile loss value of that quantile is as follows:
[0031] Calculate the difference between the true critical nutrient retention rate label and the quantile prediction value. If the difference is positive, the quantile loss value is equal to the product of the quantile value and the difference. If the difference is negative or zero, the quantile loss value is equal to the product of a difference coefficient obtained by subtracting the quantile value and the difference obtained by subtracting the true label value from the quantile prediction value.
[0032] During the online prediction phase, the current enhanced sequence samples acquired in real time are input into the trained temporal hybrid network to obtain the corresponding three quantile prediction values; the median quantile prediction value is directly used as the mean of the key nutrient retention rate prediction.
[0033] Based on the predicted values of the lower quantile, median quantile, and upper quantile, a measure of prediction uncertainty is calculated. The calculation process is as follows:
[0034] First, calculate the first median value, which is the higher quantile predicted value minus the lower quantile predicted value; then, calculate the absolute value of the median quantile predicted value; finally, divide the first median value by the absolute value to obtain the first result, which is called the nominal prediction interval width.
[0035] Next, the second intermediate value is calculated, which is the predicted value of the higher quantile minus twice the predicted value of the median quantile, plus the predicted value of the lower quantile. Then, the second intermediate value is squared to obtain the second intermediate result, called the interval symmetry deviation. After that, a preset weight adjustment coefficient is multiplied by the interval symmetry deviation to obtain the second result, called the weighted symmetry deviation penalty term.
[0036] Finally, the first result is added to the second result, and the sum is the measure of prediction uncertainty.
[0037] The median quantile prediction value is used as the mean prediction of the retention rate of key nutrients, and the calculated prediction uncertainty measure is used as the corresponding prediction uncertainty. These are then encapsulated into a binary prediction result containing the prediction mean and the prediction uncertainty.
[0038] In a preferred embodiment, the specific construction process of forming the current processing state representation vector and the robust optimization strategy upon which the reinforcement learning decision model is based in the processing parameter decision module includes:
[0039] The initial composition data of the current raw materials is denoted as the raw material composition vector, the real-time equipment parameter set is denoted as the equipment parameter vector, the particle size index of the target product is denoted as the particle size target value, the predicted mean of the retention rate of key nutrients in the binary prediction results from the nutrient retention prediction module is denoted as the prediction mean, and its prediction uncertainty is denoted as the uncertainty value; wherein, the real-time equipment parameter set includes a real-time particle size monitoring value measured in real time by an online particle size sensor deployed at the discharge end of the grinding unit;
[0040] The raw material composition vector, equipment parameter vector, particle size target value, predicted mean, and uncertainty value are spliced and fused together. Before splicing, the raw material composition vector, equipment parameter vector, and particle size target value are normalized respectively, while the predicted mean and uncertainty value are kept at their original values. Then, the normalized raw material composition vector, normalized equipment parameter vector, normalized particle size target value, predicted mean, and uncertainty value are connected in this order to form a current processing state representation vector.
[0041] The robust optimization strategy embedded in the reinforcement learning decision model is evaluated through a multi-objective weighted synthesis function. The output of this function is a comprehensive regret value, which is obtained by weighted summation of three regret values: nutrient retention target regret value, granularity proximity target regret value, and energy consumption control target regret value.
[0042] The calculation method for the regret of nutrient retention target is as follows:
[0043] V1. Calculate the product of the uncertainty value and a preset risk aversion coefficient that is greater than or equal to zero to obtain the risk compensation amount;
[0044] V2. Subtract the risk compensation from the predicted mean to obtain a pessimistic estimate of nutrient retention.
[0045] V3. Calculate the difference between the preset nutrient retention target value and the pessimistic estimate of nutrient retention; if the difference is positive, take the difference as the nutrient retention target regret value; if the difference is zero or negative, take zero as the nutrient retention target regret value.
[0046] The calculation method for the granularity approaching the target regret is as follows: calculate the absolute value of the difference between the real-time monitoring value of the granularity and the target value of the granularity, and then divide this absolute value by the target value of the granularity. The quotient is the granularity approaching the target regret.
[0047] The calculation method for the energy consumption control target regret is as follows: The real-time device parameter set contains a real-time energy consumption value that reflects the current total energy consumption. A preset benchmark energy consumption value is used as the denominator, and the real-time energy consumption value is used as the numerator. The ratio between the two is calculated, and the resulting ratio is the energy consumption control target regret.
[0048] Finally, the output value of the multi-objective weighted synthesis function, i.e. the comprehensive regret value, is equal to the sum of the regret value of the nutrient retention target multiplied by a preset nutrient target weight coefficient, the regret value of the granularity approximation target multiplied by a preset granularity target weight coefficient, and the regret value of the energy consumption control target multiplied by a preset energy consumption target weight coefficient.
[0049] In a preferred embodiment, the reinforcement learning decision model calculates a set of processing parameter adjustments to be executed based on an embedded robust optimization strategy, specifically:
[0050] First, a decision calculation is performed: the current processing state representation vector is input into the reinforcement learning decision model, which processes the data through its internal policy network and outputs an original action vector.
[0051] Then, during the training and online decision exploration phase of the reinforcement learning decision model, an adaptive exploration mechanism is introduced. This mechanism uses the uncertainty value as the basis for dynamically adjusting the exploration intensity, specifically:
[0052] When the reinforcement learning decision model outputs the original action vector, the variance of the superimposed random exploration noise is proportional to the uncertainty value. That is, the uncertainty value is used as input, multiplied by a preset exploration coefficient, and then added to a preset minimum noise variance. The result is the variance of the superimposed random exploration noise in the current decision cycle.
[0053] Finally, post-processing and output are performed: the original action vector obtained from the decision calculation step is post-processed to generate the set of processing parameter adjustment amounts to be executed. The post-processing specifically includes:
[0054] First scaling sub-step: Multiply the value of each dimension of the original motion vector by the absolute value of the maximum allowable adjustment amount of the corresponding processing parameter to be adjusted, to obtain the actual physical adjustment amount of each parameter;
[0055] The second feasibility verification sub-step: Based on the preset physical lower limit and upper limit of each processing parameter to be adjusted, the new parameter value obtained by superimposing the actual physical adjustment amount onto the current real-time equipment parameter value is verified and truncated.
[0056] The third smoothness constraint sub-step: calculate the rate of change of the difference between the actual physical adjustment amount obtained in the current decision cycle and the actual physical adjustment amount of the corresponding parameter used in the previous decision cycle;
[0057] After the above-mentioned first scaling sub-step, second feasibility verification sub-step, and optional third smoothness constraint sub-step, the resulting set containing the actual physical adjustment amounts of all processing parameters to be adjusted is the set of processing parameter adjustment amounts to be executed.
[0058] In a preferred embodiment, the instruction optimization and collaborative control module performs rolling time-domain optimization calculations by combining the dynamic characteristic constraints of each actuator of the grinding mill with a preset global optimization objective, specifically including:
[0059] A linear parameter time-varying state space model of each actuator of the grinding unit is established as a prediction model. The establishment of the linear parameter time-varying state space model is based on offline system identification of the historical operating data and step response test of actuators such as feed inverter, cutting motor driver, grinding pressure servo valve, and classifier inverter.
[0060] The specific process of generating a collaborative reference trajectory is as follows: receiving the set of processing parameter adjustment amounts to be executed from the processing parameter decision module, and denoting this set as the expected adjustment amount set; obtaining the actual measured values of the operating parameters of each device at the current moment, wherein the actual measured values are derived from the data collected in real time by the sensors deployed on the grinding unit, and forming the current output vector;
[0061] A coordination timing matrix is introduced, where each element defines the start time offset of an actuator in the grinding unit to begin responding to its corresponding single desired adjustment in the set of desired adjustment amounts, in order to describe the sequential coordination timing of the actions of multiple mechanisms.
[0062] Based on the collaborative time series matrix, a corresponding shape function is defined for each expected adjustment in the expected adjustment set. The shape function is an S-shaped function with the future discrete time point and the starting time offset corresponding to the expected adjustment as independent variables. The shape functions corresponding to all expected adjustment quantities constitute a shape function vector.
[0063] The generation calculation of the cooperative reference trajectory is as follows: For each future discrete time point in the future prediction time domain, firstly, each expected adjustment amount in the expected adjustment amount set is multiplied element-wise with the corresponding position of the shape function output value at that time point in the shape function vector to obtain an intermediate adjustment amount vector after time and temporal modulation; then, this intermediate adjustment amount vector is added element-wise to the current output vector, and the resulting vector is the cooperative reference trajectory point at that future time point; by traversing all future discrete time points in the future prediction time domain, the complete cooperative reference trajectory can be obtained;
[0064] Based on the linear parameter time-varying state-space model and the cooperative reference trajectory, a rolling time-domain optimization problem is constructed. In each control cycle, the rolling time-domain optimization problem is transformed into a constrained quadratic programming problem online and solved to obtain the optimal control input vector sequence in the future control time domain.
[0065] In a preferred embodiment, generating a sequence of coordinated control instructions to drive each actuator in the future short time domain, and distributing the coordinated control instructions corresponding to the current time to the corresponding actuators, specifically includes:
[0066] First, the sequence extraction and instruction distribution are performed: the first control input vector in the optimal control input vector sequence obtained by solving the rolling time domain optimization problem is extracted as the current moment's coordinated control instruction to be issued in the current control cycle; at the same time, the complete optimal control input vector sequence is cached; the current moment's coordinated control instruction is distributed to the corresponding actuator in the grinding unit through the corresponding communication interface and protocol.
[0067] Next, feedback correction and rolling optimization are performed: When the next control cycle arrives, the latest measured values of the equipment operating parameters from the sensors are received and recorded as the latest measured values; using the latest measured values, the system state vector estimate of the linear parameter time-varying state space model is updated through the state estimator; and using this updated system state vector estimate as the new initial state, combined with the latest received set of processing parameter adjustments to be executed, the steps of generating the cooperative reference trajectory, constructing and solving the rolling time-domain optimization problem are re-executed to achieve closed-loop feedback correction, thereby generating and issuing new cooperative control commands for the current moment;
[0068] The process of sequence extraction and instruction distribution, feedback correction and rolling optimization is performed in each control cycle.
[0069] The beneficial effects of this invention are as follows: This scheme realizes the real-time generation and precise issuance of control commands. By extracting the first command of the optimization sequence, real-time performance is ensured. At the same time, the cached sequence provides a hot start initial point for the next cycle, which significantly improves the optimization solution efficiency. In each control cycle, the state estimate is updated based on the latest measurement value, and the cooperative trajectory and optimization problem are re-planned from this starting point, forming a closed-loop feedback correction mechanism. This rolling optimization process transforms the continuous adjustment target into smooth, cooperative, and strictly satisfying actual driving commands that meet various dynamic and safety constraints, ensuring the stability of the processing process and the control quality. Attached Figure Description
[0070] Figure 1 This is a flowchart of the method of the present invention;
[0071] Figure 2 This is a block diagram of the system structure of the present invention. Detailed Implementation
[0072] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0073] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0074] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0075] Example 1
[0076] This embodiment provides, for example Figure 1-2 The automated integrated processing system for mixed grain nutritional formula powder shown includes: a signal sensing and fusion module, a nutrient retention prediction module, a processing parameter decision module, and an instruction optimization and collaborative control module, wherein;
[0077] Signal sensing and fusion module: In response to the processing start command of the grinding unit, it synchronously collects raw signals from multiple types of sensors deployed in the key physical action area of the grinding unit, and performs synchronization, alignment and feature extraction processing on the raw signals to generate a spatiotemporally aligned multidimensional fusion feature matrix;
[0078] Nutrient retention prediction module: Receives a multi-dimensional fusion feature matrix from the signal sensing and fusion module, and inputs the multi-dimensional fusion feature matrix into the deep time series prediction model. The deep time series prediction model performs time series analysis and outputs a binary prediction result containing the predicted mean of the retention rate of key nutrients and its corresponding prediction uncertainty.
[0079] Processing parameter decision module: It is used to integrate the current raw material initial composition data, real-time equipment parameter set, target product particle size index and binary prediction results from the nutrient retention prediction module to form the current processing state representation vector. The current processing state representation vector is then input into the reinforcement learning decision model, which calculates a set of processing parameter adjustment amounts to be executed based on the embedded robust optimization strategy.
[0080] Instruction optimization and collaborative control module: It is used to receive the set of processing parameter adjustment amounts to be executed from the processing parameter decision module, and with the set as the set target, combined with the dynamic characteristic constraints of each actuator of the grinding unit and the preset global optimization target, it performs rolling time domain optimization calculation, thereby generating a sequence of collaborative control instructions to drive each actuator in the future short time domain, and distributing the collaborative control instructions corresponding to the current moment to the corresponding actuator of the grinding unit.
[0081] In this embodiment, it is specifically noted that in the signal sensing and fusion module, in response to the processing start command, a synchronous acquisition command is sent to multiple sensors deployed in the key physical action area of the grinding unit, and the original signals returned by each sensor according to a unified clock are received synchronously. The unified clock preferably adopts a network synchronization clock implemented by the IEEE 1588 precision time protocol to ensure that the timestamp deviation of all sensor data acquisition is less than 1 microsecond. The key physical action area includes the cutting zone inlet, the side wall of the impact crushing chamber, and the grinding pair gap outlet. The multiple sensors include an acoustic emission sensor for capturing transient stress waves of particle crushing, a vibration accelerometer for monitoring rotor imbalance and collision characteristics, an infrared thermal imager for sensing the apparent temperature distribution of the material, and a multispectral camera for acquiring the apparent spectral information of the material. The original signal is defined by the sensor number, signal type identifier, and unified clock timestamp.
[0082] The specific operations for synchronizing, aligning, and extracting features from the original signal are as follows:
[0083] Each original signal is preprocessed, including denoising, detrending, and amplitude normalization. For denoising, wavelet threshold denoising is used for acoustic emission and vibration signals, and sliding median filtering is used for temperature and spectral signals. Amplitude normalization linearly scales the amplitude of each signal to the [-1,1] interval to obtain the preprocessed signals.
[0084] Based on the preset sensor physical location coordinates and material flow rate information, spatiotemporal alignment calculations are performed on each preprocessed signal, specifically:
[0085] First, the measurement points of each sensor are mapped to a unified virtual principal axis coordinate line of the material flow to assign them spatial coordinates. Then, the comprehensive time delay compensation of each signal from the acquisition point to the preset common reference surface is calculated. The comprehensive time delay compensation includes the propagation delay of the signal in the medium and the movement delay of the material from the sensor point to the common reference surface. Then, the timestamp of each preprocessed signal is subtracted from its corresponding comprehensive time delay compensation. Finally, the time delay compensation is used to perform time shift correction on each preprocessed signal to generate a spatiotemporal aligned signal corresponding to each sensor.
[0086] For the spatiotemporally aligned signals corresponding to acoustic emission sensors and vibration accelerometers, a continuous wavelet transform method is used to extract the time-frequency domain energy distribution values at multiple preset scales from the spatiotemporally aligned signals. The mother wavelet of the continuous wavelet transform is preferably the Morlet wavelet, with 8 preset scales covering the main frequency band from 50Hz to 10kHz. The multiple preset scales cover the main frequency components corresponding to macroscopic particle crushing to microscopic friction, thereby comprehensively capturing the mechanical energy conversion and release characteristics during the crushing process.
[0087] For the spatiotemporal alignment signals corresponding to infrared thermal imagers and multispectral cameras, statistical moments including mean, standard deviation, skewness, and kurtosis, as well as the rate of change of the spatiotemporal alignment signal within a sliding time window, are extracted from the spatiotemporal alignment signals. The length of the sliding time window can be set according to the residence time of the material at the corresponding station, for example, set to one-fifth to one-half of the residence time, to balance the trend characterization capability of the features with real-time requirements. All the above-mentioned time-frequency domain energy distribution values, statistical moments, and rates of change extracted from each spatiotemporal alignment signal are collectively defined as multiple feature values extracted from that signal. All feature values extracted from a spatiotemporal alignment signal are arranged in a predetermined order to form a feature vector, wherein the position of each feature value in the vector is uniquely determined by the sensor type and feature type that generated the feature value.
[0088] The feature vectors extracted from the spatiotemporally aligned signals of all sensors at the same time are concatenated to form an initial total feature vector. The dimension of this initial total feature vector is equal to the total number of feature values of all sensors.
[0089] The specific process for generating a spatiotemporally aligned multidimensional fusion feature matrix is as follows:
[0090] Calculate the covariance matrix of the initial total eigenvector within a sliding time window. The length of the sliding time window, T_window, is preferably set to 2 seconds to balance computational real-time performance and statistical stability.
[0091] The covariance matrix is decomposed into eigenvalues. The eigenvectors corresponding to the first set number of largest eigenvalues are selected to form a principal component projection matrix with the number of rows equal to the set number and the number of columns equal to the dimension of the initial total eigenvectors. The set number M is determined by the cumulative variance contribution rate threshold. The preferred threshold is set to 85%, that is, the minimum M is selected such that the sum of the first M eigenvalues accounts for no less than 85% of the sum of all eigenvalues.
[0092] An adaptive diagonal weight matrix is constructed. This matrix is a square matrix with the number of rows and columns equal to the dimension of the initial total eigenvector, and only the elements on the main diagonal have non-zero values. The value of each diagonal element of the adaptive diagonal weight matrix is determined by the eigenvalue in the initial total eigenvector at the same index as that diagonal element. Specifically, it is dynamically calculated using a preset nonlinear mapping function based on the signal-to-noise ratio (SNR) estimate of that eigenvalue within the nearest time window and the Mahalanobis distance calculated from the eigenvector containing that eigenvalue and the corresponding eigenvector in the historical baseline state. The SNR estimate is obtained by calculating the ratio of signal power to noise power for that eigenvalue within the nearest time window. The noise power can be estimated using the residual energy after high-pass filtering or wavelet thresholding denoising. The baseline state feature vector is obtained by statistical analysis (such as calculating the mean vector) of feature vector data collected by the system during stable operation under calibration conditions. The preset nonlinear mapping function, such as a variant of the Sigmoid function family, is designed with the following principle: when the signal-to-noise ratio estimate is higher than the first set threshold and the Mahalanobis distance is lower than the second set threshold, a larger weight value (close to 1) is output, indicating that the feature value is reliable and within the normal range of variation, and should be enhanced; when the signal-to-noise ratio estimate is too low or the Mahalanobis distance is too high, a smaller weight value (close to 0) is output to suppress unreliable feature components that may be introduced by noise or abnormal operating conditions. Through this adaptive weighting mechanism, the confidence of each feature component can be dynamically evaluated and adjusted before fusion, thereby significantly improving the robustness and information purity of the fused features.
[0093] The initial total eigenvector, principal component projection matrix, and adaptive diagonal weight matrix are input into the adaptive weighted fusion function for calculation. The calculation process is as follows: First, each element on the main diagonal of the adaptive diagonal weight matrix is multiplied by the element with the same index in the initial total eigenvector to obtain an intermediate vector. This step completes the real-time, differentiated weighting of each component of the initial total eigenvector. Then, the principal component projection matrix is transposed to obtain the transposed principal component projection matrix. Finally, the transposed principal component projection matrix is multiplied by the intermediate vector to obtain a fusion feature vector with a set number of rows. This step uses the principal component projection matrix to map the weighted high-dimensional features to a low-dimensional principal component space that can best reflect the cross-sensor collaborative change pattern. This fusion feature vector is a column vector of the multi-dimensional fusion feature matrix at the corresponding time point. This fusion feature vector not only reduces the data dimensionality but also strengthens the key information strongly correlated with the processing state and weakens noise and interference.
[0094] The fused feature vectors from continuous time series are stacked in chronological order to generate a spatiotemporally aligned multidimensional fused feature matrix. This matrix serves as a standardized data structure for the entire perception layer output, providing high-quality and highly consistent input for the downstream nutrient retention prediction module. This fundamentally improves the accuracy and reliability of predicting complex process states based on multi-source heterogeneous sensor data.
[0095] In this embodiment, it should be specifically noted that in the nutrient retention prediction module, the deep time-series prediction model is a time-series hybrid network based on quantile regression, and its construction and input processing specifically include:
[0096] Receive a multi-dimensional fusion feature matrix from the signal sensing and fusion module, where each column of the multi-dimensional fusion feature matrix corresponds to a fusion feature vector at a time point; obtain the initial key nutrient content vector of the raw material from the current processing batch information, and use a preset fixed time length as a sliding window to sequentially extract multiple consecutive fusion feature vectors at time points from the multi-dimensional fusion feature matrix to form a dynamic sequence.
[0097] The contents of the initial key nutrient content vector are copied according to the number of time points contained in the sliding window to form a static feature sequence with the same length in the time dimension as the dynamic sequence.
[0098] The dynamic sequence and the static feature sequence are merged along the data dimension to form an enhanced sequence sample;
[0099] The temporal hybrid network based on quantile regression includes an encoder and quantile output heads. The encoder consists of a multi-scale temporal convolutional network and a multi-head self-attention mechanism connected sequentially. It is used to encode features of augmented sequence samples, extract multi-scale local features and global dependencies, and output a context encoding vector. The multi-scale temporal convolutional network can be set to three to five layers, and its kernel size can be set in the range of 3x3 to 5x5. It is used to extract local dependency features at different time scales from augmented sequence samples. The number of heads of the multi-head self-attention mechanism can be set to four to eight. It is used to capture the global long-range dependencies between different feature dimensions in the augmented sequence samples.
[0100] The quantile output head is connected to the output of the encoder. It is configured to receive the context encoding vector and the initial key nutrient content vector, and output in parallel the quantile prediction values of the retention rates of the three key nutrients corresponding to the preset lower quantile, median quantile and upper quantile. The preset lower quantile, median quantile and upper quantile can be 0.1, 0.5 and 0.9 respectively, corresponding to the 10%, 50% (median) and 90% probability quantiles respectively.
[0101] The training process of the deep temporal prediction model and the generation process of the binary prediction results specifically include:
[0102] During the training phase, a historical dataset containing augmented sequence samples and their corresponding real key nutrient retention rate labels is used. The temporal hybrid network is trained by optimizing a quantile loss function. The value of the quantile loss function is obtained by multiplying the quantile loss values corresponding to the lower, median, and upper quantiles respectively by preset lower, median, and upper weight coefficients, and then summing them. The lower, median, and upper weight coefficients can be set within the ranges of 0.2 to 0.3, 0.5 to 0.6, and 0.2 to 0.3, respectively, to ensure that the model accurately predicts the central trend (median) while also having good calibration capabilities for the upper and lower limits of the prediction interval. For any one of the preset lower, median, or upper quantiles, the value of that preset quantile is denoted as the corresponding quantile value, and its corresponding quantile prediction value is the predicted value under that quantile value. The calculation rule for the quantile loss value of that quantile is as follows:
[0103] Calculate the difference between the true critical nutrient retention rate label and the quantile prediction value. If the difference is positive, the quantile loss value is equal to the product of the quantile value and the difference. If the difference is negative or zero, the quantile loss value is equal to the product of a difference coefficient after subtracting the quantile value and the difference obtained by subtracting the true label value from the quantile prediction value. The goal of training is to enable the quantile output head to accurately predict the retention rate value under different probability quantiles.
[0104] During the online prediction phase, the current enhanced sequence samples acquired in real time are input into the trained temporal hybrid network to obtain the corresponding three quantile prediction values; the median quantile prediction value is directly used as the mean of the key nutrient retention rate prediction.
[0105] Based on the predicted values of the lower quantile, median quantile, and upper quantile, a measure of prediction uncertainty is calculated. The calculation process is as follows:
[0106] First, calculate the first median value, which is the higher quantile predicted value minus the lower quantile predicted value; then, calculate the absolute value of the median quantile predicted value; finally, divide the first median value by the absolute value to obtain the first result, which is called the nominal prediction interval width.
[0107] Next, the second intermediate value is calculated, which is the high quantile predicted value minus twice the median quantile predicted value, plus the low quantile predicted value. Then, the second intermediate value is squared to obtain the second intermediate result, called the interval symmetry deviation. After that, a preset weight adjustment coefficient is multiplied by the interval symmetry deviation to obtain the second result, called the weighted symmetry deviation penalty term. The preset weight adjustment coefficient can be set in the range of 0.1 to 0.2 to balance the relative contributions of the nominal prediction interval width and the interval symmetry deviation in the prediction uncertainty measurement. When the prediction interval is obviously asymmetrical, the weight adjustment coefficient can amplify the influence of the interval symmetry deviation, so that the prediction uncertainty measurement can more accurately reflect the shape uncertainty of the prediction distribution.
[0108] Finally, the first result is added to the second result, and the sum is the measure of prediction uncertainty.
[0109] The median quantile prediction is used as the mean of the key nutrient retention rate prediction, and the calculated prediction uncertainty measure is used as the corresponding prediction uncertainty. These are encapsulated into a binary prediction result containing the prediction mean and the prediction uncertainty. This binary prediction result is then passed to the downstream processing parameter decision module as part of the module's state representation vector construction. The binary prediction result provides the downstream decision-making process with complete information that includes both the target prediction value and the quantification of the reliability of the prediction value. This enables the processing parameter decision module to adopt a more aggressive optimization strategy when the prediction confidence is high and to switch to a more conservative and robust strategy when the prediction uncertainty is high, thereby achieving an intelligent closed loop of perception, prediction, and decision-making.
[0110] In this embodiment, it is specifically necessary to explain the construction process of the current processing state representation vector and the robust optimization strategy on which the reinforcement learning decision model is based in the processing parameter decision module, including:
[0111] The initial composition data of the current raw materials is denoted as the raw material composition vector. The initial composition data of the current raw materials comes from the key nutrient content data obtained by sampling and analyzing the grain raw materials through the raw material composition detection unit before feeding the current batch of raw materials, or from the pre-stored standardized composition database of the raw materials of this batch. The real-time equipment parameter set is denoted as the equipment parameter vector. The real-time equipment parameter set comes from the operating parameters collected in real time by the sensors deployed on each key actuator of the grinding unit (such as the feeding frequency converter, cutting motor driver, grinding pressure servo valve, and classifier frequency converter), including but not limited to speed, current, power, pressure, and valve opening. The target product particle size index is denoted as the particle size target value. The target product particle size index comes from the product formula requirements or order process parameters issued by the upper-level production management system. The predicted mean value of the key nutrient retention rate in the binary prediction result from the nutrient retention prediction module is denoted as the prediction mean value, and its prediction uncertainty is denoted as the uncertainty value. Among them, the real-time equipment parameter set includes a real-time particle size monitoring value measured in real time by an online particle size sensor deployed at the discharge end of the grinding unit.
[0112] The raw material composition vector, equipment parameter vector, particle size target value, predicted mean, and uncertainty value are concatenated and fused. Before concatenation, the raw material composition vector, equipment parameter vector, and particle size target value are normalized separately, while the predicted mean and uncertainty value are kept at their original values. Then, the normalized raw material composition vector, normalized equipment parameter vector, normalized particle size target value, predicted mean, and uncertainty value are concatenated in this order to form a current processing state representation vector. The normalization process can use the min-max normalization method to linearly map the values in each vector or scalar to the interval between zero and one, thereby eliminating dimensional differences and ensuring that the reinforcement learning decision model can learn the influence of each state component equally. Since the predicted mean and uncertainty value are dimensionless ratios or relative values, they are kept at their original values and included in the concatenation.
[0113] The robust optimization strategy embedded in the reinforcement learning decision model is evaluated through a multi-objective weighted synthesis function. The output of this function is a comprehensive regret value, used to assess the long-term overall effect of adjusting specific parameters under a given processing state. A smaller value indicates a better overall effect. This comprehensive regret value is obtained by weighted summation of three regret values: nutrient retention target regret, granularity approach target regret, and energy consumption control target regret. The preset weight coefficients for nutrient, granularity, and energy consumption targets can be configured according to the process requirements of different products. For example, in high-nutrient retention products, the nutrient target weight coefficient can be set to 0.5 to 0.7, the granularity target weight coefficient to 0.2 to 0.3, and the energy consumption target weight coefficient to 0.1 to 0.2, to achieve balanced optimization among multiple objectives.
[0114] The calculation method for the regret of nutrient retention target is as follows:
[0115] V1. Calculate the product of the uncertainty value and a preset risk aversion coefficient that is greater than or equal to zero to obtain the risk compensation amount. The risk aversion coefficient is usually set between 0.5 and 2.0 to adjust the system's conservatism in the face of prediction uncertainty. The larger the coefficient, the more pessimistic the system's estimate of nutrient retention rate when making decisions, and the more robust the strategy.
[0116] V2. Subtract the risk compensation from the predicted mean to obtain a pessimistic estimate of nutrient retention.
[0117] V3. Calculate the difference between the preset nutrient retention target value and the pessimistic estimate of nutrient retention. If the difference is positive, take the difference as the nutrient retention target regret value; if the difference is zero or negative, take zero as the nutrient retention target regret value. This calculation logic is equivalent to taking the larger of zero and the difference as the nutrient retention target regret value. This calculation method ensures that when the pessimistic estimate of nutrient retention reaches or exceeds the target value, the regret value of the target item is zero, encouraging the reinforcement learning decision model to prioritize meeting nutrient retention requirements. When the estimate is lower than the target, a positive regret value is generated, driving the reinforcement learning decision model to optimize.
[0118] The calculation method for the granularity approaching the target regret is as follows: calculate the absolute value of the difference between the real-time granularity monitoring value and the granularity target value, and then divide this absolute value by the granularity target value. The quotient is the granularity approaching the target regret.
[0119] The calculation method for the energy consumption control target regret is as follows: The real-time equipment parameter set contains a real-time energy consumption value that reflects the current total energy consumption. A preset benchmark energy consumption value is used as the denominator, and the real-time energy consumption value is used as the numerator. The ratio between the two is the energy consumption control target regret. The benchmark energy consumption value can be set as the standard energy consumption of the equipment under rated operating conditions or the recent historical average energy consumption. This regret calculation method means that the greater the actual energy consumption exceeds the benchmark, the greater the penalty.
[0120] Ultimately, the output value of the multi-objective weighted synthesis function, namely the comprehensive regret value, is equal to the regret value of the nutrient retention target multiplied by a preset nutrient target weight coefficient, plus the regret value of the granularity approximation target multiplied by a preset granularity target weight coefficient, plus the sum of the regret value of the energy consumption control target multiplied by a preset energy consumption target weight coefficient.
[0121] The reinforcement learning decision model calculates a set of adjustment values for the processing parameters to be executed based on the embedded robust optimization strategy, specifically:
[0122] First, the decision calculation is performed: the current processing state representation vector is input into the reinforcement learning decision model, which processes it through its internal policy network and outputs a raw action vector. Each dimension of the raw action vector corresponds to the proposed adjustment direction and magnitude of a processing parameter to be adjusted in the normalized space. The policy network is usually a deep neural network composed of fully connected layers, whose parameters are obtained through offline training and can map high-dimensional state representations into continuous parameter adjustment proposals.
[0123] Then, during the training and online decision exploration phase of the reinforcement learning decision model, an adaptive exploration mechanism is introduced. This mechanism uses the uncertainty value as the basis for dynamically adjusting the exploration intensity, specifically:
[0124] When this reinforcement learning decision model outputs the original action vector, the variance of the random exploration noise superimposed on it is proportional to the uncertainty value. That is, the uncertainty value is taken as input, multiplied by a preset exploration coefficient, and then added to a preset minimum noise variance. The result is the variance of the random exploration noise superimposed on the current decision cycle. The exploration coefficient can be set between 0.1 and 0.3, and the minimum noise variance can be set between 0.01 and 0.05 to ensure that even when the prediction uncertainty is extremely low, the decision still retains basic exploration ability and avoids the policy from getting trapped in local optima. Thus, the higher the uncertainty value, the larger the added exploration noise variance, and the wider the range of exploration in the parameter space of this reinforcement learning decision model. Conversely, the smaller the exploration noise variance, the more the reinforcement learning decision model tends to use known effective policies.
[0125] Finally, post-processing and output are performed: the original action vectors obtained from the decision calculation steps are post-processed to generate a set of adjustment values for the processing parameters to be executed. The post-processing specifically includes:
[0126] The first scaling sub-step: Multiply the value of each dimension of the original motion vector by the absolute value of the maximum allowable adjustment amount of the corresponding processing parameter to be adjusted, to obtain the actual physical adjustment amount of each parameter; the absolute value of the maximum adjustment amount needs to be preset according to the physical performance and process safety range of each actuator (such as feed frequency converter, cutting motor driver, grinding pressure servo valve, classifier frequency converter). For example, the maximum absolute value of the cutting machine speed adjustment amount may be ±100 revolutions per minute.
[0127] The second feasibility verification sub-step: Based on the preset physical lower limit and upper limit values of each processing parameter to be adjusted, the new parameter value obtained by superimposing the actual physical adjustment amount onto the current real-time equipment parameter value is verified and truncated to ensure that each new parameter value is within the allowable range defined by its corresponding physical lower limit and upper limit value. This step is crucial to prevent the command from exceeding the limit. If the calculated new parameter value exceeds the upper limit, it is truncated to the upper limit value; if it is lower than the lower limit, it is truncated to the lower limit value, thereby ensuring the absolute safety of the control command.
[0128] The third smoothness constraint sub-step: Calculate the rate of change of the difference between the actual physical adjustment amount obtained in the current decision cycle and the actual physical adjustment amount of the corresponding parameter used in the previous decision cycle, and ensure that the rate of change does not exceed the maximum rate of change limit preset for the parameter; the rate of change is obtained by dividing the absolute value of the difference between the current adjustment amount and the adjustment amount of the previous cycle by the length of the decision cycle; the maximum rate of change limit is set according to the mechanical inertia of the equipment and the requirements of process stability, in order to avoid the command change being too drastic in adjacent control cycles, thereby protecting the equipment and maintaining the stability of the production process;
[0129] After the above-mentioned first scaling sub-step, second feasibility verification sub-step, and optional third smoothness constraint sub-step, the resulting set containing the actual physical adjustment amounts of all processing parameters to be adjusted is a set of processing parameter adjustment amounts to be executed. This set is a structured data body that clearly lists the specific adjustment values of each processing parameter to be adjusted (such as main roller speed, feeder frequency, classifier frequency) within the current control cycle, and can be directly sent to the underlying actuator.
[0130] In this embodiment, it is specifically necessary to explain that in the instruction optimization and collaborative control module, the rolling time-domain optimization calculation is performed by combining the dynamic characteristic constraints of each actuator of the grinding unit with the preset global optimization target, specifically including:
[0131] A linear parameter time-varying state-space model for each actuator of the grinding mill is established as a prediction model. This model is built upon offline system identification based on historical operating data and step response tests of actuators such as the feed inverter, cutting motor driver, grinding pressure servo valve, and classifier inverter. System identification can employ subspace identification or prediction error methods. By applying step control signals of different amplitudes to each actuator and collecting their dynamic response data, the state-space model parameters at different operating points (e.g., different speeds and loads) are identified. Then, interpolation or fitting methods are used to establish the functional relationship between the model parameters and scheduling parameters, thus forming a complete linear parameter time-varying model. This model describes the dynamic relationship from control commands to equipment responses, and its specific structure is as follows:
[0132] The linear parameter time-varying state-space model comprises the following components: a system state vector, whose elements include dynamic variables within the actuators such as the speed of each motor, winding current, and valve opening; a control input vector, whose elements include frequency setpoints, voltage setpoints, or digital setpoint commands sent to each actuator; a system output vector, whose elements include measurable equipment operating parameters, including actual speed and working pressure; a scheduling parameter, which is a real-time measurable parameter reflecting the current operating point and load rate of the system; the scheduling parameter can be selected as the ratio of the current load current of the main grinding motor to the set speed, or the ratio of the total system power to the no-load power, and its value varies continuously between zero and one, used to adjust the model parameters online so that the prediction model can adapt to the current operating conditions of the system; a measurable disturbance vector, used to characterize known external disturbances, such as load disturbances introduced by changes in raw material hardness; and an output matrix, a constant matrix that maps the system state vector to the system output vector.
[0133] The dynamic relationship of the linear parameter time-varying state-space model is as follows: at any discrete time point k, the system state vector at the next time k+1 is equal to the system matrix at the current time k multiplied by the system state vector at the current time k, plus the control input matrix at the current time k multiplied by the control input vector at the current time k, and finally the measurable disturbance vector at the current time k; at the same time, the system output vector at the current time k is equal to the output matrix multiplied by the system state vector at the current time k; where the specific values of the system matrix and the control input matrix are determined by the values of the scheduling parameters at the current time k.
[0134] This linear parametric time-varying state-space model can more accurately describe the weak coupling relationship between actuators and the dynamic characteristics of the system under varying operating conditions.
[0135] The specific process of generating the collaborative reference trajectory is as follows: receive the set of processing parameter adjustment amounts to be executed from the processing parameter decision module, and denote this set as the expected adjustment amount set; obtain the actual measured values of the operating parameters of each device at the current moment. The actual measured values are derived from the data collected in real time by the sensors deployed on the grinding unit, which constitute the current output vector;
[0136] A coordination timing matrix is introduced, where each element defines the start time offset of an actuator in the grinding unit to begin responding to its corresponding single desired adjustment in the set of desired adjustment, in order to describe the sequential coordination timing of the actions of multiple mechanisms.
[0137] Based on the coordinated timing matrix, a corresponding shape function is defined for each expected adjustment quantity in the set of expected adjustment quantities. This shape function is an S-shaped function with the future discrete time point and the starting time offset corresponding to the expected adjustment quantity as independent variables, and its function value smoothly transitions between zero and one. The specific form of the S-shaped function can be the Sigmoid function or its variant. Its key parameter, rise time, can be set according to the mechanical inertia of the corresponding actuator. For example, for a grinding motor with high inertia, the rise time of the shape function of the corresponding adjustment quantity can be set to 100 to 200 milliseconds. For a fast-responding feeding frequency converter, the rise time can be set to 20 to 50 milliseconds to ensure smooth and coordinated action. The shape functions corresponding to all expected adjustment quantities constitute a shape function vector.
[0138] The generation calculation of the cooperative reference trajectory is as follows: For each future discrete time point in the future prediction time domain, firstly, each expected adjustment amount in the expected adjustment amount set is multiplied element-wise with the corresponding position in the shape function vector at that time point to obtain an intermediate adjustment amount vector that has undergone time and temporal modulation; then, this intermediate adjustment amount vector is added element-wise to the current output vector, and the resulting vector is the cooperative reference trajectory point at that future time point; by traversing all future discrete time points in the future prediction time domain, the complete cooperative reference trajectory can be obtained.
[0139] Based on a linear parameter time-varying state-space model and a cooperative reference trajectory, a rolling time-domain optimization problem is constructed, the specific structure of which is as follows:
[0140] The optimization variable is defined as a sequence of control input vectors at consecutive future discrete time points of the future control time domain length, where the control time domain length is a preset positive integer;
[0141] The objective function is to minimize a comprehensive cost, which is a weighted sum of the following four component costs:
[0142] The first component is the tracking error cost, which is calculated as follows: Within the predicted time domain length, for each time point, calculate the difference vector between the system output vector predicted by the linear parameter time-varying state-space model and the cooperative reference trajectory at that time point. Then, calculate the weighted square norm of this difference vector, where the weighting matrix is called the output error weight matrix. The output error weight matrix is usually a diagonal matrix, and the values of its diagonal elements are set relative to the control accuracy requirements of each output variable (such as rotational speed and pressure). For variables that require precise tracking (such as the classifier rotational speed corresponding to the finished particle size), the corresponding weight coefficient can be set to 0.5 to 1. For variables that allow for a certain tracking error (such as the feeding speed), the corresponding weight coefficient can be set to 0.1 to 0.3. Finally, sum these weighted square norm values for all time points.
[0143] The second component is the control increment penalty cost, which is calculated as follows: Within the future control time domain, for each time point, calculate the change vector between the control input vector at that time point and the control input vector at the previous time point. Then, calculate the weighted square norm of this change vector, where the weighting matrix is called the control increment weight matrix. The control increment weight matrix is also a diagonal matrix, and its element values are used to limit the rate of change of each control quantity (such as frequency setpoint and valve position command). The larger the value, the stricter the restriction on the rate of change. Typical values are between 0.01 and 0.1 to prevent frequent large-amplitude movements of the actuator. Finally, sum these weighted square norm values for all time points.
[0144] The third component is the control energy penalty cost, which is calculated as follows: Within the future control time domain, for each time point, calculate the weighted square norm of the control input vector at that time point. The weighting matrix is called the control weight matrix. The control weight matrix is used to indirectly optimize energy consumption. Its element values are proportional to the rated power or typical energy consumption of each actuator. For example, the weight coefficient corresponding to a high-power grinding motor can be set to 0.1, while the weight coefficient corresponding to a low-power valve can be set to 0.01. Finally, sum these weighted square norm values for all time points.
[0145] The fourth sub-item is the cost of cooperative error penalty, which is calculated as follows: Within the predicted time domain length, for each time point, firstly, based on the preset ideal cooperative relationship between each actuator, calculate the expected cooperative output vector for that time point; then calculate the difference vector between the system output vector predicted by the linear parameter time-varying state-space model and the expected cooperative output vector; next, calculate the weighted square norm of this difference vector, where the weighting matrix is called the cooperative weight matrix; the off-diagonal elements of the cooperative weight matrix can be set to zero, and the diagonal elements are used to penalize the cooperative deviation between the outputs of different actuators. For example, the cooperative error weight used to force the feeding speed and the main grinding speed to maintain a certain proportional relationship can be set to 0.2 to 0.5; finally, sum these weighted square norm values for all time points.
[0146] The constraints of this optimization problem include:
[0147] First, control input amplitude constraints: the value of each element in the control input vector at each time point in the future control time domain must be between the preset minimum control command value and the maximum control command value allowed by the actuator corresponding to that element; the minimum and maximum control command values are determined by the actuator's datasheet or safe operating procedures, for example, the frequency setpoint limit of the frequency converter is 0 to 50 Hz, and the control current of the servo valve is 4 to 20 mA;
[0148] Second, control input change rate constraint: the value of each element in the control input change vector at each time point in the future control time domain must be between the preset minimum change rate and maximum change rate allowed by the actuator corresponding to that element; the minimum and maximum change rates are used to protect the equipment and prevent sudden changes in commands, and their values are set according to the maximum acceleration or maximum flow rate change rate of the actuator. For example, the motor speed change rate can be limited to ±100 revolutions per second.
[0149] Third, system output safety constraints: for each element in the system output prediction vector at each time point in the future prediction time domain, the value must be between the preset lower and upper safety limits allowed by the corresponding device operating parameters.
[0150] Fourth, state dynamics equality constraints: that is, the system state vector and the control input vector must follow the dynamic relationship defined by the time-varying state-space model with linear parameters;
[0151] In each control cycle, the rolling time-domain optimization problem is transformed online into a constrained quadratic programming problem and solved to obtain the optimal control input vector sequence in the future control time domain.
[0152] Generate a sequence of coordinated control instructions to drive each actuator in the short-term future, and distribute the coordinated control instructions corresponding to the current time to the corresponding actuators, specifically including:
[0153] First, the sequence extraction and instruction distribution are performed: The first control input vector is extracted from the optimal control input vector sequence obtained by solving the rolling time-domain optimization problem, serving as the current moment's collaborative control instruction to be issued in the current control cycle. Simultaneously, this complete optimal control input vector sequence is cached for use in the next control cycle when solving a new rolling time-domain optimization problem, providing an initial guess for the optimization algorithm to start warmly, thus improving solution efficiency. Specifically, when constructing the optimization problem in the next cycle, the cached sequence is shifted backward by one time step (i.e., discarding the first element and adding an element identical to the last element or a zero element at the end) as the initial iteration point for the decision variables. This significantly reduces the iteration count of optimization solvers such as the interior-point method or the effective set method, shortening the solution time by 30% to 50%, meeting real-time control requirements. The current moment's collaborative control instruction is then distributed to the corresponding actuators in the grinding unit via the corresponding communication interface and protocol. These actuators include the feeder frequency converter, the cutting motor driver, the grinding pressure servo valve, and the classifier frequency converter.
[0154] Next, feedback correction and rolling optimization are performed: When the next control cycle arrives, the latest measured values of the equipment operating parameters from the sensors are received and recorded as the latest measured values. Using these latest measured values, the system state vector estimate of the linear parameter time-varying state space model is updated through the state estimator. The state estimator can use a Kalman filter or an extended Kalman filter, whose process noise covariance matrix and measurement noise covariance matrix can be obtained through historical data identification. This is used to obtain the optimal estimate of the internal state of the system (such as motor torque and unknown disturbances) in the presence of measurement noise and model errors. The updated system state vector estimate is then used as the new initial state. Combined with the latest received set of processing parameter adjustments to be executed, the steps of generating the cooperative reference trajectory, constructing and solving the rolling time-domain optimization problem are re-executed to achieve closed-loop feedback correction, thereby generating and issuing new cooperative control commands for the current moment. The control cycle can be set according to the process response speed, with a typical value of fifty to two hundred milliseconds. Within this cycle, the entire process of state estimation, optimization solution, and command issuance must be completed.
[0155] The process of sequence extraction and instruction distribution, feedback correction and rolling optimization is carried out in each control cycle, thereby continuously transforming the adjustment target of the upper-level decision into smooth, coordinated actual driving instructions that meet all constraints.
[0156] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0157] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0158] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0159] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0160] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0161] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0162] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. An automated integrated processing system for mixed grain nutritional formula powder, characterized in that, Specifically comprising: The signal perception and fusion module, the nutrient retention prediction module, the processing parameter decision module, and the instruction optimization and collaborative control module, wherein The signal perception and fusion module: in response to a processing start instruction of the grinding machine group, synchronously collecting original signals from multiple types of sensors arranged in key physical action areas of the grinding machine group, and performing synchronous, alignment, and feature extraction processing on the original signals to generate a multi-dimensional fusion feature matrix that is spatio-temporally aligned; The nutrient retention prediction module: receiving the multi-dimensional fusion feature matrix from the signal perception and fusion module, and inputting the multi-dimensional fusion feature matrix into a deep time series prediction model, which performs time series analysis and outputs a binary prediction result containing a predicted mean value of key nutrient retention rate and its corresponding prediction uncertainty; The processing parameter decision module: used for fusing current raw material initial ingredient data, real-time equipment parameter set, target product particle size index, and the binary prediction result from the nutrient retention prediction module to form a current processing state representation vector, and inputting the current processing state representation vector into a reinforcement learning decision model, which calculates a set of processing parameter adjustment amounts to be executed according to an embedded robust optimization strategy; The instruction optimization and collaborative control module: used for receiving the set of processing parameter adjustment amounts to be executed from the processing parameter decision module, and taking the set as a set target, combining dynamic characteristic constraints of each execution mechanism of the grinding machine group and a preset global optimization target, performing a rolling horizon optimization calculation to generate a collaborative control instruction sequence for driving each execution mechanism in a future short time domain, and distributing the collaborative control instruction corresponding to the current time to the corresponding execution mechanism of the grinding machine group.
2. The integrated processing system for automated production of nutritional coarse cereal powder according to claim 1, characterized in that: In the signal perception and fusion module, in response to a processing start instruction, a synchronous collection instruction is sent to multiple types of sensors arranged in key physical action areas of the grinding machine group, and original signals returned by each sensor according to a unified clock are synchronously received, wherein the key physical action areas include a cutting zone inlet, an impact crushing chamber side wall, and a gap outlet between grinding pairs, and the multiple types of sensors include acoustic emission sensors for capturing particle breakage transient stress waves, vibration accelerometers for monitoring rotor imbalance and collision characteristics, infrared thermal imagers for perceiving material apparent temperature distribution, and multispectral cameras for acquiring material apparent spectral information.
3. The integrated processing system for automated production of nutritional coarse cereal flour according to claim 2, characterized in that: The specific operation of synchronously, aligning, and feature extracting the original signals is as follows: Each original signal is preprocessed including noise reduction, detrending, and amplitude normalization to obtain each preprocessed signal; Based on the preset sensor physical position coordinates and material flow rate information, the spatio-temporally aligned signals corresponding to each sensor are calculated to generate spatio-temporally aligned signals corresponding to each sensor; For the spatio-temporally aligned signals corresponding to the acoustic emission sensors and the vibration accelerometers, continuous wavelet transform method is used to extract time-frequency energy distribution values at multiple preset scales from the spatio-temporally aligned signals; The statistical moments including mean, standard deviation, skewness, kurtosis, and the rate of change of the time-space alignment signal in a sliding time window are extracted from the time-space alignment signal corresponding to the infrared thermal imager and the multispectral camera; all the above-mentioned time-frequency energy distribution values, statistical moments and rate of change extracted from each time-space alignment signal are collectively defined as a plurality of characteristic values extracted from the signal; All the characteristic values extracted from one time-space alignment signal are arranged in a predetermined order to form a characteristic vector; The characteristic vectors extracted from the time-space alignment signals corresponding to all sensors at the same time are spliced to form an initial total characteristic vector.
4. The integrated processing system for automated processing of nutritional flour according to claim 3, wherein: The specific process of generating the multi-dimensional fusion feature matrix in space-time alignment is: Calculate the covariance matrix of the initial total characteristic vector in a sliding time window; Perform eigenvalue decomposition on the covariance matrix, select the characteristic vectors corresponding to the first set number of maximum eigenvalues to form a principal component projection matrix with a row number equal to the set number and a column number equal to the dimension of the initial total characteristic vector; An adaptive diagonal weight matrix is constructed, which is a square matrix with a row number and a column number equal to the dimension of the initial total characteristic vector, and only the elements on the main diagonal have non-zero values; wherein the value of each diagonal element of the adaptive diagonal weight matrix is determined by the characteristic value at the same position in the initial total characteristic vector as the diagonal element, which is dynamically calculated based on the signal-to-noise ratio estimate value of the characteristic value in the adjacent time window, and the Mahalanobis distance calculated from the characteristic vector containing the characteristic value and the corresponding characteristic vector in the historical reference state, through a preset nonlinear mapping function; The initial total characteristic vector, the principal component projection matrix and the adaptive diagonal weight matrix are input into the adaptive weighted fusion function for calculation, and the calculation process is as follows: first, multiply each element on the main diagonal of the adaptive diagonal weight matrix with the element at the same position in the initial total characteristic vector to obtain an intermediate vector; then, transpose the principal component projection matrix to obtain the transposed principal component projection matrix; finally, perform matrix multiplication operation on the transposed principal component projection matrix and the intermediate vector to obtain a fusion feature vector with a row number equal to the set number, which is a column vector of the multi-dimensional fusion feature matrix at the corresponding time point; The fusion feature vectors on the continuous time sequence are stacked in time sequence to finally generate the multi-dimensional fusion feature matrix in space-time alignment.
5. The integrated processing system for automated processing of nutritional flour according to claim 4, wherein: In the nutrition retention prediction module, the deep time series prediction model is a time series hybrid network based on quantile regression, and its construction and input processing specifically include: Receive the multi-dimensional fusion feature matrix from the signal perception and fusion module; obtain the initial key nutrient content vector of the raw material from the current processing batch information, and sequentially cut the fusion feature vectors at continuous time points from the multi-dimensional fusion feature matrix to form a dynamic sequence with a preset fixed time length as a sliding window; Copy the contents of the initial key nutrient content vector according to the number of time points contained in the sliding window to form a static feature sequence consistent with the dynamic sequence in the time dimension; Merge the dynamic sequence and the static feature sequence in the data dimension direction to form an enhanced sequence sample; The quantile regression-based time series hybrid network includes an encoder and a quantile output head, the encoder is composed of a multi-scale time series convolution network and a multi-head self-attention mechanism connected in sequence, and is configured to encode the enhanced sequence sample, extract multi-scale local features and global dependency relationships, and output a context encoding vector; The quantile output head is connected to the output end of the encoder, and is configured to receive the context encoding vector and the initial key nutrient content vector, and output three quantile prediction values of the key nutrient retention rate corresponding to the preset low quantile, median quantile and high quantile in parallel.
6. The integrated processing system for automated production of nutritional flour according to claim 5, wherein: The training of the deep time series prediction model and the generation process of the binary prediction result specifically include: In the training stage, the time series hybrid network is trained by optimizing a quantile loss function using a historical data set containing enhanced sequence samples and their corresponding real key nutrient retention rate labels. The value of the quantile loss function is obtained by adding the low weight coefficient, the median weight coefficient and the high weight coefficient to the respective quantile loss values of the low quantile, the median quantile and the high quantile. For any one of the preset low quantile, the median quantile or the high quantile, the value of the preset quantile is denoted as the corresponding quantile value, and the corresponding quantile prediction value is the predicted value at the quantile value. The calculation rule of the quantile loss value for the quantile is as follows: Calculate the difference between the real key nutrient retention rate label and the quantile prediction value. If the difference is positive, the quantile loss value is equal to the product of the quantile value and the difference. If the difference is negative or zero, the quantile loss value is equal to the product of the difference coefficient obtained by subtracting the quantile value from one and the difference between the quantile prediction value and the real label value; In the online prediction stage, the current enhanced sequence sample obtained in real time is input into the trained time series hybrid network to obtain the corresponding three quantile prediction values. The median quantile prediction value is directly used as the predicted mean value of the key nutrient retention rate. Based on the low quantile prediction value, the median quantile prediction value and the high quantile prediction value, the prediction uncertainty measure is calculated, and the calculation process is as follows: First, calculate the first intermediate value, which is the high quantile prediction value minus the low quantile prediction value. Then, calculate the absolute value of the median quantile prediction value. Then, divide the first intermediate value by the absolute value to obtain the first result, which is called the nominal prediction interval width. Second, calculate the second intermediate value, which is the high quantile prediction value minus twice the median quantile prediction value, plus the low quantile prediction value. Then, square the second intermediate value to obtain the second intermediate result, which is called the interval symmetric deviation degree. Then, multiply a preset weight adjustment coefficient by the interval symmetric deviation degree to obtain the second result, which is called the weighted symmetric deviation penalty term. Finally, the first result is added to the second result, and the sum value is the prediction uncertainty measure; The median quantile prediction value is taken as the prediction mean of the key nutrient retention rate, and the calculated prediction uncertainty measure is taken as the corresponding prediction uncertainty, and a binary tuple prediction result containing the prediction mean and the prediction uncertainty is packaged.
7. The integrated processing system for automated processing of nutritional flour according to claim 6, wherein: In the processing parameter decision module, the current processing state representation vector is formed, and the specific construction process of the robust optimization strategy relied on by the reinforcement learning decision model includes: The current raw material initial component data is recorded as a raw material component vector, the real-time equipment parameter set is recorded as an equipment parameter vector, the target product particle size index is recorded as a particle size target value, the key nutrient retention rate prediction mean in the binary tuple prediction result from the nutrient retention prediction module is recorded as a prediction mean, and its prediction uncertainty is recorded as an uncertainty value; wherein the real-time equipment parameter set includes a particle size real-time monitoring value measured by an online particle size sensor deployed at the discharge end of the grinding machine group in real time; The raw material component vector, the equipment parameter vector, the particle size target value, the prediction mean and the uncertainty value are spliced and fused, wherein before splicing, the raw material component vector, the equipment parameter vector and the particle size target value are normalized respectively, the prediction mean and the uncertainty value are kept as original values, then the normalized raw material component vector, the normalized equipment parameter vector, the normalized particle size target value, the prediction mean and the uncertainty value are connected in this order to form a current processing state representation vector; The robust optimization strategy embedded in the reinforcement learning decision model has an evaluation mechanism realized by a multi-objective weighted comprehensive function, and the output value of the multi-objective weighted comprehensive function is a comprehensive regret value, which is obtained by weighted summation of a nutrient retention target regret, a particle size proximity target regret and an energy consumption control target regret; wherein, The calculation method of the nutrient retention target regret is: V1, calculate the product of the uncertainty value and a preset risk aversion coefficient greater than or equal to zero to obtain a risk compensation amount; V2, subtract the risk compensation amount from the prediction mean to obtain a nutrient retention pessimistic estimate value; V3, calculate the difference between the preset nutrient retention target value and the nutrient retention pessimistic estimate value; if the difference is positive, take the difference as the nutrient retention target regret; if the difference is zero or negative, take zero as the nutrient retention target regret; The calculation method of the particle size proximity target regret is: calculate the absolute value of the difference between the particle size real-time monitoring value and the particle size target value, and then divide the absolute value by the particle size target value, and the quotient is the particle size proximity target regret; The calculation method of the energy consumption control target regret is: the real-time equipment parameter set includes a real-time energy consumption value reflecting the current total energy consumption, a preset benchmark energy consumption value is taken as the denominator, and the real-time energy consumption value is taken as the numerator, and the ratio of the two is calculated, and the ratio is the energy consumption control target regret; Finally, the output value of the multi-objective weighted comprehensive function, i.e. the comprehensive regret value, is equal to the sum of the product of the nutrient retention target regret by a preset nutrient target weight coefficient, the product of the particle size proximity target regret by a preset particle size target weight coefficient, and the product of the energy consumption control target regret by a preset energy consumption target weight coefficient.
8. The integrated processing system for automated production of nutritional flour according to claim 7, wherein: The reinforcement learning decision model calculates a set of processing parameter adjustment amounts to be executed according to an embedded robustness optimization strategy, specifically: First, perform decision calculation: input the current processing state feature vector into the reinforcement learning decision model, which processes it through its internal policy network to output an original action vector; Then, in the training and online decision exploration phase of the reinforcement learning decision model, an adaptive exploration mechanism is introduced, which uses the uncertainty value as the basis for dynamically adjusting the exploration intensity, specifically: When the reinforcement learning decision model outputs the original action vector, the variance of the random exploration noise superimposed is proportional to the uncertainty value, i.e. the uncertainty value is multiplied by a preset exploration coefficient and then added to a preset minimum noise variance to obtain the variance of the random exploration noise superimposed in the current decision cycle; Finally, perform post-processing and output: post-process the original action vector obtained in the decision calculation step to generate the set of processing parameter adjustment amounts to be executed, which includes specifically: First scaling sub-step: multiply the value of each dimension of the original action vector by the maximum adjustment amount absolute value allowed for the corresponding parameter to be adjusted to obtain the actual physical adjustment amount of each parameter; Second feasibility check sub-step: check and truncate the new parameter value obtained by superimposing the actual physical adjustment amount on the current real-time device parameter value according to the preset physical lower and upper limit values of each parameter to be adjusted; Third smoothness constraint sub-step: calculate the difference change rate between the actual physical adjustment amount obtained in the current decision cycle and the corresponding parameter actual physical adjustment amount used in the previous decision cycle; After the above first scaling sub-step, second feasibility check sub-step, and optional third smoothness constraint sub-step, the set containing all the specific actual physical adjustment amounts of the parameters to be adjusted is obtained, which is the set of processing parameter adjustment amounts to be executed.
9. The integrated processing system for automated production of nutritional flour according to claim 8, wherein: In the instruction optimization and collaborative control module, the dynamic characteristic constraints of each actuator of the grinding unit and the preset global optimization objective are combined to perform rolling horizon optimization calculation, specifically including: A linear parameter time-varying state space model of each actuator of the grinding unit is established as a prediction model, and the establishment of the linear parameter time-varying state space model is based on offline system identification of historical operation data and step response test of the feed variable frequency drive, cutting motor drive, grinding pressure servo valve, and grading machine variable frequency drive; The specific process of generating the cooperative reference trajectory is that: receiving the set of machining parameter adjustment amounts to be executed from the machining parameter decision module, denoted as a set of expected adjustment amounts; obtaining actual measurement values of operating parameters of each device at the current time, which are derived from data collected by sensors deployed on the grinding machine group in real time, and constitute a current output vector; A cooperative timing matrix is introduced, each element of which defines a start time offset of an actuator in the grinding machine group to start responding to a single expected adjustment amount corresponding to it in the set of expected adjustment amounts, to describe the timing of the actions of multiple mechanisms; Based on the cooperative timing matrix, a corresponding shape function is defined for each expected adjustment amount in the set of expected adjustment amounts, which is a S-shaped function with future discrete time points and the start time offset corresponding to the expected adjustment amount as independent variables, and all shape functions corresponding to the expected adjustment amounts constitute a shape function vector; The generation of the cooperative reference trajectory is calculated as follows: for each future discrete time point in the future prediction time domain, first, each expected adjustment amount in the set of expected adjustment amounts is multiplied by the shape function output value at the time point in the corresponding position of the shape function vector to obtain an intermediate adjustment amount vector modulated by time and timing; then, the intermediate adjustment amount vector is added to the current output vector element by element, and the resulting vector is the cooperative reference trajectory point at the future time point; by traversing all future discrete time points in the future prediction time domain, the complete cooperative reference trajectory can be obtained; Based on the linear parameter time-varying state space model and the cooperative reference trajectory, a rolling horizon optimization problem is constructed, which is converted into a constrained quadratic programming problem at each control period and solved to obtain the optimal control input vector sequence in the future control time domain.
10. The integrated processing system for automated production of nutritional flour according to claim 9, wherein: The generation of the cooperative control instruction sequence for driving each actuator in the future short time domain and the distribution of the cooperative control instruction corresponding to the current time to the corresponding actuator include: First, sequence extraction and instruction distribution: the first control input vector in the optimal control input vector sequence obtained by solving the rolling horizon optimization problem is extracted as the cooperative control instruction at the current time to be issued in the current control period; at the same time, the complete optimal control input vector sequence is cached; the cooperative control instruction at the current time is distributed to the corresponding actuator in the grinding machine group through the corresponding communication interface and protocol; Then, feedback correction and rolling optimization are performed: when the next control cycle comes, the latest device operating parameter measurements from sensors are received, denoted as latest measurements; the latest measurements are used to update the system state vector estimation of the linear parameter-varying state-space model by a state estimator; and the updated system state vector estimation is used as the new initial state to re-perform the steps of the collaborative reference trajectory generation, the rolling horizon optimization problem construction and solving, in combination with the latest received set of processing parameter adjustment amounts to be executed, to achieve closed-loop feedback correction, thereby generating and issuing new current-time collaborative control instructions; The sequence extraction and instruction distribution, feedback correction and rolling optimization process is performed in each control cycle.