Method for predicting micro-contextual risk causal relationships of infant formula production sequence polynucleotide sequences

By extracting similar sequence segments and eliminating interfering information during the milk powder production process, a micro-scenario model was established, which solved the problems of accuracy in milk powder quality control and causal relationship mining, and realized accurate prediction of finished milk powder quality and real-time risk monitoring of the production process.

CN115796343BActive Publication Date: 2026-05-12BEINMATE (HANGZHOU) FOOD RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEINMATE (HANGZHOU) FOOD RES INST CO LTD
Filing Date
2022-11-17
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing milk powder sequence prediction methods are not very accurate in multi-step prediction, cannot accurately determine key control points and process parameters, and cannot explore the causal relationships between multiple time series, resulting in crude and inaccurate milk powder quality control.

Method used

By extracting similar sequence segments from the infant formula production process, eliminating interfering information, conducting causal analysis, establishing relevant micro-scenes, mining key control points and parameters, and using sliding time slices to expand the dataset, clustering, transfer entropy calculation, and ridge regression models for prediction.

Benefits of technology

It improves the accuracy of milk powder finished product quality prediction, enables real-time monitoring and adjustment of production processes, enhances product quality, reduces the impact of interfering information, and improves the accuracy of multi-step prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796343B_ABST
    Figure CN115796343B_ABST
Patent Text Reader

Abstract

The application provides a method for predicting the micro-scene risk causal relationship of multiple sequences in an infant formula production sequence, comprising: S1, selecting sequences in process indicators of infant formula, extracting sequence data, forming M different source sequence data as a data set D; S2, performing stationary processing on the sequence data in the data set D, expanding the data set D using a sliding time slice, forming a new data set Data, and the data set Data containing target sequences and variable sequences. The application extracts similar sequence segments and excludes interference sequence segments from the related sequences of raw material indicators, production process parameters and product indicators of infant formula, selects the most relevant sequence segment to the target prediction indicator through causal analysis, establishes a related micro-scene for prediction analysis, mines out the key control points and key process control parameters that have the greatest impact on the quality of the infant formula product, and finds out the relatively optimal value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of quality and safety risk monitoring of infant formula milk powder, specifically to a method for predicting the causal relationship of micro-scenario risks in a multivariate sequence of infant formula milk powder production. Background Technology

[0002] Besides the influence of raw materials and processing methods, identifying key processes and controlling appropriate process parameters are crucial for the quality of infant formula products. Currently, most infant formula companies manage process control parameters by setting intermediate averages under national standards; a few companies, based on past research experience, use limited single-factor experiments to determine critical control points and related process control parameters. While these methods can easily and easily ensure product quality, they are relatively crude, unable to precisely identify critical control points, and cannot set relatively optimal process control parameters.

[0003] Compared to existing milk powder sequence prediction algorithms:

[0004] (1) Unit time series forecasting method: ARIMA model, autoregressive moving average model, first transforms the non-stationary time series into a stationary time series, and then regresses the dependent variable only on its lagged values ​​and the present and lagged values ​​of the random error term. This method is only applicable to the forecasting of a single series and can perform multi-step forecasting.

[0005] (2) LSTM: Long Short-Term Memory Network, suitable for learning and predicting single long-term sequences; currently, it has a high accuracy in single-step prediction, but it is not accurate for multi-step prediction.

[0006] (3) Multivariate time series forecasting method (VAR model): namely, vector autoregression model, is a commonly used economic model. The VAR model uses all historical information in the model to regress all variables on several subsequent time windows. It can perform multi-step forecasting and risk forecasting, but it cannot uncover the causal relationship between the various series.

[0007] (4) Traditional statistical regression method (Naive). This method uses all variable sequences as x and the shifted target sequence as y for prediction. It is simple to operate and can perform multi-step prediction, but its accuracy is low.

[0008] For milk powder time series, there are extensive and complex relationships among multiple factors, including interconnections and causal relationships, and the interactions between these complex factors are not clear; at the same time, these time series contain various complex information, among which a lot of interfering information affects the accuracy of prediction.

[0009] Currently, there are two major shortcomings in the forecasting methods designed and researched: in multi-step forecasting of time series, the accuracy is not high, so we need to reduce the influence of interfering information and improve the accuracy; in the forecasting and analysis of multivariate time series, existing methods cannot uncover the causal relationships between the various series. For example, multivariate time series forecasting methods can only predict risks but cannot explain the causes of anomalies. Summary of the Invention

[0010] To address the shortcomings of existing technologies, the present invention aims to provide a method for predicting the causal relationships of micro-scenario risks in the production sequence of infant formula milk powder. This invention targets the relevant sequences of raw material indicators, production process parameters, and finished product indicators of infant formula milk powder. It extracts similar sequence segments and eliminates interfering sequence segments. Through causal analysis, it selects the sequence segments most relevant to the target prediction indicator, establishes relevant micro-scenarios for predictive analysis, scientifically identifies the key control points and key process control parameters that have the greatest impact on the quality of the finished milk powder, and finds their relatively optimal values.

[0011] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:

[0012] A method for predicting the causal relationships of micro-scenario risks in a multivariate sequence of infant formula production, characterized by:

[0013] The prediction method includes the following steps:

[0014] S1. Select the sequences from the main process indicators of infant formula milk powder, extract the sequence data, and form M sequence data from different sources, which become dataset D;

[0015] S2. Stabilize the sequence data in dataset D, and expand dataset D using a sliding time slice to form a new dataset Data, which contains the target sequence and the variable sequence.

[0016] S3. Clustering is performed by extracting the most distinguishable subsequences from the variable sequence as the feature sequences of that segment, denoted as Rs sequences; and the extracted Rs sequences are then clustered, denoted as Clu; Clu m,i The cluster center line representing the i-th cluster of variables m, and defined by Clu m,i A cluster number is denoted as cnt(Clu) to uniquely identify a category. m ); Clu m,i This refers to the Rs sequence obtained above;

[0017] S4. Randomly combine the Clus of different variables m to form a sequence cube (Scube), where Scubek represents the k-th sequence cube, and each sequence is denoted as S. v,tv represents the v-th variable sequence, and t represents the t-th subsequence in the variable sequence;

[0018] S5. Filter the extracted sequence cubes (Scubes) as described above. Filter each Scube to obtain the relevant micro-scenes M-scenes. For each M-scene, define a scene transition entropy to describe the causal impact of the scene on the target indicator sequence.

[0019] S6. Perform similarity judgment, match the most relevant micro-scenes, and use the micro-scene model to predict risks.

[0020] Further: In step S2, the dataset D is expanded by determining a sampling time slice T, the length of which is the total time of the milk powder production process. Then, the time slice is divided into time slices with a length of 60 minutes, every 10 minutes. Within the time slice T, samples are taken T / 60 minutes times, generating T / 60 minutes of time series subsequences, forming the dataset Data. The number of sequences for each variable is denoted as N, with a total of M variables. Each sequence is denoted as S. m,n (0 < m <= M, 0 < n <= N), where m represents the m-th variable and n represents the n-th sequence of the m-th variable. The length of each sequence is 60 min.

[0021] Furthermore: In step S3, the clustering operation steps are as follows: First, define an OrderRank array, which records a subsequence s and the multi-source time series dataset Data = {S} in ascending order. 1,1 S 1,2 S 1,3 S 1,N Each time series S in} 1,i The maximum distance among all subsequences, sdist(s, S). 1,i Next, we need to calculate the value of the split point dt in the OrderRank array; the split point dt is the average distance between two adjacent values ​​in the OrderRank array; each split point dt can divide the OrderRank into two subsets, that is, divide the entire dataset Data into D subsets. A With D B ,in:

[0022] sdist(s, S1)<dt, S1∈D A ;sdist(s, S1)>dt, S1∈D B #

[0023] The value of the split point dt on the OrderRank array is evaluated by the gap value, which is calculated as follows:

[0024] gap = μ B -σ B -(μ A +σ A )#

[0025] Where μ A and μ B Mean(sdist(s, D)) represents the expression meaning(sdist(s, D) A )) and mean(sdist(s,D) B )), σ A and σ B It means std(sdist(s, D) A )) and std(sdist(s,D B In this context, "mean" represents the average value, and "std" represents the variance. The optimal split point is the dt value where the largest gap value exists. A larger gap value indicates a higher variance. A With D B The wider the interval, the better the division effect.

[0026] Furthermore: In step S3, for the same variable, by calculating the sequence similarity between all sequences and each Clum,i sequence, the category with the smallest sequence similarity value is selected, and all subsequences in the dataset Data are assigned to their respective categories.

[0027] Furthermore: Using transfer entropy to calculate the causality between the variable sequence and the target sequence, the modified formula for calculating transfer entropy is as follows:

[0028]

[0029] Let represent the transfer entropy of variables A, B, and C with respect to variable X, and step represent the offset.

[0030] Where A, B, and C are different variable names, and k represents the sequence length. Let X represent a sequence of length k in the variable X at time t;

[0031] When T > 0, it means that A, B, and C have a causal relationship with X. However, time series generally have many interfering factors. Therefore, further screening is performed on the basis of T > 0, selecting the top 70% to ensure that each set of sequences in the relevant micro-scene has a strong causal relationship.

[0032] Furthermore, in step S6, for the variable sequences recorded by all instruments and equipment at a certain moment in the milk powder processing, a sequence set of length len = 60 min is taken forward, and the four sequences for base powder protein content, total bacterial count, moisture content, and vitamin content are removed, ultimately producing M-4 variable sequences of length 60 min, each sequence denoted as preS. m (0 < m <= M), where M represents the Mth variable; classify all variable sequences by calculating the Clu of each sequence and its corresponding variable. m The sequence similarity of i is used to determine the category to which the variable sequence belongs;

[0033] preClu m =mix(DTW(Clu) m,i ,preS m ))(0<i≤cnt(Clu m ))

[0034] Select the Clu with the lowest sequence similarity value. m,i As the category of the sequence, the m-4 preClus are combined, and the scene with the largest scene transition entropy in M-scene that exceeds the entropy threshold is selected, that is, the scene with the highest similarity and the most relevance.

[0035] Further: In step S6, after using the micro-scene fitting method of multi-granularity space, a training set is constructed for the scene pattern, and ridge regression is used for prediction; the N process indicator sequences and target indicator sequences of the same time period are vertically spliced ​​with the target indicator sequence of the translation step (step = 10min), and the sequence groups of all time periods of the scene are horizontally spliced; the last row is used as the y value, and the rest are used as the X value, and the ridge regression model is used for prediction.

[0036] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0037] (1) This invention targets the relevant sequences of raw material indicators, production process parameters, and finished product indicators of infant formula milk powder. It extracts similar sequence segments and eliminates interfering sequence segments. Through causal analysis, it selects the sequence segments most relevant to the target prediction indicators, establishes relevant micro-scenarios for predictive analysis, and scientifically identifies the key control points and key process control parameters that have the greatest impact on the quality of the finished milk powder, finding their relatively optimal values. By establishing a scientific and reasonable relationship between raw materials and processes, the prediction accuracy of protein content, vitamin content, total bacterial count, and moisture content in infant formula milk powder can be improved, helping companies enhance the product quality of infant formula milk powder.

[0038] (2) Since milk powder may undergo multiple heating processes during production and processing, various factors will have a comprehensive impact on the final product at different stages. This invention constructs relevant micro-scenes for the milk powder sequence and finds the most similar scene. This eliminates a large number of interfering information segments and only performs local analysis and prediction, greatly improving the accuracy of multi-step prediction. This invention mines the correlation between data sequences in these milk powder processes and the causal relationship with the data sequences of milk powder target components, that is, it mines out the main factors affecting the component indicators of the final milk powder product, including the sequence that causes the influence and the approximate time when the phenomenon occurs, which is conducive to process improvement. Combining the two advantages, this invention can continuously predict the milk powder target component indicators for a period of time in the future and mine the causes of their influence through real-time monitoring data during the process, and perform risk anomaly prediction during the production process, which is convenient for timely adjustment of raw materials or optimization of production processes. Attached Figure Description

[0039] Figure 1 This is a flow chart of the wet-process milk powder production process for infant formula milk powder of the present invention;

[0040] Figure 2 This is a sequence cubic modeling diagram in an embodiment of the present invention;

[0041] Figure 3 These are micro-scene matching and micro-scene-based prediction graphs in this embodiment of the invention. Detailed Implementation

[0042] To enable those skilled in the art to better understand the technical solutions of the present invention, preferred embodiments of the present invention are described below in conjunction with specific examples. However, it should be understood that the accompanying drawings are for illustrative purposes only and should not be construed as limiting the present invention. For better illustration of this embodiment, some components in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable that some well-known structures and their descriptions may be omitted in the drawings for those skilled in the art. The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limiting the present invention.

[0043] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this should not be construed as limiting the present invention.

[0044] like Figure 1 As shown, a method for predicting the causal relationships of micro-scenario risks in a multivariate sequence of infant formula production is described. The prediction method includes the following steps:

[0045] S1. Select the sequences from the main process indicators of infant formula milk powder, extract the sequence data, and form M sequence data from different sources, which become dataset D;

[0046] S2. Perform a stationary processing on the sequence data in the dataset D, and use a sliding time slice to augment the dataset D to form a new dataset Data, where the dataset Data contains a target sequence and a variable sequence;

[0047] S3. Cluster by extracting the most distinguishable subsequence in the variable sequence as the feature sequence of this segment of the sequence, denoted as the Rs sequence; and cluster the extracted Rs sequences, denoted as Clu; Clu m,i represents the i-th class clustering center line of the m variable, and is uniquely identified by Clu m,i to identify a category, and the number of clusters for each variable is denoted as cnt(Clu m ); Clu m,i is the Rs sequence obtained above;

[0048] S4. Randomly combine the Clus of different variables m to form a sequence cube (Scube), where Scubek represents the k-th sequence cube, and each sequence in it is denoted as S v,t , v represents the v-th variable sequence, and t represents the t-th subsequence in the variable sequence;

[0049] S5. Perform a screening process on the above-extracted sequence cube (Scube); screen each Scube, and finally what is obtained is the relevant micro-scene M-scene; for each M-scene, define a scene transfer entropy to describe the causal impact of this scene on the target index sequence;

[0050] S6. Perform a similarity judgment, match the most relevant micro-scene, and use the model of the micro-scene for risk prediction.

[0051] In the step S2, after the basic data processing of the dataset D, perform a stationary processing on the sequence and use a sliding time slice to augment the dataset. After multiple experiments and analyses, determine a sampling time slice T, and the length of this T is generally the time of the total milk powder production process. Then, cut the time slice according to a time window of 60 min in length and every 10 min. Such processing has relatively high accuracy and is convenient for analysis. Then, in the time slice of T time, sample T / 60 min times, generate T / 60 min time series subsequences, form a new dataset Data, and record the number of sequences of each variable as N, with a total of M variables. Each sequence is denoted as S m,n (0 < m <= M, 0 < n <= N), m represents the m-th variable, n represents the n-th sequence of the m-th variable, and the length of each sequence is 60 min. Take the base powder protein content, total number of colonies, moisture, and vitamin content sequences as the target sequences, and the remaining sequences as the variable sequences.

[0052] In step S3, clustering is performed by extracting the most distinguishable subsequence of a certain segment of the sequence as the feature sequence of that segment, denoted as the Rs sequence. The specific implementation is as follows:

[0053] For each set of variable sequences, perform the following operation (taking n=1 as an example):

[0054] First, we define the OrderRank array, which records the subsequence s and the multi-source time series dataset Data = {S} in ascending order. 1,1 S 1,2 S 1,3 S 1,N Each time series S in} 1,i The maximum distance among all subsequences, sdist(s, S). 1,i ).

[0055] Next, we need to calculate the value of the split points dt in the OrderRank array. A split point dt is the average distance between two adjacent values ​​in the OrderRank array. Each split point dt can divide the OrderRank array into two subsets, that is, divide the entire dataset Data into D subsets. A With D B ,in:

[0056] sdist(s, S1)<dt, S1∈D A ;sdist(s, S1)>dt, S1∈D B #

[0057] The split point dt on the OrderRank array is as follows: Figure 2 As shown, its value is assessed by the gap value, which is calculated using the following formula:

[0058] gσp=μ B -σ B -(μ A +σ A )#

[0059] Where μ A and μ B Mean(sdist(s, D)) represents the expression meaning(sdist(s, D) A )) and mean(sdist(s,D) B )), σ A and σ B It means std(sdist(s, D) A )) and std(sdist(s,D B In this context, "mean" represents the average value, and "std" represents the variance. The optimal split point is the dt value where the largest gap value exists. A larger gap value indicates a higher variance.A With D B The wider the interval, the better the division effect.

[0060] In step S3, an iterative approach is used to extract the Rs sequences. Starting with the first time series in the total dataset Data, a sliding window is used to extract all its subsequences as Rs sequences. The OrderRank array for each subsequence is calculated, and the gap formula is applied to calculate the corresponding gap value and split point dt. Then, all gap values ​​are sorted, and the subsequence corresponding to the largest gap value is the one with the best distinguishing effect, and it is added to the Rs set. Finally, time series similar to the extracted Rs sequence are separated from the total dataset Data. This process is repeated iteratively until an Rs set that can distinguish all sequences is obtained.

[0061] In step S3, for the same variable, the relationship between all sequences and each Clu is calculated. m,i Sequence similarity is used to classify the sequences into the categories with the lowest similarity values, assigning all subsequences in the dataset `Data` to their respective categories. Sequence similarity is a method for measuring the similarity between two time series of different lengths, and its calculation method is as follows:

[0062] First, assume there are two sequences Q and C, with lengths n and m respectively;

[0063] Construct a matrix D of size n*m, with matrix elements d. i,j =dist(q i c j ), where dist represents the distance calculation function, using Euclidean distance.

[0064] Searching for d in matrix D 11 to d nm The shortest path is achieved through dynamic programming.

[0065] D(i,j)=d(i,j)+min{D(i-1,j-1),D(i-1,j),D(i,j-1)}

[0066] D(i, j) represents the sum of paths in the current state, which ultimately equals the similarity between the two sequences.

[0067] From matrix D, from d 11 to d nm The shortest path is used as the similarity between sequences Q and C.

[0068] In the milk powder production process, the impact of minor changes in various indicators is negligible. Therefore, only the variable sequences that cause the main impact and have a strong causal relationship are considered. Generally, the number of these sequences is not large. Considering that selecting too many sequences would greatly increase the computational complexity, the Clus of different variables m are randomly combined to form a sequence cube (Scube). Scubek represents the k-th sequence cube, and each sequence in it is denoted as Sv, t, where v represents the v-th variable sequence and t represents the t-th subsequence in the variable sequence. Since the sequence segments of the same milk powder production process indicators in the same sequence cube have similar trends of change and fluctuation, their impact on the target indicator is also similar. Therefore, a sequence cube composed of Clus of several indicators can be described as a scenario in the milk powder production process. By separately mining the causal relationship between these scenarios and the target milk powder indicator, relevant micro-scenarios are constructed. Micro-scenario construction and matching are as follows. Figure 3 As shown.

[0069] In step S5, relevant micro-scenario modeling is performed based on risk causality:

[0070] Relevant micro-scenes, with sequence cubes exhibiting strong causal relationships, can effectively describe the scenarios causing anomalies in the target indicator sequence and explain the reasons. Since the number of sequence cubes obtained in steps S1-S4 is enormous, and the specific time series contained within each sequence cube are quite disorganized, they first need to be sorted and filtered. Secondly, to obtain relevant micro-scenes, the Scubes extracted in S4 undergo further filtering.

[0071] First, the specific sequences contained in each sequence cube are time-aligned. For time periods where some variable sequences are missing, all sequence fragments corresponding to these time periods are removed. For the remaining time periods, the strength of the causal relationship between the variable sequence group and its future target indicator sequence is calculated, and weak sequence groups are discarded.

[0072] The determination of causal relationships is achieved through entropy. Transfer entropy quantifies the information transfer between multiple variables. When a system has a large number of variables, this method of calculating transfer entropy can effectively eliminate redundant causal relationships between variables and uncover direct causal relationships, thus achieving both accuracy and simplification in determining causality. We have modified the basic transfer entropy to measure the change in entropy of a sequence of m variables at the same time in a sequence cube (hereinafter referred to as a set of sequences, such as A, B, C) to the result sequence in the next time time interval. This change is used to determine whether the sequence in that time interval has a strong causal relationship with the target sequence. The modified transfer entropy calculation formula is as follows:

[0073]

[0074] Let represent the transfer entropy of variables A, B, and C with respect to variable X, and step represent the offset.

[0075] Where A, B, and C are different variable names, and k represents the sequence length. Let X represent a sequence of length k in the variable X at time t.

[0076] When T>0, it means that A, B, and C have a causal relationship with X. However, time series generally have many interfering factors. Therefore, we further screen the data based on T>0, selecting the top 70% to ensure that each set of sequences in the relevant micro-scene has a strong causal relationship.

[0077] Each Scube is filtered to obtain the relevant micro-scenes (M-scenes). An M-scene describes a set of indicator sequences that have similar causal effects on the target indicator sequence during the milk powder processing. To improve prediction accuracy, a micro-scene is discarded if it contains too few sequences. For each M-scene, a scene transition entropy is defined to describe the causal impact of that scene on the target indicator sequence.

[0078] The following is pseudocode for modeling related micro-scenes based on risk causality:

[0079]

[0080]

[0081] In step S6, the micro-scene fitting method in multi-granularity space...

[0082] 1) Similarity assessment: Matching the most relevant scenarios.

[0083] For a given moment in the milk powder processing, considering all the variable sequences recorded by instruments at a specific point in time, a sequence set of length len = 60 min is taken forward. Excluding sequences related to base powder protein content, total bacterial count, moisture content, and vitamin content, there are a total of M-4 variable sequences of length 60 min. Each sequence is denoted as preSm (0 < m <= M), where M represents the Mth variable. All sequences need to be categorized. This is done by calculating the CLUs of each sequence and its corresponding variable. m,i The sequence similarity is used to determine the category to which the sequence belongs.

[0084] preClu m =mix(DTW(Clu) m,i ,preS m ))(0<i≤cnt(Clu m ))

[0085] Select the Clu with the lowest sequence similarity value. m,i As the category of the sequence, the m-4 preClus are combined, and the scene with the largest scene transition entropy in M-scene that exceeds the entropy threshold is selected, that is, the scene with the highest similarity and the most relevance.

[0086] 2) Use the model of this micro-scene for prediction.

[0087] After utilizing a multi-granularity space micro-scene fitting method, a training set is constructed from the scene patterns, and ridge regression is used for prediction. The method for constructing the training set is as follows: Figure 3 As shown in the diagram, the N process indicator sequences and the target indicator sequence within the same time period are vertically concatenated with the target indicator sequence shifted by one step (step = 10 min). The sequences from all time periods in this scenario are then horizontally concatenated. The last row is used as the y-value, and the rest as the X-values, which are then used for prediction using a ridge regression model. The modeling process is as follows: Figure 3 As shown.

[0088] A penalty term is added to the sum of squared residuals in traditional regression methods to prevent the least squares estimation results from varying too much.

[0089] 3) If a suitable micro-scene cannot be found, then a global model is used for prediction.

[0090] In step 1), if there are no scenes in M-scene that exceed the entropy threshold, then VAR global prediction is performed, using the detection index sequence set of each instrument and equipment that has not been segmented during the process as the training set for prediction. Here, the threshold for each micro-scene is taken as the average value of the TE set of each micro-scene in step S5.

[0091] 4) Risk Prediction

[0092] This prediction method can not only predict target sequences, but also predict risks such as outliers during the milk powder manufacturing process. The prediction method is the same as predicting target sequences: during the milk powder manufacturing process, using existing data, the target sequence for each subsequent step is predicted, and this sequence is compared with normal standards to issue a risk warning, allowing for timely remedial action or termination of the process.

[0093] At the same time, based on the predicted values ​​and the combined characteristics in the micro-scenario, the relative optimal values ​​of the raw material indicators and process parameters can be found, which can be used to guide the company's future production operation management and help the company improve the product quality of infant formula.

[0094] Based on the description and accompanying drawings of this invention, those skilled in the art can readily create or use the method for predicting the micro-scenario risk causality of multivariate sequences in infant formula production sequence according to this invention, and can produce the positive effects described in this invention.

[0095] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications or equivalent changes made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for predicting the causal relationships of micro-scenario risks in a multivariate sequence of infant formula production, characterized by: The prediction method includes the following steps: S1. Select sequences from the process indicators of infant formula milk powder, extract sequence data, and form M sequence data from different sources, which become dataset D; S2. Stabilize the sequence data in dataset D, and expand dataset D using a sliding time slice to form a new dataset Data, which contains the target sequence and the variable sequence. S3. Clustering is performed by extracting the most distinguishable subsequences from the variable sequence as the feature sequences of that variable sequence, denoted as Rs sequences; and the extracted Rs sequences are then clustered, denoted as Clu; Clu m,i The cluster center line representing the i-th cluster of variables m, and defined by Clu m,i A cluster number is denoted as cnt(Clu) to uniquely identify a category. m ); Clu m,i This refers to the Rs sequence obtained above; S4. Randomly combine the Clus of different variables m to form a sequence cube Scube, where Scubek represents the k-th sequence cube, and each sequence is denoted as S. v,t v represents the v-th variable sequence, and t represents the t-th subsequence in the variable sequence; S5. Filter the extracted sequence cubes Scubes; after filtering each Scube, the relevant micro-scenes M-scenes are finally obtained; for each M-scene, a scene transition entropy is defined to describe the causal influence of the scene on the target index sequence. S6. Perform similarity judgment, match the most relevant micro-scenes, and use the micro-scene model to predict risks. In step S6, for the variable sequences recorded by all the instruments and equipment at a certain moment in the milk powder process, a sequence set with a length of len = 60 min is taken forward, and 4 sequences of the protein content of the base powder, the total number of colonies, the moisture, and the vitamin content are removed. Finally, M - 4 variable sequences with a length of 60 min are generated, and each sequence is denoted as preS m , where 0 < m <= M, and M represents the Mth variable; all the variable sequences are classified, and by calculating the sequence similarity between each sequence and each Clu m,i of the corresponding variable, the category to which the variable sequence belongs is judged; ; Select the Clu with the lowest sequence similarity value. m,i As the category of the sequence, the m-4 preClu are combined, and the scene with the largest scene transition entropy and exceeding the entropy threshold is selected from M-scene, that is, the scene with the highest similarity and the most relevance. In step S6, after using the micro-scene fitting method of multi-granularity space, a training set is constructed for the scene pattern, and ridge regression is used for prediction; the N process indicator sequences and target indicator sequences of the same time period are vertically spliced ​​with the target indicator sequence of step=10min step, and the sequence groups of all time periods of the scene are horizontally spliced; the last row is used as the y value, and the rest are used as the X value, and the ridge regression model is used for prediction.

2. The method for predicting the micro-scenario risk causality of a multivariate sequence in the infant formula production sequence according to claim 1, characterized in that: In the step S2, the method for expanding the data set D is to determine a sampling time slice T, the length of which is the time of the total milk powder production process, and then divide the time slice at intervals of 10 minutes with a length of 60 minutes; in the time slice of T time, sample T / 60 minutes, generate T / 60 time series subsequences, form the data set Data, record the number of sequences of each variable as N, with a total of M variables, and each sequence is denoted as S m,n , where 0 < m <= M, 0 < n <= N, m represents the m-th variable, n represents the n-th sequence of the m-th variable, and the length of each sequence is 60 minutes.

3. The method for predicting the micro-scenario risk causality of a multivariate sequence in the infant formula production sequence according to claim 1, characterized in that: In step S3, the clustering operation steps are as follows: First, define an OrderRank array, which records the subsequence s and the multi-source time series dataset in ascending order. Each time series The maximum distance among all subsequences Next, we need to calculate the value of the split point dt in the OrderRank array; the split point dt is the average distance between two adjacent values ​​in the OrderRank array; each split point dt can divide the OrderRank into two subsets, that is, divide the entire dataset Data into two subsets. and ,in: ; The value of the split point dt on the OrderRank array is evaluated by the gap value, which is calculated as follows: ; in and express and , and express and Mean represents the average, and std represents the variance; the optimal split point is the dt where the largest gap value exists; the larger the gap value, the better. and The wider the interval, the better the division effect.

4. The method for predicting the micro-scenario risk causality of a multivariate sequence in the infant formula production sequence according to claim 1, characterized in that: In step S3, for the same variable, by calculating the sequence similarity between all sequences and each Clum,i sequence, the category with the smallest sequence similarity value is selected, and all subsequences in the dataset Data are assigned to their respective categories.

5. The method for predicting the micro-scenario risk causality of a multivariate sequence in the infant formula production sequence according to claim 1, characterized in that: The modified formula for calculating the causality between a variable sequence and a target sequence using transfer entropy is as follows: ; Let represent the transfer entropy of variables A, B, and C with respect to variable X, and step represent the offset. , , , ; Where A, B, and C are different variable names, and k represents the sequence length. Let X represent a sequence of length k in the variable X at time t; when This indicates that A, B, and C have a causal relationship with X. However, time series data have many interfering factors, therefore... Based on this, further screening is conducted, selecting the top 70% to ensure that each sequence in the relevant micro-scene has a strong causal relationship.