Brand-new advanced multi-dimensional low-correlation time series data analysis and prediction method
By using distributed computing and quantum encryption and other technical means in biofermentation technology to process multi-dimensional low-correlation time series data, the problem of insufficient accuracy and security of data analysis in traditional technologies is solved, and more efficient and reliable data prediction and processing is achieved.
Patent Information
- Application Number
- CN202510081308.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional biofermentation technology is difficult to achieve high-density growth of E. coli, and the lack of mechanisms to effectively process low correlation data, resulting in limited analysis accuracy and depth.
The new advanced multi-dimensional low-correlation time series data analysis and prediction method is adopted, and the multi-dimensional low-correlation time series data is processed through distributed storage and computing technology, combining encryption algorithms based on quantum key distribution and symmetric encryption, and the prediction model is constructed and distributed training is used to adopt the neural network model with adaptive structural evolution.
It improves the accuracy of data analysis and prediction, data processing efficiency and data transmission security, and can more comprehensively capture the intrinsic relationships between data, especially low correlation characteristics, which enhances the precision of data processing and the reliability of prediction results.
Smart Images

Figure CN119940641A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the fields of biological fermentation, artificial intelligence and data processing technology, and specifically to a new high-level multidimensional low-correlation time series data analysis and prediction method. Background Art
[0002] With the advancement of biotechnology, recombinant Escherichia coli has become an important host for the production of recombinant protein drugs, amino acids, enzyme preparations and other biological products. These products are widely used in the pharmaceutical, food, chemical and other industries. Escherichia coli has become one of the preferred microorganisms in the biopharmaceutical field due to its clear genetic background, short growth cycle and simple culture conditions. However, traditional fermentation methods often make it difficult to achieve high-density growth of Escherichia coli, limiting the yield and quality of the final product.
[0003] Traditional technologies have shortcomings. First, these frameworks often lack effective processing mechanisms for low-correlation data, which makes it impossible to accurately capture the low-correlation characteristics between data during the processing process, thus affecting the accuracy and depth of the analysis. Second, traditional distributed computing frameworks may have deficiencies in data segmentation and distribution, making it difficult to ensure the consistency and integrity of data during distributed processing. In addition, traditional encryption methods and model training strategies may have problems such as inefficiency and insufficient security when processing large-scale data, and it is difficult to meet the current high requirements for data processing.
[0004] In summary, traditional technologies have shortcomings and are difficult to meet the high accuracy and high reliability requirements of modern data analysis. Therefore, it is particularly important to develop a new advanced multidimensional low-correlation time series data analysis and prediction method. Summary of the invention
[0005] The purpose of the present invention is to make up for the shortcomings of the prior art and to provide a new advanced multidimensional low-correlation time series data analysis and prediction method, which can realize the effective processing and analysis and prediction of multidimensional low-correlation time series data through technical means. The method not only takes into account the seasonal factors, autocorrelation and outlier correction factors of the data, but also adopts distributed storage and computing technology, as well as an encryption algorithm based on quantum key distribution and symmetric encryption, thereby improving the accuracy of data analysis and prediction, data processing efficiency and data transmission security.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a new advanced multi-dimensional low-correlation time series data analysis and prediction method, the specific steps of which are:
[0007] S1, data preprocessing and segmentation start: preprocess the data in the fermentation tank, assuming that the data dimension is n, the time series length is m, for each dimension d (d = 1, 2, ..., n) data sequence X d =x d1 ,x d2 ,…,x dm ), calculate its trend component T d =(t d1 ,t d2 ,…,t dm ), using a trend separation algorithm based on local weighted regression in K(u) is the kernel function, and Gaussian kernel function is selected h is the bandwidth parameter. In the biological fermentation data, considering that the temperature and pH parameters may have seasonal changes during the fermentation process, the seasonal period is set to p. When m / p is greater than the preset threshold TH, σ d is the standard deviation of the dimension d data. After separating the trend component, we get the residual sequence R d =(r d1 ,r d2 ,…,r dm ), where r dt =x dt -t dt Based on the residual sequence, the segmentation algorithm ISA is used to construct the segmentation index SI ij Used to measure the suitability of the split between dimensions i and j in and are the mean of the residual sequence of dimension and j, MI ij is the mutual information between dimensions i and j, φ ij To adjust the weight, its value is based on the skewness S of the residual sequence of dimensions i and j i and S j Sure, Split the original data into k sub-datasets, each with a unique ID s (s=1,2,…,k);
[0008] S2, computing node data processing: Each computing node in the computing node cluster receives a sub-data set. Each node has a data receiving module, a data processing module, a result storage module and a communication module. In the data processing module, a low correlation detection is performed on the received sub-data set, and a low correlation measurement algorithm LMA is used. where r i t and rj t are the residual data values of dimensions i and j at time t, ω ij and λ ij To adjust the parameters, according to the kurtosis K of the residual sequence of dimensions i and j i and K j Sure, Calculate the low correlation coefficient between dimensions. For low correlation dimensions, use the feature extraction algorithm FEA. FEA is based on a dynamic projection matrix DP. Where V ij Eigenvectors of the covariance matrix based on the residual sequence constructed for dimensions i and j, ψ ij is the weight, and LMA is measured according to the low correlation ij Sure Extract feature vectors and store the results in a result storage module;
[0009] S3, information collaboration and integration between nodes: Through the collaboration mechanism between nodes, an information sharing protocol ISP is designed. The protocol stipulates that the information format is IF and the transmission encryption method is ETM. An encryption algorithm based on quantum key distribution and symmetric encryption is used. The key generated by quantum key distribution is first used to encrypt the data, and then the encrypted data is encrypted again using the symmetric encryption algorithm. The distributed consensus-based collaborative processing algorithm DCPA collects information from each node and selects the quality of the node data according to the quality of the node data. n , through data integrity C n 、AccuracyA n And the amount of data D n Comprehensive assessment, Representative data R n , by calculating the Mahalanobis distance MD between the node data and the overall data in the feature space n OK, R. n =1-MD n / MD max , where MD max is the maximum value of the Mahalanobis distance between all node data and the overall data, as well as the reliability of the node in the cluster RE n , based on the task completion accuracy AC of the node's historical operation n and the average response time RT n Sure Determine the weight W of each node information, Where N is the total number of nodes, and the information of each node is weighted and summed to obtain the globally unified low-correlation feature representation GLFR;
[0010] S4, prediction model construction and distributed training: select the prediction model, and adopt the low-correlation data analysis model for biological fermentation. The model framework is mainly divided into three layers. The first layer performs time series segmentation on the preprocessed data, maintains the time sequence of the original data, separates the data of the fermentation period, fermentation period and decay period, analyzes the fermentation indicators of each stage in sections, extracts the trend and seasonal characteristics and combines them with the original data to form new data. In the financial market data, analyze the trend and seasonal characteristics of market data at different stages, the changing trend of stock prices before and after the financial reporting season, and in the meteorological data, analyzes the changing rules of temperature and precipitation in different seasons;
[0011] The second layer of new data enters the newly constructed GRU architecture, and the data obtained after operation is subjected to multi-scale feature fusion, which not only improves the robustness of the model to noise and outliers, but also prevents the loss of associated feature points due to low correlation in the later model. The processed data is then analyzed by the newly constructed temporal convolutional network TCN.
[0012] The third layer uses the data error after the previous analysis to learn the relevant characteristics using meta-learning technology, conduct meta-training and model tuning, further optimize the model, help quickly adapt to the task data set and make predictions. In the biological fermentation data, the threshold parameter θ is determined according to the characteristic distribution of data at different stages of the fermentation process. 1 and θ 2 , the initial connection weights CW of the hidden layer neurons ij (0) Determine the feature importance of GLFR based on the low correlation feature representation, and calculate the information entropy IE of each feature in GLFR f , F is the number of features of GLFR. A distributed training method is used to distribute the training data to each computing node, and the node calculates the gradient G of the model parameters. n , the global gradient GG is obtained by synergistic aggregation and averaging among nodes, and the model parameters are updated accordingly. The adaptive learning rate adjustment strategy ALRAS is used during training. ALRAS adjusts the learning rate LR according to the training error E and the number of training rounds r, LR = LR 0 ×(1+ε×E) -τ×r , where LR 0 is the initial learning rate, ε and τ are adjustment parameters, which are determined according to the complexity of the data and the convergence of the model. NV is the noise variance of the data, r m ax is the maximum number of training rounds. In the biological fermentation data, the value of NV is determined according to the noise level of the fermentation data;
[0013] S5, result integration and output: After each computing node completes the prediction, it sends the prediction result to the central node. The central node uses the DCIFCA consistency check algorithm based on density clustering and isolation forest to identify data dense areas using the density clustering algorithm. For data points that are not in dense areas, the isolation forest algorithm is used to calculate their abnormality scores AS. p , where h(x p ) is the height of the i-th tree in the isolated forest at data point p. The results are tested, and after removing outliers, they are integrated to generate the final prediction report or data file, which is output in the form of a visual chart or data report.
[0014] Furthermore, in the data preprocessing and segmentation startup step, the adjustment mechanism of the kernel function bandwidth parameter h in the local weighted regression-based trend separation algorithm LWTSA also considers the seasonal factor of the data. Assuming that the seasonal period of the data is p, when m / p is greater than the preset threshold TH, This can make the trend separation algorithm more accurate when processing data with obvious seasonality, improve the effectiveness of data preprocessing, and provide a more reliable basis for subsequent data segmentation.
[0015] Furthermore, in the low correlation measurement algorithm LMA, the adjustment parameter ω i j and λ ij The determination of further considers the autocorrelation of the residual sequence of dimensions i and j, and calculates the autocorrelation coefficient ACF of the residual sequence of dimensions i and j i (t) and ACF j (t), t is the time lag, taking the maximum autocorrelation coefficient and but The low correlation measurement algorithm can more comprehensively reflect the intrinsic relationship of the data and improve the accuracy of low correlation detection.
[0016] Furthermore, in the feature extraction algorithm FEA, the covariance matrix V of the residual sequence is ij In the construction process, for data with outliers, an outlier correction method based on the local outlier factor LOF is used to calculate the LOF value of each data point. For data points whose LOF value is greater than the preset outlier threshold AT, the mean of their neighborhood data points is used for correction. By correcting the outliers, the accuracy of covariance matrix construction is improved, thereby improving the effect of feature extraction.
[0017] Furthermore, in the distributed consensus collaborative processing algorithm DCPA, the representativeness of the data R n The Mahalanobis distance MD is calculated nThe calculation of adopts the dynamic covariance matrix. The covariance matrix is dynamically updated according to the changes of data in different time windows. Assuming the time window size is w, the covariance matrix is recalculated every w data points. This can reflect the dynamic changes of data more timely, make the evaluation of data representativeness more accurate, and optimize the determination of node information weight. In biological fermentation data, as the fermentation process proceeds, the characteristics of the data may change. Using the dynamic covariance matrix to calculate the Mahalanobis distance can better adapt to such changes, improve the rationality of determining the node information weight, and thus improve the performance of the entire system.
[0018] Furthermore, in the neural network model ANN-ASE based on adaptive structural evolution, the initial connection weights CW of the hidden layer neurons are ij (0) Determine the feature importance of GLFR based on the low correlation feature representation, and calculate the information entropy IE of each feature in GLFR f , F is the number of features of GLFR. This determination basis enables the model to better utilize data features during initial training, accelerate model convergence, and improve training efficiency. In biological fermentation data, different features have different importance to the prediction results. Determining the initial connection weights based on feature importance can enable the model to focus on key features more quickly and improve the learning effect of the model.
[0019] Furthermore, in the adaptive learning rate adjustment strategy ALRAS, the determination of the adjustment parameters ∈ and τ also considers the noise level of the data. Let the noise variance of the data be NV, r max The maximum number of training rounds is determined by such a factor, which enables the learning rate adjustment strategy to better adapt to data with different noise levels and improve the stability and accuracy of model training. In biological fermentation data, noise may come from factors such as sensor measurement errors. By adjusting the learning rate considering the noise level, the model can better cope with noise interference during training and improve the reliability of prediction.
[0020] Furthermore, in the DCIFCA consistency check algorithm based on density clustering and isolation forest, the density threshold DT in the density clustering algorithm is dynamically determined according to the dimension of the data and the number of data points. Assume that the data dimension is n and the number of data points is N. p , This determination method enables the consistency check algorithm to effectively identify abnormal prediction results on data of different scales and dimensions, and improve the reliability of result integration. In biological fermentation data, the data scale and dimension may be different in different batches or under different fermentation conditions. Dynamically determining the density threshold can ensure the effectiveness of the algorithm and guarantee the accuracy of the final prediction results.
[0021] Furthermore, in the computing node data processing step, the result storage module adopts a distributed storage architecture based on blockchain to store data, and stores data in different blocks of blockchain according to feature vectors, statistical information, etc., and utilizes the non-tamperable and traceable characteristics of blockchain to ensure the security and integrity of data, facilitate data auditing and management, and provide more reliable protection for data storage in the entire analysis and prediction process. In biological fermentation data, the security and integrity of data are crucial for subsequent analysis and decision-making. The application of blockchain technology can effectively prevent data tampering and loss, while facilitating data traceability and auditing, and improving the efficiency and reliability of data management.
[0022] Compared with the existing technology, this new advanced multi-dimensional low-correlation time series data analysis and prediction method has the following beneficial effects:
[0023] 1. This method uses distributed storage and computing technology to divide the data into multiple sub-datasets and assign them to different computing nodes for processing, thereby realizing parallel computing and efficient data processing. At the same time, this method also designs an encryption algorithm based on quantum key distribution and symmetric encryption to ensure the security of data transmission. In the prediction model construction and distributed training stages, this method adopts a neural network model based on adaptive structural evolution, which can dynamically adjust the neuron activation function and connection weight according to the characteristic distribution of the input data, thereby improving the adaptability and training efficiency of the model. In addition, this method also adopts an adaptive learning rate adjustment strategy and a consistency verification algorithm based on density clustering and isolation forest, which further optimizes the model training and result integration process and improves the overall data processing and prediction efficiency.
[0024] 2. This method can effectively separate trend components and residual sequences from multidimensional low-correlation time series data by adopting a trend separation algorithm based on local weighted regression, and then perform low-correlation detection and feature extraction on the residual sequence. Through the low-correlation measurement algorithm and feature extraction algorithm, this method can more comprehensively capture the intrinsic relationship between data, especially the low-correlation characteristics, thereby improving the accuracy of data analysis and prediction. In addition, this method also takes into account the seasonal factors, autocorrelation and outlier correction factors of the data, further enhancing the precision of data processing and the reliability of prediction results.
[0025] Other advantages, objectives and features of the present invention will be set forth in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 This is a process operation diagram of a new advanced multi-dimensional low-correlation time series data analysis and prediction method. DETAILED DESCRIPTION
[0028] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0029] Embodiment 1
[0030] Collect multidimensional time series data from the stock market, including stock price, trading volume, turnover rate, price-earnings ratio and other dimensions. The data dimension n = 5, the time series length m = 200 (assuming the data of the past 200 trading days), for each dimension of the data series, use the local weighted regression based trend separation algorithm (LWTSA) to calculate the trend component. For example, for the stock price dimension, calculate its trend component T d Here, the kernel function uses the Gaussian kernel function, and the bandwidth parameter h is based on the standard deviation σ of the data. d Calculation, since stock data has a certain seasonality (such as the annual financial report season), assume that the seasonal cycle p = 12 (using the monthly cycle to approximately consider the seasonality), when m / p = 200 / 12>5 (assuming the preset threshold TH = 5), After calculating the trend component, we get the residual sequence R d Then, based on the residual sequence, the segmentation algorithm (ISA) is used to construct the segmentation index SI ij , the adjustment weight is determined according to the mean, mutual information and skewness of the residual sequence in each dimension The original data is divided into k=3 sub-datasets, each of which has a unique ID s (s=1,2,3), for example, according to the relationship between stock price and trading volume, turnover rate and other dimensions, the data is divided into different subsets for subsequent processing in different computing nodes.
[0031] Each computing node in the computing node cluster receives a sub-data set. Each node has a corresponding module. In the data processing module, a low correlation detection is performed on the received sub-data set. The low correlation measurement algorithm (LMA) is used to calculate the low correlation coefficient between each dimension. The parameter ω is adjusted.ij and λ ij According to the kurtosis K of the residual sequence of dimensions i and j i and K j Determine, while considering the autocorrelation of the residual sequence, calculate the autocorrelation coefficient ACF i (t) and ACF j (t), take the maximum autocorrelation coefficient and To further correct ω ij and λ ij For example, for the two dimensions of stock price and price-earnings ratio, the low correlation coefficient is determined by calculating various parameters of their residual sequences. For the low correlation dimension, the feature extraction algorithm (FEA) is used to construct the covariance matrix V based on the residual sequence. ij In the process, for data with outliers (such as abnormal fluctuations in stock prices under the influence of certain special events), an outlier correction method based on the local outlier factor (LOF) is adopted to calculate the LOF value of each data point. For data points with LOF values greater than the preset abnormal threshold AT, the mean of their neighborhood data points is used for correction. After extracting the feature vector, the result is stored in the result storage module. The storage adopts a distributed storage architecture based on blockchain, and the data is stored in different blocks of the blockchain according to feature vectors, statistical information, etc., to ensure that the data cannot be tampered with and is traceable.
[0032] Through the inter-node collaboration mechanism, the information sharing protocol (ISP) is designed, the information format IF and the transmission encryption method ETM are specified, and the encryption algorithm based on quantum key distribution and symmetric encryption is used to ensure data security transmission. Based on the distributed consensus collaborative processing algorithm (DCPA), the information of each node is collected and the node data quality Q is evaluated. n Consider the integrity of data n 、AccuracyA n And the amount of data D n , calculate the data representativeness R n Mahalanobis distance MD n The calculation uses a dynamic covariance matrix, which is recalculated every w = 20 data points, while taking into account the reliability of the nodes in the cluster RE n (According to the task completion accuracy AC of the node's historical operation n and the average response time RT n Determine), thereby determining the weight W of each node information n The weighted sum of each node information is used to obtain a globally unified low-correlation feature representation GLFR.
[0033] The neural network model based on adaptive structural evolution (ANN-ASE) is selected as the prediction model, whose input layer receives GLFR and the hidden layer neuron activation function AF i Dynamically adjust according to the characteristic distribution of the input data. For example, the sigmoid function is used for the data subset with small variance, the tanh function is used for the data subset with moderate variance, and the relu function is used for the data subset with large variance. The threshold parameter θ 1 and θ 2 The connection weights CW between neurons are determined according to the characteristics of stock data. ij The initial connection weight CW is dynamically adjusted according to the gradient changes of the data during training. ij (0) According to the feature importance of GLFR, calculate the information entropy IE of each feature in GLFR f , (F is the number of features of GLFR), a distributed training method is used to distribute the training data to each computing node, and the node calculates the gradient G of the model parameters n , the global gradient GG is obtained through collaborative aggregation and averaging among nodes and the model parameters are updated. The adaptive learning rate adjustment strategy (ALRAS) is used in the training process to adjust the parameters ∈ and τ to consider the noise level of the data (assuming that the stock data noise variance NV=0.05).
[0034] After each computing node completes the prediction, it sends the prediction result to the central node, which uses the consistency check algorithm based on density clustering and isolation forest (DCIFCA) for verification. The density threshold DT in the density clustering algorithm is based on the data dimension n = 5 and the number of data points N. p =200 dynamically determined, For data points that are not in dense areas, the isolation forest algorithm is used to calculate their anomaly score AS p , after removing outliers, the final stock market forecast report is generated through integration and output to investors or analysts in the form of visual charts (such as stock price trend forecast charts, trading volume forecast charts, etc.) or data reports.
[0035] Embodiment 2
[0036] Collect multidimensional time series data from the meteorological station, such as temperature, air pressure, humidity, wind speed, and wind direction. The data dimension n=5, the time series length m=1000 (assuming the data of the past 1000 observation times), use the LWTSA algorithm to calculate the trend component of the data sequence of each dimension, the kernel function is the Gaussian kernel function, and the bandwidth parameter h is calculated considering the seasonality of meteorological data (such as seasonal changes in temperature and air pressure). Assume that the seasonal cycle p=24 (consider the changes within a day with an hour as the cycle, which is approximate to the seasonal cycle). When m / p=1000 / 24>40 (assuming the preset threshold TH=40), Get the residual sequence R d Then, the ISA algorithm is used based on the residual sequence to determine the segmentation index SI according to the characteristics of the residual sequence in each dimension. ij , split the data into k = 4 sub-datasets, each sub-dataset is identified as ID s (s=1,2,3,4).
[0037] The computing node receives the sub-dataset and uses the LMA algorithm to detect low correlation in the data processing module and adjust the parameter ω ij and λ ij Determined by the kurtosis of the dimensional residual sequence and taking into account the autocorrelation (there may be complex autocorrelation relationships between different dimensions of meteorological data), for example, between the temperature and pressure dimensions, the autocorrelation coefficient correction ω of the residual sequence is calculated ij and λ ij , use the FEA algorithm to extract features for low correlation dimensions and construct the covariance matrix V ij When performing the analysis, the LOF method is used to correct abnormal values (such as abnormal meteorological data caused by sensor failure), and the feature vectors are extracted and stored in the result storage module of the blockchain distributed storage architecture.
[0038] The ISP protocol is used to achieve inter-node collaboration, and data is transmitted in a secure and encrypted manner. Based on the DCPA algorithm, the node data quality Q is evaluated. n (comprehensive data completeness, accuracy, and data volume), calculate data representativeness R n (Mahalanobis distance is calculated using a dynamic covariance matrix, updated every w = 50 data points) and node reliability RE n , determine the weight W n Get GLFR, select ANN-ASE model, the input layer receives GLFR, and the hidden layer neuron activation function is based on the variance of the input data and the threshold θ 1 ,θ 2 (Determined according to the statistical characteristics of meteorological data) Dynamic adjustment, connection weight CW ij Dynamic update, the initial weight is determined by the GLFR feature importance (calculation of information entropy IE f) determines that the node calculates the gradient G in distributed training n , summarize the global gradient GG to update the model, train it using the ALRAS strategy, and adjust the parameters ∈ and τ to consider the noise level of meteorological data (assuming the noise variance NV = 0.03).
[0039] After the calculation node predicts, it sends the result to the central node, which uses the DCIFCA algorithm to check. The density threshold DT is based on the data dimension n=5 and the number of data points N. p =1000, and generate a meteorological data forecast report after processing the abnormal values, which is provided to the meteorological forecast department or relevant researchers in the form of visual charts (such as temperature, air pressure and other element forecast trend charts) or data reports.
[0040] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A new advanced multi-dimensional low-correlation time series data analysis and prediction method, characterized by: The specific steps of this method are: S1, data preprocessing and segmentation start: preprocess the data in the fermentation tank, assuming that the data dimension is n, the time series length is m, for each dimension d (d = 1, 2, ..., n) data sequence X d =(x d1 ,x d2 ,…,x dm ), calculate its trend component T d =(t d1 ,t d2 ,…,t dm ), using a trend separation algorithm based on local weighted regression in K(u) is the kernel function, and Gaussian kernel function is selected h is the bandwidth parameter. In the biological fermentation data, considering that the temperature and pH parameters may have seasonal changes during the fermentation process, the seasonal period is set to p. When m / p is greater than the preset threshold TH, σ d is the standard deviation of the dimension d data. After separating the trend component, we get the residual sequence R d =(r d1 ,r d2 ,…,r dm ), where r dt =x dt -t dt Based on the residual sequence, the segmentation algorithm ISA is used to construct the segmentation index SI ij Used to measure the suitability of the split between dimensions i and j in and are the mean of the residual sequence of dimension and j, MI ij is the mutual information between dimensions i and j, φ ij To adjust the weight, its value is based on the skewness S of the residual sequence of dimensions i and j i and S j Sure, Split the original data into k sub-datasets, each with a unique ID s (s=1,2,…,k); S2, computing node data processing: Each computing node in the computing node cluster receives a sub-data set. Each node has a data receiving module, a data processing module, a result storage module and a communication module. In the data processing module, a low correlation detection is performed on the received sub-data set, and a low correlation measurement algorithm LMA is used. where r i t and r j t are the residual data values of dimensions i and j at time t, ω ij and λ ij To adjust the parameters, according to the kurtosis K of the residual sequence of dimensions i and j i and K j Sure, Calculate the low correlation coefficient between dimensions. For low correlation dimensions, use the feature extraction algorithm FEA. FEA is based on a dynamic projection matrix DP. Where V ij Eigenvectors of the covariance matrix based on the residual sequence constructed for dimensions i and j, ψ ij is the weight, and LMA is measured according to the low correlation ij Sure Extract feature vectors and store the results in a result storage module; S3, information collaboration and integration between nodes: Through the collaboration mechanism between nodes, an information sharing protocol ISP is designed. The protocol stipulates that the information format is IF and the transmission encryption method is ETM. An encryption algorithm based on quantum key distribution and symmetric encryption is used. The key generated by quantum key distribution is first used to encrypt the data, and then the encrypted data is encrypted again using the symmetric encryption algorithm. The distributed consensus-based collaborative processing algorithm DCPA collects information from each node and selects the quality of the node data according to the quality of the node data. n , through data integrity C n 、AccuracyA n And the amount of data D n Comprehensive assessment, Representative data R n , by calculating the Mahalanobis distance MD between the node data and the overall data in the feature space n OK, R. n =1-MD n / MD max , where MD max is the maximum value of the Mahalanobis distance between all node data and the overall data, as well as the reliability of the node in the cluster RE n , based on the task completion accuracy AC of the node's historical operation n and the average response time RT n Sure Determine the weight W of each node information, Where N is the total number of nodes, and the information of each node is weighted and summed to obtain the globally unified low-correlation feature representation GLFR; S4, prediction model construction and distributed training: select the prediction model, and adopt the low-correlation data analysis model for biological fermentation. The model framework is mainly divided into three layers. The first layer performs time series segmentation on the preprocessed data, maintains the time sequence of the original data, separates the data of the fermentation period, fermentation period and decay period, analyzes the fermentation indicators of each stage in sections, extracts the trend and seasonal characteristics and combines them with the original data to form new data. In the financial market data, analyze the trend and seasonal characteristics of market data at different stages, the changing trend of stock prices before and after the financial reporting season, and in the meteorological data, analyzes the changing rules of temperature and precipitation in different seasons; The second layer of new data enters the newly constructed GRU architecture, and the data obtained after operation is subjected to multi-scale feature fusion, which not only improves the robustness of the model to noise and outliers, but also prevents the loss of associated feature points due to low correlation in the later model. The processed data is then analyzed by the newly constructed temporal convolutional network TCN. The third layer uses the data error after the previous analysis to learn the relevant characteristics using meta-learning technology, conduct meta-training and model tuning, further optimize the model, help quickly adapt to the task data set and make predictions. In the biological fermentation data, the threshold parameters θ1 and θ2 are determined according to the characteristic distribution of data at different stages of the fermentation process, and the initial connection weight CW of the hidden layer neurons is ij (0) Determine the feature importance of GLFR based on the low correlation feature representation, and calculate the information entropy IE of each feature in GLFR f , F is the number of features of GLFR. A distributed training method is used to distribute the training data to each computing node, and the node calculates the gradient G of the model parameters. n , the global gradient GG is obtained by collaborative aggregation and averaging among nodes, and the model parameters are updated accordingly. The adaptive learning rate adjustment strategy ALRAS is used during training. ALRAS adjusts the learning rate LR according to the training error E and the number of training rounds r, LR = LR0 × (1 + ε × E) -τ×r , where LR0 is the initial learning rate, ε and τ are adjustment parameters, which are determined according to the complexity of the data and the convergence of the model. NV is the noise variance of the data, r m ax is the maximum number of training rounds. In the biological fermentation data, the value of NV is determined according to the noise level of the fermentation data; S5, result integration and output: After each computing node completes the prediction, it sends the prediction result to the central node. The central node uses the DCIFCA consistency check algorithm based on density clustering and isolation forest to identify data dense areas using the density clustering algorithm. For data points that are not in dense areas, the isolation forest algorithm is used to calculate their abnormality scores AS. p , Where h(x p ) is the height of the i-th tree in the isolated forest at data point p. The results are tested, and after removing outliers, they are integrated to generate the final prediction report or data file, which is output in the form of a visual chart or data report.
2. According to claim 1, a new advanced multi-dimensional low-correlation time series data analysis and prediction method is characterized in that: In the data preprocessing and segmentation start-up steps, the adjustment mechanism of the kernel function bandwidth parameter h in the trend separation algorithm LWTSA based on local weighted regression also considers the seasonal factor of the data. Assuming that the seasonal period of the data is p, when m / p is greater than the preset threshold TH, 3. According to claim 1, a new advanced multi-dimensional low-correlation time series data analysis and prediction method is characterized in that: In the low correlation measurement algorithm LMA, the adjustment parameter ω i j and λ ij The determination of further considers the autocorrelation of the residual sequence of dimensions i and j, and calculates the autocorrelation coefficient ACF of the residual sequence of dimensions i and j i (t) and ACF j (t), t is the time lag, taking the maximum autocorrelation coefficient and but This enables the low correlation measurement algorithm to more comprehensively reflect the intrinsic relationship of the data.
4. According to claim 1, a new advanced multi-dimensional low-correlation time series data analysis and prediction method is characterized in that: In the feature extraction algorithm FEA, based on the covariance matrix V of the residual sequence ij In the construction process, for data with outliers, an outlier correction method based on the local outlier factor LOF is used to calculate the LOF value of each data point. For data points whose LOF value is greater than the preset outlier threshold AT, the mean of their neighborhood data points is used for correction.
5. According to claim 1, a new advanced multi-dimensional low-correlation time series data analysis and prediction method is characterized in that: In the distributed consensus collaborative processing algorithm DCPA, the representativeness of the data R n The Mahalanobis distance MD is calculated n The calculation uses a dynamic covariance matrix, which is dynamically updated according to the changes in data in different time windows. Assuming the time window size is w, the covariance matrix is recalculated every w data points.
6. According to claim 1, a new advanced multi-dimensional low-correlation time series data analysis and prediction method is characterized in that: In the neural network model ANN-ASE based on adaptive structural evolution, the initial connection weights CW of the hidden layer neurons are ij (0) Determine the feature importance of GLFR based on the low correlation feature representation, and calculate the information entropy IE of each feature in GLFR f , F is the number of features of GLFR.
7. The new advanced multi-dimensional low-correlation time series data analysis and prediction method according to claim 1 is characterized in that: In the adaptive learning rate adjustment strategy ALRAS, the determination of the adjustment parameters ∈ and τ also considers the noise level of the data. Let the noise variance of the data be NV, r max is the maximum number of training rounds.
8. The new advanced multi-dimensional low-correlation time series data analysis and prediction method according to claim 1 is characterized in that: In the DCIFCA consistency check algorithm based on density clustering and isolation forest, the density threshold DT in the density clustering algorithm is dynamically determined according to the dimension of the data and the number of data points. Suppose the data dimension is n and the number of data points is N. p , 9. The new advanced multi-dimensional low-correlation time series data analysis and prediction method according to claim 1 is characterized in that: In the computing node data processing step, the result storage module adopts a distributed storage architecture based on blockchain to store data, and stores the data in different blocks of the blockchain according to feature vectors, statistical information, etc., making use of the tamper-proof and traceable characteristics of the blockchain.