Industrial time series data correlation analysis method based on Spark platform
By designing an incremental mining model for industrial timing data on the Spark platform, the problem of insufficient computing performance and storage space in the existing technology is solved, and the rapid analysis and real-time update of industrial timing data is realized, and good scalability and fault tolerance are achieved.
Patent Information
- Application Number
- CN202211014856.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-08-23
AI Technical Summary
The prior art has insufficient computing performance and storage space when processing industrial big data, and does not have system fault tolerance and scalability, which cannot meet the distributed characteristics and real-time requirements of industrial timing data.
Based on the Spark platform, an incremental mining model for industrial timing data is designed. By converting the original data set into RDD and grouping it, the correlation of each mode is calculated using the MS-Ecalt algorithm and PW-MinLSH, real-time processing and update of incremental data is realized.
It effectively reduces the time complexity of the algorithm, improves mining efficiency, realizes rapid analysis and real-time update of industrial timing data, and has good scalability and fault tolerance.
Smart Images

Figure CN115455075B_ABST
Abstract
Description
Technical Field
[0001] The invention provides an industrial time series data correlation analysis method based on a Spark platform, belonging to the technical field of big data mining. Background Art
[0002] In the context of the industrial Internet, industrial time series data comes from a wide range of sources and has the characteristics of large volume, multi-source, continuous sampling, low value density, and strong dynamics. The analysis and processing of multi-source time series has attracted more and more attention. Industrial data acquisition platforms often include various modular collaborative sensor equipment groups, which are more widely used in prediction and fault detection. The multivariate time series correlation between different sensors can effectively explore the mutual influence between industrial time series data with diverse modes and variable working conditions, thereby realizing in-depth mining of the relevant mechanisms of industrial time series data.
[0003] For current industrial big data analysis, most of them use traditional single-node local computing, but the software / hardware configuration of a single computer device is always limited, and the data processing capacity is always objectively limited by physical conditions. At the same time, single-node processing cannot meet the distributed characteristics of data. The contradiction between the traditional single-node model and the large amount, real-time and distribution of industrial big data is increasing, and many problems have been exposed. For example, when the traditional model drives a large amount of data, it will face insufficient computing performance and storage space, high cost of special equipment, and lack of system fault tolerance and scalability. It cannot even independently undertake a certain order of magnitude of processing tasks and cannot adapt to the distributed collection characteristics of the data itself.
[0004] Existing algorithms can only mine the time series data that has been generated, but cannot mine the subsequent real-time data, so they are not scalable and real-time. This results in the inability to update the results in a timely manner. At the same time, traditional mining methods require frequent recalculation of the entire data, which is unacceptable for the time loss of industrial flow data, and its processing method is not suitable for the real-time processing requirements of dynamic data. From the current research, the research on incremental data mining has not been fully carried out. The data types involved are mostly concentrated on spatial data, and the theory of incremental mining has not been applied to industrial data, which undoubtedly limits the application of incremental data mining in industry. Therefore, a new model is needed to efficiently solve this problem.
[0005] With the rapid rise of the big data era, traditional single-node local computing can no longer cope with the analysis and processing of distributed massive data, which requires a new analytical approach to quickly mine industrial data. Since a large number of sensors in industrial machines are constantly feeding back huge multi-source time series, each time series generated by a sensor can be analyzed by a computing node to analyze the relationship behind it. Therefore, multi-source time series in industry are more suitable for parallel computing.
[0006] Parallel computing not only avoids the shortcomings of the traditional single-node model, but also can use general commercial hardware to provide sufficient and scalable computing performance and storage space for the overall system, ensuring system performance and reducing costs. At the same time, it ensures that the computer storage cluster composed of multiple computing nodes will not cause the entire system to fail and data to be lost due to the failure of a single node, thereby improving the system's fault tolerance and availability. The Spark parallel platform is fully capable of handling real-time, distributed stream data processing and complex calculations.
[0007] Therefore, the present invention makes full use of the distributed computing capabilities of the Spark platform to design an incremental mining model of association rules for industrial time series data, which can well adapt to the processing and analysis tasks of industrial big data. Summary of the invention
[0008] In order to overcome the deficiencies in the prior art, the present invention aims to solve the technical problem of providing an improved method for industrial time series data correlation analysis based on the Spark platform.
[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is: an industrial time series data correlation analysis method based on the Spark platform, comprising the following steps:
[0010] S1: Convert the original data set into RDD and divide it into different groups;
[0011] S2: Use the MS-Ecalt algorithm to obtain the frequent partial periodic pattern set RDD of each group respectively;
[0012] S3: The frequent partial periodic pattern sets of different groups are combined to generate candidate set information RDD, and the frequent partial periodic pattern set RDD of the original data set is obtained. Then, PW-MinLSH is used to calculate the correlation of each pattern component and the label of each Hash Bucket and obtain the relevant partial periodic pattern set in each Hash Bucket, which is recorded as the original result;
[0013] S4: Convert the newly added data set into RDD and divide it into different groups;
[0014] S5: Use the MS-Ecalt algorithm to obtain the RDD of the frequent partial periodic pattern sets of each newly added group and use PW-MinLSH to calculate the label of each Hash Bucket and obtain the relevant partial periodic pattern set in each Hash Bucket, which is recorded as the incremental result. Through the Hash Bucket label, the original result and the incremental result are merged according to the incremental update strategy to obtain the updated relevant partial periodic pattern set RDD. The relevant partial periodic pattern sets that meet the requirements are screened out through the set threshold to achieve incremental mining.
[0015] The S1 specifically includes:
[0016] S1.1: Multiple processes are set up in parallel based on the Spark platform. Each process contains a sensor group. Each sensor group contains n sensors to provide real-time feedback of various data in the process. Each sensor is regarded as a node. n computing nodes are configured in the workstation, which are recorded as CM1, CM2, ..., CMn respectively. The generated data is recorded as the initial data set.
[0017] S1.2: Read each initial data set through textFile and convert it into RDD, and divide it into different groups. During the running of the Spark job, the initial data set is divided into different groups in real time through the RDD application process and RDD eviction system;
[0018] The steps of the RDD application process in step S1.2 are as follows:
[0019] Apply for storage memory on the heap and determine whether the requested memory is greater than the available memory. If the requested memory is less than the available memory, the application is successful and the memory is allocated.
[0020] When the requested memory is larger than the available memory, first evict the heap blocks to release enough memory, then determine whether the off-heap space is larger than the released memory. When the off-heap space is larger than the released memory, modify the on-heap objects, store the data off-heap, release memory space, and update the available memory. When the off-heap space is less than the released memory, first determine whether the useDisk value is True. When the useDisk value is True, serialize and store it on disk, delete the on-heap objects, release memory space, and update the available memory. When the useDisk value is not True, delete the on-heap objects, release memory space, and update the available memory.
[0021] The RDD eviction system in step S1.2 is as follows:
[0022] Determine whether to cache to local memory by determining the current checkpoint and task completion status. RDDs before the checkpoint and successfully submitted tasks are not cached.
[0023] Checkpoint judgment: The RDD logically located before the checkpoint is no longer valuable for caching. The dependency relationship of the RDD is used to determine the generation relationship between the cached data block and the checkpoint RDD, and to determine whether the current RDD has a sub-checkpoint.
[0024] Task determination: When all partitions of an RDD belong to completed tasks, it is determined that the RDD no longer needs to be cached;
[0025] Strategy for insufficient space: Since no persistence function is used to modify the storage level, this part of the cached RDD will be cleared first when the non-heap space is insufficient.
[0026] In step S2, the MS-Ecalt algorithm is used for each group to obtain the mining result of the frequent partial periodic pattern set of the group, wherein the mining steps of the MS-Ecalt algorithm are as follows:
[0027] S2.1: Traverse the database once, convert the horizontal data format into the vertical data format, and obtain the corresponding periodic occurrence set according to the timestamp of the occurrence position of pattern X, which is recorded as By comparing the elements in the collection, if any element value exceeds the set , the element appears as an invalid cycle, delete it from the set, and repeat the above steps until All elements in , and the final set is recorded as ;
[0028] S2.2: Calculation The modulus length of the set is the period value corresponding to the mode, denoted as ,like The value exceeds the set , then pattern X is called a frequent partial periodic pattern, and the length of pattern X is k, which is called a frequent k pattern;
[0029] S2.3: By taking the intersection of the TID sets of the frequent k patterns, calculate the corresponding k+1 item set, and determine whether it is a frequent partial periodic pattern according to steps S2.1-S2.2;
[0030] S2.4: Repeat steps S2.1-S2.3 until all frequent partial cycle patterns in each process are mined.
[0031] The specific steps of step S3 are as follows:
[0032] S3.1: Use the partial periodicity value (PS), support (sup), weight (weight), density rate (dr) and average periodicity (apr) of the mined frequent partial periodic patterns to build the initial matrix Input Matrix (IM). The matrix obtained after t random permutations of IM is called Signature Matrix (SM).
[0033] S3.2: Split the SM horizontally into some blocks, denoted as bands, each band contains r rows in the SM;
[0034] S3.3: For each band, calculate and process the hash value so that the hash value becomes the tag of the pre-set Hash Bucket, and then match each band with the Hash Bucket;
[0035] S3.4: Calculate the probability that the patterns corresponding to the two bands are mapped to the same Hash Bucket, thereby obtaining the relevant partial periodic pattern.
[0036] The incremental update strategy in step S5 is:
[0037] When incremental data arrives, the incremental data set is converted into RDD and divided into corresponding groups. Steps S1-S4 are used to mine the relevant partial periodic patterns in the incremental data, and the Hash Bucket corresponding to each pattern in the incremental data is calculated. The original data and incremental data are merged according to the buckets with the same corresponding Hash Bucket values to obtain the final updated result and realize incremental mining.
[0038] When the original data and incremental data are merged according to the buckets with the same corresponding Hash Bucket value, the following two situations may occur:
[0039] The modes are the same. For such results, you only need to update the parameter information, and no other operations are required.
[0040] The mode is different. For this result, just regard it as a new mode without updating the parameter information.
[0041] The beneficial effects of the present invention compared with the prior art are as follows: the correlation analysis method for industrial time series data proposed by the present invention based on the Spark platform can effectively characterize and describe the mutual influence, dependence or related information and knowledge between manufacturing data such as product data, process data, equipment data, system operation data, etc., and is an effective means and way to reveal the correlation laws between manufacturing data, so that the relationship generated by these data can be analyzed in a timely manner, which effectively reduces the time complexity of the algorithm and improves the mining efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The present invention will be further described below in conjunction with the accompanying drawings:
[0043] Figure 1 It is an overall architecture diagram of a parallel incremental model shown in an embodiment of the present invention;
[0044] Figure 2 This is a flow chart of RDD memory application after adding non-heap storage in the present invention;
[0045] Figure 3 A diagram of an embodiment of the incremental update process of the present invention;
[0046] Figure 4 Flowchart of the RDD conversion process of the present invention. DETAILED DESCRIPTION
[0047] like Figures 1 to 4 As shown, in order to update the results in time for subsequent adjustments, the present invention combines the designed incremental mining model with Spark parallel computing. The model can update the new data generated in the future to the existing time series database in time, so that it can well deal with the processing and analysis tasks of industrial big data. At the same time, in order to reduce data redundancy and the space required for storing data when using Spark parallel computing, the application process and eviction system of RDD in the traditional Spark parallel algorithm are improved. This method can effectively reduce the algorithm's time and space complexity, give the algorithm scalability and real-time performance, thereby improving mining efficiency in industrial data and achieving the purpose of quickly updating results.
[0048] The following takes incremental mining of relevant partial periodic patterns as an example to describe the process steps of the present invention in detail. Figure 1The figure shows the overall architecture of the parallel incremental model of the incremental mining of relevant partial periodic patterns of the present invention, including two processes in a certain industrial production link, each process includes a sensor group and a sensor group includes n sensors to feedback various data in the process in real time. A computing node is assigned to each sensor, and the frequent partial periodic patterns are mined in parallel in each node by using the frequent item set mining algorithm (MS-Ecalt) through n computing nodes, and the various parameter information of each pattern is obtained. These parameters are input into the PW-MinLSH algorithm as input matrices, and the correlation between each pattern and the corresponding Hash Bucket are calculated. When the incremental data arrives, it is no longer necessary to traverse the entire database again, but only to traverse the incremental data once to obtain the corresponding Hash bucket of the incremental data set, and perform incremental mining through the Hash bucket to achieve the purpose of quickly updating data. The specific steps of each step are as follows.
[0049] 1. Consider each sensor as a node, configure n computing nodes in the workstation, denoted as CM1, CM2, …, CMn, and the data generated by them is denoted as the initial data set.
[0050] 2. Read each initial data set through textFile and convert it into RDD, and divide it into different groups (the application process and eviction system of the new RDD are as follows).
[0051] a) Improvement of RDD application process
[0052] During the running of Spark jobs, RDD calculation and eviction occur frequently. Due to the memory limit of Java heap space, each transformation operation (Transformation) calculates a new RDD that needs to apply for sufficient memory in the heap space. In Spark's memory management, there are memory limits for each area. When the memory is insufficient, the blocks originally in the heap space (storage blocks in units of RDD partitions) need to be evicted. After the storage method is improved, the new storage method is added to the original memory application process. Figure 2 As shown in the figure, when the requested heap space is insufficient, the evicted memory blocks will be cached as a new storage mode in the local memory first, which significantly improves the reuse speed compared to storing them on disk.
[0053] b) Improvement of RDD eviction system
[0054] Different from the traditional method of deciding whether to discard RDD directly by judging the storage level only, the improved processing mechanism of the present invention needs to make additional judgments on whether RDD is worth caching. It determines whether to cache to local memory by determining the current checkpoint and the completion of the task. RDDs before the checkpoint and successfully submitted tasks are not cached.
[0055] i: Checkpoint judgment
[0056] Checkpoints are used to handle the situation where there is a long dependency between RDDs at a certain stage. In order to prevent the high recalculation cost caused by block loss, the checkpoint mechanism is used to reduce the recalculation cost. Logically, the RDD before the checkpoint is no longer valuable for caching. The dependency relationship between the cached data blocks and the checkpoint RDD is used to determine the generation relationship between the cached data blocks and the checkpoint RDD, and to determine whether the current RDD has a sub-checkpoint (if there is a sub-checkpoint RDD, it is considered that the RDD does not need to be cached, regardless of special circumstances).
[0057] ii: Judgment of the task to which it belongs
[0058] When the iterator of RDD is called, the present invention adds a new data structure to record the tasks run by the RDD for comparison with the completed tasks. When all partitions of an RDD belong to completed tasks, it is determined that the RDD no longer needs to be cached.
[0059] iii: Strategy when there is insufficient space
[0060] Since the persistence function is not used to modify the storage level, this part of the cached RDD will be cleared first when the non-heap space is insufficient. Compared with the persisted read data, its life cycle is short and its retention priority is lower.
[0061] 3. Use the MS-Ecalt algorithm to obtain the frequent partial periodic pattern set of each group. The specific mining steps of the MS-Ecalt algorithm are as follows:
[0062] a) Traverse the database once and convert the horizontal data format into the vertical data format. According to the timestamp of the position where pattern X appears, the corresponding period occurrence set is obtained, which is recorded as By comparing the elements in the collection, if any element value exceeds the set , the element appears as an invalid cycle and is deleted from the set. Repeat the above steps until All elements in , the final set is recorded as .
[0063] b) Calculation The modulus length of the set is the period value corresponding to the mode, recorded as ,like The value exceeds the set , then pattern X is called a frequent partial periodic pattern. The length of pattern X is k, which is called a frequent k-pattern.
[0064] c) By taking the intersection of the TID sets of the frequent k patterns, the corresponding k+1 item sets are calculated, and according to the above steps, it is determined whether it is a frequent partial periodic pattern.
[0065] d) Repeat the above steps until all frequent partial cycle patterns in each process are discovered.
[0066] 4. Combine different mining results to get the original result originResult (<m,itemsets,ps,sup,weight,dr,apr> ) and calculate the correlation between each mode through PW-MinLSH. The specific steps are as follows:
[0067] a) Use the partial periodicity values (PS), support (sup), weight (weight), density ratio (dr) and average periodicity (apr) of the frequent partial periodic patterns mined above to build the initial matrix Input Matrix (IM). The matrix obtained after performing t random permutations on IM is called Signature Matrix (SM).
[0068] b) Split the SM horizontally into some blocks, called bands, where each band contains r rows in the SM. It should be noted that bands in the same column belong to the same process.
[0069] c) For each band, calculate the hash value (here the md5 algorithm is used to calculate the hash value). These hash values need to be processed to become the tags of the pre-set Hash Bucket, and then match these bands with the Hash Bucket.
[0070] d) If the bands in the same horizontal direction of two processes are mapped to the same hash value, the patterns corresponding to the two bands are mapped to the same Hash Bucket, which means that the patterns here are considered to be similar enough. For any band of any two processes, the probability that the two band values are the same is , is the similarity of the results mined in the process. In other words, the probability that these two bands are different is 1- The SM composed of these two processes has a total of b bands, and the probability that these b bands are different is Therefore, the probability of at least one of these b bands being the same is .
[0071] Through the above steps to calculate the Hash Bucket corresponding to each mode, the probability It is the probability that the modes corresponding to the two bands are mapped to the same Hash Bucket. ,when When , the probability that two patterns are mapped to the same Hash Bucket is , thus obtaining the relevant partial periodic pattern (Relevant Partial Periodic Pattern) .
[0072] 5. Filter the results and store them in RPPP_List (<itemset,minpr> ).
[0073] 6. When incremental data arrives, convert the incremental data set into RDD and divide it into corresponding groups. Use the above steps to mine the relevant partial periodic patterns in the incremental data and calculate the Hash Bucket corresponding to each pattern in the incremental data. Merge the original data and incremental data according to the buckets with the same corresponding Hash Bucket value to obtain the final updated result and realize incremental mining. The following two situations may occur during the merge:
[0074] i: The modes are the same. For this result, you only need to update the parameter information and no other operations are required.
[0075] ii: Different modes. For this result, just treat it as a new mode without updating parameter information.
[0076] The following example shows how to perform incremental updates by taking two hash buckets with the original data set and the incremental data set as examples. Figure 3 For the relevant periodic mode For example, both Hash Buckets have this pattern, which meets the situation of i, so only some period support information needs to be updated, and the result is But for the relevant periodic pattern For example, there is no identical pattern in the corresponding Hash Bucket, so it meets the situation ii and can be regarded as a new pattern. The remaining Hash Buckets can get the final result after such incremental updates.
[0077] 7. Use the MS-Ecalt algorithm to obtain the frequent partial periodic pattern set of each group, and merge different mining results to obtain the incremental result incrementalResult (<m,itemsets,ps,sup,weight,dr,apr> ) and calculate the correlation between each pattern through PW-MinLSH to obtain the updated relevant partial periodic pattern set RDD, and filter out the relevant partial periodic pattern set that meets the requirements through the set threshold, and the mining process ends.
[0078] The RDD conversion process of the present invention is as follows Figure 4 As shown in the figure, stage0 converts the data generated by each sensor into RDD and divides it into different groups; stage1 uses the MS-Ecalt algorithm for different groups to obtain the frequent partial periodic pattern set RDD of each group; stage2 combines the frequent partial periodic pattern sets of different groups to generate candidate set information RDD, obtains the frequent partial periodic pattern set RDD of the original data set, and then uses PW-MinLSH to calculate the label of each Hash Bucket and obtain the relevant partial periodic pattern set in each Hash Bucket (recorded as the original result); stage3 converts the newly added data set into RDD and divides it into different groups; stage4 uses the MS-Ecalt algorithm to obtain the frequent partial periodic pattern set RDD of each newly added group and uses PW-MinLSH to calculate the label of each Hash Bucket and obtain the relevant partial periodic pattern set in each Hash Bucket (recorded as the incremental result), and through the Hash Bucket label, the original result and the incremental result are merged according to the incremental update strategy to obtain the updated relevant partial periodic pattern set RDD.
[0079] Advanced manufacturing production lines use various intelligent data collection tools to achieve real-time recording and perception of production operation status and operating environment. Industrial time series data based on real-time sensor data has gradually become the main form of current manufacturing resource perception and access. In addition to the massive, distributed, high-dimensional, and heterogeneous characteristics of big data, industrial time series data also presents complex strong real-time and multi-temporal and spatial correlation characteristics (that is, there are some rules within industrial time series data, such as periodicity, trend, correlation, and linear drift, etc.); at the same time, the multi-process manufacturing process that is prevalent in complex product manufacturing activities has posed new challenges to finding the propagation rules and mutual influence of various processing elements in the manufacturing process. The complex correlation between factors affecting the manufacturing process is difficult to be effectively discovered based on qualitative experience. Discovering the complex mechanism relationship between multi-dimensional time series data from massive industrial time series data is an important problem faced by manufacturing companies. The present invention proposes a correlation analysis method for industrial time series data based on the Spark platform, which can effectively characterize and describe the mutual influence, dependence or related information and knowledge between manufacturing data such as product data, process data, equipment data, system operation data, etc., and is an effective means and way to reveal the correlation law between manufacturing data. Therefore, the relationship generated by these data can be analyzed in a timely manner, effectively reducing the time complexity of the algorithm and improving the mining efficiency.
[0080] Regarding the specific structure of the present invention, it should be noted that the connection relationship between the various component modules adopted in the present invention is definite and feasible. Except for special instructions in the embodiments, the specific connection relationship can bring about corresponding technical effects, and solve the technical problems raised by the present invention without relying on the execution of corresponding software programs. The components, modules, models of specific components, the connection methods between each other, and the conventional use methods and expected technical effects brought about by the above-mentioned technical features, except for specific instructions, all belong to the disclosed contents in patents, journal articles, technical manuals, technical dictionaries, and textbooks that can be obtained by technical personnel in this field before the application date, or belong to the existing technologies such as conventional technologies and common knowledge in this field, and there is no need to elaborate, so that the technical solution provided in this case is clear, complete, and feasible, and the corresponding physical products can be reproduced or obtained according to the technical means.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The industrial time series data correlation analysis method based on the Spark platform is characterized by: The steps include: S1: Convert the original data set into RDD and divide it into different groups; S2: Use the MS-Ecalt algorithm to obtain the frequent partial periodic pattern set RDD of each group respectively; S3: The frequent partial periodic pattern sets of different groups are combined to generate candidate set information RDD, and the frequent partial periodic pattern set RDD of the original data set is obtained. Then, PW-MinLSH is used to calculate the correlation of each pattern component and the label of each Hash Bucket and obtain the relevant partial periodic pattern set in each Hash Bucket, which is recorded as the original result; The specific steps of step S3 are as follows: S3.1: The initial matrix Input Matrix (IM) is constructed using the partial periodicity value (PS), support (sup), weight (weight), density rate (dr) and average periodicity (apr) of the mined frequent partial periodic patterns. The matrix obtained after t random permutations of IM is called Signature Matrix (SM). S3.2: Split the SM horizontally into some blocks, denoted as bands, each band contains r rows in the SM; S3.3: For each band, calculate and process the hash value so that the hash value becomes the tag of the pre-set Hash Bucket, and then match each band with the Hash Bucket; S3.4: Calculate the probability that the patterns corresponding to the two bands are mapped into the same Hash Bucket, thereby obtaining the relevant partial periodic patterns; S4: Convert the newly added data set into RDD and divide it into different groups; S5: Use the MS-Ecalt algorithm to obtain the RDD of the frequent partial periodic pattern sets of each newly added group and use PW-MinLSH to calculate the labels of each Hash Bucket and obtain the relevant partial periodic pattern sets in each Hash Bucket, which are recorded as incremental results. Through the Hash Bucket labels, the original results and the incremental results are merged according to the incremental update strategy to obtain the updated relevant partial periodic pattern set RDD. The relevant partial periodic pattern sets that meet the requirements are screened out through the set threshold to achieve incremental mining.
2. The method for analyzing the correlation of industrial time series data based on the Spark platform according to claim 1 is characterized in that: The S1 specifically includes: S1.1: Multiple processes are set up in parallel based on the Spark platform. Each process contains a sensor group. Each sensor group contains n sensors to provide real-time feedback of various data in the process. Each sensor is regarded as a node. n computing nodes are configured in the workstation, which are recorded as CM1, CM2, ..., CMn respectively. The generated data is recorded as the initial data set. S1.2: Read each initial data set through textFile and convert it into RDD, and divide it into different groups. During the running of the Spark job, the initial data set is divided into different groups in real time through the RDD application process and RDD eviction system.
3. The method for analyzing the correlation of industrial time series data based on the Spark platform according to claim 2 is characterized in that: The steps of the RDD application process in step S1.2 are as follows: Apply for storage memory on the heap and determine whether the requested memory is greater than the available memory. If the requested memory is less than the available memory, the application is successful and the memory is allocated. When the requested memory is larger than the available memory, first evict the heap blocks to release enough memory, then determine whether the off-heap space is larger than the released memory. When the off-heap space is larger than the released memory, modify the on-heap objects, store the data off-heap, release memory space, and update the available memory. When the off-heap space is less than the released memory, first determine whether the useDisk value is True. When the useDisk value is True, serialize and store it on disk, delete the on-heap objects, release memory space, and update the available memory. When the useDisk value is not True, delete the on-heap objects, release memory space, and update the available memory.
4. The method for analyzing the correlation of industrial time series data based on the Spark platform according to claim 2 is characterized in that: The RDD eviction system in step S1.2 is as follows: Determine whether to cache to local memory by determining the current checkpoint and task completion status. RDDs before the checkpoint and successfully submitted tasks are not cached. Checkpoint judgment: The RDD logically located before the checkpoint is no longer valuable for caching. The dependency relationship of the RDD is used to determine the generation relationship between the cached data block and the checkpoint RDD, and to determine whether the current RDD has a sub-checkpoint. Task determination: When all partitions of an RDD belong to completed tasks, it is determined that the RDD no longer needs to be cached; Strategy for insufficient space: Since no persistence function is used to modify the storage level, this part of the cached RDD will be cleared first when the non-heap space is insufficient.
5. The method for analyzing the correlation of industrial time series data based on the Spark platform according to claim 1 is characterized in that: In step S2, the MS-Ecalt algorithm is used for each group to obtain the mining result of the frequent partial periodic pattern set of the group, wherein the mining steps of the MS-Ecalt algorithm are as follows: S2.1: Traverse the database once, convert the horizontal data format into the vertical data format, and obtain the corresponding periodic occurrence set according to the timestamp of the occurrence position of pattern X, which is recorded as By comparing the elements in the collection, if any element value exceeds the set , the element appears as an invalid cycle, delete it from the set, and repeat the above steps until All elements in , the final set is recorded as ; S2.2: Calculation The modulus length of the set is the period value corresponding to the mode, denoted as ,like The value exceeds the set , then pattern X is called a frequent partial periodic pattern, and the length of pattern X is k, which is called a frequent k pattern; S2.3: By taking the intersection of the TID sets of the frequent k patterns, calculate the corresponding k+1 item set, and determine whether it is a frequent partial periodic pattern according to steps S2.1-S2.2; S2.4: Repeat steps S2.1-S2.3 until all frequent partial cycle patterns in each process are mined.
6. The method for analyzing the correlation of industrial time series data based on the Spark platform according to claim 1 is characterized in that: The incremental update strategy in step S5 is: When incremental data arrives, the incremental data set is converted into RDD and divided into corresponding groups. Steps S1-S4 are used to mine the relevant partial periodic patterns in the incremental data, and the Hash Bucket corresponding to each pattern in the incremental data is calculated. The original data and incremental data are merged according to the buckets with the same corresponding Hash Bucket values to obtain the final updated result and realize incremental mining.
7. The method for analyzing the correlation of industrial time series data based on the Spark platform according to claim 6 is characterized in that: When the original data and incremental data are merged according to the buckets with the same corresponding Hash Bucket value, the following two situations may occur: The modes are the same. For such results, you only need to update the parameter information, and no other operations are required. The mode is different. For this result, just regard it as a new mode without updating the parameter information.
Citation Information
Patent Citations
Dead nozzle compensation
CA2508141A1
Cardinality estimation method aiming at streaming big data
CN106709001A