A method for mining periodic patterns under variable gap conditions
By introducing pruning methods of character gaps and gaps in the pattern matching technology, the problem of traditional pattern matching technology taking time and insufficient accuracy in processing large-scale data is solved, and efficient and accurate pattern matching effect is achieved.
Patent Information
- Application Number
- CN202211434751.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-11-16
AI Technical Summary
After the introduction of character gap constraints, traditional pattern matching technology leads to long-term and insufficient accuracy, especially when processing huge meteorological data, it is difficult to meet the needs of high efficiency and precision.
The periodic mode mining method under variable gap conditions is adopted, and the pattern matching process is optimized to improve efficiency and accuracy through character gap constraints and pruning methods that appear gap constraints.
It significantly improves the efficiency and accuracy of pattern matching, and can process large-scale data in a short time, meeting efficient and accurate pattern matching needs.
Smart Images

Figure CN115687362B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of electrical digital data processing, and in particular relates to a periodic pattern mining method under variable gap conditions. Background Art
[0002] With the advent of the big data era, pattern matching technology (also known as string matching technology) has been widely used in important fields such as bioinformatics, air quality monitoring, data stream mining, information retrieval and filtering, and network intrusion detection. In recent years, with the development of sensor network technology, progress in the field of bioinformatics, and the popularization of the Internet, people's ability to obtain data has significantly improved, and the amount of data has shown explosive growth. Therefore, seeking more efficient pattern matching algorithms has become the focus of renewed attention in the academic community.
[0003] Daily weather forecasts require the collection of a large amount of meteorological observation data, the processing of the data, and the identification of patterns, which are often reflected in periodic data patterns.
[0004] Early studies focused on pattern matching with a single, fixed-length wildcard. Later researchers gradually explored pattern matching with multiple, variable-length wildcards: is a fixed-length wildcard pattern, where is a wildcard character, each The symbol can match 1 arbitrary character; while the pattern T[2,4]A[3,6]CT is a pattern with multiple wildcards of variable length. T and A can match 2 to 4 uncertain characters, and A and CT can match 3 to 6 uncertain characters. Subsequent research has been extended to pattern mining conditions such as general gap conditions, one-time conditions, overlapping conditions, and approximate matching conditions.
[0005] The concept of sequence pattern mining is to find high-frequency items in the sequence whose support is greater than a specified threshold, that is, frequent patterns. The research on frequent patterns focuses on the distribution frequency of their appearance in the target sequence string, and cannot characterize the distribution position of their appearance in the entire target sequence string: the regularity of the distribution position of their appearance is also one of the important properties of pattern matching, and its importance is particularly prominent in the fields of biological genetic material research, information retrieval and analysis, etc.
[0006] Pattern matching refers to searching for a subsequence in a relatively long sequence string S that is identical or similar to a relatively short pattern P, where the sequence string S and the pattern P must use the same alphabet.
[0007] During the mining process, candidate patterns are determined based on the number and type of characters and the given character spacing. As the number of determined characters in the pattern increases, the number of candidate patterns will grow exponentially in the process of pattern growth. If they are calculated one by one directly, the overhead will be too high. For huge meteorological data, general pattern matching technology takes a very long time, cannot achieve precise results, and cannot meet some special pattern matching requirements. Summary of the invention
[0008] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method for mining periodic patterns under variable gap conditions, which solves the problems of long time consumption and insufficient precision caused by introducing character gap constraints in traditional pattern matching through character gap constraints and appearance gap constraint pruning methods.
[0009] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0010] The present invention provides a method for mining periodic patterns under variable gap conditions, comprising the following steps:
[0011] S1. Obtain a data sequence string S obtained by a meteorological data sensor, and construct a hash table HT of the data sequence string S using a character set Σ;
[0012] S2. Define a candidate pattern P, and construct an intermediate product of the candidate pattern P based on the hash table HT;
[0013] S3, determine whether there is a character char in the process of constructing the intermediate product of the candidate pattern P j If the character gap pruning condition is met, then go to step S4, otherwise go to step S5;
[0014] S4, determining that the candidate pattern P is a non-periodic pattern, and performing character gap pruning on the candidate pattern P, and returning to step S2;
[0015] S5. Repeat steps S2 to S4 to obtain all candidate patterns P that satisfy the gap constraint, and record, store and output each candidate pattern P that satisfies the gap constraint, thereby completing the periodic pattern mining under the variable gap condition.
[0016] Furthermore, the expression of the data sequence string S is as follows:
[0017] S=s 0 s 1 s 2 …S i S i+1 …s n-2 s n-1 ,s i ∈Σ
[0018] Among them, s 0 Indicates the 0th data sequence character, s n-1 Indicates the n-1th data sequence character, s i Represents the i-th data sequence character, i is a natural number, Σ represents the symbol set, and n represents the length of the data sequence string S.
[0019] Furthermore, the key key of the hash table HT stores characters, and the value value of the hash table HT stores all occurrence positions of such characters in the data sequence string.
[0020] Furthermore, the candidate pattern P is expressed as follows:
[0021] P=p 0 [Min 0 ,Max 0 ]p 1 [Min 1 ,Max 1 ]p 2 [Min 2 ,Max 2 ]…
[0022] p j [Min j ,Max j ]…p n′-2 [Min n′-2 ,Max n′-2 ]p n′-1 ,Max j ≥Min j
[0023] Among them, p 0 The character representing the 0th position of candidate pattern P, p j Represents the j-th character of candidate pattern P, [Min j ,Max j ] represents the character spacing constraint between the jth character and the j+1th character, Min j Indicates the minimum distance between the jth character and the j+1th character, Max j Indicates the maximum distance between the jth character and the j+1th character, p j+1 -p j -1 is the character spacing, Min j and Max j are all integers, and n′ is a natural number greater than or equal to 3.
[0024] Furthermore, the step S2 comprises the following steps:
[0025] S21, extract the 0th data sequence character P in the hash table HT 0 All positions of Value p0 ;
[0026] S22, respectively search for the position of the character in the 0th data sequence at [position-Min CG ,position+Max CG ] The first data sequence character P in the interval 1 , and construct several first intermediate products Occurrence (P 0 P 1 ), completing the first cycle, where position represents the position of the 0th data sequence character, Min CG Indicates the minimum adjacent appearance distance, Max CG Indicates the maximum adjacent occurrence distance;
[0027] S23, find each first intermediate product Occurrence (P 0 P 1 ) within the character spacing constraint range of the last data sequence character 2 , and construct the second intermediate product Occurrence (P 0 P 1 P 2 ), completing the second cycle;
[0028] S24, repeating the method of step S23 until no intermediate product is produced in a certain cycle.
[0029] Furthermore, the expression of the judgment condition of the character gap pruning condition is as follows:
[0030]
[0031] in, indicates existence, ST indicates that, Represents the character char in the data sequence string S j The number of occurrences, l s represents the length of the data sequence string S, n″ represents the number of determined characters in the candidate pattern P, Max OG Indicates that the maximum value of the gap range occurs.
[0032] Furthermore, the character gap pruning comprises the following steps:
[0033] A1. Define the first candidate mode P a The first character char exists in x , and the first character char x Located in the first data sequence string Sx The first position on position(char x ) y , then define the adjacent second character char x+1 Match character spacing constraint interval;
[0034] The second character char x+1 The expression for matching character spacing constraints is as follows:
[0035] [position(char x ) y -Min CG ,position(char x ) y -Max CG ]
[0036] A2、For position(char x ) y -Min CG ≤0 or position(char x ) y -Max CG ≤0, then position(char x ) y -Min CG or position(char x ) y -Max CG The value of is 0;
[0037] A3、For position(char x ) y -Min CG >n″′-1-position(char x ) y ,get Then any position(char x ) y The occurrence of the first or subsequent data sequence directly skips the character gap pruning judgment, where n″′ represents the first data sequence string S x Length;
[0038] A4, for the second character char x+1 The character char cannot be found within the matching character spacing constraint. x+1 , then the first candidate mode P a and the first candidate pattern P a Any super pattern of does not satisfy the character spacing constraint. aand the first candidate pattern P a Prune any super-pattern of .
[0039] Furthermore, the definitions appearing in step A3 are as follows:
[0040] There exists a position index sequence I of length m such that it satisfies the occurrence condition constraint, then the position index sequence I is an occurrence of the candidate pattern P in the data sequence string S;
[0041] The expression of the position index sequence I is as follows:
[0042] I= 0 ,i 1 ,…,i m-2 ,i m-1 >
[0043] Among them, i 0 Indicates the index sequence element at position 0, i m-1 represents the index sequence element at position m-1; the expression of the occurrence condition constraint is as follows:
[0044]
[0045] i j-1 ≠i j
[0046] Min j-1 ≤i j-1 -i j -1≤Max j-1
[0047] 0≤j≤m-1,0≤i-(j-1)≤n-1
[0048] in, represents the j-th position index of the character in the i-th data sequence, p j Indicates the jth character, i j-1 Indicates the index sequence element at position j-1, i j Indicates the j-th index sequence element, Min j-1 Indicates the minimum distance between the j-1th character and the jth character, Max j-1 Indicates the maximum distance between the character at position j-1 and the character at position j.
[0049] Furthermore, the super mode in step A4 is defined as follows:
[0050] Define the first candidate pattern P of length l a The second candidate pattern P with a length of l+m′ b , if the first candidate mode P a The characters in any position in the second candidate pattern P correspond to the same characters in order. b The first candidate mode P a The second candidate mode P b The second candidate pattern P b The first candidate mode P a The super mode;
[0051] The first candidate mode P a and the second candidate pattern P b The expression is as follows:
[0052] P a =a 0 a 1 …a l-1
[0053] P b =b 0 b 1 …b l b l+1 …b l+m′-2 b l+m′-1
[0054] Among them, a 0 Represents the first candidate mode P a The character at position 0, a l-1 Represents the first candidate mode P a The l-1th character, b 0 Represents the second candidate mode P b The character at position 0, b l+m′-1 Represents the second candidate mode P b The character at the l+m′-1th position, where l is a non-zero natural number and m′ is a positive integer greater than or equal to 1.
[0055] Furthermore, the definition of the gap constraint in step S5 is as follows:
[0056] If the appearance gap OG between any adjacent appearances of the candidate pattern P satisfies the appearance gap constraint, then the candidate pattern P is a periodic pattern;
[0057] The expression for the gap constraint is as follows:
[0058] [Min OG ,Max OG ]
[0059] Among them, Min OG Indicates the minimum value of the gap range.
[0060] The beneficial effects of the present invention are as follows: the present invention provides a method for mining periodic patterns under variable gap conditions, which includes a pruning method with character gap constraints and occurrence gap constraints; in traditional pattern matching problems, introducing character gap constraints can make problem solving more flexible, but it will also increase the difficulty of the problem. The present invention introduces occurrence gap constraints on the basis of character gap constraints, which can make pattern matching technology more efficient and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 The present invention is a flowchart of the steps of a method for mining periodic patterns under variable gap conditions in an embodiment of the present invention.
[0062] Figure 2 are all occurrences of a given pattern a[0,2]c in sequence S in an embodiment of the present invention.
[0063] Figure 3 Schematic diagram of the detailed process of pattern mining of sequence S under the conditions of determining the number of characters N=2, character gap and occurrence gap [0,2] in an embodiment of the present invention. DETAILED DESCRIPTION
[0064] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0065] Example 1
[0066] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides a periodic pattern mining method under variable gap conditions, comprising the following steps:
[0067] S1. Obtain a data sequence string S obtained by a meteorological data sensor, and construct a hash table HT of the data sequence string S using a character set Σ;
[0068] The expression of the data sequence string S is as follows:
[0069] S=s 0 s 1 s 2 …S i S i+1 …s n-2 s n-1 ,s i ∈Σ
[0070] Among them, s 0Indicates the 0th data sequence character, s n-1 Indicates the n-1th data sequence character, s i represents the i-th data sequence character, i is a natural number, Σ represents the symbol set, and n represents the length of the data sequence string S;
[0071] The key key of the hash table HT stores characters, and the value value of the hash table HT stores all the positions where the characters of this type appear in the data sequence string;
[0072] S2. Define a candidate pattern P, and construct an intermediate product of the candidate pattern P based on the hash table HT;
[0073] The step S2 comprises the following steps:
[0074] S21, extract the 0th data sequence character P in the hash table HT 0 All locations
[0075] S22, respectively search for the position of the 0th data sequence character at [position-Min CG ,position+Max CG ] The first data sequence character P in the interval 1 , and construct several first intermediate products Occurrence (P 0 P 1 ), completing the first cycle, where position represents the position of the 0th data sequence character, Min CG Indicates the minimum adjacent appearance distance, Max CG Indicates the maximum adjacent occurrence distance;
[0076] S23, find each first intermediate product Occurrence (P 0 P 1 ) within the character spacing constraint range of the last data sequence character 2 , and construct the second intermediate product Occurrence (P 0 P 1 P 2 ), completing the second cycle;
[0077] S24, repeat the method of step S23 until no intermediate product is produced in a certain cycle.
[0078] The expression of the candidate pattern P is as follows:
[0079]
[0080] Among them, p 0The character representing the 0th position of candidate pattern P, p j Indicates the j-th character of candidate pattern P, [Min j ,Max j ] represents the character spacing constraint between the jth character and the j+1th character, Min j Indicates the minimum distance between the jth character and the j+1th character, Max j Indicates the maximum distance between the jth character and the j+1th character, p j+1 -p j -1 is the character spacing, Min j and Max j are all integers, n′ is a natural number greater than or equal to 3;
[0081] S3, determine whether there is a character char in the process of constructing the intermediate product of the candidate pattern P j If the character gap pruning condition is met, then go to step S4; otherwise, go to step S5;
[0082] The expression of the judgment condition of the character gap pruning condition is as follows:
[0083]
[0084] in, indicates existence, ST indicates that, Represents the character char in the data sequence string S j The number of occurrences, l s represents the length of the data sequence string S, n″ represents the number of determined characters in the candidate pattern P, Max OG Indicates the maximum value of the gap range;
[0085] S4, determining that the candidate pattern P is a non-periodic pattern, and performing character gap pruning on the candidate pattern P, and returning to step S2;
[0086] The character gap pruning comprises the following steps:
[0087] A1. Define the first candidate mode P a The first character char exists in x , and the first character char x Located in the first data sequence string S x The first position on position(char x ) y , then define the adjacent second character char x+1 Match character spacing constraint interval;
[0088] The second character char x+1The expression for matching character spacing constraints is as follows:
[0089] [position(char x ) y -Min CG ,position(char x ) y -Max CG ]
[0090] A2、For position(char x ) y -Min CG ≤0 or position(char x ) y -Max CG ≤0, then position(char x ) y -Min CG or position(char x ) y -Max CG The value of is 0;
[0091] A3、For position(char x ) y -Min CG >n″′-1-position(char x ) y ,get Then any position(char x ) y The occurrence of the first or subsequent data sequence directly skips the character gap pruning judgment, where n″′ represents the first data sequence string S x Length;
[0092] The definitions appearing in step A3 are as follows:
[0093] There exists a position index sequence I of length m such that it satisfies the occurrence condition constraint, then the position index sequence I is an occurrence of the candidate pattern P in the data sequence string S;
[0094] The expression of the position index sequence I is as follows:
[0095] I= 0 ,i 1 ,…,i m-2 ,i m-1 >
[0096] Among them, i 0 Indicates the index sequence element at position 0, im-1 Represents the index sequence element at position m-1;
[0097] The expression of the occurrence condition constraint is as follows:
[0098]
[0099] i j-1 ≠i j
[0100] Min j-1 ≤i j-1 -i j -1≤Max j-1
[0101] 0≤j≤m-1,0≤i-(j-1)≤n-1
[0102] in, represents the j-th position index of the character in the i-th data sequence, p j Indicates the jth character, i j-1 Indicates the index sequence element at position j-1, i j Indicates the j-th index sequence element, Min j-1 Indicates the minimum distance between the j-1th character and the jth character, Max j-1 Indicates the maximum distance between the character at position j-1 and the character at position j;
[0103] A4, for the second character char x+1 The character char cannot be found within the matching character spacing constraint. x+1 , then the first candidate mode P a and the first candidate pattern P a Any super pattern of does not satisfy the character spacing constraint. a and the first candidate pattern P a Prune any super-pattern of
[0104] The super mode in step A4 is defined as follows:
[0105] Define the first candidate pattern P of length l a The second candidate pattern P with a length of l+m′ b , if the first candidate mode P a The characters in any position in the second candidate pattern P correspond to the same characters in order. b The first candidate mode P a The second candidate mode P b The second candidate pattern P b The first candidate mode P aThe super mode;
[0106] The first candidate mode P a and the second candidate pattern P b The expression is as follows:
[0107] P a =a 0 a 1 …a l-1
[0108] P b =b 0 b 1 …b l b l+1 …b l+m′-2 b l+m′-1
[0109] Among them, a 0 Represents the first candidate mode P a The character at position 0, a l-1 Represents the first candidate mode P a The l-1th character, b 0 Represents the second candidate mode P b The character at position 0, b l+m′-1 Represents the second candidate mode P b The character at position l+m′-1 of , where l is a non-zero natural number and m′ is a positive integer greater than or equal to 1;
[0110] S5, repeating steps S2 to S4, obtaining all candidate patterns P that satisfy the gap constraint, and recording, storing and outputting each candidate pattern P that satisfies the gap constraint, completing the periodic pattern mining under the variable gap condition;
[0111] The definition of the gap constraint in step S5 is as follows:
[0112] If the appearance gap OG between any adjacent appearances of the candidate pattern P satisfies the appearance gap constraint, then the candidate pattern P is a periodic pattern;
[0113] The expression for the gap constraint is as follows:
[0114] [Min OG ,Max OG ]
[0115] Among them, Min OG Indicates the minimum value of the gap range.
[0116] Example 2
[0117] In another embodiment of the present invention, for huge meteorological data, the general pattern matching technology consumes a very long time, cannot achieve the precise purpose, and cannot meet some special pattern matching requirements;
[0118] The present invention does not specify a specific target mode, but uses characters in a character set without repetition, such as a sequence string S=s with a length of n. 0 s 1 s 2 …S i S i+1 …s n-2 s n-1 ,s i ∈Σ, where Σ is the character set, |Σ| represents the size of the symbol set. Given a character set Σ of length m, there will be a candidate pattern P of length n (n≤m) composed of Σ. possibilities;
[0119] Define a pattern P with n characters. The spacing constraints between characters in P are all [Min,Max]. Then n+(n-1)*Max is called the maximum possible length of pattern P.
[0120] Given a length of l s The target ordered sequence string S, a character char in S j The number of occurrences of Set the gap range to [Min OG ,Max OG ], given a candidate pattern P with n characters, the maximum possible length of P is n+(n-1)*Max OG . Then character char j The occurrence gap restriction is not satisfied, so the candidate pattern P does not satisfy the occurrence gap constraint, and the candidate pattern P is not a periodic pattern. For the candidate pattern P, if the pattern appears only once or not in the sequence string S, then this pattern is not a periodic pattern, and this pruning method is called occurrence gap pruning.
[0121] like Figure 2 As shown, given a sequence string S = s 0 s 1 s 2 s 3 s 4 s 5 s 6 s 7 s 8 s 9 s 10 =taatcctgatc, pattern string P = p 0 [min1,max1]p 1=a[0,2]c, where [0,2] is the character gap, 0 represents the minimum character gap, and 2 represents the maximum character gap; a[0,2]c means there can be zero to two wildcards between a and c That is, a and c can match 0 to 2 characters. There are 4 occurrences that meet this character constraint, namely <1,4>, <2,4>, <2,5>, and <8,10>. It is further found that <1,4> and <2,4> both use s in the same position. 4 "c" appears in <2,4> and <2,5>, both using s in the same position 2 "a" in
[0122] The method of the present invention can well solve the problem of processing large-volume sequence data. Assuming that each character corresponds to a parameter in the meteorological data, given a pattern string p 1 =a[0,2]c represents the regular characteristics of rain in weather data, p 2 =t[0,3]c represents lightning characteristics, and its occurrence times and proportions in the weather data sequence are used to reflect rainfall, lightning intensity, etc.
[0123] Through the present invention, it is possible to actively find the regular features that need to be understood. The implementation principle is to actively give the character gaps and appearance gaps of the pattern. For example, if a user wants to understand the sales of a company's product and compare it with similar products of other companies, then the company's product pattern p can be found in the market transaction data. 3 , and the other party's product model p 4 Dig in, calculate the proportion, and the results are clear at a glance;
[0124] The beneficial effects of the present invention are:
[0125] (1) The method of the present invention innovatively proposes the technical concept of appearance gap, and proposes two technologies, character gap pruning and appearance gap pruning, in combination with character gap constraint conditions. In the process of pattern matching with general gap constraints, the pattern matching efficiency can be improved to a high degree, solving the problem of difficulty in solving due to gap constraints.
[0126] (2) The present invention performs periodic pattern matching and gap pruning on the sequence string after character gap pruning, and obtains the required pattern after multiple processing, which can well adapt to large-capacity data processing, greatly shorten the mining time, and simplify the calculation;
[0127] (3) The constraint conditions of the present invention include the number of characters determined by the pattern, the character spacing and the occurrence of the spacing. The relevant matching conditions can be set according to the actual situation, and the target pattern that meets the user's needs can be found accurately and quickly.
[0128] Example 3
[0129] In a practical example of the present invention, let the sequence string S be the meteorological data after collection and processing, and in order to mine the periodic regular pattern, the process of mining the periodic pattern under the condition that the character gap constraint is [0,2] and the appearance gap constraint is [0,2] is demonstrated;
[0130] like Figure 3 As shown, the number of determined characters of a given candidate pattern is N. When N=2, the character spacing constraint [0,2] is first satisfied, and there are 27 occurrences of the determined number of characters: ta<0,1>, ta<0,2>, tt<0,3>, aa<1,2>, at<1,3>, ac<1,4>, at<2,3>, ac<2,4>, ac<2,5>, tc<3,4>, tc<3,5>, tt<3,6>, cc<4,5>, ct<4,6>, cg<4,7>, ct<5,6>, cg<5,7>, ca<5,8>, tg<6,7>, ta<6,8>, tt<6,9>, ga<7,8>, gt<7,9>, gc<7,10>, at<8,9>, ac<8,10>, tc<9,10>;
[0131] The initial 27 occurrences are pruned for character gaps: There are two characters t and a in ta<0,1>. Starting from t at position 0, the character gap constraint is [0,2]. In [0-0,0-2], because 0-2=(-2), the gap constraint for matching a at position 1 is [0,0]. In sequence S, the closest to a at position 1 is at position 2, and the distance is 0, which meets the condition. Then, judging from a at position 1, the gap constraint for matching a at position 2 is [0,1]. However, under the gap constraint, the closest a to a at position 2 in sequence S is at position 8, so it does not meet the condition, so ta<0,1> is pruned by character gaps;
[0132] Judging by the appearance of aa<1,2>, starting from the 1-bit a, the character gap constraint for matching 2-bit a is [0,1]. In sequence S, the 8-bit a is closest to the 2-bit a, and the distance is 5, which does not meet the condition. According to the super pattern theorem described above, the patterns composed of 1-bit a do not meet the condition, and aa<1,2>, at<1,3>, and ac<1,4> will be pruned by the character gap.
[0133] When ac<2,4> appears, starting from a at position 2, the character gap constraint for matching t is [0,2]. In sequence S, the closest to t at position 3 is t at position 6, with a distance of 2, which meets the condition. Then, c at position 4 is used for judgment, and the gap constraint for matching c at position 5 is [5-0,5-2]=[3,5]. The c closest to c at position 5 is at position 10, with a distance of 4, which meets the condition. Therefore, ac<2,4> passes the character gap pruning judgment;
[0134] tt<3,6> appears, and the matching gap constraint of the 7-bit g is 6>11-1-6, so tg<6,7>, ta<6,8>, tt<6,9>, ga<7,8>, gt<7,9>, gc<7,10>, at<8,9>, ac<8,10>, and tc<9,10> skip the character gap constraint;
[0135] Taking the above occurrence as an example, according to the pruning judgment principle, the occurrences that can be pruned through character gaps are: ta<0,2>, ac<2,4>, ac<2,5>, tg<6,7>, ta<6,8>, tt<6,9>, ga<7,8>, gt<7,9>, gc<7,10>, at<8,9>, ac<8,10>, tc<9,10>;
[0136] Periodic mode and gap pruning judgment: According to the formula Make a judgment, is a character in S j The number of occurrences of the given gap in this example is [Min OG ,Max OG ] is [0,2];
[0137] The pattern length determined in this embodiment is 2, and the length of sequence S is 11, so In sequence S, the character that appears less than 2.5 times is g, and the occurrences containing g are tg<6,7>, ga<7,8>, gt<7,9>, and gc<7,10>. At this time, there are still ta<0,2>, ac<2,4>, ac<2,5>, ta<6,8>, tt<6,9>, at<8,9>, ac<8,10>, and tc<9,10>. Among these occurrences, only ac and ta patterns appear 2 times or more, and the rest appear only once, which does not meet the definition of periodic pattern. Therefore, only ac<2,4>, ac<2,5>, ac<8,10>, ta<0,2>, and ta<6,8> meet the definition of periodic pattern.
[0138] Gap judgment: among the occurrences of ac<2,4>, ac<2,5>, and ac<8,10>, the gap between ac<2,5> and ac<8,10> is 8-5-1=2, which meets the requirement of [0,2] character gap, while the gap between ta<0,2> and ta<6,8> is 6-2-1=3, which does not meet the requirement of [0,2]. Therefore, the last ac is the periodic pattern to be found.
[0139] The mining ends when the character gap constraint is [0,2] and the number of determined pattern characters N=2 with a gap constraint [0,2].
[0140] When the number of characters is determined to be 3, under the condition that the character gap constraint is [0,2] and the gap constraint is [0,2], the process of mining periodic patterns on S is as follows:
[0141] Given a candidate pattern with a definite number of characters N, such as N = 3, first satisfying the character spacing constraint [0,2], there are 63 definite characters: taa<0,1,2>, tat<0,1,3>, tac<0,1,4>, tat<0,2,3>, tac<0,2,4>, tac<0,2,5>, ttc<0,3,4>, ttc<0,3,5>, ttt<0,3,6>, aat<1,2,3>, aac<1,2,4>, aac<1,2,5>, atc <1,3,4>, atc<1,3,5>, att<1,3,6>, acc<1,4,5>, act<1,4,6>, acg<1,4,7>, atc<2,3,4>, atc<2,3,5>, att<2,3,6 >, acc<2,4,5>, act<2,4,6>, acg<2,4,7>, act<2,5,6>, acg<2,5,7>, aca<2,5,8>, tcc<3,4,5>, tct<3,4,6>, tcg<3 ,4,7>, tct<3,5,6>, tcg<3,5,7>, tca<3,5,8>, ttg<3,6,7>, tta<3,6,8>, ttt<3,6,9>, cct<4,5,6>, ccg<4,5,7>, cca<4,5,8>, ctg<4,6,7>, cta<4,6,8>, ctt<4,6,9>, cga<4,7,8>, cgt<4,7,9>, cgc<4,7,10>, ctg<5,6,7>, cta<5, 6,8>, ctt<5,6,9>, cga<5,7,8>, cgt<5,7,9>, cgc<5,7,10>, cat<5,8,9>, cac<5,8,10>, tga<6,7,8>, tgt<6,7,9> , tgc<6,7,10>, tat<6,8,9>, tac<6,8,10>, ttc<6,9,10>, gat<7,8,9>, gac<7,8,10>, gtc<7,9,10>, atc<8,9,10>;
[0142] Perform character spacing judgment on the above occurrences:
[0143] Since ta<0,1>, tt<0,3>, aa<1,2>, at<1,3>, ac<1,4>, at<2,3>, tc<3,4>, tc<3,5>, tt<3,6>, cc<4,5>, ct<4,6>, cg<4,7>, ct<5,6>, cg<5,7>, ca<5,8> are pruned by character gaps based on N=2, the super patterns taa<0,1,2>, tat<0,3>, ac<1,4>, at<2,3>, tc<3,4>, tc<3,5>, tt<3,6>, cc<4,5>, ct<4,6>, cg<4,7>, ct<5,6>, cg<5,7>, ca<5,8> are all pruned by character gaps. ,1,3>, tac<0,1,4>, tat<0,2,3>, ttc<0,3,4>, ttc<0,3,5>, ttt<0,3,6>, aat<1,2,3>, aac<1,2,4>, a tc<1,3,4>, atc<1,3,5>, att<1,3,6>, acc<1,4,5>, act<1,4,6>, acg<1,4,7>, atc<2,3,4>, atc<2,3, 5>, att<2,3,6>, acc<2,4,5>, acg<2,4,7>, tcc<3,4,5>, tct<3,4,6>, tcg<3,4,7>, tct<3,5,6>, tcg< 3,5,7>, tca<3,5,8>, ttg<3,6,7>, tta<3,6,8>, ttt<3,6,9>, cct<4,5,6>, ccg<4,5,7>, cca<4,5,8>, ctg<4,6,7>, cta<4,6,8>, ctt<4,6,9>, cga<4,7,8>, cgt<4,7,9>, cgc<4,7,10>, ctg<5,6,7>, cta<5,6,8>, ctt<5,6,9>, cga<5,7,8>, cgt<5,7,9>, cgc<5,7,10>, cat<5,8,9>, cac<5,8,10> will be pruned directly by the character gap.The remaining tac<0,2,4>, tac<0,2,5>, aac<1,2,5>, act<2,4,6>, act<2,5,6>, acg<2,5,7>, aca<2,5,8>, tga<6,7,8>, tgt<6,7,9>, tgc<6,7,10>, tat<6,8,9>, tac<6,8,10>, ttc<6,9,10>, gat<7,8,9>, gac<7,8,10>, gtc<7,9,10>, atc< 8,9,10>, according to the character gap pruning principle, the only ones that pass are: tac<0,2,4>, tac<0,2,5>, acg<2,5,7>, aca<2,5,8>, tga<6,7,8>, tgt<6,7,9>, tgc<6,7,10>, tat<6,8,9>, tac<6,8,10>, ttc<6,9,10>, gat<7,8,9>, gac<7,8,10>, gtc<7,9,10>, atc<8,9,10>;.
[0144] Periodic mode and gap pruning judgment:
[0145] Periodic mode and gap pruning judgment: According to the formula Make a judgment, is a character in S j The number of occurrences of the given gap in this example is [Min OG ,Max OG ] is [0,2];
[0146] The pattern length determined in this embodiment is 2, and the length of sequence S is 11, so In sequence S, the character that appears less than 1.56 times is g, and the occurrences containing g are acg<2,5,7>, tga<6,7,8>, tgt<6,7,9>, tgc<6,7,10>, gat<7,8,9>, gac<7,8,10>, gtc<7,9,10>, and tac<0,2,4>, tac<0,2,5>, aca<2,5,8>, tat<6,8,9>, tac<6,8,10>, ttc<6,9,10>, atc<8,9,10>. Among these occurrences, only the tac pattern appears more than 2 times, and the rest appear only once, which does not meet the definition of the periodic pattern. Therefore, only tac<0,2,4>, tac<0,2,5>, and tac<6,8,10> meet the definition of the periodic pattern.
[0147] Gap judgment: among the occurrences of tac<0,2,4>, tac<0,2,5>, and tac<6,8,10>, the gap between tac<0,2,4> and tac<6,8,10> is 6-4-1=1, and the gap between tac<0,2,5> and tac<6,8,10> is 6-5-1=0, which is consistent with the gap between [0,2]. Therefore, the last tac is the periodic pattern we are looking for.
[0148] The mining ends when the character gap constraint is [0,2] and the number of determined pattern characters N=3 with the gap constraint [0,2].
[0149] In the mining demonstration of determining the number of characters N=2, the character gap constraint [0,2] is satisfied in the sequence S, and the number of occurrences of the determined characters is 27 in total. After the character gap pruning of the present invention, 9 eligible patterns are still left, while after the occurrence gap pruning process, only 2 eligible patterns are left; in the mining demonstration of N=3, the character gap constraint [0,2] is satisfied in the sequence S, and the number of occurrences of the determined characters is 63 in total. After the character gap pruning of the present invention, 14 eligible patterns are still left, while after the occurrence gap pruning process, only 2 eligible patterns are left.
[0150] It can be seen that with the change of the length of the data sequence and the character constraints, the objects to be processed by periodic pattern mining will show an exponential growth. For meteorological data, which is already at an astronomical level, the objects to be processed are unimaginably large. If only a general pruning method is used, only the result after the first step in the process demonstrated by the present invention can be barely obtained. The data sample is still very large, and its periodicity cannot be intuitively seen, and it cannot meet the real needs of the user. However, by using the method of the present invention, a very intuitive and small number of patterns can be obtained after two prunings, and its high efficiency is unmatched by other pattern matching methods.
Claims
1. A method for mining periodic patterns under variable gap conditions, It is characterized in that The steps include: S1. Obtain a data sequence string S obtained by a meteorological data sensor, and construct a hash table HT of the data sequence string S using a character set Σ; S2. Define a candidate pattern P, and construct an intermediate product of the candidate pattern P based on the hash table HT; S3, determine whether there is a character char in the process of constructing the intermediate product of the candidate pattern P j If the character gap pruning condition is met, then go to step S4, otherwise go to step S5; S4, determining that the candidate pattern P is a non-periodic pattern, and performing character gap pruning on the candidate pattern P, and returning to step S2; S5. Repeat steps S2 to S4 to obtain all candidate patterns P that satisfy the gap constraint, and record, store and output each candidate pattern P that satisfies the gap constraint, thereby completing periodic pattern mining under variable gap conditions.
2. The method for mining periodic patterns under variable gap conditions according to claim 1, It is characterized in that The expression of the data sequence string S is as follows: S=s 0 s 1 s 2 …S i S i+1 …s n-2 s n-1 ,s i ∈Σ Among them, s 0 Indicates the 0th data sequence character, s n-1 Indicates the n-1th data sequence character, s i Represents the i-th data sequence character, i is a natural number, Σ represents the symbol set, and n represents the length of the data sequence string S.
3. The method for mining periodic patterns under variable gap conditions according to claim 2, It is characterized in that The key key of the hash table HT stores characters, and the value value of the hash table HT stores all the positions where the characters of this type appear in the data sequence string.
4. The method for mining periodic patterns under variable gap conditions according to claim 3, It is characterized in that The expression of the candidate pattern P is as follows: P=p 0 [My 0 ,Max 0 ]p 1 [My 1 ,Max 1 ]p 2 [My 2 ,Max 2 ]…p j [My j ,Max j ]…p n′-2 [My n′-2 ,Max n′-2 ]p n′-1 ,Max j ≥Min j Among them, p 0 The character representing the 0th position of candidate pattern P, p j Indicates the j-th character of candidate pattern P, [Min j ,Max j ] represents the character spacing constraint between the jth character and the j+1th character, Min j Indicates the minimum distance between the jth character and the j+1th character, Max j Indicates the maximum distance between the jth character and the j+1th character, p j+1 -p j -1 is the character spacing, Min j and Max j are all integers, and n′ is a natural number greater than or equal to 3.
5. The method for mining periodic patterns under variable gap conditions according to claim 4, It is characterized in that The step S2 comprises the following steps: S21, extract the 0th data sequence character P in the hash table HT 0 All locations S22, respectively search for the position of the character in the 0th data sequence at [position-Min CG ,position+Max CG ] The first data sequence character P in the interval 1 , and construct several first intermediate products Occurrence (P 0 P 1 ), completing the first cycle, where position represents the position of the 0th data sequence character, Min CG Indicates the minimum adjacent appearance distance, Max CG Indicates the maximum adjacent occurrence distance; S23, find each first intermediate product Occurrence (P 0 P 1 ) within the character spacing constraint range of the last data sequence character 2 , and construct the second intermediate product Occurrence (P 0 P 1 P 2 ), completing the second cycle; S24, repeat the method of step S23 until no intermediate product is produced in a certain cycle.
6. The method for mining periodic patterns under variable gap conditions according to claim 4, It is characterized in that The expression of the judgment condition of the character gap pruning condition is as follows: in, indicates existence, ST indicates that, Represents the character char in the data sequence string S j The number of occurrences, l s represents the length of the data sequence string S, n″ represents the number of determined characters in the candidate pattern P, Max OG Indicates that the maximum value of the gap range occurs.
7. The method for mining periodic patterns under variable gap conditions according to claim 4, It is characterized in that The character gap pruning comprises the following steps: A1. Define the first candidate mode P a The first character char exists in x , and the first character char x Located in the first data sequence string S x The first position on position(char x ) y , then define the adjacent second character char x+1 Match character spacing constraint interval; The second character char x+1 The expression for matching character spacing constraints is as follows: [position(char x ) y -Min CG ,position(char x ) y -Max CG ] A2、For position(char x ) y -Min CG ≤0 or position(char x ) y -Max CG ≤0, then position(char x ) y -Min CG or position(char x ) y -Max CG The value of is 0; A3、For position(char x ) y -Min CG >n″′-1-position(char x ) y ,get Then any position(char x ) y The occurrence of the first or subsequent data sequence directly skips the character gap pruning judgment, where n″′ represents the first data sequence string S x Length; A4, for the second character char x+1 The character char cannot be found within the matching character spacing constraint. x+1 , then the first candidate mode P a and the first candidate pattern P a Any super pattern of does not satisfy the character spacing constraint. a and the first candidate pattern P a Prune any super-pattern of .
8. The method for mining periodic patterns under variable gap conditions according to claim 7, It is characterized in that The definitions appearing in step A3 are as follows: There exists a position index sequence I of length m such that it satisfies the occurrence condition constraint, then the position index sequence I is an occurrence of the candidate pattern P in the data sequence string S; The expression of the position index sequence I is as follows: I=<i 0 ,i 1 ,…,i m-2 ,i m-1 > Among them, i 0 Indicates the index sequence element at position 0, i m-1 Represents the index sequence element at position m-1; The expression of the occurrence condition constraint is as follows: s ij =p j i j-1 ≠i j My j-1 ≤i j-1 -in j -1≤Max j-1 0≤j≤m-1,0≤i-(j-1)≤n-1 Among them, s ij represents the j-th position index of the character in the i-th data sequence, p j Indicates the jth character, i j-1 Indicates the index sequence element at position j-1, i j Indicates the j-th index sequence element, Min j-1 Indicates the minimum distance between the j-1th character and the jth character, Max j-1 Indicates the maximum distance between the character at position j-1 and the character at position j.
9. The method for mining periodic patterns under variable gap conditions according to claim 8, It is characterized in that The super mode in step A4 is defined as follows: Define the first candidate pattern P of length l a The second candidate pattern P with a length of l+m′ b , if the first candidate mode P a The characters in any position in the second candidate pattern P correspond to the same characters in order. b The first candidate mode P a The second candidate mode P b The second candidate pattern P b The first candidate mode P a The super mode; The first candidate mode P a and the second candidate pattern P b The expression is as follows: P a =a 0 a 1 …a l-1 P b =b 0 b 1 …b l b l+1 …b l+m′-2 b l+m′-1 Among them, a 0 Represents the first candidate mode P a The character at position 0, a l-1 Represents the first candidate mode P a The l-1th character, b 0 Represents the second candidate mode P b The character at position 0, b l+m′-1 Represents the second candidate mode P b The character at the l+m′-1th position, where l is a non-zero natural number and m′ is a positive integer greater than or equal to 1.
10. The method for mining periodic patterns under variable gap conditions according to claim 9, It is characterized in that The definition of the gap constraint in step S5 is as follows: If the appearance gap OG between any adjacent appearances of the candidate pattern P satisfies the appearance gap constraint, then the candidate pattern P is a periodic pattern; The expression for the gap constraint is as follows: [My OG ,Max OG ] Among them, Min OG Indicates the minimum value of the gap range.
Citation Information
Patent Citations
High average utility sequence pattern mining method under non-overlapping condition
CN111475551A
Method and system for mining generalized sequential patterns in a large database
US5742811A