Method for accurately mining time interval correlation mode

By introducing the foreground-free sequence filtering pruning and depth-first search pruning strategies in time interval-related mode mining, the problem of time events being treated as point-of-time processing is solved, and low-complexity and efficient TIRP mining is achieved.

CN120336393APending Publication Date: 2025-07-18TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174933.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, time events are treated as point-of-time processing, ignoring the duration of the event, resulting in high computational complexity and high resource consumption for frequent mode mining. No research has applied accurate query to time interval-related mode (TIRP) mining.

Method used

The promising sequence filtering (USFP) strategy is used to filter out sequences that are not promising, combining depth-first search and promising extension mode pruning (UQPP) strategy to reduce connection operations and optimize mining process.

Benefits of technology

It reduces the time for unpromising sequence exploration, reduces the computational complexity and resource consumption, and improves mining efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336393A_ABST
    Figure CN120336393A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data mining, in particular to a method for accurately mining a time interval correlation mode, which is used for solving the problems that a time event is only processed as a time point and the duration of the event is ignored in the prior art, and also solving the problems of high calculation complexity, large resource consumption and the like caused by comprehensively mining all modes. The method comprises the following steps: firstly, adopting a foreground-free sequence filtering pruning (USFP) strategy to reduce the exploration time of a sequence without a foreground, then performing search extension on a frequent time interval mode based on depth-first search, and applying a designed unpromising extension mode pruning (UQPP) strategy and a unpromising query mode pruning (UEPP) strategy in the search extension so as to realize the search extension of the sequence without the foreground. Connection operation is reduced, and the mining process is effectively optimized. Experimental results on a real data set and a synthetic data set show that the method of the scheme has excellent data mining performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This case relates to the field of data mining, and particularly to a method for accurately mining time interval-related patterns in a time interval sequence database. Background Art

[0002] With the rapid development of information technology and the wide application of sensing devices, a large amount of raw data is being continuously generated. By deeply exploring and analyzing the information behind these data through data mining, people can be helped to obtain valuable knowledge.

[0003] Frequent pattern mining was initially proposed to identify patterns that frequently occur in transaction databases, that is, those patterns whose occurrence times exceed a specified threshold. In this case, all items in a transaction are usually considered unordered and are considered to occur simultaneously. However, behavioral data is usually collected over time and space, and sequential data with a time order can more accurately reflect human activity characteristics. Sequential pattern mining believes that <A, B> and <B, A> are two different sequential patterns. For example, in a hospital, the situation where a patient tells a doctor that they cough and then have a fever is different from the situation where they have a fever and then cough, and their diagnosis and treatment plans will also be different. Currently, a variety of algorithms have been proposed to efficiently extract frequently occurring sequential patterns from sequential databases. Sequential pattern mining has been widely applied in many fields such as anomaly detection, autonomous driving, and recommendation systems.

[0004] Existing studies usually represent each event through timestamps and assume that all events either occur simultaneously or sequentially, but ignore the important fact that events occur over a specific duration. For example, a person washes up from 8:00 to 8:20 in the morning, then has breakfast from 8:20 to 8:40, and at the same time answers a phone call that lasts for 10 minutes during breakfast. If only the timestamp method is used, it is impossible to clearly determine whether the phone call occurred during breakfast or after breakfast. To solve this problem, time interval-related patterns (TIRP) were proposed, which provide a richer representation of events by including the duration, occurrence order, and time relationship of events. The goal of frequent TIRP mining is to identify patterns in a time interval sequence database whose support exceeds a specified threshold. Compared with sequential pattern mining, TIRP mining is more complex and requires not only calculating the durations of multiple events but also determining the time relationships between events, such as the case where event A intersects with event B, or event A occurs before event B. Summary of the Invention

[0005] The purpose of this case is to propose an efficient method for mining target time interval related patterns (TIRP) to overcome the deficiencies in the prior art where time events are only processed as time points and the event duration is ignored, and to solve problems such as high computational complexity and large resource consumption caused by comprehensively mining all patterns. The specific technical solutions are as follows.

[0006] In a first aspect, this case proposes a method for accurately mining time interval related patterns. The method includes the following steps: Based on a given time interval sequence database, construct a horizontal database HD, and arrange each time interval sequence in order in the horizontal database HD; Based on the horizontal database HD, identify all frequent S-TIRP patterns of length 1 and store them in the frequent pattern set SF. The S-TIRP is a set aggregated by multiple time interval related patterns sharing the same event A, and the time interval related patterns in this set are denoted as Take the number of time interval sequences that match as the vertical support. If the vertical support exceeds the set value, then is a frequent S-TIRP pattern; sequentially obtain a frequent S-TIRP pattern from SF, perform a depth-first search on the event sequence in the pattern for search expansion, and apply a non-promising expansion pattern pruning strategy and a non-promising query pattern pruning strategy during the search expansion to reduce join operations; at the end of the search expansion, take the time interval related patterns with a vertical support greater than or equal to the set value as the target time interval related patterns; in the non-promising expansion pattern pruning strategy, if the last event p of the current pattern and the current query event are combined into an event pair, and the vertical support of this pair of events is less than the set value, then terminate the expansion; in the non-promising query pattern pruning strategy, if the last event p and the event q to be expanded are combined into an event pair, and the vertical support of this pair of events is less than the set value, then terminate the expansion.

[0007] In an implementation of the above technical solution, based on the horizontal database HD, to identify all frequent S-TIRP patterns of length 1, the steps include: Based on the query event sequence, filter out the time interval sequences in the horizontal database HD that do not contain this query event sequence to generate an updated database HD'; scan the horizontal database HD' to identify all frequent S-TIRP patterns.

[0008] In an implementation of the above technical solution, to arrange each time interval sequence in order, the steps include: Represent the time interval sequence S using symbolic time intervals I, S = {I1, I2,..., I m}, where the subscript of I represents the symbol time interval identifier, I = {start, end, e}, start represents the start time of event type e, end represents the end time of event type e, and I.start, I.end, and I.e represent the values in the above triple respectively; for two symbol time intervals I i and I j , if I i < I j , then I i , I j satisfies (Ii.start < I j .start), or when I i .start is equal to I j .start, I i .end < I j .end, or when I i .start is equal to I j .start and I i .end is equal to I j .end, I i .e < I j .e.

[0009] In an implementation of the above technical solution, in the hopeless expansion mode pruning strategy and in the hopeless query mode pruning strategy, the vertical support of paired events is obtained based on the pattern support matrix, the pattern support matrix is constructed based on the horizontal database HD', and the element value in the pattern support matrix is the vertical support of the S-TIRP pattern formed by a pair of events.

[0010] In an implementation of the above technical solution, during the search expansion by performing a depth-first search on the event sequence in the pattern, record the matching progress between the current pattern and the query event sequence to determine whether the query event sequence has been fully matched.

[0011] In an implementation of the above technical solution, before performing a depth-first search on the event sequence in the pattern for search expansion, establish a vertical database for each S-TIRP pattern in SF. The vertical database contains a list of sequence IDs and is extended to include the position, start time, end time, and relevant symbol time intervals of the event in the sequence ID, so that during the search expansion by performing a depth-first search on the event sequence in the pattern, the symbol time interval of the sequence and the time interval of the pattern can be directly obtained, and by constraining the symbol time interval and the time interval of the pattern, meaningless patterns can be avoided from being generated.

[0012] In one implementation of the above technical solution, when arranging each time interval sequence in order, it includes: setting a noise value ∈≥0. If the time points t i and t j of two events satisfy |t i -t j |≤∈, then the time points of these two events are equal, denoted as t i =t j ; if (t i -t j )>∈, it is expressed as t i <t j .

[0013] In one implementation of the above technical solution, the constructed vertical database is sorted in descending order according to the number of sequences containing the S-TIRP pattern.

[0014] In the second aspect, this case proposes a computer-readable storage medium storing a computer program that can be loaded and executed by a processor to perform any of the above methods.

[0015] The beneficial technical effects of this case: In this solution, events do not need to be processed only as time points during mining, which can meet the time relationship requirements of events and can also consider the time interval requirements of patterns. This solution can reduce the exploration time for sequences without prospects, reduce join operations during exploration, and has the advantages of low computational complexity, small resource consumption, and high mining efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 、 one The overall flow schematic diagram in one implementation.

[0018] Figure 2 、 one The schematic diagram of the time interval sequence database exemplified in one implementation.

[0019] Figure 3 、 one The time interval mutual relationship and corresponding conditions in one implementation.

[0020] Figure 4 、 one The vertical database schematic diagram of the time interval pattern in one implementation.

[0021] Figure 5 , one Schematic diagram of the structure of PSM in a certain implementation manner.

[0022] Figure 6 , one Schematic diagram of the TaTIRP flow chart in a certain implementation manner.

[0023] Figure 7 , one Comparison chart of running time results under different minSup conditions in a certain implementation manner. Specific implementation manner

[0024] For the TIRP mining task, many studies have proposed different solutions. Initially, the temporal relationship between different events was usually simply assumed to be a complementary relationship of two events. For example, event A occurring before event B is equivalent to event B occurring after event A. However, this method has obvious limitations: when more than two events are involved, it is difficult to accurately determine the temporal relationship between the first event and the third event. For example, if A occurs before B and B occurs before C, it is easy to conclude that A occurs before C. But if A occurs before B and at the same time C also occurs before B, the relationship between A and C cannot be determined. To solve this problem, the TPrefixSpan algorithm was proposed to discover non-ambiguous temporal patterns. Subsequently, Papapetrou proposed a noise tolerance method to solve the ambiguity problem in Allen temporal relationships. For example, it can be considered that a meeting at 10:00 is equivalent to a meeting at 10:02. On this basis, KarmaLego introduced ε-tolerance and utilized the transitive property in Allen temporal relationships. Then, VertTIRP further clarified the definition of temporal relationships and adopted a vertical representation method, effectively utilizing the transitive property of temporal relationships. Recently, FastTIRP significantly reduced the frequency of join operations by introducing the pairwise support pruning (PSP) method, thus improving the computational efficiency. However, these solutions still have the problem of low efficiency when dealing with large-scale databases.

[0025] Since the inefficiency mainly stems from generating a large number of worthless sequence patterns during the mining process. For this reason, precise query, as a customized mining technology, has emerged, aiming to more precisely meet user needs. For example, precise pattern query can discover all relevant frequent patterns when a customer lists a pencil and a book in the shopping list. Currently, some research methods for precise query have been proposed, such as precise sequence pattern query, precise continuous sequence pattern mining, precise high-utility itemset query, and precise high-utility sequence query. However, to our knowledge, no research has applied precise query to the mining of time interval-related patterns (TIRP).

[0026] Based on this, precise query is introduced in the time interval related pattern mining in this solution, see Figure 1 , and the steps include: first, adopt the unpromising sequence filtering and pruning (USFP) strategy to reduce the exploration time of sequences without prospects, and then search and expand frequent time interval patterns based on depth-first search. In the search and expansion, apply the designed unpromising expansion pattern pruning (UQPP) strategy and unpromising query pattern pruning (UEPP) strategy to reduce join operations and effectively optimize the mining process.

[0027] This solution can have broad application prospects. For example, by collecting customer intentions and analyzing continuous monitoring data of patients, such as blood glucose, blood pressure, electrocardiogram, etc., personalized treatment plans and preventive measures can be formulated by identifying time interval related patterns in the disease process. In the financial market, by analyzing user preferences and high-frequency data, such as trading volume, price fluctuations, etc., high-frequency trading algorithms can be optimized and trading returns can be improved by discovering time interval related patterns. Similarly, by analyzing the activity data of users within different time intervals, user behavior patterns can be mined to provide personalized content recommendations and advertisement pushes, thereby enhancing the marketing effect.

[0028] The terms involved in this solution are introduced as follows.

[0029] Definition 1: Assume there is a set of event types E = {e1, e2,..., e n}, and they are arranged in lexicographical order, denoted as <. The symbolic time interval is defined as the triple I = {start, end, e}. In this scenario, start represents the start time of event e, and end represents the end time of event e.

[0030] For the sake of simplicity, the terms I.start, I.end, and I.e are used to represent the values in the above triple respectively. In addition, we tend to use an approximate ∈ method to handle noise instead of using an exact operator to compare two time intervals.

[0031] Definition 2: Assume a given noise value ∈ ≥ 0. If two time points t i and t j satisfy |t i - t j | ≤ ∈, then these two time points are equal, denoted as t i = t j . In addition, if (t i - t j ) > ∈, it is expressed as t i < t j . These symbols enable us to arrange time interval events in order.

[0032] Definition 3: Let a finite sequence of time intervals S = {I1, I2, …, I m}, which consists of m different symbol time intervals and is arranged in < order. If I i < I j , then one of the following conditions is satisfied:

[0033] (I i .start < I j .start) V (I i .start = I j .start ∧ I i .end < I j .end)

[0034] V (I i .start = I j .start ∧ I i .end = I j .end ∧ I i .e < I j .e)

[0035] where: I i .e < I j .e can be arranged in alphabetical order.

[0036] Definition 4: The time interval sequence database D consists of a set of time interval sequences, denoted as D = <S1, S2, …, S n >>. Each sequence S i is assigned a unique identifier i.

[0037] For example, as Figure 2 shown, the time interval sequence database consists of five sequences. It should be noted that the same event in the sequence cannot occur at the same time point. In sequence S1, event A starts at time 5 and ends at time 12, then starts at time 14 and ends at time 20; event B starts at time 2 and ends at time 10, then starts at time 12 and ends at time 20; event C starts at time 12 and ends at time 18; event D starts at time 8 and ends at time 18. According to the sorting rules specified in Definition 3, S1 can be represented as {(2, 10, B), (5, 12, A), (8, 18, D), (12, 18, C), (12, 20, B), (14, 20, A)}. Similarly, the original database is converted into a horizontal database HD arranged in < order, as shown in Table 1.

[0038] Table 1 Horizontal Database

[0039] SID Time-interval sequence Sl (2, 10, B), (5, 12, A), (8, 18, D), (12, 18, C), (12, 20, B), (14, 20, A) S2 (2, 16, B), (8, 10, C), (12, 14, C), (14, 18, A), (18, 20, D) S3 (2, 6, A), (2, 8, C), (11, 13, A), (11, 15, D), (14, 19, B), (15, 19, A), (16, 19, D) S4 (2, 15, A), (6, 13, C), (6, 13, D) S5 (0, 2, A), (3, 9, C), (5, 13, A), (13, 16, B), (15, 20, D), (17, 20, B)

[0040] Definition 5: We define the temporal relationship between any pair of symbolic time intervals I i and I j as r(I i , I j ). There are a total of eight temporal relationships, namely: before (b), meet (m), overlap (o), contain (c), finish (f), equal (e), start (s), and left contain (l). See Figure 3 the schematic relationships and corresponding conditions. In the figure, ∈ is the given noise value.

[0041] To better describe these relationships, several constraints are used, such as minGap, maxGap, minDura, and maxDura. Among them, minGap and maxGap are used to limit the interval between two symbolic time intervals when their relationship is defined as "before"; while minDura and maxDura require that the time interval of the pattern be within a specific range. The specific conditions and their schematic diagrams for each temporal relationship are shown in Figure 3 . The red dashed box in the figure represents the noise range. Since the duration of all patterns must be within a certain range, this range is not listed separately in the figure.

[0042] Definition 6: Assume that in this example, the set of all event types is Φ = {A, B, C, D}, and the set of all relationships is Γ = {b, m, o, c, f, e, s, l}, where E and R are ordered subsets of Φ and Γ respectively. A Temporal Interval Related Pattern (TIRP) is represented as X = (E, R), where R = {r(I1, I2), r(I1, I3), …, r(I2, I3), …, r(I k-1 , I k ), R ij represents the relationship between the i-th and j-th events in E.

[0043] For example, assume that minGap is set to 0, maxGap is set to 5, minDura is set to 0, and maxDura is set to 20. In sequence S1, TIRP(BA, o) appears at {(2, 10, B), (5, 12, A)}, and (BA, b) appears at {(2, 10, B), (14, 20, A)}. Additionally, TIRP(BAC, obm) appears in sequence S1, denoted as {(2, 10, B), (5, 12, A), (12, 18, C)}. According to the maxGap constraint, TIRP(CB, b) does not appear in S3 because (B.end - C.start) is greater than 5. In this example, even if two events are the same, their relationships may vary within the same sequence or different sequences. For instance, in sequence S1, the temporal relationship between event B and A is both "overlap" and "before".

[0044] Definition 7: We aggregate multiple interval-related patterns that share the same event E into a set, called S-TIRP, and the interval-related patterns of this set are denoted as

[0045] For example, (AB, o) and (AB, b) can both be called (AB, _). Taking (AB, _) as a new pattern, that is, S-TIRP classifies patterns with different relationships but the same event sequence as one type of pattern. However, in this case, it is not considered that (AB, o) and (ABC, obm) belong to the same S-TIRP, nor that (AB, o) and (BA, b) belong to the same S-TIRP. Definition 8: Suppose there exists a TIRP X = (E, R) and a sequence of time intervals S = {I1, I2,..., I m}, when TIRP X is included in sequence S, we consider X to match S. In other words, for any event e i and e j in E, there exist I k and I h in sequence S such that e i = I k .e, e j = I h .e, and r(I k , I h ) = R ij .

[0046] In the actual mining process, when matching interval-related patterns with a sequence of time intervals, both events and relationships need to be considered simultaneously, which can be very strict. To discover more interesting patterns, we propose a looser matching definition.

[0047] Definition 9: Let there be an S-TIRP where E = {e1, e2,..., e s}, and a sequence of time intervals S = {I1, I2,..., I m}. We consider to be matched with S if and only if for 1 ≤ v ≤ s, there exists k v such that 1 ≤ k1 ≤ k2 ≤... ≤ k v ≤ m, and

[0048] As shown in the above example, TIRP(BA, o) is matched with sequences S1 and S2. Additionally, S-TIRP is also matched with S1 and S2.

[0049] Definition 10: The vertical support (VSup) of S-TIRP is defined as the number of sequences that are matched . If exceeds the parameter minSup × |D|, where minSup is the percentage of the database size set by the expert and |D| is the size of the given database, then is considered a frequent S-TIRP pattern. If the event in E is 1, then is a frequent S-TIRP pattern of length 1. The horizontal support (HSup) of is defined as the number of times

[0050] matches a certain time interval sequence S. Continuing with the above example, the vertical support of is

[0051] In comparison with the horizontal database, the vertical database provides a more functional representation for S-TIRP, which is a tabular form containing multiple S-TIRP attributes.

[0052] Definition 11: The vertical database of S-TIRP contains a list of sequence IDs and is extended to include event IDs, start and end times, source time intervals, and the corresponding relationships, all of which are recorded in the SID sequence. For ease of explanation, Figure 4 shows an example vertical database.

[0053] In sequence S1, eid represents the position 5 of the last symbol time interval of and The minimum and maximum values of the start time and end time. According to the relational condition, their relationship is represented as s. The source time interval records the symbolic time intervals that make up the

[0054] To better meet customer needs, instead of mining a large number of patterns, it is better to explore more precise target patterns through the query patterns provided by users.

[0055] Definition 12: Given a query event sequence qes = {e1, e2,..., e k}}, the target S-TIRP satisfies and where |D| is the size of the given database.

[0056] The purpose of target time interval related pattern mining is to mine a set of target S-TIRPs that match the query event sequence qes.

[0057] Given a time interval sequence database D, a query event sequence qes, a specified minimum support minSup, and some optional constraint conditions such as minGap, maxGap, minDura, maxDura, ∈. The goal is to identify the complete set of target S-TIRPs that satisfy and

[0058] For example, in Table 1, assume the query event sequence is qes = {A, C}, and the minimum support minSup is set to 0.4. The size of the given time interval sequence database is 5. At the same time, set minGap = 0, maxGap = 5, minDura = 0, maxDura = 20, and ∈ = 0. According to these conditions, a total of eight S-TIRPs are mined, which are:

[0059] Generally, there are two ways of sequence pattern growth, namely S-extension and I-extension. S-extension is used when new events occur at different times, while I-extension is used when multiple events occur simultaneously. In the current context, since the defined sorting rule strictly arranges each symbolic time interval, there are no identical symbolic time intervals. Therefore, this paper only adopts S-extension to extend patterns.

[0060] Definition 13: S-extension, also known as sequence extension, is to generate a new event sequence by appending a new event after the last event of the sequence pattern X, denoted as

[0061] ​Definition 14: The event corresponding to the matching position in the query event sequence is called the current query event, denoted as qe. During the pattern growth process, we use a flag bit match to track the matching position of the current qe. When the extended event e does not match qe, the match value is 0. On the contrary, when the event e matches qe, the match value is incremented by 1. If match reaches the length of the query event sequence, it means that the sequence has been fully matched.

[0062] Since the events of the target S-TIRP must contain the query event sequence qes, the following pruning strategy is recommended.

[0063] The first strategy: (Unpromising Sequence Filtering Pruning strategy, USFP): For the query event sequence qes, any time interval sequence S in the time interval sequence database D that does not match should be discarded. Sequences lacking cannot contribute to generating the target S-TIRP. By eliminating these non-contributing sequences, the TaTIRP method not only reduces memory consumption but also optimizes processing performance. Suppose the size of the time interval sequence database D is lower than the minimum support threshold minSup × |D|, then it means that frequent target S-TIRPs cannot be discovered.

[0064] In the running example, for the query event sequence qes = {A, C}, we find that only sequence S2 does not contain qes. Therefore, we can directly remove S2 from the database, which will greatly reduce the search space for pattern growth. FastTIRP adopts a data structure called Pattern Support Matrix (PSM), which stores the co-occurrence information between two S-TIRPs to minimize the number of required join operations. Here, we continue to use this structure to further optimize the algorithm.

[0065] Definition 15: For each pair of event types e1 and e2 ∈ E in the time interval sequence database D, if e1 is temporally before e2, the Pattern Support Matrix (PSM) will save a triple, where

[0066] Figure 5 shows the construction process of the PSM structure, where sequence S2 has been removed. Using this structure, the following two pruning strategies are designed. The second strategy: Unpromising Extended Pattern Pruning strategy (UEPP): If the support degree of the last event p of the S-TIRP and the upcoming extended event q does not reach the support threshold, for example, PSM(p, q) < minSup × |D|, then it is redundant to construct a more complex pattern with the pattern of event q. Due to the downward closure property of support, S-TIRP(pq) -will not be a frequent pattern.

[0067] When constructing the PSM structure, FastTIRP requires that the durations of two S-TIRPs are less than maxDura, and the interval between them is less than maxGap. However, it is found in the actual implementation that considering the interval limit in the initial stage will prune some candidate patterns that should be frequent patterns. For example, assume a sequence contains four symbol time intervals: (0, 20, A), (0, 3, B), (5, 8, C), and (15, 20, C). It can be found that the S-TIRP (CC) - has an interval of 7, exceeding maxGap. However, the pattern (ABCC) - is still qualified in the final result. This phenomenon is inconsistent with the downward closure property of support. The reason is that the start time and end time of (AB) - and (ABC) - are both 0 and 20, and the extension conditions between them are satisfied. Therefore, the pattern (ABCC) - is still reasonable. Based on the above example, we decide to remove the interval limit when constructing the PSM structure.

[0068] The third strategy: (Useless Query Pattern Pruning Policy, UQPP): If in the S-TIRP, the support of the last event p and the current query event qe is lower than the threshold, for example, PSM(p, qe) < minSup × |D|, then the event p cannot be used for pattern extension.

[0069] Although both the second strategy and the third strategy are based on the PSM structure, there are differences between them. The second strategy focuses on expanding events, while the third strategy focuses on the current query event. For example, assume the query event sequence qes = {C, B} and minSup is set to 0.4. According to Strategy 1, since the sequence S2 does not contain qes, s2 can be directly removed. In addition, assume the current event is D and the event to be expanded is C. If according to PSM(D, C) < minSup × |D| (the second strategy), the expansion can stop. If the last event of the current pattern is B and the current query event qe is C, then since PSM(B, C) < minSup × |D| (the third strategy), there is no need to continue expanding this pattern.

[0070] The following clearly and completely describes how to apply precise query to the mining of time interval related patterns (TIRP) in this case. Obviously, the described implementation manners are only part of the implementation manners of this case, rather than all of them.

[0071] See Figure 6, a method for accurately mining time interval related patterns, denoted as TaTIRP, includes the following steps.

[0072] Step S1, Input data preparation.

[0073] The input data includes: a time interval sequence database D, a query event sequence qes provided by the user, the required minimum support minSup, and some optional parameters, such as the noise tolerance ∈, the minimum interval minGap between two time points, the maximum interval maxGap, the minimum duration minDura and the maximum duration maxDura between two S-TIRPs. Among them, the set minimum support and time interval limit can accurately define the range of the mined target patterns.

[0074] Step S2, Scan the time interval sequence database D to create a horizontal database HD.

[0075] This horizontal database contains multiple sequences, and each sequence consists of symbolic time intervals sorted in the defined order <.

[0076] Step S3, Apply the USFP strategy to the horizontal database HD to filter out the sequences that do not contain the query event sequence qes, thereby generating an updated horizontal database HD'.

[0077] Here, in order to prepare for the subsequent pruning strategy, a pattern support matrix can be constructed based on HD'. Traverse all time interval sequences in the dataset. For each time interval sequence, check all event pairs that appear in the sequence and calculate their vertical support to obtain the pattern support matrix. Through the pattern support matrix, patterns that do not require pattern expansion can be directly judged.

[0078] Step S4, Scan the updated horizontal database HD' to identify all frequent S-TIRPs and store them in the frequent pattern set SF.

[0079] Step S5, Construct a vertical database VD for each frequent S-TIRP and sort the vertical database in descending order according to the number of sequences containing the S-TIRP.

[0080] In the vertical database, the position, start time, end time, time relationship, and related symbolic time intervals of each event eid in the sequence S sid need to be recorded. In this way, the key information can be directly used for pattern expansion, avoiding multiple scans of the sequence database.

[0081] Step S6, Perform a depth-first search (DFS) for each S-TIRP P in SF to expand the frequent pattern and identify the target pattern through pattern expansion.

[0082] Taking the horizontal database in Table 1 as an example, for the horizontal database shown in Table 1 which consists of five sequences S1 to S5, the query event sequence is AB. Sequences S2 and S4 that do not contain AB can be deleted, and the remaining S1, S3, and S5 form HD'.

[0083] In HD', the time interval related patterns include (A, _), (B, _), (C, _), (D, _), and their vertical supports are all 3. That is to say, S1, S3, and S5 all contain the time interval related patterns (A, ), (B, ), (C, ), (D, _).

[0084] Taking the extended time interval related pattern (A, _) as an example, it can be extended to (AA, _), (AB, _), (AC, _), (AD, _), and the support counts are all 3. For further extension, assuming the extension of (AA, ), the vertical support count of (AAA, ) is 1. If the threshold is 2, then (AAA, ) does not meet the condition and there is no need to extend it further. Continue to judge whether (AAB, ) meets the condition, and so on.

[0085] In one implementation, in order to identify the matching position of the query event sequence qes, a variable match is introduced to record the matching progress of the current pattern with the query event sequence.

[0086] For each S-TIRP P in SF, in order to extend to a longer S-TIRP, TaTIRP will perform a depth-first search (DepthFirstSearch) process to achieve pattern extension.

[0087] In one implementation, the input of the depth-first search process includes the prefix S-TIRP P, the query event sequence qes, a set SF containing all extensible frequent S-TIRPs, the required minimum support minSup, and the matching position match of the query event sequence. Denote the last event of P as lastEventOfP, and the current query event as qe.

[0088] Combined Figure 6 , the depth-first search process is as follows:

[0089] First, check whether lastEventOfP is equal to qe. If they are equal, then increase match by 1; otherwise, match remains 0.

[0090] Second, evaluate whether this S-TIRP has matched all events in the query event sequence. If match is equal to the length of the query event sequence, output the current pattern P.

[0091] If the vertical support after combining lastEventOfP and qe is less than or equal to minSup × |D|, the expansion operation of S-TIRP P is terminated according to the UQPP strategy (the third strategy).

[0092] If the vertical support after combining lastEventOfP and qe is greater than minSup × |D|, for each pattern F in SF, assign the value of match to newMatch to avoid missing a match during the depth-first search process.

[0093] For each pattern F in SF, if the value of PSM(lastEventOfP, F) is less than the threshold minSup × |D|, then according to the UEPP strategy (the second strategy), any pattern extended from these two patterns cannot become a frequent pattern.

[0094] For each pattern F in SF, if the value of PSM(lastEventOfP, F) is greater than or equal to the threshold minSup × |D|, then a vertical database can be constructed for the extended pattern and further determine whether the vertical support meets the frequent pattern condition. If it meets, then it can be further extended.

[0095] For the extensible pattern F, determine whether the current query event qe is equal to the pattern F; if they are equal, increment newMatch by 1. This process will continue until all target patterns are mined.

[0096] In practical applications, the time interval between two S-TIRPs may be too long or too large, resulting in a decrease in the significance of the mining results; therefore, we use maxGap and maxDura to constrain the size of the time interval.

[0097] In this case, the efficiency of the proposed method is verified and analyzed through experiments, mainly focusing on the running time. To evaluate the correctness and effectiveness of the designed strategy, we developed two variants of the TaTIRP algorithm. First, we adopted FastTIRP as the baseline algorithm and applied post-processing techniques to extract the target patterns corresponding to the query event sequence. The algorithm combining the baseline algorithm and the post-processing technique is called FastTIRP*. Since the second strategy has been verified to be effective in FastTIRP, in subsequent experiments, all variants of TaTIRP default to include the strategy. TaTIRP1 represents using Strategy 1 without using the third strategy, and TaTIRP2 represents using the third strategy without using the first strategy. TaTIRP12 applies all strategies simultaneously. In summary, there are five algorithms for comparison. All algorithms are implemented in Java, and the experiments are conducted on a computer equipped with a 64-bit Windows 10 operating system, 16GB of memory, and a 12th-generation Core i7-12700 processor.

[0098] Six datasets were used in this experiment, including four real-world datasets (ASL, Hepatitis, Diabetes, and Smarthome) and two synthetic datasets (DS 1 and DS2). These datasets have their own characteristics, which contribute to comprehensively evaluating the effectiveness of the algorithms from different perspectives.

[0099] · The American Sign Language dataset (ASL) contains a large number of videos demonstrating various gestures and actions of American Sign Language. These videos are usually performed by native ASL users to ensure the authenticity of sign language expressions. The annotations in the dataset mark the start and end times of each gesture, facilitating in-depth analysis of the fluency and temporal characteristics of sign language.

[0100] · The Hepatitis dataset is mainly used for medical and health research, especially in the analysis of hepatitis-related data. This dataset records information such as the treatment situation and results of patients, their survival status, survival time, occurrence of complications, and prognosis assessment.

[0101] · The Diabetes dataset is mainly used to analyze diabetes-related data and records the treatment plans and effects of patients, including information such as drug use, insulin injection, and blood glucose monitoring.

[0102] · The Smarthome dataset is a comprehensive dataset of a home automation system, covering multiple aspects of a smart home environment, such as sensor data, user interactions, environmental conditions, and device status.

[0103] · DS1 is a synthetic time interval dataset generated by the FTDPMiner-EP algorithm, containing 1,000 sequences, each sequence consists of 20 symbolic time intervals, and there are 100 event types in total.

[0104] · DS2 is also a synthetic time interval sequence dataset, containing 100,000 sequences, each sequence consists of 10 symbolic time intervals, and there are 100 event types in total.

[0105] By adjusting the minSup parameter, six datasets are experimented to evaluate the efficiency of the proposed algorithm. The experimental results are as Figure 7As shown. Generally speaking, as minSup gradually increases, the running times of the five comparison algorithms on all datasets will gradually decrease. It can be observed that FastTIRP consumes the longest running time on all datasets, although FastTIRP* adds a post-processing operation on its basis. This is because FastTIRP outputs a larger number of patterns than FastTIRP*, and the time required for transmitting the output results is longer than the time required for judging whether it is the target TIRP pattern. As can be seen from the figure, the running time of TaTIRP1 is shorter than that of TaTIRP2. The reason is that TaTIRP1 removes some time interval sequences at the initial stage of the algorithm, which not only helps to exclude some frequent patterns and their vertical databases that are irrelevant to the query event sequence earlier, but also reduces the exploration time for unpromising sequences when expanding the remaining frequent patterns. TaTIRP1 considers the availability of the query sequence events from an overall perspective, while TaTIRP2 judges whether the combination of the current expanded pattern and the query event is frequent and can form part of the target pattern from a local perspective during the pattern expansion process. Taking the ASL dataset as an example, the differences in running times between the algorithms are very obvious. FastTIRP and FastTIRP* show the slowest and nearly equal running times, followed by TaTIRP1 and TaTIRP2 with slightly shorter running times, while TaTIRP12 shows the shortest running time, indicating that the performance of the proposed algorithm has been significantly optimized when all strategies are applied. It should be noted that some of the running time curves do not show a linear decline. This is because when minSup is low, the change in minSup has little impact on the number of generated target time interval patterns. In dense datasets, since each sequence contains a large number of time intervals, the running times of TaTIRP1 and TaTIRP12 are almost the same and almost overlap in the figure. For example, in the Hepatitis and DS1 datasets, the red and green curves almost completely coincide, which is mainly due to the dominant role of Strategy 1. Through the description of the above implementation manners, those skilled in the art can clearly understand that the method of the present disclosure can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits, etc. However, in more cases for the present disclosure, software program implementation is a better implementation manner.

[0106] Although the embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, and all of these fall within the scope of protection of the present disclosure.

Claims

1. A method for accurately mining time interval-related patterns, characterized in that, The method includes the following steps: Based on a given time interval sequence database, construct a horizontal database HD, and arrange each time interval sequence in order in the horizontal database HD; Based on the horizontal database HD, identify all frequent S-TIRP patterns of length 1 and store them in the frequent pattern set SF. The S-TIRP is a set aggregated by multiple time interval-related patterns sharing the same event A, and the time interval-related patterns of this set are denoted as Take the number of time interval sequences that match as the vertical support. If the vertical support exceeds the set value, then is a frequent S-TIRP pattern; Successively obtain a frequent S-TIRP pattern from SF, perform depth-first search on the event sequence in the pattern for search expansion, and apply a non-promising expansion pattern pruning strategy and a non-promising query pattern pruning strategy during the search expansion to reduce join operations; At the end of the search expansion, regard the time interval-related patterns with vertical support greater than or equal to the set value as the target time interval-related patterns; In the non-promising expansion pattern pruning strategy, if the last event p of the current pattern and the current query event are combined into an event pair, and the vertical support of this pair of events is less than the set value, then terminate the expansion; In the non-promising query pattern pruning strategy, if the last event p and the event q to be expanded are combined into an event pair, and the vertical support of this pair of events is less than the set value, then terminate the expansion.

2. The method according to claim 1, wherein Based on the horizontal database HD, identify all frequent S-TIRP patterns of length 1. The steps include: Based on the query event sequence, filter out the time interval sequences in the horizontal database HD that do not contain the query event sequence to generate an updated database HD'; Scan the horizontal database HD' to identify all frequent S-TIRP patterns.

3. The method according to claim 1, wherein Arrange each time interval sequence in order. The steps include: The time interval sequence S is represented by the symbol time interval I, S = {I1, I2,..., I m}, the subscript of I represents the symbol time interval identifier, I = {start, end, e}, start represents the start time of the event type e, end represents the end time of the event type e, I.start, I.end and I.e respectively represent the values in the above triple; For two symbol time intervals I i and I j , if I i <I j , then I i 、I j satisfy (I i .start < I j .start), or when I i .start is equal to I j .start, I i .end < I j .end, or when I i .start is equal to I j .start and I i .end is equal to I j .end, I i .e < I j .e.

4. The method according to claim 1, wherein In the non-promising expansion pattern pruning strategy and in the non-promising query pattern pruning strategy, the vertical support of the paired events is obtained based on a pattern support matrix. The pattern support matrix is constructed based on the horizontal database HD'. The element value in the pattern support matrix is the vertical support of the S-TIRP pattern formed by a pair of events.

5. The method according to claim 1, wherein During the depth-first search for search expansion of the event sequence in the pattern, record the matching progress of the current pattern and the query event sequence to determine whether the query event sequence has been fully matched.

6. The method according to claim 1, characterized in that Before performing the depth-first search for search expansion of the event sequence in the pattern, establish a vertical database for each S-TIRP pattern in SF. The vertical database contains a list of sequence IDs and is extended to include the position, start time, end time, and related symbolic time interval of the event in the sequence ID, so that during the depth-first search for search expansion of the event sequence in the pattern, the symbolic time interval of the sequence and the time interval of the pattern can be directly obtained, and by constraining the symbolic time interval and the time interval of the pattern, meaningless patterns can be avoided from being generated.

7. The method according to claim 1, wherein When arranging each time interval sequence in order, it includes: setting a noise value ∈≥0. If the time points t i and t j satisfy |t i -t j |≤∈, then the time points of these two events are equal, denoted as t i =t j ; if (t i -t j )>∈, then it is expressed as t i <t j .

8. The method according to claim 7, characterized in that, Sort the constructed vertical database in descending order according to the number of sequences containing the S-TIRP pattern.

9. A computer-readable storage medium, characterized in that: There is stored a computer program that can be loaded and executed by a processor to perform any one of the methods in claims 1 to 8.