Parallelization motif transfer mode identification method and system in large-scale sequence diagram
Through the parallelized motif transfer pattern recognition method, the time area division strategy and PTMT algorithm are used to solve the problems of high computational complexity and large memory usage in large-scale timing chart processing, and efficient and accurate motif transfer counting is achieved, supporting ordinary server deployment and significantly improving processing speed.
Patent Information
- Application Number
- CN202510549562.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing methods have high computational complexity, huge memory usage when dealing with large-scale timing charts, and low utilization of multi-core resources, making it difficult to deploy on ordinary servers, and there are problems of redundant calculations and repeated counting.
The parallelized motif transfer pattern recognition method is adopted, and the timing chart is divided into independent regions through the time zone division strategy (TZP), and efficient calculation and precise counting are performed in combination with the three-stage framework. Multi-threaded parallel calculation is used to eliminate redundancy and graph isomorphic calculation is eliminated through hash encoding.
It significantly improves counting efficiency and scalability, reduces memory usage, supports ordinary server deployment, accelerates processing speed by 50 times, and has completely correct statistical results.
Smart Images

Figure CN120067652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pattern recognition, and in particular, to a method and system for parallel motif transfer pattern recognition in large-scale time series graphs. Background Art
[0002] A time series graph is a dynamic network structure whose edges and nodes change over time, and can truly reflect the dynamic evolution of the interaction relationships in complex systems. As a frequently occurring subgraph pattern in time series graphs, the transfer process of motifs reveals the dynamic characteristics of the network structure and is crucial for applications such as anomaly detection and behavior prediction.
[0003] Although motif analysis has important value, existing methods face the following challenges when dealing with large-scale time series graphs: High computational complexity, the time complexity of the existing best method (TMC) algorithm, resulting in excessive time consumption when processing billions of edges. In addition, the memory occupancy is huge. Existing methods require more than 300GB of memory to store all candidate motif transfer trajectories for most publicly available datasets, making it difficult to be deployed on ordinary servers. And existing methods rely on global synchronization, with an acceleration ratio of only 9.2% under 32 threads, resulting in low utilization of multi-core resources. Moreover, simple time window partitioning leads to repeated counting of motif transfers. For example, the triangle closure process across windows is counted multiple times by different threads. Summary of the Invention
[0004] To solve the above-mentioned problems, the present invention provides a method and system for parallel motif transfer pattern recognition in large-scale time series graphs.
[0005] In a first aspect, a method for parallel motif transfer pattern recognition in large-scale time series graphs provided by the present invention adopts the following technical solutions: A method for parallel motif transfer pattern recognition in large-scale time series graphs includes: Obtain the original time series graph; Perform data preprocessing on the obtained original time series graph and configure parameters; Perform execution area division based on the TZP algorithm; Perform multi-threaded parallel computing based on the PTMT algorithm to eliminate redundancy; Result parsing.
[0006] Further, the obtaining of the original time series graph includes obtaining the original time series graph =(V, E, T), where each edge e = (u, v, t) represents the interaction between node u and node v at timestamp t. The data file type is a text file, and the data content consists of several lines. Each line contains three columns of numbers representing the digital serial number corresponding to the first vertex, the serial number corresponding to the second vertex, and the timestamp of each temporal edge. The set of temporal edges formed by all lines completely describes the temporal graph.
[0007] Furthermore, the data preprocessing and parameter configuration for the obtained original temporal graph include unifying the unit of timestamp t in seconds to avoid cross-granularity errors; sorting in ascending order of timestamp to ensure the consistency of subsequent partitioning and processing; and removing abnormal timestamps and duplicate edges.
[0008] Furthermore, the data preprocessing and parameter configuration for the obtained original temporal graph include setting time constraints, maximum transfer step size, expansion factor, and number of threads respectively. Among them, the time constraint δ determines the time window for motif transfer, and the maximum transfer step size limits the number of motif conversions and determines the upper limit of the number of different types of motif edges in the final result. The expansion factor ω corresponds to the interval length determination parameter in the TZP partitioning algorithm.
[0009] Furthermore, the execution area partitioning based on the TZP algorithm includes calculating the time span of the growing area, defining the time range of the current growing area and extracting the edges that meet the time range, and defining the time range of the boundary area . By updating the time range of the boundary area until it covers the entire time axis, the partition set is output .
[0010] Furthermore, the multi-threaded parallel calculation based on the PTMT algorithm is used to eliminate redundancy, including allocating each in the partition set to an independent thread, and traversing the edges in the current growing area in chronological order. By updating the thread-local motif conversion candidate set ; collecting all threads' , after merging in the order of the growing area, each motif transfer instance is encoded as a unique string using hash coding, and the occurrence times of each encoding are counted through a hash table, and the global frequency table is output .
[0011] Further, the result parsing includes analyzing the transition relationship between digital string motifs through a pattern parsing algorithm and calculating the ratio of specific motif transitions; among them, extract the prefix of the motif, that is, the substring after removing the last two digits, traverse the PrefixMap, classify and count by prefix, and calculate the number of motif transitions and non - transitions; in the ratio calculation stage, calculate the ratio of a certain motif transition: If querying all, output the transition ratios of all motifs; if querying a specific motif, output the detailed transition situation of this motif.
[0012] In a second aspect, a parallel motif transition pattern recognition system in a large - scale time - series graph includes: A data acquisition module, configured to acquire the original time - series graph; A configuration module, configured to perform data pre - processing on the acquired original time - series graph and configure parameters; A region division module, configured to perform execution region division based on the TZP algorithm; A calculation module, configured to perform multi - thread parallel calculation based on the PTMT algorithm to eliminate redundancy; An analysis module, configured to perform result analysis.
[0013] In a third aspect, the present invention provides a computer - readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the parallel motif transition pattern recognition method in a large - scale time - series graph.
[0014] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer - readable storage medium, the processor is used to implement each instruction; the computer - readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the parallel motif transition pattern recognition method in a large - scale time - series graph.
[0015] In summary, the present invention has the following beneficial technical effects: The present invention divides the time - series graph into independently processable regions through a time - region division strategy (TZP), and combines a three - stage framework to achieve efficient calculation and accurate counting. A large - scale time - series graph motif counting strategy based on topological constraints is disclosed, which realizes efficient pruning through matrix operations, combines time - window division and parallel calculation, significantly improves the counting efficiency and scalability, and is applicable to large - scale dynamic graph analysis scenarios such as social networks and biological networks.
[0016] The present invention only takes 2,923 seconds to process 120 million edges under 32 threads in a large public dataset, accelerating 50 times compared with the traditional method. The peak memory occupancy is reduced from 320GB to 92GB, supporting the deployment on ordinary servers. The cross-region duplicate counting is eliminated through the TZP strategy, and the statistical results are completely correct. The streaming processing adapts to dynamic data, and the peak memory is controlled within 150GB. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of a method for parallel motif transfer pattern recognition in a large-scale temporal graph according to Embodiment 1 of the present invention; Figure 2 It is a schematic diagram of the temporal graph according to Embodiment 1 of the present invention; Figure 3 It is a schematic diagram of the TZP strategy according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The present invention will be further described in detail below with reference to the accompanying drawings.
[0019] Term Explanation: Temporal Graph: A dynamic network structure composed of nodes and edges that change over time, with time stamps attached to the edges.
[0020] Motif Transfer: The evolution process of a specific subgraph pattern (motif) over time in a temporal graph.
[0021] TZP Strategy (Temporal Zone Partitioning): A partitioning strategy that divides a temporal graph into a growth zone and a boundary zone to ensure unique counting of cross-region motif transfers.
[0022] Deterministic Encoding: Mapping motif instances to unique strings to avoid graph isomorphism calculations.
[0023] Embodiment 1 Referring to Figure 1 , a method for parallel motif transfer pattern recognition in a large-scale temporal graph in this embodiment includes: Obtaining the original temporal graph; Performing data preprocessing on the obtained original temporal graph and configuring parameters; Performing execution region division based on the TZP algorithm; Performing multi-threaded parallel computing based on the PTMT algorithm to eliminate redundancy; Result parsing.
[0024] Specifically, it includes the following steps: S1. Obtain interaction data such as social networks, 1. System architecture, Data collection layer: Consists of distributed crawlers or platform open APIs, and can simultaneously and parallelly pull multi-source logs such as Weibo, Twitter, Zhihu, etc.
[0025] Message queue layer: Uses Apache Kafka for buffering, and the Topics are divided into raw_interactions (raw) and clean_interactions (after cleaning).
[0026] Data cleaning layer: The consumer program subscribes to raw_interactions, and after formatting, deduplication, and exception removal, it writes to clean_interactions.
[0027] Storage layer: The cleaned data is landed in HDFS (Parquet files) or NoSQL (such as HBase, MongoDB) for sub-library and sub-table storage.
[0028] Scheduling and monitoring: Uses Airflow to trigger full-scale pull and incremental cleaning tasks regularly, and Prometheus + Grafana monitors the throughput, latency, and error rate of each node.
[0029] 2. Field definitions and formats, Required fields: u (initiator interaction user ID, Int64), v (interacted user ID, Int64), t (Unix timestamp, Int64, unit seconds, UTC).
[0030] Optional fields: type (interaction type, Int8, 1 = like, 2 = comment, 3 = private message), content_id (content identifier, String), geo (JSON object, lat / lng), app_version (client version, String).
[0031] As an example JSON: {"u":1001,"v":2002,"t":1679788800,"type":2,"content_id":"post_84321","geo":{"lat":31.23,"lng":121.47},"app_version":"8.4.2"} 3. Collection process, 1) API pull: Use the platform batch interface, and pull in pages when t > t_last_max, with 100,000 items per page.
[0032] 2) Distributed crawler: Based on Scrapy + Redis, multi - machine collaboration; sharding of user accounts and topic keywords; 3) Cleaning logic: Time verification (discard records where t > current UTC + 60s or t < 0); deduplication within the same second (keep the earliest one); formatted output of the quadruple (u, v, t, type), with type = 0 by default; write to clean_interactions.
[0033] S2. Construct a time - series graph based on the acquired data, 1. Data loading, Use Spark to read HDFS Parquet: val df = spark.read.parquet("hdfs: / / ... / clean_interactions / date=*") Select fields u, v, t and convert to a DataFrame.
[0034] 2. Global sorting and micro - batches, df_sorted = df.orderBy("t") to complete global time - series sorting; Slice t by hour and write to multiple small files for subsequent parallel processing.
[0035] 3. Design of node and edge tables, Node table nodes: CREATE TABLE nodes (node_id BIGINT PRIMARY KEY, first_seen BIGINT,last_seen BIGINT, degree INT); Edge table edges: CREATE TABLE edges (edge_id BIGINT PRIMARY KEY, src_id BIGINT, dst_idBIGINT, ts BIGINT, type TINYINT, INDEX(ts), INDEX(src_id,dst_id)); The edge table is partitioned and stored by ts.
[0036] 4. Graph construction process, 1) Node storage in the database: After deduplicating all u and v, write them to nodes, with first_seen = MIN(t) and last_seen = MAX(t).
[0037] 2) Edge warehousing: Write edges in batches and update nodes.degree in real time.
[0038] 3) Verification: Check that the minimum / maximum ts in edges is consistent with the corresponding field in nodes.
[0039] 5. Graph enhancement Add a prev_ts field to the edge table to calculate the time difference between the current edge and the previous edge of the same source node; Count the number of interactions of (u,v) within the statistical window for reference during pruning; Add labels in combination with user portraits (registration duration, number of fans, activity, etc.).
[0040] 6. Example pseudocode (Spark+GraphFrame) from graphframes import GraphFrame nodes = df.select("u").union(df.select("v")).distinct().withColumnRenamed("u","id") edges = df_sorted.selectExpr("u as src","v as dst","t as ts","type") graph = GraphFrame(nodes, edges) 7. Output file Time-series edge set .txt: # u v t 1001 2002 1679788800 1001 2003 1679788815 2003 2002 1679788900 S3. Perform data preprocessing on the constructed time-series graph Input: Original time-series graph , where each edge represents the interaction between node u and node v at timestamp t..
[0041] Among them, unify the unit of timestamp t in seconds (or milliseconds) to avoid cross-granularity errors. Sort in ascending order of timestamp to ensure the temporal consistency of subsequent partitioning and processing. Eliminate abnormal timestamps (such as future times or negative values) and duplicate edges (only keep one of multiple interactions of the same node pair at the same timestamp).
[0042] Output: Normalized temporal edge set .
[0043] S4. Configure parameters Time constraint : Dynamically set according to the application scenario. For example: Bitcoin transaction network: = 60 seconds (capture consecutive transactions within 1 minute); Social network interaction: = 10 minutes (analyze short-term user sessions).
[0044] Maximum transfer step size : Control the maximum number of transfers according to requirements, such as .
[0045] Expansion factor ω: For sparse data, you can choose to increase the parameter value of ω to reduce the number of partitions; for dense data, you can reduce the parameter value of ω to improve the parallel granularity.
[0046] Number of threads #thread: Set according to hardware resources, such as 32 threads (it is necessary to ensure that the number of threads ≤ the number of physical CPU cores).
[0047] S5. Perform region division (TZP algorithm) Divide the temporal graph into: 1. Growth Zone (Growth Zone): Divide the temporal graph into multiple independently processed regions along the time axis, and its time span is: Among them, ω is the thread expansion factor (used to determine the time span of the divided regions), δ is the time constraint (a parameter in the motif transfer process, determining the time limit for motif conversion), is the maximum transfer step size (a parameter in the motif transfer process, determining the step limit for motif conversion).
[0048] Design intention: Balance the partition granularity and parallel efficiency by adjusting ω.
[0049] 2. Boundary Zone (Boundary Zone): Set an overlapping region between adjacent growth regions, and the time span is: Uniqueness guarantee: The edge set of any motif transfer instance completely belongs to a single growth region or its boundary region.
[0050] Lemma 1 (Uniqueness Proof): Suppose has a time span of ≤ ; if spans multiple regions, its first edge must belong to a certain growing region , and subsequent edges are classified into 's boundary region due to 's coverage . Therefore, is only counted once by .
[0051] As a further implementation method, Input: Sorted edge set , time constraint , maximum transfer step size , expansion factor , number of threads .
[0052] Algorithm process: 1. Initialization: Starting time .
[0053] 2. TZP policy partitioning: Calculate the time span of the growing region: .
[0054] Define the time range of the current growing region as , and extract all edges that satisfy . Define the time range of the boundary region as , where . Enumerate the temporal edges of the current non-added temporal edge region from front to back , and add the temporal edges that satisfy the current or or time constraint to the corresponding region to form a subset of the edge set. Update , and repeat until the entire time axis is covered
[0055] Output: Partition set .
[0056] For a more detailed description, first, initialize the starting time to the earliest timestamp of all edges in the temporal edge set . Then enter the loop. As long as the current starting time does not exceed the latest timestamp in the edge set, perform the following operations: Calculate the end time of the current growing region , whose value is the sum of the products of the start time and the preset parameters , and , that is . Then, extract all the edges from the edge set whose timestamps are within the range of to form the current growth region . Subsequently, define the boundary region according to the formula , whose time range is to ensure the unique attribution of cross-region motif transfer. After adding the generated growth region and the boundary region to the partition set Q, update the start time to the current for processing the next time window. Loop and repeat the above steps until the entire time axis is covered, and finally return the set Q containing all partitions to complete the dynamic partitioning of the time series graph.
[0057] S6. Execute the parallel computing framework (PTMT), where (1) the growth region is expanded in parallel, Input: the growth region assigned to thread k , set the current transfer candidate set M k as an empty set.
[0058] Process: Traverse the edge set in chronological order, retrieve whether the current edge and each in have common vertices, and update the candidate and transfer by judging whether the constraints such as are satisfied. And the following two constraints need to be satisfied. Time constraint: the interval between adjacent edges . Step size constraint: the transfer step size . When and the motif in i have common vertices and satisfy the time and step size constraints, add to the motif i and complete one transfer; if e does not meet the above conditions, then directly add to to complete the update of . Finally, output the local transfer candidate set .
[0059] (2) Overlap-aware aggregation, Principle: When merging the results of each thread, since our partitioning method ensures that duplicate elements only appear between two adjacent partitions, we use the formula and combine it with the region partitioning strategy to eliminate duplicate counting or omission caused by ordinary parallel partitioning.
[0060] Example: If a certain transfer is counted as 10 and 8 respectively in and , but is double-counted as 5 in , then its global value is 10 + 8 - 5 = 13.
[0061] Process: Since there is no data dependency between non-adjacent partitions, overlap elimination can be efficiently completed by automatically allocating threads. The specific implementation method is as follows: In the first stage, obtain each local motif transfer candidate set . For adjacent candidate sets i and i + 1, use the formula to obtain the correct result and reduce it to the final set to obtain the global motif conversion set .
[0062] (3) Deterministic encoding, Encoding rule: After obtaining the global motif conversion set , map various in it to a fixed-length string. Its formal description is as follows: where f(Ai): V → N is a continuous mapping function for node IDs (mapping node names to a number for easy encoding), Ai is all nodes in the Motif in chronological order, and ⊕ is string concatenation. This method aims to avoid the complexity O(n!) of graph isomorphism calculation and achieve O(1) hash duplicate checking.
[0063] (4) Dynamic load balancing and optimization, Adaptive partitioning: Dynamically adjust the ω value according to data density: As one of the hyperparameters of the method, ω supports incremental processing of tens of billions of edges by controlling the rolling partition length of the time window, reducing the memory peak by 80%.
[0064] As a further implementation method, (1) Growth region parallel expansion, Thread allocation: Allocate each in to an independent thread, and the number of threads is set by the parameter .
[0065] Candidate generation: Traverse the edges in chronological order in, and for each edge , call the edge expansion conversion function: The pseudocode is as follows: Update the thread-local transformation candidate set .
[0066] (2)Overlap-aware aggregation, Collect all threads' , and merge them in the order of growing regions. For adjacent regions and , apply the formula to correct the duplicate count: Finally, obtain the global transformation set.
[0067] (3)Deterministic encoding, First, use hash encoding to encode each motif transition instance into a unique string.
[0068] For example, for the motif sequence ⟨(A,B,1:00),(B,C,1:20)⟩, it is mapped to: Code(⟨(A,B,1:00),(B,C,1:20)⟩)=f(A)⊕f(B)⊕f(B)⊕f(C)="0112" where f is the mapping from node ID to consecutive integers (e.g., A→0, B→1, C→2).
[0069] Subsequently, obtain the global frequency table M: Count the occurrences of each encoding through a hash table, and the output format is: encoding: count such as 01010203: 1896181, 01010101: 289241, Among them, in the PTMT algorithm, in the first stage, the temporal edges are divided into multiple regions and processed in parallel. Through the try_to_transit(e,δ, , ) function, under the δ time constraint, Expand candidate motifs under the constraints of the maximum path length and node continuity conditions; in the second stage, the candidate motif sets generated in each region are merged into a global set , and duplicates generated by overlapping boundary regions are eliminated through a conflict resolution mechanism; in the third stage, the motif transfer process after merging is encoded into a fixed-length string, and finally a frequency mapping table is returned .
[0070] Step S7. Result parsing and application, In practical applications, this method parses motif (digit string) transfer data by given query parameters and calculates its transfer ratio. The specific implementation methods include the format of input data, query methods, algorithm execution processes, and the structure of the final output.
[0071] First, the input data format is a text file (such as the global frequency table M obtained in step 4), and each line contains a motif (digit string) and its occurrence times. For example: 011212 10 011213 5 010102 7 0112 18 Among them, 011212 represents the motif, and 10 represents its occurrence times. There may be shared prefixes among the motifs in the data. For example, 011212 and 011213 have the same prefix 0112, which indicates possible motif transfers.
[0072] Query parameter settings: The user specifies the following parameters: <input_file>: The path of the input data file.
[0073] <query>: Query for a specific motif prefix, or use "all" to query all motifs.
[0074] [min_length max_length] (optional): Specify the minimum and maximum lengths of the motif, with the default range being 4 to 8.
[0075] For example, execute the command: . / motif_analysis dataset.txt 0112 4 8 This means querying for the motif transition of the prefix 0112 in dataset.txt, and restricting the motif length to be between 4 and 8.
[0076] Algorithm execution process: Data preprocessing: Read dataset.txt, parse the motifs and occurrence counts, and extract the motif prefixes (remove the last two digits, e.g., 011212 becomes 0112). Record the motif transition data PrefixMap["0112"]["011212"] = 10.
[0077] Accumulate the total transition count for this prefix CountMap["0112"] += 10.
[0078] Query the motif transition ratio: If querying "all", traverse all prefixes in PrefixMap. If querying a specific motif, such as 0112, then only calculate PrefixMap["0112"]. Calculate the ratio of each motif's transition occurrence: Output result: Generate the result file result_dataset.txt, with an example as follows: Motif: 0112 011212: 10 (66.67%) 011213: 5 (33.33%) Total Evolved: 15 (83.33%) If there is no transition for the motif, then output "No data found for prefix: <query>".
[0079] Final output: The result finally output by this method is the ratio of motif transfer, which helps users analyze the evolutionary relationship between motifs. For example, users can query the global motif statistics through "all", or specify a query to query the transfer situation of specific motifs. The obtained analysis results can be used in the following fields: High-frequency pattern recognition: Bitcoin trading network: For example, the path "010203" of Motif 0102 accounts for 70.75%, which is marked as multi-account collaborative trading and triggers the risk trading identification rule. If the path frequency of a certain account "010101" exceeds the threshold (such as 100 times within 1 minute), it is determined as a self-loop risk trading.
[0080] Wikipedia collaboration network: For example, the path "010102" of Motif 0101 accounts for 30.72%, which is identified as multi-person relay editing and marked as a high-quality collaborative entry. When the frequency of the path "011213" of Motif 0112 suddenly increases, the administrator is automatically notified to review potential editing conflicts.
[0081] S8. Final output and application scenario implementation, Output the global motif frequency table (CSV), prefix transfer report (TXT / JSON), and real-time API ( / motif / report?prefix=0102).
[0082] Typical applications: anti-fake account detection, interest recommendation, hot topic monitoring.
[0083] In the Wikipedia user discussion page interaction data (WikiTalk dataset), set hour, , analyze the motif transfer pattern of user editing behavior, and reveal the following key collaboration and conflict dynamics: 1. Motif 0101 (unilateral dominant interaction) High-frequency paths: "010101" (34.75%): A single user continuously edits the same discussion page, which may be for content maintenance or controversial modification.
[0084] "010102" (30.72%): After user A edits user B's discussion page, user B quickly responds to user C, forming a chain interaction.
[0085] Anomaly detection: If the frequency of "010101" is abnormally high (for example, a single user accounts for more than 50%), it may involve malicious screen swiping or destructive editing.
[0086] Collaboration analysis: The "010102" path reflects relay discussions between users, which is common when multiple people collaborate to improve entry content.
[0087] 2. Motif 0102 (Triangle Collaboration Network) Dominant Path: "010203" (70.75%): The continuous editing of users A→B→C forms a triangular information flow, which may be a collaborative revision of controversial content by multiple people.
[0088] "010223" (1.51%): The closed-loop interaction of users A→B→C→A reflects the reaching of consensus or escalation of conflict in the discussion.
[0089] Community governance: High-frequency triangular paths can identify core editor groups and assist administrators in allocating permissions.
[0090] Conflict warning: If the closed-loop path "010223" is accompanied by an edit rollback, it may indicate the outbreak of an edit war.
[0091] 3. Motif 0121 (Complex Collaboration Level) Core Path: "012123" (28.28%): Chain propagation of users A→B→C→D, which may be topic diffusion or cross-community collaboration.
[0092] "012131" (14.37%): The nested closed loop of users A→B→C→A, reflecting that core users dominate the discussion agenda.
[0093] In the editing of a certain popular entry, the "012123" path accounts for more than 40%, indicating efficient collaboration across user groups; if the "012131" path surges in a short period of time, it may indicate that a small number of users monopolize the right to speak and need to intervene to balance.
[0094] 4. Motif 0112 (Conflicting Editing Mode) Abnormal path: "011213" (27.14%): Repeated modifications by user A→B→A→C may be an editing war or content tug-of-war.
[0095] "011212" (14.80%): The cyclic confrontation of users A→B→A→B indicates a continuous conflict between the two parties.
[0096] Detection rules: If the proportion of the "011213" path in a certain discussion page exceeds 30%, the administrator review will be automatically triggered; Implement real-time monitoring of the "011212" path to prevent malicious damage.
[0097] As a further implementation method, To further improve the efficiency of the large-scale temporal motif transfer discovery system, the present invention designs and introduces a pre-pruning module that runs before the main computing process of the system. This module combines structural constraints and frequency constraints to pre-screen motif transfer paths in advance, thereby significantly reducing the scale of the motif candidate set and accelerating the discovery process of high-frequency motif transfers. The method proposes a structural pruning idea and combines it with a time-aware mechanism to form a pruning framework for the motif transfer scenario.
[0098] The pruning strategy includes three types of mechanisms: structural pruning, frequency pruning, and time pruning, which are specifically as follows: 1. Structural Pruning This part is based on the transfer relationship between and analyzes whether it meets specific structural constraints (such as triangle closure, star center) to eliminate structurally invalid combinations in advance.
[0099] Construct the adjacency matrix of the static transfer graph between motifs , where A[i][j] = 1 indicates that can evolve into ; For the triangle closure structure, construct a pruning matrix: used to identify the mutual transfer reachability between three points; For the star structure, construct a pruning matrix: S used to identify the multi-point transfer mode of the shared center ; If the value of the transfer pair (i, j) in the above matrix is 0, it means that it does not meet the target structure and can be directly excluded to avoid participating in the main process calculation.
[0100] 1.1 Pruning of triangle motifs, Definition of the adjacency matrix: Let the static adjacency matrix of the temporal graph be , and use to represent the element in the i-th row and j-th column of the matrix, where represents the number of edges from node to .
[0101] Define the symmetric adjacency matrix of an undirected graph as: where is the transpose of, representing a two-way edge.
[0102] Triangle topology matrix : The element of represents the number of paths of length 2 between nodes and .
[0103] is the Hadamard product (element-wise multiplication), and it is required that there must be a direct edge connection at the end of the path to ensure the formation of a triangle.
[0104] Pruning rule (Rule 1): If , then nodes and cannot form any triangle motifs.
[0105] Proof logic: If and only if there exists at least one node such that forms a closed triangle. If , then there is no such path.
[0106] 1.2 Pruning of star and pair motifs Star topology matrix : Represents that there is a two-way edge between nodes and .
[0107] Pruning rule (Rule 2): The conditions for node to be a star core or a pair motif node: 1. , satisfying any of the following conditions: (There is a two-way edge) ( to has multiple edges) ( to having multiple edges) 2. Node degree .
[0108] Proof logic: The star motif requires the central node to be connected to at least 3 edges, and the paired motif requires multiple interactions between nodes.
[0109] 1.3 Pruning algorithm implementation (Algorithm 3: getTCM) 1. Calculate the matrices and .
[0110] 2. Traverse all nodes : If , mark as a pruning node (unable to participate in triangle motifs).
[0111] If and and and the degree < 3, mark as a pruning node (unable to participate in star / paired motifs).
[0112] 3. Return the pruning mapping table Map M and Map S , guiding subsequent calculations to skip irrelevant nodes.
[0113] 2. Frequency-based Pruning, To avoid processing sporadic or extremely low-frequency motif transfer paths, the system sets a frequency lower threshold (e.g., minimum occurrence times ≥ 5). If the frequency of a certain motif or its associated transfer pair is lower than this threshold, the corresponding path is removed from the transfer graph, thereby reducing redundant candidates.
[0114] This mechanism can be approximately regarded as a "dimensionality reduction" operation in a high-dimensional sparse space, compressing the motif space into a subspace with stronger statistical stability and improving the quality of candidate sampling.
[0115] 3. Integrate into the system preprocessing stage, Step description: 1. Input preprocessing: Input the timing diagram , time constraint , partition length .
[0116] 2. Temporal Partitioning: Generate a set Q of overlapping temporal partitions.
[0117] 3. Partition-Level Pruning: For each partition : a. Extract subgraphs .
[0118] b. Apply Algorithm 1 (getTCM) to calculate the pruning mapping tables Map M and Map S .
[0119] c. Delete the nodes marked for pruning and retain the candidate node sets , where and are the sets of nodes to be pruned obtained from the pruning mapping tables Map M and Map S respectively.
[0120] 4. Output: The candidate node sets and subgraphs of each partition for use by the subsequent motif counting module.
[0121] Mathematical Tools and Complexity: Adjacency matrix power operation is used for path counting, and the Hadamard product is used for constrained closed structures. The complexity of the pruning phase is (matrix multiplication), but can be reduced to through sparse matrix optimization. After partitioning, the scale of each subgraph is reduced, further reducing the computational amount.
[0122] System Integration Advantages: Reduce the number of candidate nodes and the search space for subsequent motif enumeration. Achieve parallelization through temporal partitioning to improve the overall throughput. Can customize pruning rules for specific motif types (such as triangles, stars) to flexibly adapt to system requirements.
[0123] As Figure 2 shown, the temporal interaction between nodes is demonstrated through a timing diagram example, where A, B, C, D represent vertices, and 3:10, 10:10, etc. represent the timestamps corresponding to the directed timing edges.
[0124] As Figure 3 shown, the schematic diagram of the TZP strategy.
[0125] Example 2 This example provides a parallel motif transition pattern recognition system in large-scale temporal graphs, including: A data acquisition module, configured to: A computer-readable storage medium stores a plurality of instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the method.
[0126] A terminal device includes a processor and a computer-readable storage medium. The processor is configured to implement each instruction; the computer-readable storage medium is configured to store a plurality of instructions, and the instructions are adapted to be loaded and executed by the processor to perform the method.
[0127] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.< / query> < / query>
Claims
1. A parallel motif transfer pattern recognition method in a large-scale time series graph, characterized in that: include: Get the original timing diagram; Perform data preprocessing on the acquired original time series diagram and configure parameters; Execute area division based on TZP algorithm; Multi-threaded parallel computing based on PTMT algorithm to eliminate redundancy; Result analysis.
2. The parallelized motif transfer pattern recognition method in a large-scale time sequence graph according to claim 1, characterized in that: The obtaining of the original timing diagram includes obtaining the original timing diagram =(V,E,T), where each edge e=(u,v,t) represents the interaction between node u and node v at timestamp t. The data file type is a text file, and the data content is a number of lines. The three columns of numbers in each line represent the numerical sequence number corresponding to the first vertex of each timing edge, the sequence number and timestamp corresponding to the second vertex. The timing edge set composed of all rows completely describes the timing graph.
3. The parallelized motif transfer pattern recognition method in a large-scale time sequence diagram according to claim 2, characterized in that: The data preprocessing and parameter configuration of the acquired original time series graph include unifying the time stamp t into seconds to avoid cross-granularity errors; arranging in ascending order of the time stamp to ensure that the subsequent partitions are consistent with the processing time series; Eliminate timestamp anomalies and duplicate edges. The timestamp unit is unified and guaranteed by the initial TXT text acquisition stage, and the deletion of the same lines is completed to achieve the required task of removing duplicate edges; the ascending order of timestamps is implemented by the C++ programming language using the system library function sort(); the elimination of abnormal time values is achieved through the judgment statement in the program.
4. The parallelized motif transfer pattern recognition method in a large-scale time sequence graph according to claim 3, characterized in that: The data preprocessing and parameter configuration of the acquired original timing diagram include setting time constraints, maximum transfer step length, expansion factor and number of threads respectively, wherein the time constraints δ Determine the time window for motif transfer and the maximum transfer step length Limit the number of motif conversions and determine the upper limit of the number of different motif edges in the final result, the expansion factor ω Corresponds to the interval length determination parameter in the TZP partitioning algorithm.
5. The parallelized motif transfer pattern recognition method in a large-scale time sequence graph according to claim 4, characterized in that: The execution area division based on the TZP algorithm includes calculating the time span of the growth area and defining the current growth area. The time range is extracted and the edges that meet the time range are defined, and the boundary area is defined time range, by updating the boundary region The time range is increased until the entire time axis is covered, and the partition set Q is output.
6. The parallelized motif transfer pattern recognition method in a large-scale time sequence graph according to claim 5, characterized in that: The multi-threaded parallel computing based on the PTMT algorithm is used to eliminate redundancy, including partitioning the set Each of Assign to independent threads and traverse the current growth area in chronological order The edges in the thread are transformed into candidate sets by updating the thread-local motif ; Collect all threads After merging in the order of growth areas, each motif transfer instance is encoded into a unique string using hash coding, and the number of occurrences of each code is counted through the hash table to output a global frequency table .
7. The parallelized motif transfer pattern recognition method in a large-scale time sequence graph according to claim 6, characterized in that: The result analysis includes analyzing the transfer relationship between digital string motifs through a pattern analysis algorithm and calculating the transfer ratio of a specific motif; wherein, the prefix of the motif is extracted, that is, the substring with the last two digits removed, the PrefixMap is traversed, and classification statistics are performed by prefix, and the number of motifs that have been transferred and those that have not been transferred is calculated; in the ratio calculation stage, the transfer ratio of a certain motif is calculated: If you query all, the transfer ratios of all motifs will be output; if you query a specific motif, the detailed transfer information of the motif will be output.
8. A parallelized motif transfer pattern recognition system in a large-scale time series graph, characterized in that: include: The data acquisition module is configured to acquire the original timing diagram; A configuration module is configured to perform data preprocessing on the acquired original timing diagram and configure parameters; The region partitioning module is configured to perform region partitioning based on the TZP algorithm; The computing module is configured to perform multi-threaded parallel computing based on the PTMT algorithm to eliminate redundancy; The parsing module,is configured to,parse the results.
9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the method according to claim 1 .
10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the method as claimed in claim 1 .
Citation Information
Patent Citations
Global maximization of time limit revenues by a travel provider
CA2793186A1
High-resolution remote sensing image segmentation method based on inter-scale mapping
CN104361589A
Graph embedding learning method based on graph primitives
CN111581445A
Network motif-based local high-order community discovery method and device
CN113870042A
Cross-layer walk community detection method based on motif perception
CN117036079A