Data running state analysis method and system for multi-dimensional data association mining

CN122818282APending Publication Date: 2026-09-25SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611043111.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

基于时序数据库的事件索引匹配方法侧重于将事件按时间顺序排列成序列后进行相似度计算,其底层模型假设事件之间存在明确的前后继关系且路径唯一;基于状态转移特征提取的关联分析方法虽然能够处理一定的分支结构,但其分支数量通常被预设为有限且已知,难以应对事件链内部节点出度不确定、分岔层级不可预知的复杂情形

Benefits of technology

[0020]综上所述,本申请提供的一种面向多维数据关联挖掘的数据运行状态分析方法通过构建混合事件图并引入超边结构,能够完整保留事件间的多对一汇聚关系与一对多分岔关系,为后续识别非规则化网状事件链提供准确的拓扑基础;通过对事件状态节点平均驻留时间的变点检测并辅以相对时间偏移矩阵的扩展约束,可以准确定位时序异常的事件节点,进而抽取包含多个起点或多个终点的候选子图;在此基础上,通过计算融合起点节点数、终点节点数和内部节点分支熵的多起点-多终点显著度并以此判定非规则事件链,能够自动发现组织运行中传统线性分析方法难以捕捉的多因汇聚与多果发散型事件传导路径,弥补了现有技术对网状事件链识别能力的缺失;进一步地,将事件链显著度作为增强权重附加至人员节点与事件节点之间并生成增强型多维关联图,可以将非规则事件链中隐含的高阶关联关系显式融入全局关联强度矩阵,使得下游分析能够充分利用事件链挖掘结果;最终,通过综合节点带权特征向量中心性、属性漂移程度以及事件链与频繁模式库的相似度进行状态判定,可以在区分罕见正常模式与真实异常模式的基础上输出状态判定结果,用以实现对组织运行中非规则化事件链的自动识别与准确状态评估,有效降低异常漏报和误判风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818282A_ABST
    Figure CN122818282A_ABST
Patent Text Reader

Abstract

The application relates to a data running state analysis method and system facing multi-dimensional data correlation mining, which comprises the following steps: time alignment and entity alignment are performed on multi-dimensional heterogeneous event data to generate a standardized event tuple set and an entity library; a mixed event graph is constructed by taking an event state as a node, a time sequence causal relationship as a directed edge and a convergence and divergence relationship as a hyperedge, and an average residence time sequence and a relative time offset matrix are calculated; a candidate node is obtained by performing a change point detection on the average residence time sequence, a constraint is expanded based on the relative time offset matrix, a candidate subgraph is extracted, a multi-start point-multi-end point significance is calculated, and an irregular event chain is determined according to the significance; an enhanced weight of the event chain significance is attached between personnel nodes and event nodes to generate an enhanced multi-dimensional correlation graph and a correlation strength matrix; and a node running state score is calculated based on the enhanced multi-dimensional correlation graph and the correlation strength matrix, an event chain state is determined in combination with a similarity in a frequent pattern library, and a result is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to a data operation status analysis method and system for multidimensional data correlation mining. Background Technology

[0002] In the field of organizational operations analysis and process mining, event correlation analysis based on multidimensional data is an important means of identifying operational patterns and discovering potential anomalies. These methods typically collect event records from multiple data sources, such as business systems, personnel management, and resource scheduling. By analyzing the temporal sequence, causal transmission, and entity participation relationships between events, they construct correlation chains between events, thereby assessing the health of organizational operations. However, in real-world organizational scenarios, the triggering and transmission of events often do not proceed linearly in a single direction. For example, a project delivery delay may be triggered by delays in early decision-making from multiple departments, and after several parallel or sequential sub-tasks, it ultimately manifests different consequences such as timeouts or quality deviations at multiple delivery nodes. This complex transmission pattern of "multiple causes converging and multiple effects diverging" gives the event chain in organizational operations a naturally multi-starting point, multi-end point, and internally multi-branched network topology.

[0003] Existing event association analysis methods typically treat event chains as linear progressions or finite branching structures, assuming events have a single beginning and propagate in a single direction. Event indexing and matching methods based on time-series databases focus on arranging events into sequences in chronological order and then calculating similarity. Their underlying models assume clear successor relationships between events and unique paths. While association analysis methods based on state transition feature extraction can handle certain branching structures, the number of branches is usually pre-set to be finite and known, making it difficult to handle complex situations where the out-degree of nodes within the event chain is uncertain and the bifurcation level is unpredictable. Furthermore, association mining methods based on knowledge graphs often use binary relationships (node-edge) for modeling when handling multi-source data fusion, lacking effective means of expressing higher-order combinations of three or more nodes acting on the same goal. When faced with irregular event chains that commonly occur in organizational operations, characterized by multiple convergence points, multiple intermediate bifurcations, and multiple divergence points, these methods often only identify local fragments while missing complete causal transmission paths, leading to distorted operational status assessment results and inaccurate identification of root causes of anomalies. Summary of the Invention

[0004] Based on this, the purpose of this invention is to provide a data operation status analysis method and system for multidimensional data association mining that can perform implicit association mining and automatic identification of irregular multi-starting-multi-ending-point event chains in organizational operation.

[0005] The objective of this invention is achieved through the following solution:

[0006] In a first aspect, the present invention provides a data operation status analysis method for multidimensional data association mining, comprising the following steps:

[0007] S1: Perform time and entity alignment on the event data in the collected multidimensional heterogeneous data, generate a standardized event tuple set with a source entity identifier set, a target entity identifier set and an attribute vector for each event, and extract personnel entities, resource entities and time period entities to form a standardized entity library;

[0008] S2: Construct a hybrid event graph for the standardized event tuple set, with event states as nodes, temporal causal relationships as directed edges, and many-to-one convergence relationships and one-to-many bifurcation relationships as hyperedges, and calculate the average dwell time series of each event state node and the relative time offset matrix between state nodes.

[0009] S3: Accumulate and detect change points in the average residence time series to obtain event nodes with time anomalies as candidate nodes. Extend the constraints on the subgraphs where the candidate nodes are located based on the relative time offset matrix. Extract candidate subgraphs containing at least two start points or at least two end points from the mixed event graph. Calculate the multi-start point-multi-end point saliency based on the number of start point nodes, the number of end point nodes, and the internal node branch entropy in each candidate subgraph. Event subgraphs with saliency exceeding the preset graph saliency threshold are judged as irregular event chains and stored in the irregular event chain set.

[0010] S4: The saliency of the event chains in the irregular event chain set is used as an enhancement weight to be added between personnel nodes and event nodes in the standardized entity library, and the association strength between node pairs is calculated based on the enhanced node relationship to generate an enhanced multidimensional association graph and association strength matrix;

[0011] S5: Based on the enhanced multidimensional association graph and association strength matrix, calculate the weighted eigenvector centrality of each node under the association strength matrix, and calculate the running state score of each node by combining the degree of drift of node attributes relative to the historical baseline. Based on the distribution of state scores of nodes within each irregular event chain and the similarity between the irregular event chain and the frequent patterns in the frequent pattern library, output the state determination result of the irregular event chain; whereby, the frequent pattern library is constructed from historical standardized event tuples based on the association rule mining algorithm.

[0012] Secondly, this invention provides a data operation status analysis system for multidimensional data association mining, which is configured with the following modules:

[0013] The heterogeneous data standardization module is used to perform time and entity alignment on event data in the collected multidimensional heterogeneous data, generate a standardized event tuple set with a source entity identifier set, a target entity identifier set and an attribute vector for each event, and extract personnel entities, resource entities and time period entities to form a standardized entity library.

[0014] The hybrid event graph construction module is used to construct a hybrid event graph from a set of standardized event tuples, with event states as nodes, temporal causal relationships as directed edges, and many-to-one convergence relationships and one-to-many bifurcation relationships as hyperedges, and to calculate the average dwell time series of each event state node and the relative time offset matrix between state nodes.

[0015] The irregular event chain mining module is used to accumulate and detect change points in the average dwell time series, obtain event nodes with time anomalies as candidate nodes, and extend the constraints on the subgraphs where the candidate nodes are located based on the relative time offset matrix. It extracts candidate subgraphs containing at least two start points or at least two end points from the mixed event graph. It calculates the multi-start point-multi-end point saliency based on the number of start point nodes, the number of end point nodes, and the internal node branch entropy of each candidate subgraph. Event subgraphs with saliency exceeding the preset graph saliency threshold are judged as irregular event chains and stored in the irregular event chain set.

[0016] The multidimensional association graph enhancement module is used to add the saliency of the event chain in the irregular event chain set as an enhancement weight to the personnel nodes and event nodes in the standardized entity library, and calculate the association strength between node pairs based on the enhanced node relationship to generate an enhanced multidimensional association graph and association strength matrix.

[0017] The event chain state assessment module is used to calculate the weighted eigenvector centrality of each node under the association strength matrix based on the enhanced multidimensional association graph and association strength matrix, and to calculate the running state score of each node by combining the degree of drift of node attributes relative to the historical baseline. Based on the state score distribution of nodes within each irregular event chain and the similarity between the irregular event chain and frequent patterns in the frequent pattern library, the module outputs the state judgment result of the irregular event chain. The frequent pattern library is constructed from historical standardized event tuples based on the association rule mining algorithm.

[0018] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the above-mentioned data operation status analysis methods for multidimensional data association mining.

[0019] Fourthly, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements any of the above-mentioned data operation status analysis methods for multidimensional data association mining.

[0020] In summary, the data operation status analysis method for multidimensional data association mining provided in this application, by constructing a hybrid event graph and introducing a hyperedge structure, can completely preserve the many-to-one convergence relationship and one-to-many bifurcation relationship between events, providing an accurate topological foundation for subsequent identification of irregular network event chains. By detecting the change point of the average residence time of event state nodes and supplementing it with the extended constraint of the relative time offset matrix, it can accurately locate event nodes with temporal anomalies, and then extract candidate subgraphs containing multiple start points or multiple end points. On this basis, by calculating the multi-start point-multi-end point significance by fusing the number of start point nodes, the number of end point nodes, and the internal node branch entropy, and using this to determine irregular event chains, it can automatically discover multiple events that are difficult to capture by traditional linear analysis methods in organizational operations. Because of its convergent and multi-effect divergent event propagation paths, this approach compensates for the lack of existing technologies in identifying network event chains. Furthermore, by adding event chain saliency as an enhancing weight between personnel nodes and event nodes to generate an enhanced multidimensional association graph, the high-order associations implicit in irregular event chains can be explicitly integrated into the global association strength matrix, allowing downstream analysis to fully utilize the event chain mining results. Finally, by comprehensively considering the centrality of node weighted eigenvectors, the degree of attribute drift, and the similarity between the event chain and the frequent pattern library, a state determination can be made. This allows for the output of state determination results based on distinguishing between rare normal patterns and true abnormal patterns, thereby achieving automatic identification and accurate state assessment of irregular event chains in organizational operations, effectively reducing the risk of anomaly underreporting and misjudgment.

[0021] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0022] Figure 1 A flowchart illustrating a data operation status analysis method for multidimensional data association mining provided in this application embodiment;

[0023] Figure 2 This is a schematic diagram of the structure of a data operation status analysis system for multidimensional data association mining, provided as another embodiment of this application. Detailed Implementation

[0024] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0026] In one embodiment, such as Figure 1 As shown, a data operation status analysis method for multidimensional data association mining is provided. This embodiment illustrates the method's application to a terminal. It is understood that this method can also be applied to a server, and to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0027] S1: Perform time and entity alignment on the event data in the collected multidimensional heterogeneous data, generate a standardized event tuple set with a source entity identifier set, a target entity identifier set and an attribute vector for each event, and extract personnel entities, resource entities and time period entities to form a standardized entity library.

[0028] Specifically, the system collects event records from multiple heterogeneous data sources, including business support systems, operations support systems, customer relationship management systems, human resource management systems, project management systems, and IoT monitoring devices. These data sources differ in data structure, storage format, timestamp recording methods, and entity encoding rules. The system performs time alignment on each collected raw event record, converting timestamps from different data sources into a standard time format and correcting clock drift between systems. During time alignment, the system selects a reference event co-occurring across systems as a benchmark, estimates the offset function of each data source relative to the benchmark clock, and corrects the original timestamps accordingly. Simultaneously, the system performs entity alignment, extracting personnel designations, resource designations, and time period markers from the raw records of each data source. Through multi-dimensional similarity calculations, including string edit distance, attribute feature similarity, co-occurrence relation similarity, and semantic vector similarity, different representations pointing to the same real-world entity are merged into the same entity cluster, and a globally unique identifier is assigned to each entity cluster.

[0029] After time and entity alignment, the system converts each original event record into a standardized event tuple. Each tuple contains a globally unique identifier for the event, a source entity identifier set, a target entity identifier set, a standardized timestamp, and an attribute vector. The source entity identifier set records all entities that triggered the event, the target entity identifier set records all entities affected by the event, and the attribute vector is extracted and quantized into a real-valued vector from the original record by the system according to a preset attribute template. The system organizes all standardized event tuples into a set and stores it in the event database. Simultaneously, the system extracts all occurrences of entity identifiers from the standardized event tuples, categorizes entities into personnel entities, resource entities, and time-cycle entities based on entity type, and records the globally unique identifier and basic attribute information aggregated from various data sources for each entity, forming a standardized entity database.

[0030] S2: Construct a hybrid event graph for the standardized event tuple set, with event states as nodes, temporal causal relationships as directed edges, and many-to-one convergence relationships and one-to-many bifurcation relationships as hyperedges, and calculate the average dwell time sequence of each event state node and the relative time offset matrix between state nodes.

[0031] Specifically, the system takes a standardized set of event tuples as input and constructs a hybrid event graph. First, the system performs cluster analysis on the attribute vectors of the standardized event tuples, dividing the attribute vector distribution into several clusters. Each cluster corresponds to an event state type. The system assigns a unique state node identifier to each event state and maps all event tuples belonging to the same cluster to that state node. The system then establishes directed edges for temporal causal relationships between event state nodes. For any two event state nodes, the system retrieves the event instances contained in the two states. If there exists a pair of event instances where the source entity set or the target entity set intersects and the timestamp order satisfies a sequential relationship, the system marks this node pair as a temporal causal candidate. The system employs a statistical test method based on transition entropy, using the historical occurrence sequence of the candidate predecessor state as a condition, to test whether its predictive ability for the occurrence sequence of the successor state is statistically significant. Node pairs that pass the test are added with directed edges, pointing from the predecessor state to the successor state.

[0032] Furthermore, the system constructs hyperedges to represent many-to-one convergence relationships and one-to-many bifurcation relationships. When event instances of multiple event state nodes act on the same subsequent event state node, the system constructs a convergence hyperedge, which connects multiple predecessor nodes and one successor node. When an event instance of one event state node acts on multiple subsequent event state nodes, the system constructs a bifurcation hyperedge, which connects one predecessor node and multiple successor nodes. The hybrid event graph is composed of a set of event state nodes, a set of directed edges representing temporal causal relationships, and sets of convergence and bifurcation hyperedges. The system calculates the average dwell time of each event state node, which is the average time interval between the occurrence time of each event instance in that state and the occurrence time of its causal successor event instance. The system arranges the average dwell times of each state node in order of state identifier, forming an average dwell time sequence. The system also calculates the relative time offset between any two event state nodes. For node pairs with a temporal causal relationship, the system collects all event instance pairs that satisfy the chronological order, calculates the average time difference as the offset, and stores it in the relative time offset matrix. The corresponding positions of node pairs without a causal relationship in the matrix are empty values.

[0033] S3: Accumulate and detect change points in the average residence time series to obtain event nodes with temporal anomalies as candidate nodes. Extend the constraints on the subgraphs containing the candidate nodes based on the relative time offset matrix. Extract candidate subgraphs containing at least two start points or at least two end points from the mixed event graph. Calculate the multi-start point-multi-end point saliency based on the number of start point nodes, the number of end point nodes, and the internal node branch entropy in each candidate subgraph. Event subgraphs with saliency exceeding the preset graph saliency threshold are judged as irregular event chains and stored in the irregular event chain set.

[0034] Specifically, the system performs cumulative sum and change point detection using the average residence time series as input. The system calculates the mean and standard deviation of the entire sequence, converting the average residence time of each state node into a standardized deviation value. The system initializes the upper and lower cumulative sum statistics to zero, and sequentially accumulates the difference between the standardized deviation value and a preset allowable offset, tracking positive and negative deviation accumulation respectively. When any cumulative sum statistic exceeds a preset control limit, the system determines that a structural change point has occurred at that location and marks the event state node corresponding to that change point as a candidate node for temporal anomalies. After obtaining the candidate node set, the system performs time-constrained subgraph expansion for each candidate node. Starting from the candidate node, the system performs bidirectional breadth-first expansion along the directed edges and hyperedges of the hybrid event graph. Forward expansion includes successor nodes whose offsets in the time offset matrix are within a preset time window, and backward expansion includes predecessor nodes whose offsets are within the same time window. The expansion depth is limited by a preset maximum depth. The system extracts the expanded subgraph corresponding to each candidate node from the hybrid event graph.

[0035] The system identifies the starting and ending nodes for each candidate subgraph. A starting node is defined as a state node with zero in-degree, and an ending node is defined as a state node with zero out-degree. The system selects candidate subgraphs with at least two starting nodes or at least two ending nodes, excluding the rest. The system calculates the saliency of each retained subgraph across multiple starting and ending points. The saliency comprises three components: the first is calculated from the number of starting nodes, the second from the number of ending nodes, and the third is the branch entropy calculated based on the out-degree distribution of nodes within the subgraph. The branch entropy reflects the complexity of the bifurcation structure within the subgraph. The system weights and sums these three components according to preset weights to obtain the saliency. The system compares the saliency with a preset graph saliency threshold. Subgraphs with saliencies reaching or exceeding the threshold are classified as irregular event chains. The system stores these irregular event chains in an irregular event chain set, recording the node set, edge set, super-edge set, and corresponding saliency value for each event chain.

[0036] S4: The saliency of the event chains in the irregular event chain set is used as an enhancement weight to be added between personnel nodes and event nodes in the standardized entity library, and the association strength between node pairs is calculated based on the enhanced node relationship to generate an enhanced multidimensional association graph and association strength matrix.

[0037] Specifically, the system extracts the saliency of each event chain from the set of irregular event chains and uses this saliency as an enhancement weight to map between personnel nodes and event state nodes in the standardized entity library. For each irregular event chain, the system traverses all event state nodes contained in the event chain. For each event state node, the system obtains all event instances mapped to that node, extracts personnel entity identifiers from the source entity identifier set of each event instance, and establishes an enhancement relationship between the personnel entity and the event state node. The enhancement weight is taken as the saliency of the event chain. If the same pair of personnel entities and event state nodes appears in multiple irregular event chains or multiple event instances of the same event chain, the system accumulates the corresponding saliency as the final enhancement weight for that node pair.

[0038] After adding enhanced weights, the system constructs an enhanced multidimensional association graph. The node set of this graph includes all personnel entities, resource entities, time-cycle entities in the standardized entity library, and all event state nodes in the hybrid event graph. The system establishes association edges between entities and events in this graph. The base weight of the association edge is determined by the participation frequency of the entity in the event instance, and the enhanced weight is added as an additional item to the base weight. The system also establishes association edges between entities, with the base weight of the association edge determined by the frequency of the two entities appearing together in the same event or the same time window. After obtaining the graph structure containing the direct association weights of all node pairs, the system calculates the association strength between any pair of nodes. The association strength consists of two parts: direct association strength and indirect association strength. The direct association strength is directly taken from the weight of the corresponding edge in the strong association graph. The indirect association strength is calculated by traversing all paths through intermediate nodes in the graph. The contribution of each path is the product of the association strengths of each segment on the path multiplied by the attenuation factor of the intermediate node. The attenuation factor is calculated based on the degree of the intermediate node and a preset attenuation coefficient. The system organizes the association strength of all node pairs into an association strength matrix. The rows and columns of the matrix correspond to entity nodes and event state nodes, and the matrix elements represent the association strength between node pairs.

[0039] S5: Based on the enhanced multidimensional association graph and association strength matrix, calculate the weighted eigenvector centrality of each node under the association strength matrix, and calculate the running state score of each node by combining the degree of drift of node attributes relative to the historical baseline. Based on the distribution of state scores of nodes within each irregular event chain and the similarity between the irregular event chain and the frequent patterns in the frequent pattern library, output the state determination result of the irregular event chain; whereby, the frequent pattern library is constructed from historical standardized event tuples based on the association rule mining algorithm.

[0040] Specifically, the system takes the association strength matrix as input and calculates the weighted eigenvector centrality of each node. The system uses a power-law iteration method to solve for the principal eigenvalues ​​and corresponding principal eigenvectors of the association strength matrix. The eigenvectors are initialized as all-unit vectors. Iterative matrix-vector multiplication and vector normalization operations are performed until the second norm of the vector difference between two adjacent iterations is less than the preset convergence precision, at which point the iteration terminates. The components of the obtained normalized eigenvectors represent the weighted eigenvector centrality of the corresponding node. The system also calculates the attribute drift degree of each node. Attribute vectors for each node are extracted from the standardized entity library and event tuples. The statistical mean of the node's attributes within the historical window period is used as the historical baseline. The Mahalanobis distance between the current attribute vector and the historical baseline vector is calculated. The inverse of the covariance matrix of the attribute vectors within the historical window period is introduced during the Mahalanobis distance calculation to eliminate dimensional differences and covariance effects between attribute dimensions.

[0041] Furthermore, the system combines the weighted feature vector centrality and attribute drift degree of each node, and obtains the node's operational status score by multiplying the centrality by an exponential function of the drift degree. The system pre-constructs a frequent pattern library based on historical standardized event tuples. During construction, a frequent pattern mining algorithm scans event state sequences within historical time periods, identifying event state subsequences with support not lower than a preset minimum support threshold as frequent patterns, and calculating the confidence between frequent patterns. Strong association rules that satisfy the minimum confidence threshold are stored together with their corresponding frequent patterns in the frequent pattern library. For each event chain in the set of irregular event chains, the system extracts the operational status scores corresponding to the event state nodes contained in the event chain, calculates the mean, maximum, and variance of the scores, and the proportion of nodes with scores exceeding a preset anomaly threshold, to obtain the internal anomaly severity.

[0042] Preferably, the system compares the event state sequence of the event chain with each frequent pattern sequence in the frequent pattern library using the longest common subsequence. The similarity is calculated as twice the length of the longest common subsequence divided by the sum of the lengths of the two sequences. The maximum similarity among all similarities is taken as the pattern similarity between the event chain and the frequent pattern library. The pattern deviation is calculated based on the difference between this similarity and the benchmark value. The system weights and sums the internal anomaly severity and pattern deviation according to preset weights to obtain a comprehensive status judgment score. The system classifies the event chain status into different risk levels based on the comparison between the comprehensive score and preset grading thresholds. The system outputs a structured diagnostic report, which includes a graph structure description of each irregular event chain, a comprehensive status judgment score, a risk level, the identifier of the node with the highest anomaly score and its associated entity information, and the frequent pattern sequence with the highest similarity to the event chain in the frequent pattern library, for further verification by analysts.

[0043] In summary, the data operation status analysis method for multidimensional data association mining provided in this application, by constructing a hybrid event graph and introducing a hyperedge structure, can completely preserve the many-to-one convergence relationship and one-to-many bifurcation relationship between events, providing an accurate topological foundation for subsequent identification of irregular network event chains. By detecting the change point of the average residence time of event state nodes and supplementing it with the extended constraint of the relative time offset matrix, it can accurately locate event nodes with temporal anomalies, and then extract candidate subgraphs containing multiple start points or multiple end points. On this basis, by calculating the multi-start point-multi-end point significance by fusing the number of start point nodes, the number of end point nodes, and the internal node branch entropy, and using this to determine irregular event chains, it can automatically discover multiple events that are difficult to capture by traditional linear analysis methods in organizational operations. Because of its convergent and multi-effect divergent event propagation paths, this approach compensates for the lack of existing technologies in identifying network event chains. Furthermore, by adding event chain saliency as an enhancing weight between personnel nodes and event nodes to generate an enhanced multidimensional association graph, the high-order associations implicit in irregular event chains can be explicitly integrated into the global association strength matrix, allowing downstream analysis to fully utilize the event chain mining results. Finally, by comprehensively considering the centrality of node weighted eigenvectors, the degree of attribute drift, and the similarity between the event chain and the frequent pattern library, a state determination can be made. This allows for the output of state determination results based on distinguishing between rare normal patterns and true abnormal patterns, thereby achieving automatic identification and accurate state assessment of irregular event chains in organizational operations, effectively reducing the risk of anomaly underreporting and misjudgment.

[0044] In one embodiment, S1 of the data operation status analysis method for multidimensional data association mining provided by the present invention specifically includes the following steps:

[0045] S11: Based on the data access interface, collect multi-dimensional raw data from multiple business systems, including order milestone nodes, personnel attendance records, equipment and material allocation records, and event trigger logs. Align the clock deviations of different systems through unified timestamp formatting to generate an aligned multi-dimensional event record set.

[0046] Specifically, the system establishes connections with multiple business systems through data access interfaces, including an order management system, a human resources management system, a materials management system, and a log collection system. The system collects order milestone node data from the order management system, which records the time and status information of orders reaching each key stage in the process; it collects personnel attendance records from the human resources management system, which includes personnel's check-in and check-out times, attendance status, and work group information; it collects equipment and material transfer records from the materials management system, which includes the sending department, receiving department, quantity, and time of the transfer; and it collects event trigger logs from the log collection system, which records the occurrence time and target of various business operations.

[0047] The timestamp recording methods of different business systems vary. Some systems use the Unix timestamp format, while others use date and time string formats. Clock drift also exists between different systems. The system performs a unified timestamp formatting operation on each raw event record, identifying the timestamp field format in the raw record and converting non-standard time representations to a unified standard time format according to the format type. For data sources with clock drift, the system calculates the offset of each data source relative to the base clock by selecting a reference event that co-occurs across systems, and applies the offset to the timestamp correction of all records in that data source. After completing timestamp formatting and clock drift correction, the system reassembles the fields of the same data source according to a preset event record structure, ensuring that each record includes an event identifier, event occurrence time, source entity identifier, target entity identifier, and event content description fields. The reassembled set of all records is output as an aligned multidimensional event record set.

[0048] S12: Based on the globally unique entity encoding rule, perform cross-system entity matching and disambiguation processing on the source entity identifier and target entity identifier in the multi-dimensional event record set, map different aliases of the same entity to a unified identifier, and generate a standardized event tuple set with the source entity identifier set and the target entity identifier set; wherein, the globally unique entity encoding rule is a segmented encoding rule with cross-system uniqueness obtained by orderly concatenating the entity type identifier, source system identifier, key attribute hash value and timestamp and then performing a consistent hash operation.

[0049] Specifically, the system loads a globally unique entity encoding rule, which specifies the ordered concatenation order of entity type identifier, source system identifier, key attribute hash value, and timestamp, as well as the parameters for consistent hashing. The system first extracts the original identifier string from the source entity identifier field and target entity identifier field of each record. By parsing the format and context information of the original identifier, it identifies the entity type and source system to which the identifier belongs. For each identified entity identifier, the system concatenates its entity type identifier, source system identifier, and the key attribute field extracted from the record. The current timestamp is appended to the concatenated string. The complete concatenated string is used as input to the consistent hash function. After hashing, a fixed-length hash value is output. This hash value is combined with a preset segment identifier to form a globally unique code for the entity.

[0050] During processing, the system needs to identify different aliases pointing to the same entity from different data sources. The identification method is as follows: the system extracts the original identifier and associated attributes of the same entity from different data sources, calculates the similarity of each attribute dimension, and determines that multiple identifiers with a comprehensive similarity exceeding a threshold point to the same entity. After determination, the system uniformly maps these identifiers to a generated globally unique code. For event records that have completed entity alignment, the system replaces the source entity identifier field in the original record with a standardized set of source entity identifiers, and replaces the target entity identifier field with a standardized set of target entity identifiers. The timestamp and event content description fields in the event record remain unchanged. The system organizes all event records that have undergone entity matching and disambiguation processing into a standardized event tuple set. Each standardized event tuple contains the event global identifier, standardized timestamp, source entity identifier set, target entity identifier set, and attribute vectors extracted from the original record.

[0051] S13: Extract entity identifiers from the standardized event tuple set by classifying them into personnel entities, resource entities, and time period entities, and associate them with the role tags and resource quota attributes defined in the business process to generate a standardized entity library.

[0052] Specifically, the system performs classification and extraction operations on all entity identifiers in the standardized event tuple set. The system iterates through the source and target entity identifier sets of each standardized event tuple to obtain the globally unique codes of all occurrences. Based on the segment identifier within the globally unique entity code, the system determines the entity type corresponding to that code. The segment identifier contains an entity type identifier field, which the system parses to categorize the entity into one of three types: personnel entities, resource entities, or time-cycle entities. For codes identified as personnel entities, the system retrieves the employee's name, employee ID, department, job title, and role tag information from various business systems. The role tag defines the employee's functional positioning in various business operations within the business process. The system summarizes this information into attribute records for the personnel entity.

[0053] For codes identified as resource entities, the system retrieves the resource's name, specifications, type, department, and quota attributes from the materials management system and equipment management system. The quota attributes record the resource's capacity limit, available quantity, and scheduling restrictions. The system aggregates this information into the resource entity's attribute record. For codes identified as time-cycle entities, the system retrieves the cycle's start time, end time, cycle type, and time mapping relationship with business nodes from the cycle definition table in the business system. Cycle types include financial settlement cycles, project milestone cycles, and scheduling cycles. The system aggregates this information into the time-cycle entity's attribute record. The system allocates storage space for each entity to record its globally unique code and corresponding attribute record, organizing the mapping relationship between all entity codes and attribute records into a key-value pair storage structure to form a standardized entity library. The system also establishes a reference index from the standardized event tuple set to the standardized entity library, enabling the source entity identifier set and target entity identifier set in each event tuple to be quickly associated with the corresponding entity attribute record in the entity library.

[0054] In one embodiment, S2 of the data operation status analysis method for multidimensional data association mining provided by the present invention specifically includes the following steps:

[0055] S21: Based on the source entity identifier set, target entity identifier set, and absolute timestamp of each event in the standardized event tuple set, construct an initial directed graph with event states as nodes and causal transmission and temporal order between events as directed edges. Perform loop detection and delooping on the initial directed graph to generate an acyclic directed graph.

[0056] Specifically, the system extracts event state nodes from the standardized event tuple set. Using the attribute vector of each event tuple as input, a clustering method is employed to divide the attribute vector space into multiple regions, each corresponding to an event state type. The system generates a state node for each event type and maps all event tuples whose attribute vectors fall into that region to that state node. After node generation, the system establishes directed edges between event state nodes, with the establishment rules based on a comprehensive judgment of the source entity identifier set, the target entity identifier set, and the absolute timestamp. The system traverses all event state node pairs. For each pair of predecessor and successor candidate nodes, the system searches the event tuples contained in the predecessor and successor candidate nodes. If a predecessor and successor event tuple satisfy the condition that the absolute timestamp of the predecessor event is earlier than that of the successor event, and the target entity identifier set of the predecessor event intersects with the source entity identifier set of the successor event, the system adds a directed edge between the predecessor and successor candidate nodes, with the edge pointing from the earlier event to the later event. After the system completes the above determination for all node pairs, it obtains the initial directed graph.

[0057] Furthermore, the system performs cycle detection on the initial directed graph, using a depth-first traversal strategy to mark the visit status of each node. When an edge pointing to an existing node in the current recursion stack is found during traversal, the system records all nodes and edges contained in that cycle. For each detected cycle, the system performs cycle removal processing. The cycle removal strategy calculates the average time difference between the two event state nodes connected by each directed edge in the cycle, and removes the edge with the smallest average time difference from the cycle. This strategy is based on the assumption that the edge with the smallest time difference has the weakest temporal association. The system repeats cycle detection and cycle removal operations until no cycles exist in the graph, generating an acyclic directed graph.

[0058] S22: Generate a converging superedge for multiple predecessor nodes pointing to the same successor node within the same time window in an acyclic directed graph, and generate a bifurcation superedge for multiple successor nodes generated by the same predecessor node within the same time window. All superedges together with the directed edges of the acyclic directed graph constitute a hybrid event graph.

[0059] Specifically, the system performs hyperedge generation operations based on an acyclic directed graph. The system sets a time window parameter, which defines the maximum time interval between two event state nodes that indicate they are performing the same convergence or bifurcation operation. The system scans all nodes in the acyclic directed graph. For each event state node, the system obtains the set of all predecessor nodes pointing to that node. The system groups the predecessor nodes according to the average time difference of the directed edges from each predecessor node to its successor node, grouping multiple predecessor nodes whose average time differences fall within the same time window. For each group containing at least two predecessor nodes, the system generates a convergence hyperedge. This convergence hyperedge uses all predecessor nodes in the group as its source node set and the successor node as its target node. The hyperedge signifies that the triggering of the successor event state requires the combined action of all predecessor event states in the group. For each event state node, the system also obtains the set of all successor nodes pointed to by that node.

[0060] Furthermore, the system groups successor nodes according to the average time difference corresponding to the directed edges from the current node to each successor node, grouping multiple successor nodes whose average time differences fall within the same time window. For each group containing at least two successor nodes, the system generates a bifurcation hyperedge. This hyperedge uses the predecessor node as its source node and all successor nodes in the group as its target node set. The hyperedge signifies that the triggering of the predecessor event state will lead to the concurrent propagation of successor event states in the group. After generating all convergence hyperedges and bifurcation hyperedges, the system merges all temporally causal directed edges and all hyperedges to form a hybrid event graph. In the hybrid event graph, any two nodes are allowed to have both directed edge and hyperedge relationships simultaneously.

[0061] S23: For each event state node in the mixed event graph, divide its dwell time observations into fixed time windows, calculate the average dwell time within each time window, and serialize and store it to generate an average dwell time series.

[0062] Specifically, the system calculates the average dwell time series for each event state node in the hybrid event graph. First, the system obtains all event instances mapped to each event state node and their corresponding absolute timestamps. For each event instance, the system needs to determine its dwell time on the event state node. The dwell time is calculated as follows: the system retrieves causal successor event instances reachable from the event instance along directed edges or hyperedges. Among all causal successor event instances, it selects the instance with the smallest time difference whose absolute timestamp is greater than the current event instance's timestamp. This smallest time difference is used as the dwell time observation value for the current event instance on that state node. For event instances without any causal successor event instances, the system marks the dwell time of that instance as a missing value. The system divides all dwell time observation values ​​for each event state node according to a fixed time window length. The division is based on the time window to which the absolute timestamp of each event instance belongs, grouping all dwell time observation values ​​falling within the same time window into one group.

[0063] Furthermore, the system calculates the arithmetic mean of the dwell time for each group. For groups containing no valid observations, the system marks the average dwell time of that time window as null. The system arranges the average dwell times of each time window into a sequence according to their start times, generating an average dwell time sequence for that event state node. The system repeats all the above operations for each event state node in the mixed event graph, obtaining an independent average dwell time sequence for each node. The length of each sequence depends on the number of time windows covered by the event instances contained in that node over the entire observation time span.

[0064] S24: Based on the absolute timestamps of each event state node in the hybrid event graph, calculate the average time difference between any two nodes, approximate the cumulative time offset of node pairs without direct directed edges through the shortest path, and generate a relative time offset matrix.

[0065] Specifically, the system calculates a relative time offset matrix based on a hybrid event graph. The rows and columns of the matrix correspond to all event state nodes in the hybrid event graph, and the matrix elements represent the average time offset between the corresponding node in a row and the corresponding node in a column. For cases where there are direct directed edges from one node to another in the graph, the system collects all event instance pairs associated with that directed edge. Each pair contains a predecessor event instance and a successor event instance, and the timestamp of the predecessor event instance is earlier than the timestamp of the successor event instance. The system calculates the time difference between each pair of event instances, takes the arithmetic mean of all time differences as the average time offset of the directed edge, and fills this value into the corresponding position in the matrix.

[0066] For node pairs connected by hyperedges in the graph, the system processes the relationship between the predecessor and successor node sets according to the type of hyperedge. For convergence hyperedges, the system calculates the average time offset from each predecessor node to the target node and fills it into the corresponding position in the matrix. For bifurcation hyperedges, the system calculates the average time offset from the source node to each successor node and fills it into the corresponding position in the matrix. For node pairs in a mixed event graph that are not directly connected by any directed edge or hyperedge, the system uses the shortest path accumulation method to estimate the time offset. The system traverses all directed paths from the corresponding row node to the corresponding column node in the graph. The time offset of each path is obtained by accumulating the time offsets between consecutive node pairs on the path. The system selects the path with the smallest accumulated time offset among all reachable paths and uses this smallest accumulated value as the estimated relative time offset for that node pair. For node pairs that have no reachable paths, the system marks the corresponding position in the matrix as infinity. The system organizes the time offsets of all node pairs into a two-dimensional matrix in row and column order to generate a relative time offset matrix.

[0067] In one embodiment, S3 of the data operation status analysis method for multidimensional data association mining provided by the present invention specifically includes the following steps:

[0068] S31: Based on the average dwell time series, calculate the cumulative sum and difference between the mean dwell time in each time window and the historical baseline according to the preset sliding window. When the cumulative sum and difference exceed the preset change point detection threshold, mark the corresponding event state node as a time series abnormal node and generate a candidate node set.

[0069] Specifically, the system acquires the average dwell time series for each event state node. This series is arranged in fixed time windows, with each element corresponding to the arithmetic mean of the dwell times of all event instances on that state node within a time window. The system deploys sliding windows on the series, with a window length equal to the preset number of time windows and a sliding step size of one time window unit. At each sliding window position, the system calculates the difference between the average dwell time of each time window within the window and the historical baseline. The historical baseline is determined by the weighted average of all dwell time observations of that node during its initial normal operation period. The system defines the calculation method for the cumulative difference as assigning a weight coefficient that varies with the position within the window to the dwell time deviation value of each time window. The closer the deviation value is to the current time, the larger its weight coefficient value. The system constructs a weighted cumulative difference statistic, which is obtained by summing the products of the deviation values ​​of each time window within the window and their corresponding weight coefficients.

[0070] When this statistic exceeds a preset change point detection threshold, the system marks the event state node corresponding to the end of the current sliding window as a time-series aberration node. The system performs the above change point detection operation on each event state node in the mixed event graph, collecting all event state nodes marked as time-series aberrations to form a candidate node set. The system uses the following formula to express the weighted cumulative difference during the calculation process:

[0071]

[0072] in, The number of time windows contained within the sliding window. For the first The average dwell time observation of the event state node within a time window. This is the historical baseline mean. The standard deviation is the historical baseline. This is the position number within the window. The linear weighting coefficient is assigned to the position deviation value, and this weighting coefficient increases as the position number increases. This is the weighted cumulative difference calculated at the k-th time window. In this formula... This is the dimensionless standardized deviation value, multiplied by a weighting coefficient that also has dimensionless properties. It is a dimensionless statistic, suitable for cross-node comparisons. The system will... If the value exceeds the preset change point detection threshold, it is determined that there is a temporal change point at that location.

[0073] S32: Starting from each candidate node in the candidate node set, expand in both the reverse and forward directions along the directed edges of the hybrid event graph. During the expansion process, prune branches that exceed the preset time span limit based on the relative time offset matrix, and select connected subgraphs that contain at least two source entity identifier sets and whose number of nodes is not less than the preset subgraph node number baseline to generate a candidate subgraph set.

[0074] Specifically, the system uses each candidate node in the candidate node set as the starting point for subgraph expansion. The expansion process proceeds along the directed edges of the hybrid event graph, performing both reverse and forward expansion. Reverse expansion traces the predecessor node in the reverse direction of the directed edges, while forward expansion traces the successor node in the direction of the directed edges. During the expansion process, the system applies time constraints to each branch based on the relative time offset matrix. When the cumulative relative time offset from the current node to the node to be expanded exceeds a preset time span limit, the system stops further expansion of that branch and performs pruning. Simultaneously, the system applies constraints to the expansion range based on the source entity identifier set. The system retrieves the source entity identifier set of the event tuples contained in each node during the expansion process, counts the total number of source entity identifiers already covered in the current subgraph, and reduces the priority of expansion in that direction when a node encountered during the expansion process cannot introduce a new source entity identifier. After completing the expansion, the system obtains an initial expanded subgraph set starting from each candidate node. The system performs connectivity verification on each initial expanded subgraph. The verification method is to check whether there is at least one undirected connected path between any two nodes in the subgraph, which consists of a directed edge and a hyperedge. If there is a node that cannot be connected to other nodes through any path, the system removes the node from the subgraph.

[0075] The system further filters the expanded subgraphs based on the following criteria: the size of the source entity identifier set in the subgraph is not less than a preset threshold, and the total number of nodes in the subgraph is not less than a preset baseline threshold for the number of nodes in the subgraph. The system includes all connected subgraphs that meet the filtering criteria into the candidate subgraph set. During the expansion process, the system uses the following formula to define the node expansion priority, which is used to determine the expansion order among multiple optional expansion directions:

[0076]

[0077] in, For the current expansion of the frontier node, For the adjacent nodes to be expanded, For nodes The set of source entity identifiers contained in the event tuples. This is the set of source entity identifiers already covered by the current subgraph. For nodes The number of source entity identifiers that can be added to the current subgraph. Nodes in the relative time offset matrix To the node The relative time offset, To prevent extremely small constants with a denominator of zero. The larger the value, the more it indicates that from Expand to The higher the priority, the more likely the system will select it in each round of expansion. The neighboring node with the largest value is included in the subgraph. The physical meaning of this formula is to prioritize expanding neighboring nodes that can contribute new source entity identifiers and have smaller time offsets.

[0078] S33: For each candidate subgraph in the candidate subgraph set, count the number of internal starting nodes, the number of ending nodes, and the branch entropy of internal nodes. Calculate the multi-starting-multi-ending saliency of the candidate subgraph by weighted summing of the proportion of starting nodes, the proportion of ending nodes, and the branch entropy. Candidate subgraphs with saliency exceeding the preset multi-starting-multi-ending saliency threshold are identified as irregular event chains and stored in the irregular event chain set. Extract the offset quantum matrix between internal nodes of the irregular event chain from the relative time offset matrix to generate a local time offset matrix.

[0079] Specifically, the system performs topological statistics on each candidate subgraph in the candidate subgraph set. The system traverses all event state nodes in the candidate subgraph, counting the total number of nodes with zero in-degree as the number of starting nodes, the total number of nodes with zero out-degree as the number of ending nodes, and the total number of nodes with both non-zero in-degree and non-zero out-degree as the number of internal nodes. The system also counts the total number of all directed edges in the candidate subgraph and the out-degree value of each internal node. Based on these statistics, the system calculates the multi-starting-point and multi-ending-point significance of the candidate subgraph. The system uses a weighted summation method to combine the starting point contribution, ending point contribution, and branch entropy contribution into a significance value. The starting point contribution is determined by the ratio of the number of starting nodes to the sum of the number of starting nodes and internal nodes; the ending point contribution is determined by the ratio of the number of ending nodes to the sum of the number of ending nodes and internal nodes. The branch entropy contribution is calculated based on the out-degree distribution of internal nodes. The out-degree of each internal node is divided by the total number of directed edges to obtain the out-degree probability of that node. The negative value of the product of the out-degree probabilities of all internal nodes and the base-2 logarithm is then summed to obtain the branch entropy. The system multiplies the contribution of the starting point, the contribution of the ending point, and the contribution of the branch entropy by preset weight coefficients and then sums them to obtain the saliency of the candidate subgraph.

[0080] The system classifies candidate subgraphs with saliency exceeding a preset threshold as irregular event chains and stores them in an irregular event chain set. The system extracts all relative time offsets between nodes within each irregular event chain from the relative time offset matrix. These offsets are then organized into a two-dimensional submatrix according to the topological order of the nodes in the event chain. The rows and columns of this submatrix correspond to the nodes within the event chain, and the matrix elements are the relative time offsets between corresponding node pairs. The system stores this submatrix as the local time offset matrix for that irregular event chain. The formula for calculating the saliency of a multi-start-point / multi-end-point event chain is:

[0081]

[0082] in, For multi-start-multi-endpoint significance, This represents the total number of starting nodes with an in-degree of zero in the candidate subgraph. This represents the total number of endpoint nodes with an out-degree of zero in the candidate subgraph. This represents the total number of internal nodes in the candidate subgraph that are neither the start nor the end point. Let be the total number of directed edges in the candidate subgraph. For internal nodes The out-degree is the number of successor nodes that an internal node points to. For the preset weighting coefficients, satisfy These are used to adjust the contribution of the starting point proportion, ending point proportion, and branch entropy to the significance calculation, respectively; the branch entropy term... This is used to quantify the bifurcation complexity of the internal structure of a candidate subgraph. The larger the entropy value, the richer the bifurcation structure of the candidate subgraph.

[0083] In one embodiment, S4 of the data operation status analysis method for multidimensional data association mining provided by the present invention specifically includes the following steps:

[0084] S41: Based on the association mapping relationship between personnel nodes in the standardized entity library and event state nodes in the hybrid event graph, establish initial association edges between personnel nodes and event state nodes, and assign a basic association strength to each initial association edge determined by the event type code and personnel role weight, thereby generating an initial multidimensional association graph.

[0085] Specifically, the system extracts all personnel entity nodes from the standardized entity library and all event state nodes from the hybrid event graph, using these personnel nodes and event state nodes as the initial node set of the multidimensional association graph. The system establishes initial association edges between personnel nodes and event state nodes. The rules for establishing these edges are based on the attribution relationship between personnel entities and event states in the standardized event tuple set. When there is at least one event tuple in the standardized event tuple set, and the source entity identifier set or target entity identifier set of that event tuple contains the personnel entity and the event tuple is mapped to the event state node, the system establishes an initial association edge between that personnel node and that event state node. The system calculates the basic association strength for each initial association edge. The basic association strength is jointly determined by the event type code and the personnel role weight. The event type code is the value of the event type corresponding to the event state node in the preset event type code table, and the personnel role weight is the value of the personnel entity's role label in the business process in the preset role weight table.

[0086] Furthermore, the system multiplies the event type code by the personnel role weight, and then multiplies this by the normalized ratio of the co-occurrence frequency between the personnel entity and the event state node relative to the co-occurrence frequency distribution of all personnel-event pairs, to obtain the basic association strength of the initial association edge. After the system completes the establishment of all initial association edges and the assignment of basic association strengths, it generates an initial multidimensional association graph with personnel nodes and event state nodes as the node set, and the initial association edges and their corresponding basic association strengths as the edge set. The basic association strength is calculated using the following formula:

[0087]

[0088] in, For personnel nodes With event state nodes The fundamental correlation strength between them is dimensionless. For personnel nodes The role weight is determined by the role label of the person in the business process, and the value is the corresponding value of the role in the preset weight table, with a dimension of one. Event state node The event type encoding weight is obtained by normalizing the encoding value of the event type in the preset encoding table, and its dimension is one. For personnel nodes With event state nodes The co-occurrence frequency in a set of standardized event tuples, with a dimension of one. The numerator is the maximum frequency of all personnel-event relationships with the CCP, with a dimension of one. The numerator in this formula... Applying logarithmic compression to the co-occurrence frequency makes the effect of frequency difference on intensity exhibit diminishing marginal returns, and the denominator term... As a normalization factor, the intensity value is constrained to the range of 0 to 1, ensuring The range of values ​​and and The product is consistent.

[0089] S42: For each irregular event chain in the set of irregular event chains, use the multi-start-multi-end-point salience as the enhancement weight, and attach the enhancement weight to the association edges between the event state nodes and the associated personnel nodes inside the irregular event chain. For the association edges between event state nodes and personnel nodes that are not in any irregular event chain, the basic association strength remains unchanged, and the enhanced node association relationship is generated.

[0090] Specifically, the system acquires all irregular event chains and their corresponding multi-start and multi-endpoint saliency values ​​from the set of irregular event chains. The system traverses each event chain in the set, and for all event state nodes contained in the current event chain, identifies all personnel nodes associated with these event state nodes in the initial multidimensional association graph. The system uses the multi-start and multi-endpoint saliency of the current event chain as an enhancement weight, and adds this enhancement weight to the association edges between each event state node and each associated personnel node within the current event chain. The addition method is to add an enhancement contribution on top of the basic association strength. For association edges between event state nodes and personnel nodes that are not within any irregular event chain, the system maintains their basic association strength unchanged. When adding enhancement weights, if the same personnel node and the same event state node are simultaneously in multiple irregular event chains, the system accumulates the multiple enhancement weights obtained by the personnel node and the event state node. After adding enhancement weights, the system updates the strength values ​​of the corresponding association edges in the initial multidimensional association graph and records the updated association relationship as the enhanced node association relationship. During the addition of enhancement weights, the system uses the following formula to calculate the enhancement contribution:

[0091]

[0092] in, For personnel nodes With event state nodes The enhancement contribution of the associated edges from the irregular event chain is on the order of one. To include personnel nodes and event state nodes A collection of irregular event chains. For the first The significance of multiple starting and ending points of an irregular event chain is denoted by a value between 0 and 1, with a dimension of one. This is the preset attenuation coefficient, with dimensions of one. In irregular event chains In the middle, personnel nodes To the event state node The shortest path length, measured in units of the number of directed edges, has a dimension of one. This formula uses an exponential decay function. The augmentation contribution of more distant nodes decreases with increasing path length. When the path length is zero, the augmentation contribution equals the significance of the event chain; as the path length approaches infinity, the augmentation contribution approaches zero. After the system completes the augmentation contribution calculation, it will... and The sum is used as the enhanced correlation strength.

[0093] S43: Based on the enhanced node association relationship, the association strength between any two node pairs is fused and calculated. The basic association strength between node pairs, the original edge weights of the event graph, and the enhanced weights contributed by the irregular event chains are combined to construct a two-dimensional array of association strengths containing all node pairs, generate an association strength matrix, and construct an enhanced multidimensional association graph using the association strength matrix as the adjacency matrix and the standardized entity library and event state nodes as the node set.

[0094] Specifically, the system acquires all enhanced node relationships, encompassing all associated edges between personnel nodes and event state nodes, along with their enhanced strength values. The system performs a fusion calculation on the association strength between any two nodes, including all personnel entities, resource entities, and time-cycle entities in the standardized entity library, as well as all event state nodes in the hybrid event graph. For any node pair, the system obtains the association strength contribution value from three sources: the first source is the sum of the basic association strength and the enhanced contribution directly corresponding to the node pair in the enhanced node relationships; the second source is the original edge weight assigned to the directed edge or superedge during initial construction if such an edge or superedge exists between the node pair in the hybrid event graph; and the third source is the enhanced weight contributed by the saliency of the event chain after path decay when the node pair is both within a non-regular event chain.

[0095] Furthermore, the system performs a weighted summation of the three contribution values, with the weighting coefficients adaptively determined by the system based on the rank of each contribution value relative to the distribution of contribution values ​​across all nodes. The system organizes the fused association strengths of all node pairs into a two-dimensional array according to the order of their node identifiers, generating an association strength matrix. A fixed mapping relationship exists between the row and column indices of the matrix and the node identifiers. Using the association strength matrix as the adjacency matrix and all entity nodes in the standardized entity library and all event state nodes in the hybrid event graph as the node set, the system constructs an enhanced multidimensional association graph. In this graph, an undirected edge exists between any two nodes, and the weight of the edge is the value of the element at the corresponding position in the association strength matrix. The formula for fusion calculation is as follows:

[0096]

[0097] in, For nodes With nodes The overall correlation strength between them is dimensionless. This represents the basic association strength of the node pair, and its value is only taken when the node pair is a combination of a person node and an event state node. Other types of node pairs take a value of zero and have a dimension of one. The enhancement contribution of the node pair is set to a value only when the node pair is a combination of a person node and an event state node. Other types of node pairs take a value of zero and have a dimension of one. Nodes in a mixed event graph With nodes The original edge weight is set if there is a directed edge or superedge between them; otherwise, it is zero and has a dimension of one. To contain nodes and nodes A collection of irregular event chains. For the first The salience of an irregular event chain. For nodes With nodes In irregular event chains The shortest path length within the interior. , , These are adaptive weight coefficients for the three contribution sources, with a sum of 1. The system determines the value of each coefficient based on the rank proportion of each contribution value among all node pairs; the higher the rank proportion, the greater the weight of that source. This formula, through an adaptive weighting mechanism, ensures that the dominant association types in the overall network receive higher fusion weights, thereby preserving the most representative association patterns in the data.

[0098] In one embodiment, S5 of the data operation status analysis method for multidimensional data association mining provided by the present invention specifically includes the following steps:

[0099] S51: Based on the adjacency structure of the enhanced multidimensional association graph, each node traverses its two-hop neighborhood of neighboring nodes, calculates the association weight between neighboring nodes according to the association strength matrix, calculates the weighted eigenvector centrality of each node by iteratively solving the eigenvalue equation, and calculates the drift degree according to the Jensen-Shannon divergence between the statistical distribution of each node's attributes and the historical baseline. The centrality and drift degree are weighted and fused to generate the running status score of each node.

[0100] Specifically, the system acquires the adjacency structure of the enhanced multidimensional association graph, which is defined by the association strength matrix. The system performs a two-hop neighborhood traversal operation for each target node. A two-hop neighborhood is defined as the set of all nodes reachable by moving two steps along the association edges from the target node. The system records the association paths between the target node and each two-hop neighborhood node. Each path contains an intermediate node, and the path weight is determined by the product of the association strength from the target node to the intermediate node and the association strength from the intermediate node to the two-hop neighborhood node. If multiple paths exist between the target node and the same two-hop neighborhood node, the system takes the sum of all path weights as the equivalent association weight between the target node and the two-hop neighborhood node. The system then re-fills the matrix with the equivalent association weight as a correction value at the corresponding position in the association strength matrix. After completing the two-hop neighborhood traversal of all nodes, the system obtains the modified association strength matrix. Using this matrix as the adjacency matrix, a system of linear equations is established. The solution vector of this system of equations is the weighted eigenvector centrality of each node. The solution method is to use the power iteration method to perform eigenvalue decomposition on the modified association strength matrix and extract the eigenvectors corresponding to the principal eigenvalues ​​as the centrality vectors.

[0101] The system simultaneously calculates the attribute drift degree for each node, based on the Jensen-Shannon divergence measure between the node's current attribute distribution and its historical baseline distribution. The system defines an attribute probability distribution for each node, formed by normalizing the node's attribute values ​​within the current time window. The historical baseline distribution is formed by normalizing the node's attribute values ​​during its historical normal operation periods. The system calculates the Jensen-Shannon divergence between these two probability distributions, using this divergence value as the node's attribute drift degree. The system then weights and fuses the weighted eigenvector centrality and attribute drift degree for each node, multiplying the centrality by an exponential function of the drift degree to obtain the operational status score for each node. The iterative calculation formula for weighted eigenvector centrality is as follows:

[0102]

[0103] in, For the first Nodes in the next iteration The weighted eigenvector centrality estimate is dimensionless and initially set to 1 for all nodes. For nodes The set of two-hop neighborhood nodes, that is, from Starting from the associated edge, we pass through all the nodes that can be reached by exactly two edges. To correct the nodes in the correlation strength matrix To the node The equivalent association weight is calculated as follows: ,in and These are elements in the original correlation strength matrix. Dimensions and Consistent. Let be the set of all nodes. The denominator of this formula is a normalization factor, which is the square root of the sum of squares of the weighted sum of all nodes, ensuring that the centrality vector has a unit length after each iteration. The formula for calculating the degree of attribute drift is as follows:

[0104]

[0105] in, For nodes The degree of attribute drift, with a non-negative value and a unit of knight. For nodes The attribute probability distribution within the current time window, which is composed of nodes. exist The normalized values ​​of each attribute dimension constitute the sum of the values ​​of each dimension being 1. For nodes The historical baseline attribute probability distribution is composed of the attribute normalization values ​​of the same node during its historical normal operation period. It is the arithmetic mean of two distributions. The Kullback-Leibler divergence is defined as follows: The Jensen-Shannon divergence ensures the symmetry of the divergence values ​​and avoids infinitely large values ​​by introducing a mean distribution as a reference. The fusion calculation formula for node running status scores is as follows:

[0106]

[0107] in, For nodes The operational status score is dimensionless. The node after convergence of the power iteration method The final value of the weighted eigenvector centrality is greater than zero and dimensionless. This is an exponential function of the drift degree, mapping the drift degree from the non-negative real number domain to the real number domain greater than or equal to 1. This ensures that when the drift degree is zero, the score equals the centrality, and the score increases exponentially as the drift degree increases. This fusion method allows nodes with high centrality and large drift degrees to obtain higher state scores.

[0108] S52: Based on historical standardized event tuples, a frequent pattern mining algorithm that supports superedges is used to iteratively scan the co-occurrence relationship of events within each time window, extract subgraph patterns with support higher than the preset frequent threshold, organize them into a database containing multiple event chain topology types, and generate a frequent pattern library.

[0109] Among them, the historical standardized event tuple is a standardized event tuple extracted from the set of standardized event tuples within a historical period based on a preset historical time window.

[0110] Specifically, the system extracts standardized event tuples within historical time periods from the standardized event tuple set as the historical dataset, with the historical time period defined by a preset historical time window. The system groups the event tuples in the historical dataset according to time windows, with event tuples within the same time window forming a transaction set. The system uses a frequent pattern mining algorithm that supports hyperedges to iteratively scan the transaction sets. In the first round of scanning, the algorithm counts the frequency of each single event state node in each transaction set and calculates the support of each single event state node. Support is defined as the ratio of the number of transaction sets containing that event state node to the total number of transaction sets. Event state nodes with support not lower than a preset frequency threshold are recorded as frequent itemsets. In the k-th round of scanning, where k is an integer greater than 1, the algorithm generates candidate k-itemsets based on the frequent k-1 itemsets generated in the previous round. The candidate k-itemsets contain k event state nodes and the possible hyperedge connections between them, including convergence hyperedges and bifurcation hyperedges. The system counts the frequency of each candidate k-itemset within the transaction set. This frequency count requires both the co-occurrence condition of event state nodes and the matching condition of hyperedge structures. The hyperedge structure matching condition requires that the transaction set simultaneously contain all event instances involved by the hyperedges defined in the candidate k-itemset. The system calculates the support of each candidate k-itemset and designates those with support not lower than a preset frequency threshold as frequent k-itemsets. The algorithm iterates until no new frequent k-itemsets can be generated.

[0111] Furthermore, the system stores all frequent itemsets and their corresponding hyperedge structures as graph patterns. Each graph pattern contains a set of nodes, edges, and hyperedges, as well as a record of the pattern's support in historical data. The system organizes all graph patterns into a frequent pattern library, which supports indexing and querying by graph topology type and node labels. The support calculation formula in frequent pattern mining is as follows:

[0112]

[0113] in, For graph mode Support in historical data, with a value between 0 and 1, and a dimension of one. This refers to the set of all transactions divided by time window. Representation of graph pattern All event state nodes appear in the transaction set middle. Let be the hyperedge structure matching determination function, when Each hyperedge defined in the transaction set The function returns true if a co-occurrence relationship can be found in all event instances; otherwise, it returns false. The number of transaction sets that satisfy the node inclusion condition and the hyperedge structure matching condition. This represents the total number of transaction sets. By introducing hyperedge matching conditions, this formula allows the preservation of support information for higher-order combinational relationships during frequent pattern mining, avoiding the drawback of traditional frequent itemset mining where hyperedge structures are lost.

[0114] S53: For each irregular event chain in the set of irregular event chains, extract the running status score of each node inside each irregular event chain, the total number of nodes in the chain, and the distribution of the number of start and end points. Calculate the maximum value of the node score and multiply it by an adjustment coefficient composed of the proportion of the number of edges in the chain to the total number of edges, the number of start and end points, and the edit distance similarity factor between the irregular event chain and the most similar pattern in the frequent pattern library to generate a chain anomaly index. If the chain anomaly index exceeds the preset anomaly index threshold, the irregular event chain is determined to be an abnormal event chain, and the status judgment result containing the irregular event chain identifier, anomaly type, and internal node list is output.

[0115] Specifically, the system extracts irregular event chains one by one from the set of irregular event chains as objects to be judged. The system obtains the running status scores of all event state nodes within the current irregular event chain and extracts the maximum value. Simultaneously, the system calculates the total number of directed edges within the irregular event chain, the total number of event tuples involved after mapping the nodes contained in the event chain to standardized event tuples, and the number of starting and ending nodes of the event chain. The system calculates the graph edit distance between the irregular event chain and each frequent pattern in the frequent pattern library. The graph edit distance calculation is based on the minimum cost of node replacement, edge insertion, and edge deletion operations, where hyperedges are included in the overall edit distance calculation. The replacement cost of a hyperedge is determined by the size of the symmetric difference set between the node sets connected by the hyperedge. The system takes the minimum graph edit distance between the event chain and all frequent patterns in the frequent pattern library, divides this minimum edit distance by half the sum of the total number of edges in the event chain and the total number of edges in the frequent patterns, and obtains the normalized similarity value.

[0116] The system constructs a chain anomaly index, which is obtained by multiplying the maximum node score by an adjustment coefficient. The adjustment coefficient is composed of a factor representing the proportion of edges within the chain to the total number of edges, a factor representing the number of start and end points, and a normalized similarity factor. The start and end point factor is calculated by adding one to the sum of the number of start and end nodes and then taking the reciprocal of that result. The normalized similarity factor is obtained by subtracting the normalized similarity value between the event chain and the most similar pattern in the frequent pattern library from one. The system compares the chain anomaly index with a preset anomaly index threshold. If the chain anomaly index exceeds the threshold, the irregular event chain is classified as an anomalous event chain. The system records the identifier, anomaly type label, and a list of identifiers for all event state nodes within each classified irregular event chain, organizing these into a structured output of the state determination results. The formula for calculating the chain anomaly index is as follows:

[0117]

[0118] in, For the first The chain anomaly index of a chain of irregular events, dimensionless. It represents the maximum value of the running status scores of all event state nodes within the event chain, is dimensionless, and plays a major weighting role. Let be the total number of directed edges within this event chain. The total number of directed edges in the mixed event graph is denoted by , and the ratio of to represents the proportion of the event-related edge chain in the entire graph. , , This is a preset power exponent parameter with a dimension of one, used to adjust the degree of nonlinear influence of each factor on the adjustment coefficient. and These are the starting and ending node numbers of the event chain, respectively. The reciprocal of the sum of these two numbers plus one is used as the starting and ending point quantity factor. The value decreases as the number of starting and ending points increases, and the dimension is one. This is a library for frequent patterns. non-regular event chain With frequent patterns The graph editing distance between them is measured in the number of edges. The normalization benchmark for edit distance is half the sum of the number of edges in the two graphs. This represents the normalized structural similarity between the event chain and the frequent patterns, with a value between 0 and 1. The maximum similarity score between the event chain and all patterns in the frequent pattern library is taken. The three factors in the adjustment coefficient are weighted from three dimensions: edge size ratio, endpoint sparsity, and structural rarity, so that event chains with a high edge ratio, few endpoints, and low similarity to historical frequent patterns receive a higher anomaly index.

[0119] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0120] Based on the same inventive concept, this application also provides a data operation status analysis system for implementing the data operation status analysis method for multidimensional data association mining as described above. The solution provided by this system is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the data operation status analysis system for multidimensional data association mining provided below can be found in the limitations of the data operation status analysis method for multidimensional data association mining described above, and will not be repeated here.

[0121] Preferably, such as Figure 2 As shown, this invention provides a data operation status analysis system 600 for multidimensional data association mining, which is configured with the following modules:

[0122] The heterogeneous data standardization module 610 is used to perform time alignment and entity alignment on event data in the collected multidimensional heterogeneous data, generate a standardized event tuple set with a source entity identifier set, a target entity identifier set and an attribute vector for each event, and extract personnel entities, resource entities and time period entities to form a standardized entity library.

[0123] The hybrid event graph construction module 620 is used to construct a hybrid event graph from a set of standardized event tuples, with event states as nodes, temporal causal relationships as directed edges, and many-to-one convergence relationships and one-to-many bifurcation relationships as hyperedges, and to calculate the average dwell time series of each event state node and the relative time offset matrix between state nodes.

[0124] The irregular event chain mining module 630 is used to accumulate and detect change points in the average dwell time series, obtain event nodes with time-series anomalies as candidate nodes, and extend the constraints on the subgraph where the candidate nodes are located based on the relative time offset matrix. It extracts candidate subgraphs containing at least two start points or at least two end points from the mixed event graph, calculates the multi-start point-multi-end point saliency based on the number of start point nodes, the number of end point nodes and the internal node branch entropy of each candidate subgraph, and determines the event subgraph with saliency exceeding the preset graph saliency threshold as an irregular event chain and stores it in the irregular event chain set.

[0125] The multidimensional association graph enhancement module 640 is used to add the saliency of the event chain in the irregular event chain set as an enhancement weight to the personnel nodes and event nodes in the standardized entity library, and calculate the association strength between node pairs based on the enhanced node relationship to generate an enhanced multidimensional association graph and an association strength matrix.

[0126] The event chain state assessment module 650 is used to calculate the weighted eigenvector centrality of each node under the association strength matrix based on the enhanced multidimensional association graph and association strength matrix, and to calculate the running state score of each node by combining the degree of drift of node attributes relative to the historical baseline. Based on the state score distribution of nodes within each irregular event chain and the similarity between the irregular event chain and the frequent patterns in the frequent pattern library, the module outputs the state judgment result of the irregular event chain. The frequent pattern library is constructed from historical standardized event tuples based on the association rule mining algorithm.

[0127] In one embodiment, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described data operation status analysis method for multidimensional data association mining.

[0128] In one embodiment, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described data operation status analysis method for multidimensional data association mining.

[0129] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0130] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0131] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data operation status analysis method for multidimensional data association mining, characterized in that, Includes the following steps: S1: Perform time and entity alignment on the event data in the collected multidimensional heterogeneous data, generate a standardized event tuple set with a source entity identifier set, a target entity identifier set and an attribute vector for each event, and extract personnel entities, resource entities and time period entities to form a standardized entity library; S2: Construct a hybrid event graph for the standardized event tuple set, with event states as nodes, temporal causal relationships as directed edges, and many-to-one convergence relationships and one-to-many bifurcation relationships as hyperedges, and calculate the average dwell time sequence of each event state node and the relative time offset matrix between state nodes. S3: Accumulate and detect change points in the average residence time series to obtain event nodes with time-series anomalies as candidate nodes, and extend the constraints on the subgraph where the candidate nodes are located based on the relative time offset matrix. Extract candidate subgraphs containing at least two start points or at least two end points from the mixed event graph. Calculate the multi-start point-multi-end point saliency based on the number of start point nodes, the number of end point nodes, and the internal node branch entropy in each candidate subgraph. Determine the event subgraphs with saliency exceeding the preset graph saliency threshold as irregular event chains and store them in the irregular event chain set. S4: The saliency of the event chain in the irregular event chain set is used as an enhancement weight and added to the personnel node and event node in the standardized entity library. Based on the enhanced node relationship, the association strength between node pairs is calculated to generate an enhanced multidimensional association graph and an association strength matrix. S5: Based on the enhanced multidimensional association graph and the association strength matrix, calculate the weighted eigenvector centrality of each node under the association strength matrix, and calculate the running state score of each node by combining the degree of drift of the node attributes relative to the historical baseline. Based on the state score distribution of nodes within each irregular event chain and the similarity between the irregular event chain and the frequent patterns in the frequent pattern library, output the state determination result of the irregular event chain; wherein, the frequent pattern library is constructed from historical standardized event tuples based on the association rule mining algorithm.

2. The method according to claim 1, characterized in that, S1 includes: S11: Based on the data access interface, collect multi-dimensional raw data from multiple business systems, including order milestone nodes, personnel attendance records, equipment and material allocation records, and event trigger logs. Align the clock deviations of different systems through unified timestamp formatting to generate an aligned multi-dimensional event record set. S12: Based on the globally unique entity encoding rule, perform cross-system entity matching and disambiguation processing on the source entity identifier and target entity identifier in the multi-dimensional event record set, map different aliases of the same entity to a unified identifier, and generate a standardized event tuple set with a source entity identifier set and a target entity identifier set; wherein, the globally unique entity encoding rule is a segmented encoding rule with cross-system uniqueness obtained by orderly concatenating entity type identifier, source system identifier, key attribute hash value and timestamp and then performing a consistent hash operation; S13: Extract entity identifiers from the standardized event tuple set by classifying them into personnel entities, resource entities, and time period entities, and associate them with the role tags and resource quota attributes defined in the business process for each entity to generate a standardized entity library.

3. The method according to claim 1, characterized in that, S2 includes: S21: Based on the source entity identifier set, target entity identifier set and absolute timestamp of each event in the standardized event tuple set, construct an initial directed graph with event state as nodes and causal transmission and time sequence between events as directed edges, and perform loop detection and loop removal on the initial directed graph to generate an acyclic directed graph. S22: Generate a convergence superedge for multiple predecessor nodes pointing to the same successor node within the same time window in the acyclic directed graph, and generate a bifurcation superedge for multiple successor nodes generated by the same predecessor node within the same time window. All superedges together with the directed edges of the acyclic directed graph constitute a hybrid event graph. S23: For each event state node in the hybrid event graph, divide its dwell time observations into fixed time windows, calculate the average dwell time within each time window and serialize and store it to generate an average dwell time series. S24: Based on the absolute timestamps of each event state node in the hybrid event graph, calculate the average time difference between any two nodes, approximate the cumulative time offset of node pairs without direct directed edges through the shortest path, and generate a relative time offset matrix.

4. The method according to claim 1, characterized in that, S3 includes: S31: Based on the average dwell time series, calculate the cumulative sum and difference between the mean dwell time in each time window and the historical baseline according to the preset sliding window. When the cumulative sum and difference exceed the preset change point detection threshold, mark the corresponding event state node as a time series abnormal node and generate a candidate node set. S32: Starting from each candidate node in the candidate node set, expand in the reverse and forward directions along the directed edge direction of the hybrid event graph. During the expansion process, prune branches that exceed the preset time span limit according to the relative time offset matrix, and filter connected subgraphs that contain at least two source entity identifier sets and whose number of nodes is not less than the preset subgraph node number baseline to generate a candidate subgraph set. S33: For each candidate subgraph in the candidate subgraph set, count the number of internal starting nodes, the number of ending nodes, and the branch entropy of internal nodes. Calculate the multi-starting-multi-ending point saliency of the candidate subgraph by weighted summing of the proportion of starting nodes, the proportion of ending nodes, and the branch entropy. Determine candidate subgraphs with saliency exceeding a preset multi-starting-multi-ending point saliency threshold as irregular event chains and store them in the irregular event chain set. Extract the offset quantum matrix between internal nodes of the irregular event chain from the relative time offset matrix to generate a local time offset matrix.

5. The method according to claim 4, characterized in that, The formula for calculating the significance of the multi-start point-multi-endpoint method is as follows: ; in, For multi-start-multi-endpoint significance, This represents the total number of starting nodes with an in-degree of zero in the candidate subgraph. This represents the total number of endpoint nodes with an out-degree of zero in the candidate subgraph. This represents the total number of internal nodes in the candidate subgraph that are neither the start nor the end point. Let be the total number of directed edges in the candidate subgraph. For internal nodes The out-degree is the number of successor nodes that an internal node points to. For the preset weighting coefficients, satisfy These are used to adjust the contribution of the starting point proportion, ending point proportion, and branch entropy to the significance calculation, respectively; the branch entropy term... This is used to quantify the bifurcation complexity of the internal structure of a candidate subgraph. The larger the entropy value, the richer the bifurcation structure of the candidate subgraph.

6. The method according to claim 1, characterized in that, S4 includes: S41: Based on the association mapping relationship between personnel nodes in the standardized entity library and event state nodes in the hybrid event graph, establish initial association edges between personnel nodes and event state nodes, and assign a basic association strength to each initial association edge determined by the event type code and personnel role weight, thereby generating an initial multidimensional association graph. S42: For each irregular event chain in the set of irregular event chains, the multi-start point-multi-end point salience is used as the enhancement weight. The enhancement weight is added to the association edge between the event state node and the associated personnel node inside the irregular event chain. The basic association strength is maintained unchanged for the association edge between the event state node and the personnel node that is not in any irregular event chain, thereby generating the enhanced node association relationship. S43: Based on the enhanced node association relationship, the association strength between any two node pairs is fused and calculated. The basic association strength between the node pairs, the original edge weights of the event graph, and the enhanced weights contributed by the irregular event chains are combined to construct a two-dimensional array of association strengths containing all node pairs, generate an association strength matrix, and construct an enhanced multidimensional association graph using the association strength matrix as the adjacency matrix and the standardized entity library and the event state nodes as the node set.

7. The method according to any one of claims 1-6, characterized in that, S5 includes: S51: Based on the adjacency structure of the enhanced multidimensional association graph, for each node, traverse the adjacent nodes in its two-hop neighborhood, calculate the association weight between adjacent nodes according to the association strength matrix, calculate the weighted eigenvector centrality of each node by iteratively solving the eigenvalue equation, and calculate the degree of drift according to the Jensen-Shannon divergence between the statistical distribution of each node's attributes and the historical baseline. The centrality and the degree of drift are weighted and fused to generate the running status score of each node. S52: Based on historical standardized event tuples, a frequent pattern mining algorithm that supports superedges is used to iteratively scan the co-occurrence relationship of events within each time window, extract subgraph patterns with support higher than the preset frequent threshold, organize them into a database containing multiple event chain topology types, and generate a frequent pattern library. The historical standardized event tuple is a standardized event tuple extracted from the set of standardized event tuples within a historical period based on a preset historical time window; S53: For each irregular event chain in the set of irregular event chains, extract the running status score, the total number of nodes in the chain, and the distribution of the number of start and end points of each node in each irregular event chain. Calculate the maximum value of the node score and multiply it by an adjustment coefficient composed of the proportion of the number of edges in the chain to the total number of edges, the number of start and end points, and the edit distance similarity factor between the irregular event chain and the most similar pattern in the frequent pattern library to generate a chain anomaly index. If the chain anomaly index exceeds a preset anomaly index threshold, the irregular event chain is determined to be an abnormal event chain, and a status determination result containing the irregular event chain identifier, anomaly type, and internal node list is output.

8. A data operation status analysis system for multidimensional data association mining, characterized in that, The system includes: The heterogeneous data standardization module is used to perform time and entity alignment on event data in the collected multidimensional heterogeneous data, generate a standardized event tuple set with a source entity identifier set, a target entity identifier set and an attribute vector for each event, and extract personnel entities, resource entities and time period entities to form a standardized entity library. The hybrid event graph construction module is used to construct a hybrid event graph for the standardized event tuple set, with event states as nodes, temporal causal relationships as directed edges, and many-to-one convergence relationships and one-to-many bifurcation relationships as hyperedges, and to calculate the average dwell time sequence of each event state node and the relative time offset matrix between state nodes. The irregular event chain mining module is used to accumulate and detect change points in the average dwell time series, obtain event nodes with time-series anomalies as candidate nodes, and extend the constraints on the subgraph where the candidate nodes are located based on the relative time offset matrix. It extracts candidate subgraphs containing at least two start points or at least two end points from the mixed event graph, calculates the multi-start point-multi-end point saliency based on the number of start point nodes, the number of end point nodes and the internal node branch entropy in each candidate subgraph, and determines the event subgraphs with saliency exceeding the preset graph saliency threshold as irregular event chains and stores them in the irregular event chain set. The multidimensional association graph enhancement module is used to add the saliency of the event chain in the irregular event chain set as an enhancement weight to the personnel nodes and event nodes in the standardized entity library, and calculate the association strength between node pairs based on the enhanced node relationship to generate an enhanced multidimensional association graph and an association strength matrix. The event chain state assessment module is used to calculate the weighted eigenvector centrality of each node under the association strength matrix based on the enhanced multidimensional association graph and the association strength matrix, and to calculate the running state score of each node by combining the degree of drift of the node attributes relative to the historical baseline. Based on the state score distribution of nodes within each irregular event chain and the similarity between the irregular event chain and the frequent patterns in the frequent pattern library, the module outputs the state judgment result of the irregular event chain. The frequent pattern library is constructed from historical standardized event tuples based on the association rule mining algorithm.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.