Incremental monitoring method for distributed cluster state

By uniformly representing and incrementally verifying the multimodal state event streams of distributed clusters, the problems of discovering cross-modal causal relationships and adjusting monitoring strategies are solved, realizing the perception of causal loops and dynamic synchronization of resource allocation, thus improving the efficiency and accuracy of monitoring.

CN122496438APending Publication Date: 2026-07-31SUZI INFORMATION TECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZI INFORMATION TECHNOLOGY (HANGZHOU) CO LTD
Filing Date
2026-06-11
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to directly discover and label different types of cross-modal causal relationships within a unified framework in distributed cluster operations and maintenance. They lack topology-level measurements of causal graph changes, and monitoring strategy adjustments are limited in scope and open-loop, leading to a mismatch between resource configuration and the actual state of the cluster.

Method used

By uniformly representing the multimodal state event stream of the distributed cluster, generating time-series signals carrying modal source labels, performing incremental causality checks and topology measurements, generating monitoring and scheduling signals and channel start/stop commands, dynamically planning the neighborhood expansion range, and realizing on-demand start/stop and feedback correction of cross-modal acquisition channels.

Benefits of technology

It enables the discovery and labeling of different types of cross-modal causal relationships under a unified framework, perceives the impact of key changes such as causal loops on cluster status, dynamically adjusts monitoring strategies and resource configurations, and improves the efficiency and accuracy of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496438A_ABST
    Figure CN122496438A_ABST
Patent Text Reader

Abstract

This application discloses an incremental monitoring method for distributed cluster status, belonging to the fields of artificial intelligence infrastructure operation and maintenance, intelligent chips, and distributed control in the next-generation information technology industry. By uniformly representing multimodal state event flows and performing incremental causal verification and topological measurement, it achieves the discovery and labeling of different types of cross-modal causal relationships within a unified framework. This method can measure the stability of causal structures at the topological level and perceive the impact of key changes such as causal loops on the cluster status. Based on the trend of causal structure changes, the monitoring strategy is dynamically adjusted, including the collection frequency, reporting content, and on-demand start / stop of cross-modal collection channels, with feedback corrections. This ensures that the configuration of monitoring resources is synchronized with the actual changes in the cluster status, improving the efficiency and accuracy of monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the operation and maintenance of artificial intelligence infrastructure, intelligent chips, and distributed control in the next-generation information technology industry, and in particular to an incremental monitoring method for the status of distributed clusters. Background Technology

[0002] In distributed cluster operation and maintenance, observable data encompasses multimodal state data, including continuous numerical metrics, discrete text log events, and asynchronous call chain tracing events. Understanding the causal dependencies between cross-modal state data is crucial for quickly locating the root cause of a failure and predicting its propagation path.

[0003] However, existing causal discovery schemes, when dealing with cross-modal causal relationships, typically require converting discrete or asynchronous events into representations of the same dimension as numerical indicators before analysis. This conversion process is independent of the causal testing itself, making it difficult to directly discover and label different types of cross-modal causal relationships within a unified causal graph framework. Furthermore, existing schemes only assess changes in the causal graph by adding or deleting edges, lacking means to measure the stability of the causal structure at the topological level and failing to perceive the impact of the formation and demise of causal loops on the cluster state.

[0004] Furthermore, adjustments to existing monitoring strategies only involve a single dimension: the frequency of data collection. Moreover, these are open-loop controls that lack the ability to make predictive adjustments based on changes in the causal structure and also lack feedback verification of the control effects. This makes it difficult for the configuration of monitoring resources to keep pace with changes in the actual state of the cluster. Summary of the Invention

[0005] This application provides an incremental monitoring method for the status of a distributed cluster, the technical solution of which is as follows: On the one hand, an incremental monitoring method for the status of a distributed cluster is provided, the method comprising: A unified representation is performed on the multimodal state event stream of the distributed cluster to obtain a time-series signal carrying a modality source label. The multimodal state event stream includes continuous numerical index events, discrete text log events, and asynchronous call chain tracing events. Incremental causality testing and topological measurement are performed on the time-series signal carrying modal source labels and the causal hypergraph to obtain an updated causal hypergraph, causal type label and topological change measurement. The incremental causality testing adaptively selects the testing method based on the modal source labels and adjusts the accumulation rate of causal edge confidence according to the previous round of topological change measurement. Based on the topology change metric, the causal type marker, and the modal source label, monitoring and scheduling signals and channel start / stop instructions are generated. The topology change metric drives the adjustment of the acquisition frequency and incremental reporting content. The causal type marker, combined with the modal source label, drives the on-demand start / stop of cross-modal acquisition channels. Based on the monitoring and scheduling signals and the channel start / stop instructions, the access parameters of the multimodal state event stream are corrected by feedback, and the neighborhood expansion range of the next round of incremental causality verification is dynamically planned. Attached Figure Description

[0006] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0007] Figure 1 This is a flowchart of an incremental monitoring method for the status of a distributed cluster provided in an embodiment of this application; Figure 2 This is a flowchart of another incremental monitoring method for the state of a distributed cluster provided in an embodiment of this application; Figure 3 This is a flowchart of another incremental monitoring method for the state of a distributed cluster provided in an embodiment of this application; Figure 4 This is a flowchart of another incremental monitoring method for the state of a distributed cluster provided in an embodiment of this application; Figure 5 This is a flowchart of another incremental monitoring method for the state of a distributed cluster provided in an embodiment of this application. Detailed Implementation

[0008] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0009] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0010] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0011] In the operation and maintenance of distributed clusters using related technologies, causal analysis of cross-modal state data faces challenges, as it is difficult to directly discover and label different types of causal relationships within a unified framework. Existing solutions only assess changes in causal graphs by adding or deleting edges, lacking topological stability measurements and failing to detect key changes such as causal loops. Furthermore, monitoring strategies are adjusted in a single dimension and are open-loop controls, making it difficult to make predictive adjustments based on changes in causal structure, and lacking feedback verification, resulting in a mismatch between resource configuration and the actual cluster state.

[0012] To address this issue, this application proposes an incremental monitoring method for the state of a distributed cluster. This method unifies the representation of multimodal state event streams within the distributed cluster, obtaining time-series signals carrying modality source labels. Incremental causality checks and topology measurements are then performed on the time-series signals with modality source labels and the causal hypergraph, resulting in an updated causal hypergraph, causal type labels, and topology change metrics. Based on the topology change metrics, causal type labels, and modality source labels, monitoring scheduling signals and channel start / stop commands are generated. Based on these signals and commands, the access parameters of the multimodal state event streams are corrected, and the neighborhood expansion range for the next round of incremental causality checks is dynamically planned.

[0013] For ease of understanding, the following explains some key terms in this embodiment: A distributed cluster is a system in which multiple computer nodes are interconnected through a network and work together to provide a unified service.

[0014] Multimodal state event streams refer to various types of data streams generated during the operation of a distributed cluster, including continuous numerical indicator events, discrete text log events, and asynchronous call chain tracing events. Continuous numerical indicator events include data such as CPU utilization, memory usage, and network bandwidth, which change continuously over time. Discrete text log events include system error logs and application runtime logs, which are recorded in text format. Asynchronous call chain tracing events include events reflecting internal system interactions and performance, such as inter-service call chains and request processing times.

[0015] Modal source tags are used to identify additional information about the original modal type (such as numerical indicators, text logs, call chains) of each data point in a time series signal.

[0016] A time-series signal refers to a data sequence arranged in chronological order. In this embodiment, it is a multimodal state event stream after unified characterization.

[0017] A causal hypergraph is an extended graph structure in which nodes represent entities or state variables in a cluster, and edges or hyperedges represent causal relationships between these entities or variables. Hyperedges can represent many-to-one or one-to-many causal dependencies.

[0018] Incremental causality testing refers to the process of updating and verifying local causal relationships based on an existing causal hypergraph using newly arrived time-series signals.

[0019] Topological metrics are methods for quantitatively evaluating the structural features (such as connectivity, loops, density, etc.) of causal hypergraphs.

[0020] Causality type markers are used to identify the specific type of causal relationship discovered, such as linear causality, nonlinear causality, event-driven causality, etc.

[0021] The topology change metric quantifies the degree to which the topology of a causal hypergraph changes between different points in time.

[0022] The monitoring and scheduling signals are instructions used to adjust the data acquisition frequency and reported content based on changes in the causal hypergraph.

[0023] Channel start / stop commands are used to control the opening or stopping of specific modal data acquisition channels based on the requirements of cross-modal causality.

[0024] Access parameters refer to the configuration parameters of multimodal state event streams during the acquisition and processing phases, such as sampling interval and data source address.

[0025] The neighborhood expansion range refers to the range of causal relationship exploration that extends outward from a specific node in incremental causality testing.

[0026] This embodiment provides an incremental monitoring method for the status of a distributed cluster. See [link to documentation]. Figure 1 This includes the following steps.

[0027] 101. Perform a unified representation of the multimodal state event stream of the distributed cluster to obtain time-series signals carrying modal source labels. The multimodal state event stream includes continuous numerical index events, discrete text log events, and asynchronous call chain tracing events.

[0028] Specifically, all modal data is simply merged by timestamp. For example, the timestamps of all continuous numerical metric events, discrete text log events, and asynchronous call chain tracing events are aligned, and then their values ​​or event types are directly concatenated into a long sequence. Each data point is given a simple modal identifier, such as "numerical," "log," or "call chain."

[0029] 102. Perform incremental causality testing and topological measurement on the time-series signal carrying modal source labels and the causal hypergraph to obtain the updated causal hypergraph, causal type label and topological change measurement.

[0030] The incremental causality test adaptively selects the testing method based on the modal source label and adjusts the accumulation rate of causal edge confidence based on the previous round of topological change metrics. Specifically, a universal causal testing method, such as the Granger causality test, is applied to all node pairs to determine whether a causal relationship exists between them. The test results are directly used to update the edges in the causal hypergraph. The topological metric can be easily obtained by calculating the number of added or deleted causal edges. The accumulation rate of causal edge confidence can be set to a fixed value and is not adjusted with changes in the topology.

[0031] 103. Based on topology change measurement, causal type labeling, and modality source labeling, generate monitoring and scheduling signals and channel start / stop commands.

[0032] Specifically, topology change metrics drive adjustments to the acquisition frequency and incremental reporting content, while causal type labeling combined with modal source tags drives the on-demand activation and deactivation of cross-modal acquisition channels. In detail, by setting a fixed threshold, when the topology change metric exceeds this threshold, the acquisition frequency of all modalities is uniformly increased, and a complete causal hypergraph is reported. For channel activation and deactivation, based on the causal type labeling, it is manually configured which modal channels need to be always on and which can be deactivated.

[0033] 104. Based on monitoring and scheduling signals and channel start / stop instructions, the access parameters of the multimodal state event stream are corrected by feedback, and the neighborhood expansion range of the next round of incremental causality test is dynamically planned.

[0034] Specifically, the acquisition frequency adjustment signal from the monitoring and scheduling signal is directly applied to all modal channels, and the corresponding channel is simply turned on or off according to the channel start / stop command. The neighborhood expansion range of the next round of incremental causality testing can be set to a fixed value, for example, always checking all nodes.

[0035] This embodiment provides an incremental monitoring method for distributed cluster status. By uniformly representing multimodal state event flows and performing incremental causal verification and topological measurement, it achieves the discovery and labeling of different types of cross-modal causal relationships within a unified framework. This method can measure the stability of the causal structure at the topological level and perceive the impact of key changes such as causal loops on the cluster status. Based on the trend of causal structure changes, the monitoring strategy is dynamically adjusted, including the collection frequency, reporting content, and the on-demand start and stop of cross-modal collection channels, with feedback corrections. This ensures that the configuration of monitoring resources is synchronized with the actual changes in the cluster status, improving the efficiency and accuracy of monitoring.

[0036] This application further proposes a method for unified characterization of multimodal state event streams in distributed clusters, obtaining time-series signals carrying modality source labels. See [link to relevant documentation]. Figure 2 The method includes: 201. Perform time-series alignment on the continuous numerical index event to obtain the numerical time-series components.

[0037] 202. Based on the fluctuation range of the numerical time series component within the target time period, the kernel function bandwidth for continuous reconstruction of the discrete text log event is adaptively adjusted, and continuous reconstruction is performed with the adjusted kernel function bandwidth to obtain the continuous log component.

[0038] 203. Determine the weighting coefficients based on the semantic features of the discrete text log event, and perform weighted aggregation of the arrival time intervals of the asynchronous call chain tracing event based on the weighting coefficients to obtain the continuous component of the call chain.

[0039] 204. Merge the numerical time-series component, the log continuous component, and the call chain continuous component on a unified time axis, assign the corresponding modal source label, and obtain the time-series signal carrying the modal source label.

[0040] For example, time-series alignment of continuous numerical index events to obtain numerical time-series components refers to unifying continuous numerical index events from different sources or with different sampling frequencies on the time axis, making them comparable at a common point in time. Its purpose is to provide a unified time benchmark for subsequent cross-modal data fusion and causal analysis, ensuring the accuracy and comparability of numerical indicators. For instance, interpolation methods (such as linear interpolation and spline interpolation) can be used to unify data with different sampling frequencies to a fixed sampling interval. If the original data sampling intervals are inconsistent, a unified sampling interval (such as 1 second) can be set, and then interpolation calculations can be performed on the values ​​at each time point. Alternatively, resampling techniques can be used to downsample high-frequency data or upsample low-frequency data to match the target time granularity. For example, CPU utilization data collected per second can be aggregated into an average value per minute, or memory usage data collected per minute can be extended to per second while preserving previous values.

[0041] Based on the fluctuation amplitude of the numerical time-series component within the target time period, the kernel function bandwidth for continuous reconstruction of the discrete text log event is adaptively adjusted, and continuous reconstruction is performed with the adjusted kernel function bandwidth to obtain the continuous log component. The fluctuation amplitude is an indicator that measures the drastic change of the numerical time-series component within a specific time period, reflecting the stability or activity of the cluster state. The kernel function bandwidth is a key parameter in kernel density estimation (KDE), which determines the "range of influence" or "smoothing degree" of each discrete event on the continuous signal. Adaptively adjusting the kernel function bandwidth aims to dynamically optimize the continuous reconstruction process of discrete text log events according to the real-time state of the cluster (reflected by the fluctuation of numerical indicators), thereby more accurately capturing the temporal characteristics of the log events. For example, the fluctuation amplitude can be obtained by calculating the standard deviation, variance, or the difference between the maximum and minimum values ​​of the numerical time-series component within the target time period. The adaptive adjustment of the kernel function bandwidth can be based on a mapping function, mapping the fluctuation amplitude to a bandwidth value. When the fluctuation amplitude is large, a smaller bandwidth is selected to highlight short-term event details. When the fluctuation amplitude is small, a larger bandwidth is selected to smooth long-term trends. Alternatively, the fluctuation amplitude can be characterized by calculating the mean absolute deviation or coefficient of variation of the numerical time series components within the target time period. The kernel function bandwidth can be adjusted using piecewise linear or nonlinear functions, dynamically selecting different bandwidth values ​​based on the magnitude of the fluctuation amplitude. For example, multiple fluctuation amplitude ranges can be set, with each range corresponding to a specific bandwidth adjustment strategy.

[0042] Weighting coefficients are determined based on the semantic features of discrete text log events, and these coefficients are then used to aggregate the asynchronous call chain tracing events by arrival time interval weighting, resulting in continuous call chain components. Semantic features refer to the meaning or importance inherent in discrete text log events, such as the severity level (error, warning, message, etc.), keywords, or event type. Weighting coefficients are numerical values ​​used to quantify the influence of these semantic features on the subsequent aggregation process. Determining weighting coefficients through semantic features allows for prioritizing log events with higher importance or indicating more serious problems when aggregating asynchronous call chain tracing events, thus more accurately reflecting the potential propagation paths and impacts of anomalies in the cluster. For example, semantic features can be analyzed using Natural Language Processing (NLP) techniques to extract keywords, sentiment, or predefined event patterns from the log text. Weighting coefficients can be assigned values ​​based on the severity level of the logs; for example, error logs are given high weight, warning logs medium weight, and message logs low weight. Alternatively, semantic features can also be used through machine learning models to classify log events and identify log types associated with specific failure modes. The weighting coefficients can be dynamically calculated based on the correlation strength between log events and known failure modes or their importance in historical failure analysis. For example, the weights of different log types can be determined through expert experience or historical data statistics.

[0043] The numerical time-series component, the log continuous component, and the call chain continuous component are merged on a unified time axis and assigned corresponding modality source labels to obtain the time-series signal carrying the modality source label. Merging refers to integrating preprocessed time-series data from different modalities on the same time axis according to their timestamps. The modality source label is metadata used to identify the original modality type (such as numerical metric, text log, call chain) of each time-series data point. By merging and assigning modality source labels, a unified multimodal time-series dataset is formed. This dataset is not only time-aligned but also retains the original modality information of each data point. This is crucial for subsequent cross-modal causal analysis because it allows causal testing methods to adaptively select based on the characteristics of the data modality. For example, merging can be achieved by sorting the timestamps of all components and then merging the data points with corresponding timestamps into a unified data structure. The modality source label can be directly used as a field in the data structure, storing string identifiers such as "metric", "log", or "trace". Alternatively, merging can also be achieved by creating a unified time grid and projecting the data of each component onto that grid. Modal source labels can be enumerated types or binary encoding to efficiently identify the modal attributes of each data point.

[0044] The above technical solution provides a unified representation of the multimodal state event stream of a distributed cluster, resulting in a time-series signal carrying modal source labels. This solves the problems of mismatched data characteristics across different modalities, inability to adjust representation parameters in conjunction with cluster state fluctuations, and poor time-series signal quality. For example, by aligning continuous numerical index events in time series, a unified time benchmark and accurate basis for cluster state fluctuations are provided for subsequent processing. Based on the fluctuation amplitude of numerical time-series components within the target time period, the kernel function bandwidth for continuous reconstruction of discrete text log events is adaptively adjusted. This allows the continuous reconstruction of logs to dynamically adapt to the actual state changes of the cluster. When cluster fluctuations are large, a smaller bandwidth is used to capture short-term details; when the cluster is stable, a larger bandwidth is used to integrate long-term effects. This avoids the problem of fixed-bandwidth reconstruction failing to adapt to cluster state changes, ensuring that continuous log components better reflect the actual cluster state. Weighting coefficients are determined based on the semantic features of discrete text log events, and these are used to perform weighted aggregation of arrival time intervals for asynchronous call chain tracing events. This allows the call chain aggregation results to reflect the correlation with the severity of log events, improving the relevance of the unified representation and more accurately reflecting the actual characteristics of state propagation in the cluster. Merging these components on a unified timeline and assigning modality source labels not only ensures the temporal alignment of all modality data but also preserves the modality source information of each data point. This provides high-quality foundational data for subsequent adaptive selection of causal testing methods based on modality source labels, significantly improving the accuracy of causal analysis.

[0045] This application further proposes the above-mentioned method, wherein, based on the fluctuation amplitude of the numerical time-series component within the target time period, the kernel function bandwidth for continuous reconstruction of the discrete text log event is adaptively adjusted, including: determining the variance of the numerical time-series component within the target time period as the fluctuation amplitude. When the fluctuation amplitude exceeds a preset threshold, the kernel function bandwidth for continuous reconstruction of the discrete text log event is reduced to capture the short-term time-series characteristics of the discrete text log event. When the fluctuation amplitude does not exceed the preset threshold, the kernel function bandwidth is increased to integrate the long-term time-series effects of the discrete text log event.

[0046] For example, when determining the variance of a numerical time-series component within a target time period as the amplitude of fluctuation, this variance is a statistic that measures the degree of data dispersion and can intuitively reflect the severity of fluctuations in the numerical indicator within a specific time period. Its implementation can include: collecting data points for all numerical time-series components within the target time period, calculating their mean, and then calculating the average of the sum of squares of the differences between each data point and the mean. Alternatively, a sliding window mechanism can be used to calculate the variance of the data within the window with a fixed or dynamic step size within the target time period to capture local fluctuation characteristics.

[0047] When the fluctuation amplitude exceeds a preset threshold, the bandwidth of the kernel function used for the continuous reconstruction of the discrete text log event is reduced to capture its short-term temporal characteristics. The kernel function bandwidth is a key parameter in kernel density estimation, determining the smoothing range of the kernel function around the data points. Reducing the bandwidth makes the kernel function "sharper," thereby enhancing its ability to capture local details. This can be achieved by multiplying the current kernel function bandwidth by a fixed coefficient less than 1 (e.g., 0.5 or 0.7) or subtracting a fixed value when the fluctuation amplitude exceeds the preset threshold. Alternatively, a new bandwidth value can be dynamically calculated using a predefined function (such as a logarithmic function, exponential function, or piecewise linear function) based on the degree to which the fluctuation amplitude exceeds the threshold; the larger the amplitude, the greater the bandwidth reduction.

[0048] When the fluctuation amplitude does not exceed a preset threshold, increasing the kernel function bandwidth helps to integrate the long-term temporal effects of the discrete text log events. Increasing the bandwidth makes the kernel function "flatter," thereby enhancing its ability to integrate long-term trends or cumulative effects. This can be achieved by multiplying the current kernel function bandwidth by a fixed coefficient greater than 1 (e.g., 1.2 or 1.5) or adding a fixed value when the fluctuation amplitude does not exceed the preset threshold. Alternatively, a new bandwidth value can be dynamically calculated using a predefined function based on the difference between the fluctuation amplitude and the threshold; the larger the difference (i.e., the more stable the data), the greater the bandwidth increase.

[0049] The above technical solution adaptively adjusts the kernel function bandwidth by utilizing the fluctuations of numerical time-series components, solving the problem that fixed bandwidth cannot meet the continuous log reconstruction requirements under different cluster states, and improving the accuracy of unified representation of multimodal states. The variance of the numerical time-series components within the target time period is selected as the fluctuation amplitude. Variance can intuitively reflect the overall volatility of numerical indicators, is simple to calculate, and yields reliable results, providing an accurate basis for subsequent kernel function bandwidth adjustments. Adjusting the log reconstruction bandwidth based on the fluctuations of numerical indicators allows the sensitivity of numerical indicators to changes in cluster state, ensuring that the granularity of log reconstruction matches the changing characteristics of the current cluster state. When the fluctuation amplitude exceeds a preset threshold, the kernel function bandwidth is reduced to capture the short-term temporal characteristics of discrete text log events. A fluctuation amplitude exceeding the threshold indicates that the current cluster state is changing rapidly, requiring higher log temporal resolution. A smaller kernel function bandwidth avoids over-smoothing of log features at different times, preserving short-term log change characteristics and adapting to the need for capturing log temporal information under rapidly changing cluster states. When the fluctuation amplitude does not exceed a preset threshold, the kernel function bandwidth is increased to integrate the long-term temporal effects of discrete text log events. A fluctuation amplitude within the threshold indicates that the cluster's current state is changing smoothly. A larger kernel function bandwidth can integrate the cumulative effects of logs over a longer period, avoiding the log feature fragmentation problem caused by small bandwidth, and more accurately reflecting the long-term impact of logs on the current cluster state. This adaptive adjustment mechanism allows continuous log components to more accurately reflect the true temporal characteristics of discrete text log events under the current cluster state, thus providing more accurate and unified representation data for subsequent incremental causal testing, improving the accuracy and efficiency of distributed cluster state monitoring.

[0050] This application further proposes specific steps for performing continuous reconstruction using the adjusted kernel function bandwidth to obtain continuous log components, including: grouping the discrete text log events by event type and extracting the event occurrence time sequence of each group on the time axis; performing kernel density estimation on the event occurrence time sequence of each group using the adjusted kernel function bandwidth to obtain the corresponding continuous density signal for each group; merging the corresponding continuous density signals of each group on a unified time axis and concatenating them with the corresponding semantic feature vectors to obtain the continuous log component.

[0051] For example, grouping discrete text log events by event type aims to categorize them based on their inherent attributes or content. Different event types often correspond to different system behavior patterns, error types, or operational states. Grouping ensures that subsequent processing targets semantically similar events, thus avoiding confusion of temporal patterns between different event types. For instance, using predefined log templates, regular expressions or natural language processing techniques can be used to match incoming log events with known templates, classifying them into the corresponding event type. Alternatively, for unknown or dynamically changing log patterns, unsupervised learning algorithms, such as K-Means or DBSCAN, can be used to extract features from the log text and then cluster them, grouping similar log content into one category. After grouping, it is necessary to extract the event occurrence time sequence for each group on the timeline. The purpose of this step is to obtain the precise timestamps of events from each event type group. These timestamp sequences constitute the raw data for that specific event type in the time dimension, which is crucial for understanding its occurrence pattern. Typically, these timestamps are directly parsed from the log records and arranged in chronological order. Alternatively, if the log itself does not contain an accurate timestamp, the event arrival time recorded by the event handling system can be used as the occurrence time.

[0052] Kernel density estimation is performed on the event occurrence time series of each group using the adjusted kernel function bandwidth. Kernel density estimation is a non-parametric method used to estimate the probability density function of a random variable. By applying it to the event occurrence time series, discrete event points are transformed into smooth continuous signals, reflecting the intensity or frequency of events over time. The "adjusted kernel function bandwidth" used here is adaptively determined based on the fluctuation amplitude of the numerical time series components within the target time period. It controls the smoothness of the estimated density function, thus adapting to the volatility of the underlying system state. Performing kernel density estimation on each group separately ensures that the unique temporal characteristics of each event type are preserved, avoiding dilution by other event types. Commonly used kernel functions include the Gaussian kernel function, the Epanechnikov kernel function, or the trigonometric kernel function, with the Gaussian kernel function often chosen due to its good smoothness and computational efficiency. The result of kernel density estimation is the corresponding continuous density signal for each group. These signals are numerical sequences or functions with time indices; higher values ​​indicate higher density or frequency of that specific event type at the corresponding time point, thus transforming discrete log events into a quantifiable continuous representation.

[0053] The continuous density signals of each group are merged on a unified time axis. This merging operation aims to integrate continuous density signals of different event types into a single, coherent representation. This is crucial for aligning log data with other modalities (such as numerical metrics and call chain tracing events) on a common time reference, thereby enabling comprehensive multimodal analysis. The merging can be achieved by simply summing the continuous density signal values ​​of all groups at the same time point to obtain a single overall log event density signal. Alternatively, the continuous density signals of each group can be treated as different dimensions or features, and their feature vectors can be concatenated on a unified time axis. This concatenation is then combined with the corresponding semantic feature vectors of each group. This step further enriches the purely time-based continuous density signal by introducing the semantic information of the log events themselves. Semantic features provide context and meaning beyond event frequency, helping to understand the potential impact of log data more comprehensively. Semantic feature vectors can be encoded based on predefined semantic labels (such as "error," "warning," etc.), or by using text embedding models (such as Word2Vec, BERT) to convert typical log text into high-dimensional semantic vectors, or by combining the severity level of the log. Through the above steps, the continuous component of the log is obtained. This component is a unified, continuous and multi-dimensional representation of discrete text log events. It contains both temporal dynamics (from continuous density signals) and semantic meaning (from semantic feature vectors), enabling it to be better integrated with other modal data and used for subsequent causal analysis.

[0054] The above technical solution solves the problem that failing to distinguish between different event types during continuous reconstruction of discrete text log events can lead to confusion in temporal distribution features, inaccurate representation by semantic features, and negative impacts on the accuracy of subsequent causal tests. For example, by finely grouping discrete text log events by event type, it ensures that the unique temporal patterns of each event type can be captured and processed independently, avoiding mutual interference between features of different log types. Utilizing a kernel function bandwidth that adaptively adjusts based on the fluctuation amplitude of numerical indicators, kernel density estimation is performed on the occurrence time sequences of each group of events. This allows the continuous reconstruction process to accurately match the temporal distribution characteristics of each log event and adapt to fluctuations in the overall cluster state, thereby significantly improving the accuracy and representativeness of continuous log components. Furthermore, merging the continuous density signals of each group on a unified time axis and concatenating them with the corresponding semantic feature vectors of each group not only achieves temporal alignment of log data with other modal data, but more importantly, it integrates the semantic information of log events into the continuous representation, so that the continuous components of the log have both rich temporal dynamics and semantic context, providing a more comprehensive and accurate input for subsequent cross-modal causal relationship testing, thereby improving the accuracy and effectiveness of distributed cluster status monitoring.

[0055] This application further proposes a method for determining weighting coefficients based on the semantic features of discrete text log events, comprising: identifying the severity level of the discrete text log event to obtain a severity level identifier; determining a baseline value for the weighting coefficient based on the severity level identifier; and correcting the baseline value based on the fluctuation range of the continuous numerical index event within the corresponding time period to obtain the weighting coefficient.

[0056] To better understand the above technical solutions, the key technical features involved will be explained in detail below.

[0057] The core of the step "identifying the severity level of the discrete text log event and obtaining its severity level label" lies in extracting information about the degree of impact on the system state contained in the original text log. This can be achieved in several ways. For example, a set of rules can be predefined, including keyword matching (such as "ERROR", "CRITICAL", "WARNING", "INFO", etc.) or regular expression patterns, to automatically identify and assign the corresponding severity level by matching the log content. Another approach is to use a machine learning model. By training on a large amount of log data already labeled with severity levels, the model can automatically classify new discrete text log events, thereby obtaining their severity level labels. These labels can be enumerated values ​​(such as levels 1 to 5) or more fine-grained numerical values.

[0058] The step of "determining the baseline value of the weighting coefficient based on the severity level identifier" aims to transform the identified severity level into a quantifiable initial weight. For example, a mapping table or function can be created to directly map different severity level identifiers to their corresponding baseline values. For instance, "CRITICAL" can be mapped to a higher baseline value (e.g., 10), "ERROR" to a second-highest value (e.g., 7), "WARNING" to a medium value (e.g., 3), and "INFO" to a lower value (e.g., 1). Alternatively, a non-linear function can be used so that a small increase in severity level leads to a more significant increase in the baseline value, better reflecting the potential impact of high-severity events.

[0059] The step of "adjusting the benchmark value based on the fluctuation range of the continuous numerical index event within the corresponding time period to obtain the weighting coefficient" aims to introduce dynamic information about the overall system operation status to adjust the initially determined benchmark value to better reflect reality. This fluctuation range can be measured by various statistical measures, such as calculating the standard deviation, variance, range, or coefficient of variation of the continuous numerical index event within the target time period. Methods for adjusting the benchmark value can include: when the fluctuation range is large, multiplying the benchmark value by a correction factor greater than 1 to amplify the impact of the log event; when the fluctuation range is small, using a correction factor close to or less than 1 to reduce its impact. For example, a correction function can be designed that takes the fluctuation range as input, outputs a correction coefficient, and then multiplies or adds the benchmark value to this correction coefficient to obtain the weighting coefficient.

[0060] The above technical solution enables a more accurate and dynamic determination of the weighting coefficients for discrete text log events. By identifying the severity level of discrete text log events, the inherent impact of each event can be extracted from the log itself, providing a basis for determining the weighting coefficients that aligns with actual business meaning and avoiding potential biases from relying solely on semantic features. Determining the baseline value for the weighting coefficients based on the severity level identifier ensures that log events of different severity levels receive appropriate initial weights in subsequent processing, intuitively reflecting their different impacts on the cluster state. Furthermore, by introducing the fluctuation range of continuous numerical indicator events within the corresponding time period to correct the baseline value, the weighting coefficients can dynamically adapt to the overall operating state of the cluster. When the cluster state fluctuates drastically, the weight of the log event is amplified accordingly to highlight its potential anomaly indication role. When the cluster state is stable, the weight may be moderately adjusted to avoid oversensitivity. This two-step correction mechanism ensures that the resulting weighting coefficients not only consider the severity of the log event itself but also incorporate the overall dynamic context of the cluster, thus more accurately reflecting the actual impact of discrete text log events. This significantly improves the accuracy of weighted aggregation of subsequent asynchronous call chain tracing events, enabling the continuous components of the aggregated call chain to more realistically represent the actual state change patterns of the cluster. This provides a more reliable and accurate data foundation for subsequent incremental causal testing and distributed cluster state monitoring, and solves the problem of unreasonable weighting caused by relying solely on semantic features to determine weighting coefficients.

[0061] This application further proposes a method for weighted aggregation of arrival time intervals of asynchronous call chain tracing events based on weighted coefficients to obtain continuous components of the call chain. The method includes: Extract the inter-service call relationship from the asynchronous call chain tracing event, and divide the asynchronous call chain tracing event into event groups with different call paths according to the inter-service call relationship.

[0062] For each event group, the time interval between the arrival time and the target time of each event in the event group is attenuated and corrected using the weighting coefficient as the baseline weight, so as to obtain the delay contribution value of each event to the target time.

[0063] The delay contribution of each event to the target time is accumulated by the event group, and the accumulated results of each event group are merged according to the topology of the inter-service call relationship to obtain the continuous component of the call chain.

[0064] For example, extracting inter-service call relationships from asynchronous call chain tracing events refers to identifying the dependency patterns between services in a distributed cluster. This can be achieved by parsing metadata such as parent-child span information, service names, and operation names contained in the call chain tracing events, thereby constructing a call graph between services. For example, analyzing the source and target service identifiers of each tracing event, or reconstructing the complete call path using the tracing ID and span ID. Another approach is to utilize the APIs or data models provided by distributed tracing systems (such as OpenTelemetry and Zipkin) to directly query or extract the service dependency graph.

[0065] Dividing asynchronous call chain tracing events into event groups based on inter-service call relationships means grouping asynchronous call chain tracing events belonging to the same logical call path together based on the extracted inter-service call relationships. For example, a complete business request might go through a call path: service A calls service B, and service B then calls service C, forming an "A→B→C" call path. All tracing events belonging to this path will be grouped into one event group. This can be achieved by generating a unique identifier for each unique call path and associating the event with that identifier. Alternatively, common call path templates can be predefined, and events matching these templates can be grouped into the corresponding event groups.

[0066] For each event group, using the weighted coefficient as the baseline weight, the time interval between the arrival time and the target time of each event within that group is attenuated to obtain the delay contribution value of each event to the target time. This means that when calculating the impact of each asynchronous call chain tracing event on the current target time, not only its intrinsic importance (reflected by the weighted coefficient) is considered, but also the distance between its occurrence time and the target time. For example, using an exponential or linear attenuation function, the time interval between the event's arrival time and the target time is used as input to calculate an attenuation factor. Then, this attenuation factor is multiplied by the weighted coefficient to obtain the delay contribution value of the event. The design of the attenuation function should ensure that the longer the time interval, the smaller the attenuation factor, so that earlier events contribute less to the current target time.

[0067] The latency contribution of each event to the target time is summed by event group. This means that within each predefined event group of the call path, the latency contribution values ​​calculated for all events are summed. For example, if an event group contains events e1, e2, and e3, and their latency contribution values ​​are C1, C2, and C3 respectively, then the summation result for that event group is C1 + C2 + C3. This ensures that the overall activity or impact of each call path is quantified into a single value.

[0068] The cumulative results of each event group are merged according to the topology of the inter-service call relationships to obtain the continuous components of the call chain. This means that after obtaining the cumulative contribution value of each call path event group, these cumulative results are integrated according to the actual call dependencies between services to form a continuous time-series signal that reflects the state of the entire distributed cluster call chain. For example, a service dependency graph is constructed to associate each service node with the cumulative results of its related event groups. Then, through graph traversal or aggregation algorithms, the contributions of downstream services are propagated to upstream services, or the contributions of parallel call paths are merged (such as summation, averaging, or maximization) to obtain one or a group of continuous components representing the state of the entire or key parts of the call chain.

[0069] The above technical solution addresses the asynchronous nature of asynchronous call chain tracing events and the service call topology characteristics. By combining the weighting coefficients previously determined based on the semantic features of discrete text log events, it transforms discrete asynchronous call events into unified temporal continuous components. This solution extracts inter-service call relationships from asynchronous call chain tracing events and divides events into event groups with different call paths, solving the problem of previous aggregation methods ignoring call path differences. This allows events with different service interaction logics to be processed independently, preserving the structural characteristics of the call path itself. By using the weighting coefficients as the baseline weight, the time interval between the arrival time and the target time of each event within the event group is attenuated, yielding the delay contribution value of each event to the target time. This not only incorporates the semantic importance reflected by the severity level of discrete text log events but also reasonably quantifies the actual rule that earlier-occurring call events have less impact on the current target time, ensuring the accuracy of the continuous components' reflection of the call state. The delay contribution values ​​of each event to the target time are accumulated by event group and merged according to the topology of the service call relationship. This not only preserves the state information of different call paths, but also fits the actual topology of the distributed cluster service call. This allows the continuous components of the resulting call chain to be directly aligned with the continuous components of other modalities on a unified time axis, meeting the alignment and integration requirements of unified temporal representation of multimodal states. This provides accurate and reliable input for subsequent incremental causal verification, thereby improving the accuracy and effectiveness of distributed cluster state monitoring.

[0070] This application further proposes incremental causality testing and topological measurement of the time-series signal carrying modal source labels and the causal hypergraph, resulting in an updated causal hypergraph, causal type label, and topological change measurement. See [link to relevant documentation]. Figure 3 The method includes the following steps.

[0071] 301. Based on the modality source label, for the node pairs in the local subgraph affected by the time series signal in the causal hypergraph, an incremental causal test method is adaptively selected to perform causal test, and the causal relationship between the node pairs and the causal type label are obtained.

[0072] Modal source tags are metadata used to identify the modal type (e.g., continuous numerical index events, discrete text log events, asynchronous call chain tracing events) of each data point in the multimodal state event stream of a distributed cluster. Local subgraphs refer to the parts of the causal hypergraph directly related to the current time-series signal changes. Their influence range is typically determined by analyzing the activity or trend of the time-series signal, thus avoiding unnecessary traversal and computation of the entire large causal hypergraph. Adaptive selection of causal testing methods refers to dynamically selecting the most suitable causal testing algorithm based on the modal combination type of node pairs within the local subgraph (e.g., both nodes are numerical indices, one is a numerical index and the other is a text log event). For example, for node pairs with the same modality, statistically based testing methods such as Granger causality tests or transitive entropy can be used. For cross-modal node pairs, methods based on specific domain models, such as methods based on stochastic process modeling or event sequence analysis, can be used. Incremental causality testing, on the other hand, verifies and updates causal relationships only for node pairs within a local subgraph based on an existing causal hypergraph, rather than reconstructing the entire causal graph from scratch each time, significantly improving efficiency. This step accurately identifies causal relationships between node pairs and assigns them corresponding causal type labels, such as linear causality, nonlinear causality, and event-driven causality, providing a refined basis for subsequent structural updates and monitoring scheduling.

[0073] 302. Based on the causal relationship and the causal type label, update the structure and attributes of the local subgraph, and adjust the confidence accumulation rate of the causal edges in the local subgraph according to the topological change metric of the previous round, to obtain the updated causal hypergraph.

[0074] The structural and attribute updates refer to dynamically adding new causal edges or hyperedges, deleting invalid causal edges, or modifying the attributes of existing causal edges (such as causal strength and causal type labels) in the local subgraph based on the results of incremental causal testing. For example, when a new causal relationship is discovered, a corresponding causal edge or hyperedge is created in the local subgraph and labeled with a causal type label. The confidence accumulation rate of causal edges refers to the speed at which causal relationships are confirmed or denied. Adjusting the confidence accumulation rate based on the previous round of topology changes means that when the topology of the causal hypergraph is in a stable state, a slower confidence accumulation rate is used to avoid misjudgments. When the topology changes drastically, the confidence accumulation rate can be accelerated to respond more quickly to changes in the cluster state and promptly confirm or deny causal relationships. This dynamic adjustment mechanism makes the update process of the causal hypergraph more flexible and accurate, and better adapts to the dynamic characteristics of distributed clusters.

[0075] 303. Based on the causal strength of each causal edge in the updated causal hypergraph, construct a filtering sequence, and measure the topological change of the updated causal hypergraph on the filtering sequence to obtain the topological change metric.

[0076] Causal strength is an indicator that quantifies the strength of causal relationships. It can be a statistic for causal testing, an information-theoretic measure (such as transit entropy), or a difference based on model fit. Constructing a filtering sequence involves performing a series of filtering operations on the updated causal hypergraph based on the causal strength values, generating a series of subgraphs with varying sparsity. For example, multiple causal strength thresholds can be set, and the causal hypergraph can be filtered sequentially using these thresholds, retaining causal edges with a causal strength not lower than the current threshold. This results in a series of subgraphs, which form the filtering sequence in descending order of strength thresholds. This serialized subgraph representation can reveal the structural characteristics of causal relationships at different strength levels. Measuring topological changes on the filtering sequence involves analyzing how the topological characteristics (e.g., the number of connected components, the number of loop structures, etc.) of each subgraph in the filtering sequence change with the strength thresholds to quantify the overall topological changes of the causal hypergraph. This multi-scale, multi-granularity measurement method can capture the structural changes of the causal hypergraph at different causal strength levels, including the formation and disappearance of causal loops, thus providing a more comprehensive and accurate measure of topological changes.

[0077] The above technical solutions address the shortcomings of existing methods in incremental causality testing and topology measurement, including the lack of refined execution logic, insufficient adaptability, and inaccurate measurement of topology changes. For example, by adaptively selecting a causality testing method based on modality source labels, this application can employ the most suitable testing strategy for node pairs with different modality combinations, avoiding the limitations of single testing methods in identifying cross-modal causal relationships and significantly improving the accuracy of causal relationship identification. Incremental causality testing and structural attribute updates are performed only on local subgraphs affected by time-series signals, greatly reducing computational overhead and improving update efficiency. Furthermore, the accumulation rate of causal edge confidence is dynamically adjusted based on the previous round of topology change measurement, allowing the confirmation or denial process of causality to better adapt to dynamic changes in cluster state, avoiding oversensitivity in stable states or slow response during drastic changes. This ensures that the updated causal hypergraph more accurately reflects the true causal dependencies of the cluster. Furthermore, by constructing a filtering sequence based on causal strength and measuring topological changes on it, this application can capture structural changes in causal hypergraphs at multiple scales and granularities, filter out noise interference from weak causal relationships, and perceive the formation and demise of causal loops, thus providing a more comprehensive and accurate measurement of topological changes. These improvements collectively provide more accurate, reliable, and refined inputs for subsequent monitoring and scheduling signal generation and monitoring strategy adjustment, enabling the monitoring system to respond more intelligently and efficiently to changes in the operational status of distributed clusters.

[0078] This application further proposes an incremental causal test based on modal source labels. For node pairs within local subgraphs affected by time-series signals in a causal hypergraph, an adaptive causal test method is selected to obtain the causal relationship and causal type label between node pairs. Specifically, when the modal combination type of the node pair is numerical index and numerical index, a linear causal test is performed on the node pair to obtain a linear test statistic, and a nonlinear causal test is performed on the node pair to obtain a nonlinear test statistic. When both the linear test statistic and the nonlinear test statistic satisfy the linear significance condition, it is labeled as a strong causal relationship. When only the linear test statistic satisfies the linear significance condition, it is labeled as a weak linear causal relationship. When only the nonlinear test statistic satisfies the nonlinear significance condition, it is labeled as a nonlinear causal relationship. When the modal combination type of the node pair is numerical index and discrete text log events, a stochastic process model is performed on the occurrence sequence of discrete text log events to obtain a baseline fitting model containing only historical event information. Numerical indices are introduced as external explanatory variables into the stochastic process model to obtain an extended fitting model containing numerical index information. By comparing the goodness-of-fit differences between the baseline and extended fitting models, the causal relationships between node pairs are determined and labeled as event-driven causal relationships.

[0079] For node pairs where both modality combinations are numerical indicators, this application employs a dual causality testing mechanism, simultaneously performing linear and nonlinear causality tests. Linear causality tests aim to capture direct, proportional influence relationships between variables, such as Granger causality tests or vector autoregression (VAR) model analysis. This can be achieved by constructing a predictive model where the future value of one variable is predicted by both its own historical value and the historical value of another variable, and the linear causal relationship is determined by comparing the difference in predictive performance with and without the historical value of the other variable. Another approach is a linear approximation method based on mutual information or transitive entropy. Nonlinear causality tests are used to identify more complex, non-proportional influences between variables, such as kernel function-based methods (e.g., kernel Granger causality tests), information theory methods (e.g., transitive entropy), or neural network-based predictive models. This can be achieved by constructing a nonlinear predictive model, such as using support vector regression (SVR) or long short-term memory networks (LSTM), to predict one variable and assessing whether the introduction of another variable significantly improves predictive accuracy. These two tests comprehensively assess the various causal relationships that may exist between numerical indicators, avoiding the omission of important causal information due to the limitations of a single test method.

[0080] After completing linear and nonlinear causality tests, this application refines the causal relationship labeling based on whether the obtained test statistics meet the significance criteria. The significance criteria are typically determined by comparing the results with a significance level (e.g., p-value). When both the linear and nonlinear test statistics reach significance, it indicates a significant linear and nonlinear influence between the two numerical indicators, and this is labeled as a "strong causal relationship" to reflect its multi-dimensional and high-intensity association. If only the linear test statistic meets the significance criteria, it is labeled as a "weak linear causal relationship," indicating a predominantly linear influence, but with no significant nonlinear influence. If only the nonlinear test statistic meets the significance criteria, it is labeled as a "nonlinear causal relationship," indicating a predominantly nonlinear influence, but with no significant linear influence. This hierarchical labeling helps to more accurately understand the nature and strength of the causal relationship, providing more granular information for subsequent fault diagnosis and monitoring strategy adjustments.

[0081] For cross-modal numerical indicators and discrete text log event node pairs, this application employs a specially designed causality test method. The occurrence sequence of discrete text log events is modeled as a stochastic process, using models such as Poisson regression, negative binomial regression, or autoregressive moving average (ARMA) to capture the time dependence of the log events themselves, constructing a "benchmark fit model" based solely on historical log information. The numerical indicator is then introduced as an external explanatory variable (or covariate) into the stochastic process modeling, constructing an "extended fit model" that incorporates numerical indicator information. For example, in Poisson regression, the numerical indicator is used as one of the predictor variables. The causal relationship is determined by comparing the goodness-of-fit difference between the two models. The goodness-of-fit difference can be evaluated using statistics such as the likelihood ratio test, the Akaike Information Criterion (AIC), or the Bayesian Information Criterion (BIC). If the extended fit model significantly improves the goodness-of-fit compared to the benchmark fit model, the numerical indicator is considered to have a causal influence on the occurrence of discrete text log events, and this is labeled as an "event-driven causal relationship." This method can directly process discrete event data, avoiding the information loss that may result from forced conversion to continuous data, and more accurately reveals how changes in numerical indicators drive the occurrence of log events.

[0082] The above technical solution adaptively selects the most suitable causal testing method for node pairs with different modal combinations and refines the labeling of causal relationships, thus overcoming the problems of single causal testing methods and inaccurate identification of cross-modal causal relationships in related technologies. For example, for node pairs that are both numerical indicators, performing both linear and nonlinear causal tests can comprehensively capture various complex relationships that may exist between variables, avoid missing important causal information, and subdivide the causal relationship into strong causality, weak linear causality, or nonlinear causality according to the significance of the test results, providing richer causal attribute information. For cross-modal node pairs of numerical indicators and discrete text log events, this application innovatively adopts a method based on stochastic process modeling and goodness-of-fit comparison to directly process discrete log event sequences, avoiding the information distortion that may be caused by forcing discrete events to be continuous in traditional methods. It can accurately identify the driving effect of numerical indicators on log events and label them as event-driven causal relationships. This adaptive causality test and refined causality type labeling provide an accurate and reliable basis for subsequent monitoring and scheduling based on causality type. For example, the collection frequency can be adjusted preferentially based on strong causal relationships, and cross-modal collection channels can be started and stopped as needed based on event-driven causal relationships. This enables more intelligent and efficient distributed cluster status incremental monitoring, significantly improving the accuracy of fault location and the timeliness of early warning.

[0083] This application further proposes performing a linear causality test on the node pair to obtain a linear test statistic, and performing a nonlinear causality test on the node pair to obtain a nonlinear test statistic, specifically including: For the time series of the first node in the node pair, linear prediction modeling is performed using the time series of the second node to obtain the linear prediction residual sequence.

[0084] Based on the linear prediction residual sequence and the time series of the second node, the nonlinear information residual in the linear prediction residual sequence explained by the time series of the second node is determined as the nonlinear test statistic.

[0085] Based on the goodness of fit of the linear prediction model, the linear test statistic is determined.

[0086] In this process, the time series of the first node in the node pair is used to perform linear prediction modeling with the time series of the second node, resulting in a linear prediction residual sequence. This aims to isolate the portion of the first node's time series that can be linearly explained by the second node's time series. By constructing a linear prediction model, components in the first node's time series that are linearly correlated with the second node's time series are captured and removed, thus ensuring that the remaining residual sequence mainly contains nonlinear dependencies or random noise, providing a clean input for subsequent nonlinear analysis. For example, a linear regression method based on a vector autoregression (VAR) model can be used, with the first node's time series as the dependent variable and its own historical values ​​and the historical values ​​of the second node's time series as independent variables. The difference between the predicted and actual values ​​is the linear prediction residual sequence. Alternatively, an autoregressive moving average (ARIMA-X) model with exogenous variables can be used, with the second node's time series as exogenous input, to linearly fit the first node's time series; the model residuals are the desired linear prediction residual sequence.

[0087] Based on the linear prediction residual sequence and the time series of the second node, the residual nonlinear information explained by the time series of the second node in the linear prediction residual sequence is determined as the nonlinearity test statistic. The purpose of this step is to accurately quantify the nonlinear causal contribution of the second node time series to the first node time series after excluding linear effects. By analyzing the nonlinear portion of the residual sequence that can still be explained by the second node time series, the interference of linear components on the nonlinearity test is avoided. For example, the conditional mutual information (CMI) between the second node time series and the linear prediction residual sequence can be calculated. This CMI, after controlling for the influence of their respective historical values, can reveal the residual nonlinear information between the two. Another approach is to use a kernel-based nonlinear Granger causality test method, mapping the linear prediction residual sequence and the second node time series to a high-dimensional feature space, and performing a linearity test in this space to indirectly capture the nonlinear relationship. The test statistic is the residual nonlinear information.

[0088] Based on the goodness of fit of the linear prediction model, the linearity test statistic is determined. This step directly uses the performance index of the linear prediction model to measure the strength of the linear causal relationship. The higher the goodness of fit of the model, the stronger the linear predictive ability of the second node time series on the first node time series, thus directly reflecting the significance of the linear causal association. For example, the coefficient of determination (R-squared) can be used as the linearity test statistic; a higher value indicates a stronger ability of the model to explain the variance of the first node time series, and a more significant linear causal relationship. Alternatively, within the framework of Granger causality testing, the F-statistic can be used directly. This statistic assesses the improvement in the linear predictive ability of the second node time series on the first node time series by comparing the sum of squared residuals between prediction models that include and do not include the second node time series, thus serving as the linearity test statistic.

[0089] Through the above technical solution, when performing causal testing on node pairs whose modal combinations consist of numerical indices and numerical indices, the time series of the first node is used to perform linear prediction modeling with the time series of the second node to obtain a linear prediction residual sequence. This step pre-separates the portion of the first node's time series that can be linearly explained by the second node's time series, ensuring that the residual sequence retains only information that cannot be linearly explained. Based on this linear prediction residual sequence and the time series of the second node, the residual amount of nonlinear information explained by the second node in the residual is determined as the nonlinear test statistic, thus ensuring that the nonlinear test statistic is not disturbed by linear causal information and can accurately reflect the true degree of nonlinear causal association between node pairs. Directly determining the linear test statistic based on the goodness of fit of the linear prediction model simplifies the calculation process of the statistic and ensures that its value directly corresponds to the degree of linear causal explanation, thereby accurately reflecting the true strength of the linear causal association between node pairs. Through this step-by-step processing method, this solution can avoid mutual interference between linear and nonlinear causal information in the calculation process of the statistic, ensuring the accuracy of the linear and nonlinear test statistics. This provides a reliable foundation for accurately labeling a strong causal relationship when both the linear and nonlinear test statistics satisfy the linear significance condition, a weak linear causal relationship when only the linear test statistic satisfies the linear significance condition, and a nonlinear causal relationship when only the nonlinear test statistic satisfies the nonlinear significance condition. This helps improve the accuracy of causal hypergraph updates and enhances the reliability of topology metrics, thereby optimizing the incremental monitoring of distributed cluster states.

[0090] This application further proposes a method for linear prediction residual sequence by using the time series of the first node in the node pair and the time series of the second node to perform linear prediction modeling. The specific steps include: Using the historical values ​​of the time series of the first node as the predictor variable, an autoregressive fit is performed on the current values ​​of the time series of the first node to obtain the first fitting residual.

[0091] In this autoregressive fitting, the historical value of the time series of the second node is introduced as an additional predictor variable, and the current value of the time series of the first node is subjected to joint regression fitting to obtain the second fitting residual.

[0092] The first fitting residual and the second fitting residual are differencing each other along the time axis to obtain the linear prediction residual sequence.

[0093] For example, in the step of using the historical values ​​of the first node's time series as predictor variables and performing an autoregressive fit on the current value of the first node's time series to obtain the first fitting residual, the aim is to model the inherent time dependence of the first node's time series itself. By utilizing its own historical values, a baseline prediction is established for the current value, thereby capturing the autocorrelation and autoregressive characteristics of the series. This first fitting residual represents the portion of the current value that cannot be explained by its own history, providing a reference point for subsequent consideration of external influences. For example, the autoregressive (AR) component of the Autoregressive Moving Average (ARIMA) model can be used, and the model parameters can be determined through least squares or maximum likelihood estimation to fit the relationship between the historical values ​​and the current value of the first node's time series. Alternatively, a simple linear regression model can be constructed, using the values ​​of the first node's time series at the past p times as independent variables and the current value as the dependent variable to fit the predicted value, and then the residual can be calculated.

[0094] In the autoregressive fitting step, the historical values ​​of the second node's time series are introduced as additional predictor variables. A joint regression fitting is then performed on the current value of the first node's time series to obtain the second fitting residual. This step further incorporates historical information from the second node's time series into the aforementioned self-predictive model. By performing the joint regression fitting, the amount of variance that can be additionally explained in the current value of the first node when considering both the past values ​​of the second node and the past values ​​of the first node itself is evaluated. This second fitting residual reflects the portion that remains unexplained after considering both the history of the first node and the history of the second node. For example, based on the first step's autoregressive model, the values ​​of the past q times of the second node's time series are added as additional independent variables to the regression model, predicting the current value of the first node's time series together with the historical values ​​of the first node's time series. Alternatively, the Granger causality test model can be used, which determines causality by comparing the predictive performance of autoregressive models that include and do not include external variables (i.e., the historical values ​​of the second node). Its internal mechanism is highly consistent with the joint regression fitting in this step.

[0095] The first and second fitted residuals are differencing each other along the time axis to obtain the linear prediction residual sequence. This step is crucial for separating the specific influence of the second node on the first node. By subtracting the second fitted residual, which includes the influence of the second node, from the first fitted residual, which only includes the influence of the first node itself, the resulting linear prediction residual sequence accurately represents the portion of the first node's current value that is linearly explained only by the historical values ​​of the second node, after excluding the autoregressive effect of the first node itself. This pure residual sequence is essential for accurate causal inference. For example, by using point-by-point subtraction, for each time point t in the time series, the first fitted residual (t) is subtracted from the second fitted residual (t) to obtain the linear prediction residual for that time point. Alternatively, if the residual sequence is stored in vector or matrix form, vector or matrix subtraction can be performed directly to achieve efficient differencing.

[0096] By employing the aforementioned step-by-step fitting and differencing technique, the additional residuals introduced by the second node's time series in predicting the first node can be accurately separated. For example, by performing autoregressive fitting on the first node's time series, a first fitting residual is obtained. This step eliminates the interference of the first node's own temporal correlation on subsequent residual calculations, ensuring an accurate assessment of the first node's predictive ability. Introducing historical values ​​of the second node's time series as additional predictor variables into the autoregressive fitting for joint regression fitting yields a second fitting residual. This step, while preserving the historical predictive basis of the first node, accurately reflects the change in prediction accuracy after incorporating the second node's information. Dividing the first and second fitting residuals along the corresponding time axis positions results in a linear prediction residual sequence that only includes residual changes contributed by the second node's time series, thus completely eliminating the interference of the first node's own temporal correlation. This precise residual separation mechanism provides a reliable foundation for subsequent calculation of accurate linear and nonlinear test statistics, significantly improving the accuracy of causal relationship judgment. Especially in the incremental monitoring of distributed cluster status, it can more accurately identify the causal dependencies between event streams of different modal states, avoid misjudgments caused by inaccurate residuals, and thus optimize the generation of monitoring scheduling signals and channel start / stop instructions, improving the response efficiency and resource allocation rationality of the entire monitoring system.

[0097] In some of the solutions described above in this application, incremental causality testing and topological measurement are proposed for time-series signals carrying modal source labels and causal hypergraphs to obtain updated causal hypergraphs, causal type labels, and topological change measures. This requires updating the structure and attributes of local subgraphs and adjusting the causal edge confidence accumulation rate to obtain the updated causal hypergraph. However, in this process, the original methods do not differentiate between newly discovered causal relationships and existing causal relationships, nor do they adjust the causal edge confidence accumulation rate based on the previous round of topological changes. The fixed confidence accumulation step size cannot adapt to different stable states of the topological structure, making it difficult to quickly confirm persistent causal relationships and prone to misconceptions of frequently changing causal relationships, thus affecting the accuracy and efficiency of causal hypergraph updates.

[0098] To address this, this application proposes a method for updating the structure and attributes of a local subgraph based on causal relationships and causal type labels, and adjusting the confidence accumulation rate of causal edges within the local subgraph according to the previous round of topological change metrics to obtain an updated causal hypergraph. Specifically, the method includes: when incremental causal testing discovers a new causal relationship, establishing a corresponding causal edge or hyperedge in the local subgraph and labeling it with a causal type label; when incremental causal testing confirms an existing causal edge in the local subgraph, increasing the confidence of the existing causal edge according to the confidence increment step size corresponding to the causal type label; and determining the confidence accumulation rate adjustment coefficient for each causal edge within the local subgraph based on the previous round of topological change metrics, and correcting the current round of confidence increment step size with the confidence accumulation rate adjustment coefficient to obtain the updated causal hypergraph.

[0099] For example, when an incremental causal test discovers a new causal relationship, a corresponding causal edge or hyperedge is created in the local subgraph and labeled with a causal type tag. This step aims to add the newly identified causal dependencies to the causal hypergraph in the form of a graphical structure. For instance, when an incremental causal test (e.g., through Granger causality tests, mutual information tests, or event-driven causality tests based on stochastic process models) identifies a previously unrecorded causal relationship between two nodes within the local subgraph, the system creates a new directed edge between these two nodes, representing a unidirectional causal relationship. If the causal relationship involves multiple source or target nodes, a hyperedge is created. Based on the results of the causal test (e.g., linear, nonlinear, event-driven, etc.), the newly created causal edge or hyperedge is labeled with the corresponding causal type for subsequent processing and analysis. Another implementation is that when the statistic or confidence level of the incremental causal test first reaches a significance threshold, indicating the existence of a new causal association, the system dynamically instantiates a causal edge object or hyperedge object in the in-memory representation of the local subgraph and associates it with the relevant source and target nodes. The object will contain an attribute field to store causal type markers, such as "strong causation", "weak linear causation", "non-linear causation" or "event-driven causation", which are derived directly from the output of incremental causality tests.

[0100] When incremental causality testing confirms an existing causal edge in a local subgraph, the confidence of the existing causal edge is increased according to the confidence increment step size corresponding to the causality type label. This step is used to strengthen the confidence of existing but re-validated causal relationships, reflecting their continued validity. For example, when incremental causality testing re-validates an existing causal edge in a local subgraph (i.e., the causal edge was established in a previous round), the system looks up a pre-configured confidence increment step size table based on the causality type label stored on the causal edge. Different causal types may have different confidence accumulation characteristics; for example, linear causal relationships may use a smaller step size, while nonlinear or event-driven causal relationships may use a larger step size to confirm their existence more quickly. The found step size value is added to the current confidence value of the causal edge. Another implementation is that, for existing causal edges that are re-validated, the system calls the corresponding confidence update function according to its causality type label. For example, for "strong causal relationships," a fixed, larger increment step size may be used. For "weak linear causal relationships," a smaller increment step size may be used. For "non-linear causal relationships," the increment step size can be dynamically adjusted based on the strength of the non-linear test statistic. This approach allows for more precise control over the rate of confidence accumulation for different types of causal relationships, ensuring the rationality of confidence updates.

[0101] Based on the topology change metric from the previous round, the system determines the confidence accumulation rate adjustment coefficient for each causal edge within the local subgraph, and uses this adjustment coefficient to correct the confidence increment step size for the current round, resulting in the updated causal hypergraph. This step introduces awareness of overall topology changes in the causal hypergraph to dynamically adjust the accumulation rate of causal edge confidence, thereby improving the adaptability and accuracy of causal graph updates. For example, the system obtains the topology change metric from the previous round of incremental causal testing. This metric reflects the severity of the causal hypergraph's structural changes. When the topology change metric is high, it indicates that the cluster state may be undergoing drastic changes. In this case, the system determines a confidence accumulation rate adjustment coefficient greater than 1 to accelerate the confirmation or decay of causal relationships. Conversely, when the topology change metric is low, it indicates that the cluster state is relatively stable. The system determines an adjustment coefficient close to or less than 1 to maintain or slow down the confidence accumulation rate. This adjustment coefficient is then multiplied by the confidence increment step size calculated in the current round to form a corrected increment step size, used to update the confidence of causal edges. Another implementation involves retrieving the corresponding confidence accumulation rate adjustment coefficient from a lookup table based on the interval in which the previous round of topological change measurement occurred. For example, the adjustment coefficient might be 1 in a stationary interval, 1.2 in an interval of interest, and 1.5 in an interval of significant change. This adjustment coefficient is then used to multiply and correct the baseline confidence increment step size corresponding to the causal type label, thereby obtaining the step size used to update the confidence of causal edges.

[0102] By employing the aforementioned technical solution, and by distinguishing between newly discovered and confirmed causal relationships, and dynamically adjusting the confidence accumulation rate in conjunction with the previous round of topology changes, the inflexible confidence adjustment and insufficient update accuracy during the causal hypergraph update process of the original method are resolved. When incremental causal testing discovers a new causal relationship, a corresponding causal edge or hyperedge is established in the local subgraph and labeled with the causal type. This allows the causal hypergraph to promptly capture new changes in the cluster state and retain the type information of the causal relationship, providing a more refined basis for subsequent monitoring and scheduling signal generation and channel start / stop commands. When incremental causal testing confirms an existing causal edge in the local subgraph, the confidence of the existing causal edge is increased according to the confidence increment step size corresponding to the causal type label. This differentiated confidence accumulation method allows different types of causal relationships to be confirmed at a speed more consistent with their characteristics, avoiding the inaccuracies caused by a uniform fixed step size and improving the accuracy of the causal hypergraph. Furthermore, based on the previous round of topology change measurement, the confidence accumulation rate adjustment coefficient of each causal edge within the local subgraph is determined, and this coefficient is used to correct the confidence increment step size in the current round. This mechanism enables the causal hypergraph update process to adaptively respond to dynamic changes in the cluster environment. When the cluster topology changes drastically, the confirmation or decay of causal relationships is accelerated, thereby adapting to new system behavior patterns more quickly. When the cluster topology is stable, a reasonable update speed can be maintained, avoiding over-response to noise. This significantly improves the causal hypergraph's ability to perceive changes in the state of the distributed cluster and its update efficiency, ensuring that the causal hypergraph can more accurately and timely reflect the true causal dependencies, providing a reliable foundation for subsequent monitoring, scheduling, and fault diagnosis. Overall, the scheme in this application makes the causal hypergraph update process more intelligent and adaptive, capable of coping with the complexity and dynamism of the distributed cluster environment, improving the accuracy and efficiency of causal discovery, and thus providing stronger support for the operation and maintenance management of distributed clusters.

[0103] This application further proposes adjusting the confidence increment step size in the current round using a confidence accumulation rate adjustment coefficient to obtain an updated causal hypergraph. This includes: when the previous round's topology change metric indicates that the causal hypergraph topology is in a stable state, updating the confidence of the existing causal edge using a baseline confidence increment step size. When the previous round's topology change metric indicates that the causal hypergraph topology is in a changing state, expanding the baseline confidence increment step size using the confidence accumulation rate adjustment coefficient, and updating the confidence of the existing causal edge with the expanded confidence increment step size. When the existing causal edge is a hyperedge and the confidence of the hyperedge exceeds the hyperedge confirmation threshold, performing time decay acceleration processing on the ordinary directed edges within the hyperedge.

[0104] For example, when the topology of the system's causal hypergraph is in a stable state—that is, when the previous round's topology change metric indicates a low or no significant change—this method uses a baseline confidence increment step size to update the confidence of existing causal edges. This baseline step size is typically a small, fixed value, such as 0.01 or 0.05, designed to gradually strengthen the confirmation of stable causal relationships in a conservative and smooth manner. Alternatively, the baseline step size can be determined proportionally to the difference between the current confidence and the maximum confidence of the causal edge; for example, it can be set as a fixed percentage of the remaining confidence space to ensure that the confidence smoothly approaches the upper limit.

[0105] When the topology of the system's causal hypergraph is in a changing state—that is, when the previous round's topology change metric indicates a significant change—this method uses a confidence accumulation rate adjustment coefficient to amplify the baseline confidence increment step size. This adjustment coefficient is typically calculated dynamically or obtained from a table based on the magnitude of the topology change metric. For example, when the topology change metric exceeds a certain threshold, the adjustment coefficient can be set to 1.5; when the change metric further increases, the adjustment coefficient can be set to 2.0 or higher. In this way, the amplified confidence increment step size accelerates the confirmation process of newly emerging or significantly enhanced causal relationships, enabling the causal hypergraph to adapt to and reflect changes in the system state more quickly.

[0106] When an existing causal edge in a causal hypergraph is identified as a hyperedge, and its confidence level reaches or exceeds the hyperedge confirmation threshold, it indicates that the higher-order causal relationship has been sufficiently verified. At this point, this method performs accelerated time decay processing on the ordinary directed edges that constitute the hyperedge. For example, it significantly increases the confidence decay rate of these ordinary directed edges, such as increasing their decay coefficient from the usual 0.005 to 0.05, causing their confidence to decrease rapidly. Another approach is to directly set the confidence level of these ordinary directed edges to a lower preset value, or even remove them directly from the causal hypergraph, to prevent them from continuing to interfere with the overall structure of the causal hypergraph and the expression of higher-order causal relationships.

[0107] The above technical solution adaptively adjusts the update strategy of causal edge confidence based on the changing state of the causal hypergraph topology, thereby improving the accuracy and response speed of the causal hypergraph. For example, when the causal hypergraph topology is stable, updating the confidence of existing causal edges using a baseline confidence increment step size avoids unnecessary structural fluctuations caused by overly rapid updates, while saving computational resources. When the topology is changing, the baseline confidence increment step size is expanded by adjusting the confidence accumulation rate coefficient, enabling the causal hypergraph to confirm new or enhanced causal relationships more quickly, accelerating the convergence of the causal hypergraph to the true system state, and improving the ability to perceive dynamic changes in the system. Furthermore, after a hyperedge is confirmed, time decay acceleration processing is performed on ordinary directed edges within the hyperedge, reducing the interference of low-value ordinary edges on the accuracy of the causal hypergraph structure, making the complex causal relationships represented by confirmed higher-order hyperedges more clearly prominent, thereby improving the accuracy and interpretability of the causal hypergraph in expressing causal dependencies. This scenario-specific confidence update mechanism, along with the optimized processing of ordinary edges within the hyperedge, together ensures that the causal hypergraph maintains an efficient, accurate, and hierarchical representation of causal relationships under different system states.

[0108] This application further proposes a method for constructing a filtering sequence based on the causal strength of each causal edge in the updated causal hypergraph, the method comprising: The causal strength values ​​of all causal edges in the updated causal hypergraph are extracted and sorted from highest to lowest to obtain a causal strength ranking sequence. This step aims to comprehensively acquire and organize the strength information of all causal relationships in the current causal hypergraph. Causal strength is a quantitative indicator measuring the strength of a causal relationship, calculated using various methods such as the significance level of statistical tests, the magnitude of the causal effect, or confidence levels. For example, for linear causal relationships, the causal strength can be the absolute value of the F-statistic or partial correlation coefficient of the Granger causality test. For nonlinear causal relationships, the strength is calculated based on information theory (such as transitive entropy) or kernel methods (such as kernel Granger causality). By sorting from high to low, the importance distribution of causal relationships is clearly displayed, providing a foundation for subsequent threshold setting and hierarchical filtering. In practice, all causal edges in the updated causal hypergraph are traversed, and the causal strength attribute value associated with each causal edge is obtained. Then, standard sorting algorithms (such as quicksort and mergesort) are used to sort these strength values ​​in descending order to generate a causal strength ranking sequence. Alternatively, during the causality verification phase, the causal strength values ​​can be directly stored in the causal edge attributes, and after each update of the causal hypergraph, all causal strength values ​​can be extracted and sorted through database queries or graph traversal operations.

[0109] Within the range of values ​​in the causal strength ranking sequence, a series of strength thresholds is generated from high to low with a preset step size. The purpose of this step is to systematically generate a series of thresholds for hierarchical filtering within the known distribution range of causal strength values. The preset step size determines the granularity of the threshold sequence; a smaller step size generates finer filtering levels, while a larger step size provides a coarser hierarchical division. Generating the strength threshold sequence ensures that subsequent filtering processes can cover the entire range from the strongest to the weakest causal relationship, thus comprehensively examining the topology at different strength levels. Specifically, the maximum and minimum values ​​in the causal strength ranking sequence are determined as the upper and lower limits of the value range. Starting from the maximum value, the thresholds are decreased with a fixed preset step size until they reach or fall below the minimum value, generating a series of strength thresholds. Alternatively, the step size can be dynamically determined based on the statistical distribution of causal strength values ​​(e.g., quantiles, standard deviation), or a non-uniform step size can be used, such as a smaller step size in densely distributed areas and a larger step size in sparse areas, to better capture key structural change points.

[0110] The updated causal hypergraph is sequentially edge-filtered using each intensity threshold in the intensity threshold sequence, retaining causal edges with a causal strength not lower than the current threshold. This results in subgraphs corresponding to each intensity threshold, which form the filtering sequence in descending order of the intensity threshold sequence. This step is the core of constructing the filtering sequence. By applying a series of decreasing intensity thresholds, a set of subgraphs reflecting different levels of causal strength is gradually built. Edge filtering means that for each threshold, only those causal edges with a causal strength reaching or exceeding that threshold are retained, thus forming a subgraph with a more "pure" or "significant" causal relationship. These subgraphs are arranged in descending order of intensity thresholds, forming an ordered sequence from sparse (containing only the strongest causal edges) to dense (containing all causal edges), providing a multi-scale, multi-level perspective for subsequent topological structure analysis. In specific implementation, the generated intensity threshold sequence is traversed. For each threshold in the sequence, an updated causal hypergraph is copied, and then all causal edges in the copy are traversed, removing all causal edges with a causal strength lower than the current threshold, thus obtaining a subgraph. These subgraphs are stored in order according to the threshold sequence, forming the filtering sequence. Alternatively, more efficient graph data structure operations can be used. For example, in a graph database, subgraphs that meet the filtering conditions can be directly queried and generated by setting edge weights. Alternatively, in memory, subgraph views under different thresholds can be quickly generated by marking or copying, and then organized into a filtering sequence.

[0111] The above technical solution addresses the problem of the lack of standardized methods for constructing filter sequences, which fails to provide a reliable basis for measuring topological changes. It constructs a filter sequence that conforms to the causal strength distribution through hierarchical ordered filtering, providing a clear and structured foundation for accurate measurement of topological changes. For example, it collects and sorts the strength information of all causal edges in the updated causal hypergraph, clarifying the overall distribution range of causal strength. This provides a complete data foundation for generating strength thresholds, ensuring that the generated thresholds cover all causal strength ranges. Generating ordered thresholds based on the sorted causal strength ranges ensures that the strength thresholds uniformly cover the entire causal strength range, avoiding threshold omissions or uneven density, and providing a unified and stable standard for subsequent hierarchical filtering. Filtering according to strength from high to low yields a series of ordered subgraphs, from retaining only high-intensity, reliable causal edges to retaining all causal edges. This clearly shows the topological changes of the causal hypergraph at different confidence levels, providing a hierarchical and structured foundation for comparing topological differences and obtaining accurate measurements of topological changes. It can distinguish the impact of different significant causal edges on topological changes, helping to more accurately capture the true changes in causal topology. This enables a more comprehensive and accurate assessment of the dynamic evolution of causal relationships from a structured, multi-layered perspective when measuring subsequent changes in the topology of the updated causal hypergraph, thereby improving the granularity of distributed cluster state monitoring.

[0112] This application further proposes a method for measuring the topological changes of the updated causal hypergraph on the filtering sequence, specifically including: extracting the number of connected components and the number of loop structures for each subgraph in the filtering sequence to obtain a topological feature vector corresponding to each intensity threshold; tracking the occurrence and disappearance times of the number of connected components and the number of loop structures on the filtering sequence to generate a causal topological persistence graph; and comparing the difference between the causal topological persistence graph generated in this round and the causal topological persistence graph generated in the previous round to obtain the topological change measure.

[0113] For example, in the step of extracting the number of connected components and the number of loop structures for each subgraph in the filtering sequence to obtain the topological feature vector corresponding to each intensity threshold, the number of connected components refers to the number of sets of mutually connected nodes in the subgraph. This can be extracted using graph traversal algorithms, such as breadth-first search (BFS) or depth-first search (DFS), traversing all nodes in the subgraph and recording the number of independent connected regions found during the traversal. Alternatively, connected components can be identified by analyzing the adjacency matrix of the subgraph and using matrix operations. The number of loop structures refers to the number of cyclic paths existing in the subgraph. This can be extracted using strong connected component algorithms, such as Tarjan's algorithm or Kosaraju's algorithm, which identifies strongly connected components in a directed graph. Each strongly connected component containing more than one node indicates the existence of a loop. Alternatively, loops can be identified and counted by detecting back edges during DFS traversal. Combining the extracted number of connected components and the number of loop structures forms the topological feature vector of the subgraph at a specific intensity threshold.

[0114] In the step of generating a causal topological persistence graph by tracking the occurrence and disappearance times of the number of connected components and the number of loop structures in the filtering sequence, the occurrence and disappearance times refer to the intensity thresholds of the first and second occurrences of a specific number of connected components or loop structures in the filtering sequence (i.e., the subgraph sequence arranged in descending order of causal strength). Tracking these times constructs a causal topological persistence graph, which can be represented as a two-dimensional graph. One dimension represents the causal strength threshold, and the other dimension represents the topological feature value (such as the number of connected components or the number of loop structures). Line segments or regions are used to visually represent the persistence range of a specific topological feature under different causal strengths. Alternatively, a data structure, such as a list or hash table, can be used to record each topological feature (e.g., "number of connected components is X" or "number of loop structures is Y") and its corresponding occurrence and disappearance intensity thresholds.

[0115] In the step of comparing the causal hypergraph generated in the current round with that generated in the previous round to obtain a measure of topological change, this comparison aims to quantify the degree of change in the overall topological structure of the causal hypergraph across different monitoring rounds. This comparison can be achieved by calculating a distance metric between the two causal hypergraphs. For example, using the idea of ​​edit distance, the minimum number of operations required to transform one graph into another (such as adding, deleting, or modifying feature persistence intervals) can be calculated. Alternatively, the change can be quantified by comparing the overlap, area of ​​difference, or Jaccard similarity coefficient of persistence intervals (occurrence and disappearance times) of the same topological features in the two graphs. For example, newly added persistence intervals, disappeared persistence intervals, and the degree of change in persistence intervals can be calculated. The resulting measure of topological change is a quantitative value that reflects the degree of change in the causal hypergraph topological structure between the current and previous rounds.

[0116] The above technical solution enables accurate measurement of the topological changes in the updated causal hypergraph. For example, by extracting the number of connected components and loop structures from each subgraph in the filtering sequence, the connectivity aggregation characteristics and cyclic dependency changes of the causal structure can be captured from the topological essence, overcoming the shortcomings of the original method which only counted the increase or decrease of a single edge, thus providing a more comprehensive reflection of the stability of the causal structure. By tracking the appearance and disappearance of these topological features in the filtering sequence, a causal topological persistence graph is generated, which can distinguish between stable strong causal structures and occasional weak causal structures, avoiding misjudgments of weak temporary structural changes and significantly improving the accuracy of change measurement. By comparing the causal topological persistence graph generated in this round with that generated in the previous round, the changes in the overall topological structure of the causal hypergraph can be accurately quantified, rather than just changes in a single local edge, thus accurately reflecting the overall stability of the causal structure and providing an accurate and reliable decision-making basis for subsequent monitoring and scheduling strategy adjustments.

[0117] This application further proposes a method for generating monitoring and scheduling signals and channel start / stop commands, see [link to relevant documentation]. Figure 4 Specifically, it includes: 401. Compare the topological change measure with the stationary threshold and the significant change threshold to determine the change range of the current causal hypergraph.

[0118] 402. Based on this variation range, generate a sampling frequency adjustment signal and an incremental reporting content selection signal. The sampling frequency adjustment signal is used to control the sampling interval of the multimodal state event stream, and the incremental reporting content selection signal is used to determine the incremental causal subgraph or the complete causal subgraph to be reported.

[0119] 403. Identify cross-modal causal relationships based on the causal type label, determine the modal channel to which the node involved in the cross-modal causal relationship belongs using the modal source label, and generate start / stop instructions for the modal channel.

[0120] The topology change metric is a quantitative indicator used to reflect the degree to which the structure of the causal hypergraph of a distributed cluster changes over time. This metric can be calculated based on methods such as graph edit distance, graph analysis, or persistent cohomology, and its role is to provide a macroscopic assessment of the stability of dependencies within the cluster, serving as a key basis for determining whether the cluster state has undergone significant changes. The stationarity threshold and the significant change threshold are two pre-defined critical values ​​used to classify the degree of change in the causal hypergraph topology. These thresholds can be set based on historical data analysis, expert experience, or dynamically optimized through machine learning models, providing clear criteria for classifying the cluster state. The change interval is a classification of the current state of the causal hypergraph based on the comparison between the topology change metric and the pre-defined thresholds. For example, it can be divided into "stable interval," "interval of concern," or "abnormal interval." This interval classification provides a high-level decision-making basis for subsequent differentiated monitoring strategy adjustments. The sampling frequency adjustment signal is an instruction used to instruct adjustments to the sampling frequency of each modality of data in the multimodal state event stream. This signal can be a relative adjustment (e.g., increasing by 20% or decreasing by 50%) or an absolute target sampling interval (e.g., sampling once per second or once every five seconds). Its function is to dynamically balance the granularity of data acquisition with system resource consumption. The incremental reporting content selection signal is an instruction used to determine the scope and granularity of the causal hypergraph information to be reported under different cluster states. This signal can instruct the reporting of only topology change metrics, incremental causal subgraphs, or complete causal subgraphs, optimizing data transmission bandwidth and storage resources and avoiding unnecessary data redundancy. The causal type label is an attribute attached to a causal relationship, describing the specific nature of the causal relationship, such as linear causality, nonlinear causality, or event-driven causality. This label provides richer information than simple causal existence, helping to understand causal mechanisms more accurately. Cross-modal causality refers to the causal dependencies between different modalities of data in a distributed cluster, such as changes in numerical indicators leading to the generation of text log events. Identifying such relationships is crucial for understanding the behavior of complex systems. The modality source tag is an identifier attached to each time-series signal or causal hypergraph node, indicating the modality from which the data or node originates (e.g., continuous numerical indicators, discrete text log events, or asynchronous call chain tracing events). This tag is fundamental for cross-modal analysis and control. The modality channel refers to the logical or physical path used to acquire, transmit, and process data of a specific modality. For example, there are dedicated indicator acquisition channels, log acquisition channels, and call chain tracing channels. The channel start / stop command is a control command used to dynamically enable or disable a specific modality channel. This command allows the system to enable or disable data acquisition channels as needed based on actual monitoring requirements, thereby achieving fine-grained resource management.

[0121] The above technical solution enables the output of monitoring and control commands from three dimensions—collection frequency, reported content, and collection channels—based on different state changes in distributed clusters and information on causal structure changes and cross-modal causal relationships. This allows for the reasonable matching of monitoring resource allocation and the adaptation of monitoring configuration to the actual state of the cluster. For example, the topology change metric is compared with a stability threshold and a significant change threshold to define the change range of the current causal hypergraph. Since the topology change metric accurately reflects the causal structure, i.e., the overall degree of change in cluster dependencies, dividing the range based on the degree of topology change combined with preset thresholds ensures that the range division matches the actual state changes of the cluster, avoiding subjective and blind state classification and providing an accurate basis for subsequent differentiated control. Based on the defined change range, a collection frequency adjustment signal and an incremental reporting content selection signal are generated. Given that different change ranges correspond to different monitoring needs, and more significant state changes require higher monitoring accuracy, generating corresponding adjustment commands based on the change range can reduce resource consumption when the state is stable and ensure sufficient monitoring accuracy when the state changes significantly, thus balancing monitoring resource consumption and monitoring effectiveness. Determining the reporting content based on the change range avoids reporting unnecessary full data, reducing data transmission and storage resource consumption. Cross-modal causal relationships are identified based on causal type tags, and the modal channel to which the relevant node belongs is determined by modal source tags, thereby generating start / stop commands for the corresponding channel. By utilizing causal type tags, cross-modal causal relationships related to the current state change can be accurately located. Furthermore, modal source tags accurately pinpoint the acquisition channels that need adjustment, enabling only channels related to the current change and disabling unnecessary channels. This not only ensures sufficient data for cross-modal causal analysis but also avoids resource waste caused by continuous acquisition of all channels, achieving on-demand configuration of acquisition channels. Overall, this application significantly improves the adaptability and resource utilization efficiency of the distributed cluster monitoring system through multi-dimensional and refined monitoring scheduling, ensuring appropriate monitoring granularity and data support under different operating conditions.

[0122] This application further proposes comparing the topological change metric with a stationarity threshold and a significant change threshold to determine the change interval of the current causal hypergraph. Specifically, this includes: when the topological change metric is lower than the stationarity threshold, the change interval is determined to be a stationary interval; when the topological change metric is not lower than the stationarity threshold and is lower than the significant change threshold, the change interval is determined to be an interval of interest; and when the topological change metric is not lower than the significant change threshold, the change interval is determined to be a significant change interval.

[0123] The topology change metric is an indicator that quantifies the evolution of the causal hypergraph structure of a distributed cluster over time. It reflects the changes in topological attributes such as nodes, edges, connectivity, and loops in the causal network and is a key basis for assessing the stability of the cluster state. This metric can be calculated by comparing the differences between the current causal hypergraph and the previous causal hypergraph in specific topological features (such as the number of connected components, the number of loops, and edge weight distribution). For example, it can be calculated by calculating the graph edit distance between two hypergraphs or a similarity metric based on graph embedding. Alternatively, it can be obtained by tracking the appearance and disappearance of the number of connected components and loop structures in the causal hypergraph on the filtering sequence, constructing a causal topological persistence graph, and comparing the differences between the causal topological persistence graphs in different rounds.

[0124] This stability threshold is a pre-defined value used to define the "stationary" state of changes in the causal hypergraph topology. When the measure of topology change is below this threshold, the clustered causal network is considered to be in a relatively stable state. This threshold serves as a benchmark for determining whether the cluster state requires a reduction in monitoring intensity. It is determined through historical data analysis, expert experience, or statistical methods (such as control charts). For example, it can be set as the 95th percentile of the topology change measure during historical normal operation. Alternatively, it can be determined through simulation experiments, running the system under known stable conditions and recording the maximum or average value of the topology change measure, then multiplying this value by a safety factor.

[0125] This significant change threshold is another pre-set value, higher than the stability threshold, used to define a "significant" state of change in the causal hypergraph topology. When the topology change metric reaches or exceeds this threshold, a significant change is considered to have occurred in the cluster's causal network, potentially indicating a potential fault or anomaly. This threshold serves as a benchmark for determining whether the cluster status requires increased monitoring intensity or in-depth analysis. It can also be determined through historical fault data analysis, anomaly detection algorithms, or expert experience; for example, it can be set as the average of the topology change metrics before historical faults or a high percentile. Alternatively, a fault can be injected into a controlled environment, the response of the topology change metric can be observed, and a critical value that can distinguish between normal and abnormal states can be selected.

[0126] Comparing the topology change metric with the stability threshold and the significant change threshold involves numerically comparing the currently calculated topology change metric with these two preset thresholds. Through this comparison, the system can quantify the degree of causal topology change in the current cluster and categorize it into a predefined state interval. This can be achieved by directly using conditional statements (such as if-else if-else structures) for numerical comparison, or by mapping the topology change metric to a continuous interval and then determining its corresponding interval using a lookup table or piecewise function.

[0127] Based on the comparison results, the current change range of the causal hypergraph is determined, which is the process of classifying the current state of the causal hypergraph into a "stable range," a "range of concern," or a "range of significant change." This provides a clear and hierarchical decision-making basis for subsequent monitoring and scheduling signal generation and channel start / stop commands. When the topology change metric is below the stability threshold, the system marks the current state as a stable range, indicating that the cluster causal relationship network is in a stable state, and at this time, reducing monitoring resource consumption can be considered. When the topology change metric is not lower than the stability threshold but is lower than the significant change threshold, the system marks the current state as a range of concern, indicating that the cluster causal relationship network has undergone a certain degree of change and requires continuous monitoring, but has not yet reached the level of emergency handling. When the topology change metric is not lower than the significant change threshold, the system marks the current state as a significant change range, indicating that the cluster causal relationship network has undergone significant changes, and there may be a fault or anomaly, requiring more proactive monitoring and response measures.

[0128] The above technical solution enables precise hierarchical classification of topology change metrics, providing accurate classification criteria for generating differentiated monitoring and scheduling signals. For example, by introducing a stability threshold and a significant change threshold, the degree of topology change in the causal hypergraph is divided into three levels: stable interval, interval of interest, and interval of significant change. This multi-level classification, compared to a single threshold judgment, more precisely reflects the dynamic changes in the causal topology of the distributed cluster, avoiding potential misjudgments or resource waste caused by rigid black-and-white judgments. When the topology change metric is in the stable interval, it indicates that the cluster causal network is stable, and the system can reduce the sampling frequency and reporting content accordingly, saving monitoring resources. When the topology change metric is in the interval of interest, it indicates that the cluster causal network has changed to a certain extent, and the system can maintain the basic sampling frequency and report incremental causal subgraphs to ensure continuous tracking of potential changes. When the topology change metric is in the interval of significant change, it indicates that the cluster causal network has undergone significant changes, and the system can quickly increase the sampling frequency and report the complete causal subgraph and full state data of the affected area to support rapid fault location and root cause analysis. Therefore, this application improves the adaptability and resource utilization efficiency of the monitoring system by refining the metric for topology changes to match the actual causal structure of the cluster more accurately.

[0129] This application further proposes a specific method for generating a sampling frequency adjustment signal and an incremental reporting content selection signal based on a change interval. The sampling frequency adjustment signal controls the sampling interval of the multimodal state event stream, and the incremental reporting content selection signal determines whether to report an incremental causal subgraph or a complete causal subgraph. For example, when the change interval is a stable interval, a first sampling frequency adjustment signal and a first incremental reporting content selection signal are generated. The first sampling frequency adjustment signal indicates a reduction in the sampling frequency of the multimodal state event stream, and the first incremental reporting content selection signal indicates that only topology change metric identifiers are reported. When the change interval is a region of interest, a second sampling frequency adjustment signal and a second incremental reporting content selection signal are generated. The second sampling frequency adjustment signal indicates maintaining the base sampling frequency, and the second incremental reporting content selection signal indicates reporting an incremental causal subgraph. When the change range is a significant change range, a third acquisition frequency adjustment signal and a third incremental reporting content selection signal are generated. The third acquisition frequency adjustment signal is used to indicate that the sampling frequency of the multimodal state event stream is increased, and the third incremental reporting content selection signal is used to indicate that the full state data of the complete cause-effect subgraph and the affected area are reported.

[0130] The change interval refers to classifying the current state of the causal hypergraph into a stable interval, a region of interest, or a region of significant change based on the comparison between the topology change metric and a preset threshold. A stable interval indicates that the topology change metric of the causal hypergraph is below the stability threshold, suggesting a relatively stable cluster state. The first sampling frequency adjustment signal is used to instruct a reduction in the sampling frequency of the multimodal state event stream. This can be achieved by sending a command to the data acquisition module to extend the current sampling period to a low-frequency sampling period, for example, from once per second to once every five seconds. Alternatively, it can be achieved by adjusting the time window aggregation parameter in the data stream processing pipeline, reducing the number of data points aggregated within the same time window, thereby indirectly reducing the sampling frequency. The first incremental reporting content selection signal is used to instruct the system to report only the topology change metric identifier. This can be achieved by the system sending only the topology change metric value of the current round or a Boolean flag representing a stable state to the upper-layer monitoring system. Alternatively, the system can generate a lightweight data packet containing a timestamp and the topology change metric identifier and push it to a message queue or storage system.

[0131] The attention interval indicates that the topological change metric of the causal hypergraph is not lower than the stability threshold but lower than the significant change threshold, indicating that there is a certain degree of change in the cluster state, which requires attention. The second sampling frequency adjustment signal is used to indicate maintaining the basic sampling frequency. This can be achieved by the system keeping the current sampling frequency of the data acquisition module unchanged, i.e., using the default or previous round's basic sampling frequency. Alternatively, the system can send an instruction to the data acquisition module to explicitly specify the basic sampling period, such as once per second. The second incremental reporting content selection signal is used to indicate reporting incremental causal subgraphs. This can be achieved by the system identifying newly added causal edges, deleted causal edges, and causal edges with changed attributes in the current causality check, and packaging and reporting the subgraph structure formed by these changes and their related attributes. Alternatively, the system can generate a structured data object containing changed nodes, changed edges, and their causal type markers, and transmit it through an API interface or message bus.

[0132] The significant change interval indicates that the topological change metric of the causal hypergraph is not lower than the significant change threshold, suggesting a drastic change in the cluster state and potentially a serious anomaly. The third sampling frequency adjustment signal instructs the system to increase the sampling frequency of the multimodal state event stream. This can be achieved by the system sending a command to the data acquisition module to shorten the current sampling period to a higher frequency, such as from once per second to once every 0.1 seconds. Alternatively, the system can activate additional acquisition agents or parallel data streams to acquire more data samples within the same time period. The third incremental reporting content selection signal instructs the system to report the full state data of the complete causal subgraph and affected areas. This can be achieved by the system packaging and reporting all structural information of the currently updated causal hypergraph (including all nodes, edges, and their attributes) and all original multimodal state data (such as original metrics, logs, and call chain data) directly or indirectly associated with the significantly changed nodes. Alternatively, the system can generate a compressed file containing a serialized representation of the complete causal hypergraph and snapshots of all relevant monitoring data within the affected areas, and upload it to a distributed storage or log analysis platform.

[0133] By employing the aforementioned technical solutions, differentiated data collection and reporting strategies are matched to varying degrees of change in the causal topology, addressing the issues of insufficient differentiation in existing monitoring strategy adjustments, leading to resource waste or inadequate information. When the cluster is stable, reducing the sampling frequency and reporting only topology change metrics significantly reduces unnecessary monitoring resource consumption and data transmission / storage overhead, while still maintaining macroscopic cluster stability. When the cluster enters a period of interest, maintaining the basic sampling frequency and reporting incremental causal subgraphs ensures sufficient detailed information for analysis in the early stages of change without excessive resource consumption, achieving a balance between information acquisition and resource consumption. When the cluster is in a period of significant change, increasing the sampling frequency and reporting the complete causal subgraph and full state data of the affected areas ensures rapid and comprehensive acquisition of all necessary information in the event of severe cluster anomalies, providing strong support for root cause localization and rapid response. This dynamic, on-demand monitoring strategy adjustment ensures that the allocation of monitoring resources remains highly synchronized with actual cluster state changes, significantly improving the efficiency and responsiveness of the monitoring system, avoiding resource waste and information loss, and thus optimizing the operation and maintenance management of distributed clusters.

[0134] This application further proposes a method for identifying cross-modal causal relationships based on causal type tags. It uses modal source tags to determine the modal channel to which the nodes involved in the cross-modal causal relationship belong, and generates start / stop instructions for the modal channel. Specifically, this includes: filtering causal edges marked as event-driven causal relationships or cross-modal joint causal relationships from the causal type tags to obtain a set of cross-modal causal edges; determining the modal channel type of the nodes connected to each causal edge in the set of cross-modal causal edges based on the modal source tags, obtaining a list of modal channels to be enabled; comparing the list of modal channels to be enabled with the list of currently enabled modal channels, and generating enable instructions for newly added modal channels and deactivation instructions for existing modal channels that no longer need monitoring.

[0135] The process involves filtering causal edges marked as event-driven causality or cross-modal joint causality from causal type tags to obtain a cross-modal causal edge set. This aims to focus on causal relationships that truly require cross-modal coordination, thereby excluding intramodal causal relationships and reducing the complexity of subsequent processing. In one implementation, the system can iterate through all identified causal edges and their corresponding causal type tags. If a causal edge marked as "event-driven causality" or "cross-modal joint causality" is found, it is included in the cross-modal causal edge set. In another implementation, a predefined list containing all cross-modal causal types is used, and all causal edges are type-matched. Causal edges whose types are in this predefined list are filtered out. Furthermore, the modal source tags of the source and target nodes of the causal edge can be used for judgment. If the modal source tags of the source and target nodes are different, the causal edge is considered a cross-modal causal edge, and further filtering is performed based on its causal type tag.

[0136] Based on modal source tags, the modal channel type of each node connected to a causal edge in the cross-modal causal edge set is determined, resulting in a list of modal channels to be enabled. This serves to accurately map abstract causal relationships to specific physical acquisition channels, providing a clear basis for subsequent channel activation and deactivation operations. In one implementation, for each causal edge in the cross-modal causal edge set, the system extracts the modal source tags of its source and target nodes. Based on these tags, a mapping table between modal source tags and modal channel types is queried to determine the corresponding modal channel type, and these types are deduplicated and added to the list of modal channels to be enabled. In another implementation, each node carries its modal source tag upon creation, which directly encodes its associated modal channel information. When traversing the cross-modal causal edge set, the system can directly read the modal source tags from the nodes at both ends of the causal edge, convert them to the corresponding modal channel type, and then deduplicate them to form the list of modal channels to be enabled.

[0137] The system compares the list of modal channels to be enabled with the list of currently enabled modal channels, generating enable commands for newly added modal channels and deactivation commands for existing modal channels that no longer need monitoring. The aim is to accurately identify channels that need to be enabled and disabled, avoiding redundant operations and resource waste. In one implementation, the system can calculate the difference and intersection between the list of modal channels to be enabled and the list of currently enabled modal channels. Channels in the difference that need to be enabled but are not in the enabled list will generate enable commands, while channels that are enabled but not in the enable list will generate deactivation commands. In another implementation, the system can maintain a global modal channel status table. It iterates through the list of modal channels to be enabled; if a channel is marked "stopped" in the status table, an enable command is generated and the status is updated to "enabled". It iterates through the list of currently enabled modal channels; if a channel is not in the enable list, a deactivation command is generated and the status is updated to "stopped".

[0138] The above technical solution accurately identifies causal edges requiring cross-modal collaboration from complex causal relationships and associates these causal relationships with specific acquisition channels based on modal source labels. By comparing the list of currently required modal channels with the list of actually enabled channels, the system can intelligently generate enable and disable commands, thereby avoiding unnecessary channels continuously occupying monitoring resources and ensuring that the multimodal data required for causal detection can be collected in a timely and accurate manner. This on-demand start-stop mechanism not only optimizes the configuration efficiency of monitoring resources and reduces system operating costs, but also improves the accuracy and real-time performance of causal detection by dynamically adjusting the data acquisition range, enabling the incremental monitoring method of distributed cluster status to respond more flexibly and efficiently to changes in cluster status.

[0139] This application further proposes a method based on the modal source label to determine the modal channel type of each node connected by a causal edge in the cross-modal causal edge set, thereby obtaining the list of modal channels to be enabled. Specifically, this includes: extracting the modal source label from each of the two endpoints of each causal edge in the cross-modal causal edge set; determining the modal channel type corresponding to each node based on the modal source label of each endpoint; and including the modal channel types corresponding to the two endpoints, excluding those currently enabled, into the list of modal channels to be enabled.

[0140] In this approach, for each causal edge in the cross-modal causal edge set, the modality source labels of the nodes at both ends of the edge are extracted. This aims to accurately obtain the data modality information corresponding to the source and target nodes constituting these causal relationships from the identified cross-modal causal relationships. This provides a foundation for subsequent determination of which specific modal acquisition channels need to be enabled. In one implementation, the system can maintain an internal data structure of a causal hypergraph, where each node pre-stores its corresponding modality source label, for example, by setting a "modality_source" attribute in the node object. When traversing the cross-modal causal edge set, the modality source label can be obtained by directly accessing the two node objects connected to each causal edge and reading their "modality_source" attribute. In another implementation, different namespaces or ID prefixes are assigned to nodes of different modalities during the construction phase of the causal hypergraph. For example, numerical metric node IDs begin with "metric_", and log event node IDs begin with "log_". During extraction, the modality source label can be inferred by parsing the IDs of the nodes at both ends of the causal edge.

[0141] Based on the modal source tags of the two endpoints, the modal channel type corresponding to each node is determined. This step transforms the abstract modal source tags into specific, operable modal channel types, enabling the system to identify and control the corresponding acquisition channels. In one implementation, the system internally pre-defines a static mapping table or configuration dictionary, associating known modal source tags (such as "continuous numerical indicators," "discrete text log events," and "asynchronous call chain tracing events") one-to-one with their corresponding modal channel types (such as "indicator acquisition channel," "log acquisition channel," and "call chain acquisition channel"). Once the modal source tag of a node is obtained, its modal channel type can be determined by consulting this mapping table. In another implementation, a rule-based dynamic judgment mechanism is used. For example, the rule "If the modal source tag contains 'indicator,' then the channel type is 'indicator acquisition.' If it contains 'log,' then the channel type is 'log acquisition'" can be defined. This approach allows for a more flexible definition of the correspondence between modalities and channels, especially suitable for situations where modal source tags may have variations or combinations.

[0142] The modal channel types corresponding to the two endpoints of the causal edge, excluding those currently enabled, are added to the list of modal channels to be enabled. This step is crucial for achieving "on-demand start / stop." By excluding already running channels, it ensures that the generated enable commands are incremental and non-redundant, thereby optimizing resource utilization and avoiding unnecessary system overhead. In one implementation, the system maintains a real-time "set of currently enabled modal channels." After determining all modal channel types for the two endpoints of the causal edge, these types are compared with the "set of currently enabled modal channels." Only channel types not present in the "set of currently enabled modal channels" are added to the "list of modal channels to be enabled." In another implementation, after determining each modal channel type, the current running status of the channel is obtained in real-time by calling the monitoring system's API or querying its status database. If the query result indicates that the channel is currently in a "stopped" or "inactive" state, it is added to the "list of modal channels to be enabled."

[0143] The above technical solution accurately identifies the modal channels that actually need to be activated, avoiding the problem of repeatedly including already activated modal channels in the activation list. For example, by extracting the modal source tags of the nodes at both ends of each causal edge in the cross-modal causal edge set, the targeted and efficient nature of subsequent processing is ensured, avoiding useless processing of irrelevant causal edges. Based on the modal source tags, the modal channel type corresponding to each node is determined, achieving accurate mapping between modal information and specific acquisition channels. Furthermore, by excluding already activated modal channels from the activation list, redundant activation commands are avoided, thereby reducing the operational burden of monitoring channel scheduling and preventing interference with existing normal acquisition channels. This makes the on-demand start and stop of cross-modal acquisition channels more accurate, optimizes the allocation of monitoring resources, avoids unnecessary resource waste, and improves the efficiency and stability of the entire distributed cluster state incremental monitoring method.

[0144] This application further proposes a method comprising: based on monitoring and scheduling signals and channel start / stop instructions, performing feedback correction on the access parameters of a multimodal state event stream, and dynamically planning the neighborhood expansion range for the next round of incremental causality testing. For example, see... Figure 5 The method includes the following steps: 501. Based on the acquisition frequency adjustment signal in the monitoring and scheduling signal, correct the sampling interval parameter of each modal channel in the multimodal state event stream.

[0145] 502. Based on the channel start / stop command, perform start or stop operations on the corresponding modal channel in the multimodal state event stream.

[0146] 503. Based on the topological change measurement of this round and the set of nodes involved in the newly added or changed causal edges in the updated causal hypergraph, determine the number of neighborhood layers to be expanded in the next round of incremental causality test, as the neighborhood expansion range.

[0147] The monitoring and scheduling signal is a control signal generated by the system based on the analysis results of the distributed cluster state, such as changes in the topology of the causal hypergraph and causal type markings, to guide the behavior of the monitoring system. This signal can contain various instructions, such as a sampling frequency adjustment signal and an incremental reporting content selection signal, and its function is to dynamically adjust monitoring resources and strategies. The sampling frequency adjustment signal is a component of the monitoring and scheduling signal; it indicates the direction and magnitude of the adjustment of the data sampling frequency of each modal channel in the multimodal state event stream. For example, it can indicate increasing, decreasing, or maintaining the current sampling frequency. This can be implemented by directly specifying the target sampling frequency through a numerical parameter, or by indicating a relative adjustment through an enumeration value (such as "high," "medium," or "low"). The multimodal state event stream refers to various types of state data from the distributed cluster, including continuous numerical index events, discrete text log events, and asynchronous call chain tracing events. Each modal channel is an independent data acquisition and transmission path established for different modal data types or sources in the multimodal state event stream. Each modal channel is responsible for accessing data of a specific type or source. The sampling interval parameter is a key parameter controlling the data acquisition frequency of this modal channel, determining the time granularity of data acquisition. Adjusting this sampling interval parameter means adjusting the data acquisition frequency of a specific modal channel according to the requirements of the acquisition frequency adjustment signal, to adapt to the monitoring needs of changes in cluster status. For example, when the acquisition frequency adjustment signal indicates an increase in the sampling frequency, the sampling interval parameter is decreased, thereby acquiring more data per unit time.

[0148] The channel start / stop command is a command generated by the system based on the analysis of causal relationships, especially the identification of cross-modal causal relationships, to control the data acquisition status of a specific modal channel. This command typically includes the identifier of the target modal channel and the type of operation to be performed (start or stop). The start or stop operation refers to activating or deactivating the corresponding modal channel in the multimodal state event stream according to the channel start / stop command. An start operation means initiating the data acquisition and processing flow of that modal channel, while a stop operation means pausing or terminating the data acquisition of that modal channel. This can be achieved by sending control commands to the corresponding modal acquisition agent or by modifying system configuration parameters.

[0149] This topology change metric quantifies the degree of change in the causal hypergraph structure, reflecting the differences in the causal hypergraph's topology at different points in time. This metric can be calculated based on changes in topological features such as the number of connected components and the number of loop structures. The updated causal hypergraph is the latest causal graph model reflecting the current state of causal relationships in the distributed cluster after this round of incremental causal testing. The set of nodes involved in the newly added or changed causal edge refers to all nodes connected to newly discovered causal edges or causal edges whose attributes (such as causal strength and causal type) have changed in the causal hypergraph during this round of incremental causal testing. These nodes represent the core region where causal relationships have changed. The neighborhood expansion range refers to the range of causal testing expanded outwards from the nodes involved in the newly added or changed causal edge in the next round of incremental causal testing. This range is usually measured by the "neighborhood layer number." For example, expanding the neighborhood by one layer means testing nodes directly connected to the core node, while expanding by two layers includes the neighbors of nodes connected to the core node. Determining the extent of this neighborhood extension aims to balance the comprehensiveness of causal testing with computational efficiency, ensuring that key areas of change are adequately tested while avoiding unnecessary full-map scans.

[0150] The above technical solution implements the monitoring and control commands generated in the preceding steps, achieving closed-loop control of incremental monitoring of the distributed cluster status. Simultaneously, it rationally plans the scope of the next round of incremental causal verification, thus balancing computational resource consumption and the accuracy of causal relationship discovery. For example, by correcting the sampling interval parameters of each modal channel based on the acquisition frequency adjustment signal in the monitoring scheduling signal, the parameters can be corrected according to the acquisition frequency adjustment requirements generated in the preceding steps combined with topology changes, rather than uniformly adjusting all channels. This ensures that the sampling parameters of relevant channels are adjusted only for corresponding needs, conforming to the actual needs of causal structure changes and avoiding the waste of monitoring resources caused by useless frequency adjustments. By executing start or stop operations on corresponding modal channels based on channel start / stop commands, operations can be performed according to the channel start / stop requirements obtained after identifying cross-modal causal relationships in the preceding steps, thereby enabling modal channels that require monitoring as needed and shutting down modal channels that do not require monitoring. This optimizes the allocation of monitoring resources while ensuring that no potential causal relationships are overlooked, reducing resource consumption caused by invalid data acquisition. Furthermore, the neighborhood expansion range for the next round of incremental causality testing is determined based on the topology change measurement in this round and the set of nodes involved in the newly added causal edges in the updated causal hypergraph. This range is determined by combining the actual degree of topology change in this round and the core nodes involved, rather than using a fixed range or full-scale testing. This ensures that potential relationships around the causal structure changes are fully tested, avoiding omissions of potential changes, while also avoiding unnecessary computational overhead from full-scale testing. This significantly improves the operational efficiency of incremental causality testing, thus better adapting to the dynamic adjustment needs of incremental monitoring of distributed cluster states.

[0151] This application further proposes a step for correcting the sampling interval parameters of each modal channel in a multimodal state event stream based on the sampling frequency adjustment signal in the monitoring and scheduling signal. This includes: extracting the target sampling frequency corresponding to the sampling frequency adjustment signal from the monitoring and scheduling signal; determining the modal channels associated with nodes where causal edges are newly added or whose causal edge attributes change in the current round of incremental causality testing, thus obtaining a set of modal channels that need adjustment; and correcting the sampling interval parameters of each modal channel in the set of modal channels that need adjustment to the sampling interval parameters corresponding to the target sampling frequency.

[0152] For example, when extracting the target sampling frequency corresponding to the sampling frequency adjustment signal from the monitoring and scheduling signal, multiple methods are employed. For instance, the monitoring and scheduling signal may contain an explicit numerical field that directly encodes the desired sampling frequency value, such as the number of samples per second or the sampling period duration. The system can obtain the target sampling frequency by parsing this field. Alternatively, the monitoring and scheduling signal may carry a predefined frequency level identifier. The system uses this identifier to query a frequency configuration table to obtain the specific target sampling frequency. Furthermore, the monitoring and scheduling signal may contain an adjustment factor, which is used to multiply the current base sampling frequency to calculate the target sampling frequency.

[0153] When determining the modal channels associated with nodes that have experienced new causal edges or changes in causal edge attributes during the current round of incremental causality testing, and obtaining the set of modal channels that need adjustment, a mapping table between nodes and modal channels is maintained in advance. When the system identifies nodes that have changed in the causal hypergraph, it can directly obtain the modal channels to which these nodes belong by querying this mapping table, thereby constructing the set of modal channels that need adjustment. Another approach is to embed the node identification information it is responsible for into the data acquisition agent of each modal channel. When receiving a list of changed nodes, each agent determines whether it is associated with itself and reports its own modal channel. In addition, the affected modal channels can be dynamically inferred by analyzing the modal source tags of the changed nodes and combining them with the system's modality-channel correspondence rules.

[0154] When adjusting the sampling interval parameters of each modal channel in the set of modal channels to match the sampling interval parameters corresponding to the target acquisition frequency, a command carrying the new sampling interval parameters is sent to the corresponding modal acquisition agent via a remote configuration interface or API. Upon receiving the command, the acquisition agent updates its internal sampling logic. Alternatively, the new sampling interval parameters can be written to a shared configuration center, from which each modal acquisition agent periodically pulls the latest configuration and applies it. Another option is to directly modify the internally maintained parameter tables for each modal channel in the centralized controller and push these updated parameters to the corresponding acquisition modules at the start of the next data acquisition cycle.

[0155] The above technical solution enables accurate feedback and correction of access parameters for multimodal state event streams. Extracting the target acquisition frequency from the monitoring and scheduling signals ensures that the direction of parameter adjustment matches the needs of current cluster state changes. By identifying the modal channels associated with nodes that have added causal edges or changed causal edge attributes in this round of incremental causality testing, the system can accurately locate the modal channels that truly require adjustment, avoiding unnecessary modifications to irrelevant channels and thus preventing resource waste and interference with stable channels. Adjusting the sampling interval parameters only for each modal channel in the set requiring adjustment makes the allocation of monitoring resources more rational, satisfying the monitoring needs after cluster state changes while maintaining the acquisition stability of other stable channels, significantly improving the efficiency and responsiveness of the monitoring system.

[0156] This application further proposes a method for performing start or stop operations on corresponding modal channels in a multimodal state event stream based on channel start / stop commands, specifically including: Analyze the channel start / stop commands to determine the target modal channel and operation type.

[0157] When the operation type is start operation and the target modal channel is currently in a stopped state, resume data acquisition of the target modal channel and initialize the sampling parameters of the modal acquisition agent corresponding to the target modal channel.

[0158] When the operation type is a stop operation, check whether the target modal channel is still the source or target node of other cross-modal causal edges in the current causal hypergraph. If not, perform a stop operation on the target modal channel and save a state snapshot of the target modal channel. The state snapshot is used to restore the sampling parameters when the target modal channel is reopened.

[0159] To better understand the above technical solutions, the technical features involved will be described in detail below.

[0160] Regarding the step of "parse the channel start / stop command to determine the target modal channel and operation type," this step aims to clarify the specific object and execution direction of the operation. For example, a channel start / stop command can be a structured data packet containing a clear channel identifier (e.g., channel ID) and an opcode (e.g., 0 for stop operation, 1 for start operation). The system parses this data packet, extracts the corresponding channel identifier as the target modal channel, and identifies the operation type indicated by the opcode. Alternatively, the command can be a predefined text command. The system uses pattern matching or semantic analysis to identify the name of the target modal channel and operation keywords such as "start" and "stop" from the text, thereby determining the operation type. This step is fundamental to all subsequent operations, ensuring that the entire start / stop process accurately applies to the intended modal channel and executes according to the correct logic.

[0161] Regarding the step of "when the operation type is 'start' and the target modal channel is currently in a stopped state, resuming data acquisition of the target modal channel and initializing the sampling parameters of the modal acquisition agent corresponding to the target modal channel," this step aims to efficiently and accurately reactivate a stopped modal channel. Resuming data acquisition can be achieved by sending a start signal to the acquisition agent corresponding to the target modal channel, prompting it to re-establish its connection with the data source and begin transmitting data. Alternatively, the system can directly activate the operating system-level data flow pipeline, allowing data to flow back into the processing module. Initializing sampling parameters can be done by loading the channel's default sampling parameters from the configuration library, such as sampling frequency, data format, and data filtering rules, or, more preferably, by restoring its historical sampling parameters from a previously saved state snapshot of the channel. This approach avoids redundant or invalid initialization operations on channels that are already running normally, reduces unnecessary resource consumption, and ensures that the data acquired after restarting meets the current monitoring requirements and the parameter requirements of incremental causal analysis, guaranteeing the correctness of the input data.

[0162] Regarding the step of "when the operation type is a stop operation, check whether the target modal channel is still a source or target node of other cross-modal causal edges in the current causal hypergraph. If not, perform a stop operation on the target modal channel and save a state snapshot of the target modal channel. The state snapshot is used to restore sampling parameters when the target modal channel is restarted," this step aims to prevent accidental stopping of modal channels that still have actual monitoring needs and to facilitate possible future restarts. Before performing a stop operation, the system performs a pre-verification of the target modal channel. This verification mechanism can traverse all causal edges in the currently updated causal hypergraph and check whether the modal source labels of the source and target nodes of these causal edges match the target modal channel to be stopped. If it is found that the target modal channel is still a source or target node of any cross-modal causal edge, it indicates that the channel still plays an important role in the current causal relationship analysis. In this case, a stop operation will not be performed, thus avoiding errors in incremental causal analysis due to missing data. Another checking method is for the system to maintain an active channel reference counter. Each time a channel is referenced by an edge in the causal hypergraph, the counter is incremented. Before stopping the operation, the counter is checked to see if it is zero. The system will only execute the stop operation if the verification result indicates that the target modal channel no longer participates in any cross-modal causal relationships, i.e., it is no longer a source or target node of any causal edge. The system that executes the stop operation saves a snapshot of the current state of the target modal channel. This snapshot can contain all configuration information such as the channel's sampling frequency, data filtering rules, and data aggregation strategies, and is serialized and stored in persistent storage media, such as a database or file system, and associated with the channel's unique identifier. When the modal channel needs to be restarted in the future, its original sampling parameters can be restored directly using this snapshot without reconfiguring the entire set of parameters. This significantly improves the efficiency of channel restart and ensures the consistency of data acquisition after restart, thereby guaranteeing the continuity of incremental causal analysis.

[0163] The above technical solution solves the problems of accidentally stopping useful channels and cumbersome restart parameter configuration during on-demand start / stop of modal channels. By parsing channel start / stop commands, the target modal channel and operation type can be accurately identified, ensuring the accuracy of subsequent operations. For start operations, only channels in a stopped state are restored for data acquisition and parameter initialization, avoiding repeated operations on already running channels, reducing resource waste, and ensuring the correctness of data acquired after restarting, making it meet the current monitoring needs and parameter requirements of incremental causal analysis. More importantly, for stop operations, a pre-verification mechanism based on the current causal hypergraph node role is introduced. The stop operation is only executed when the target modal channel is no longer the source or target node of any cross-modal causal edge. This fundamentally avoids accidentally stopping channels that still have monitoring needs, prevents incremental causal analysis errors caused by data loss, and thus ensures the accuracy of causal topology analysis. Furthermore, saving a snapshot of the channel's state during a stop operation allows the channel to directly restore its original sampling parameters upon future restarts, significantly improving channel restart efficiency. It also preserves the channel's previous configuration information, ensuring data consistency after restart and thus guaranteeing the continuity and accuracy of incremental causal analysis. This layered verification and state preservation mechanism enables more intelligent and accurate matching of monitoring resource configuration to actual cluster state changes, optimizing monitoring resource allocation and improving the robustness and efficiency of the entire distributed cluster incremental state monitoring method.

[0164] This application further proposes a method for determining the neighborhood expansion range for the next round of incremental causality testing. This method determines the number of neighborhood layers to be expanded in the next round of incremental causality testing based on the topological change metric of the current round and the set of nodes involved in the newly added or changed causal edges in the updated causal hypergraph. This number serves as the neighborhood expansion range. For example, the method includes the following steps: When the topological change metric in the current round is lower than the stationarity threshold, the neighborhood expansion layer number for the next round of incremental causality testing is determined as the first layer number.

[0165] When the topological change metric in this round is not lower than the stability threshold, the nodes connected to the newly added causal edges and the causal edges with attribute changes in the updated causal hypergraph are extracted to form a set of changed nodes.

[0166] Centered on each node in the set of changing nodes, the number of neighborhood layers to expand outward is determined based on the magnitude of the topological change metric exceeding the stability threshold, which serves as the neighborhood expansion range. The neighborhood expansion range increases with the magnitude of the change.

[0167] To better understand the above technical solutions, the key technical features involved will be explained in detail.

[0168] The topology change metric in this round is a quantitative indicator that measures the degree to which the structure of the distributed cluster causal hypergraph changes within the current monitoring period. It reflects the addition or deletion of causal relationships, changes in their strength, and the resulting overall changes in topological features such as connectivity and loops. This metric can be obtained in various ways, for example, by comprehensively calculating multiple dimensions such as the number of additions or deletions of causal edges in the causal hypergraph, the sum of changes in causal strength, changes in the number of connected components, and changes in the number of loop structures. Alternatively, it can be obtained by comparing the differences between the current causal hypergraph and the previous round's causal hypergraph in specific topological features (such as spectral radius and clustering coefficient). In some implementations, this metric can be obtained by constructing a filtering sequence based on the causal strength of each causal edge in the updated causal hypergraph, as described above, and then measuring the topology change of the updated causal hypergraph on the filtering sequence.

[0169] The set of nodes involved in newly added or changed causal edges in the updated causal hypergraph refers to the core region where the causal hypergraph structure has actually changed within the current incremental causality test cycle. It includes nodes connected by newly established causal edges (including ordinary causal edges and hyperedges), as well as nodes connected by causal edges whose original causal edge attributes (such as causal strength, causal type labeling, etc.) have significantly changed. This set can be constructed by traversing the results of the current round of incremental causality tests, filtering out all causal edges marked as "new" or "attribute updated," and collecting the source and target nodes of these causal edges. Alternatively, it can be constructed by comparing the edge lists of the updated causal hypergraph with those of the previous round, identifying the differing edges, and then extracting the nodes associated with these differing edges.

[0170] The neighborhood expansion range refers to the number of layers or distances explored outward from a specific node when performing incremental causality testing, encompassing nodes with causal relationships. It determines the degree of locality and computational complexity of the incremental causality test. This range can be defined as the maximum hop count of nodes reachable from the central node along causal edges. Alternatively, it can be defined as the set of nodes within the embedding space of the causal hypergraph centered at a certain distance from the central node that are less than a certain threshold.

[0171] The stationarity threshold is a critical value used to determine whether changes in the topology of a causal hypergraph are in a stable state. When the measure of topology change is below this threshold, it indicates that the overall structure of the causal hypergraph is relatively stable and has not undergone large-scale or significant changes. This threshold can be determined by analyzing historical data, statistically analyzing the distribution of topology change measures under normal cluster operation, and using a certain percentile of the distribution (e.g., 90% or 95%) as the stationarity threshold. Alternatively, it can be manually set based on expert experience or the system's tolerance for the stability of the causal structure.

[0172] The first layer refers to a small number of neighborhood expansion layers set for the next round of incremental causality testing when the causal hypergraph topology is in a stationary state. It is usually a small positive integer, such as layer 1, which only checks the neighboring nodes directly connected to the changed node. Alternatively, it can be set to layer 2, depending on the system's sensitivity to small changes, to capture more subtle local effects.

[0173] The magnitude by which the topology change metric exceeds the stationarity threshold refers to the amount by which the currently calculated topology change metric exceeds a preset stationarity threshold. It quantifies the "severity" or "abnormality" of changes in the causal hypergraph structure. This magnitude can be directly calculated as "topology change metric - stationarity threshold". Alternatively, it can be the result of this difference after some normalization or nonlinear transformation to better map it to the number of neighboring layers.

[0174] The neighborhood expansion range increases with the magnitude, which is an adaptive strategy. This means that the more drastic the topological changes in the causal hypergraph (i.e., the greater the magnitude of the topological change exceeding the stationarity threshold), the larger the neighborhood range that the incremental causality test needs to explore. This expansion can be linear, piecewise, or non-linear. For example, a piecewise function can be used, expanding one layer when the magnitude is in a certain interval, two layers when it is in another larger interval, and so on. Alternatively, a continuous function can be used, such as neighborhood layer number = f(magnitude), where f is a monotonically increasing function, such as a logarithmic or exponential function.

[0175] The above technical solution enables the dynamic and intelligent adjustment of the neighborhood expansion range for the next round of incremental causality testing based on the actual changes in the current causal hypergraph topology. When the topology change metric in this round is below the stability threshold, it indicates that the causal hypergraph is in a stable state. At this point, the neighborhood expansion layer number is determined to be a smaller first layer, thereby reducing the computational range of incremental causality testing, avoiding unnecessary recalculation of stable regions, significantly saving computational resources, and improving the overall processing efficiency of incremental monitoring.

[0176] When the topology change metric is not lower than the stationarity threshold, the system can identify a substantial change in the causal hypergraph. At this point, by extracting the nodes connected to newly added causal edges and those with attribute changes in the updated causal hypergraph, a set of changed nodes is formed, allowing for accurate localization of the changed region. Centered on these changed nodes, the number of outward expansion layers is determined based on the magnitude of the topology change metric exceeding the stationarity threshold, with the expansion range increasing with the magnitude. This adaptive adjustment mechanism ensures that when the causal structure change magnitude is small, only a moderate expansion is performed in a local area, avoiding overcomputation. Conversely, when the change magnitude is large, the neighborhood range is expanded accordingly to fully cover the areas potentially affected by the change, ensuring that incremental causal checks accurately capture complete causal structure changes and avoiding omissions due to an excessively small neighborhood range. This provides a reliable foundation for subsequently generating accurate monitoring and scheduling signals and channel start / stop commands.

[0177] In summary, the proposed solution minimizes unnecessary computational resource waste while ensuring the accuracy of causal verification, improves the efficiency and accuracy of incremental monitoring of distributed cluster states, and enables the monitoring system to respond more flexibly and intelligently to changes in cluster states.

[0178] This application further proposes determining the number of neighboring layers to expand outward, based on the magnitude of the topological change metric exceeding a stationarity threshold, as the range of the neighboring expansion. For example, the method includes: determining the difference between the current topological change metric and the stationarity threshold to obtain the threshold exceedance magnitude. This "threshold exceedance magnitude" refers to a quantitative indicator of the degree of topological change in the currently detected causal hypergraph exceeding the system's "stationary" state. Its function is to transform the abstract degree of topological change into a concrete numerical value for subsequent hierarchical judgment and decision-making. For example, it can be obtained by directly calculating the arithmetic difference between the current topological change metric and the stationarity threshold. Alternatively, a normalized difference or percentage difference can be used to calculate it to eliminate the influence of different dimensions, making the threshold exceedance magnitude more universal.

[0179] When the threshold amplitude falls within the first amplitude interval, a new neighborhood is expanded outward from each node in the changed node set. This "first amplitude interval" represents a relatively small change in the causal hypergraph topology, but it exceeds the range of a stationary state. In this case, "expanding the neighborhood" means taking the detected changed node as a starting point and including its directly connected nodes (i.e., nodes reachable in one step) in the next round of incremental causality testing. Its purpose is to ensure that, when the change is not drastic, the directly affected area can be covered with minimal computational cost. For example, the system can maintain a predefined "first amplitude interval" range; when the calculated threshold amplitude falls within this interval, a new neighborhood expansion strategy is triggered. Neighborhood expansion can be implemented using a graph traversal algorithm, starting from each node in the changed node set, traversing only its directly adjacent nodes, and adding these nodes and their associated edges to the next round of testing.

[0180] Furthermore, when the threshold amplitude falls within the second amplitude interval, a two-layer neighborhood is expanded outward from each node in the changed node set, with the lower limit of this second amplitude interval not lower than the upper limit of the first amplitude interval. This "second amplitude interval" indicates a significant change in the causal hypergraph topology, requiring broader attention. "Expanding the neighborhood by two layers" means including the nodes directly connected to the changed node, as well as those connected to these directly connected nodes (i.e., nodes reachable in two steps), in the next round of incremental causal testing. Its purpose is to address more complex and far-reaching changes in causal structure, avoiding the omission of potential associations. For example, the system can set the range of the "second amplitude interval" and ensure its lower limit is not less than the upper limit of the first amplitude interval to guarantee interval exclusivity and logical clarity. Expanding the neighborhood by two layers can also be achieved using a graph traversal algorithm, starting from each node in the changed node set and performing a two-layer depth traversal, including all two-step reachable nodes and their associated edges in the next round of testing.

[0181] The above technical solution transforms the degree to which causal structure changes exceed a stable state into a quantifiable comparative indicator, providing a clear basis for determining the number of expansion layers and avoiding subjectivity in setting the number of expansion layers. When the exceedance threshold is within the first range, only one layer of neighborhood is expanded. This covers the adjacent areas that may be affected by the change, meeting the scope requirements of incremental causal testing, without incurring unnecessary computational overhead due to excessive expansion, thus controlling the overall computational cost. When the exceedance threshold is within the second range, two layers of neighborhood are expanded to cover a larger area of ​​potentially affected regions, ensuring that the next round of incremental causal testing can completely capture all causal relationships affected by the change, avoiding omissions of causal associations and ensuring the integrity of the causal hypergraph update. By dividing the exceedance threshold into different ranges and matching the corresponding number of neighborhood expansion layers, this solution can flexibly adjust the scope of the next round of incremental causal testing according to the severity of the causal structure change, balancing computational efficiency and the completeness of coverage of causal structure changes. This solves the problem of wasted computational resources or omissions of causal relationships caused by unreasonable neighborhood expansion range settings in related technologies, significantly improving the accuracy and efficiency of distributed cluster state incremental monitoring.

[0182] This application further proposes that, after correcting the access parameters of the multimodal state event stream based on monitoring and scheduling signals and channel start / stop commands, the method further includes: obtaining the change in the number of newly added causal edges in the updated causal hypergraph before and after the execution of this round of feedback correction, as the information gain of this round of regulation. This information gain is compared with a desired gain threshold; when the information gain is lower than the desired gain threshold, the response sensitivity of the acquisition frequency adjustment in the next round of monitoring and scheduling signals is reduced. This information gain is correlated with the topology change metric; when the topology change metric indicates that the causal hypergraph is in a change range and the information gain remains lower than the desired gain threshold, the weight of the topology change metric in the generation of the next round of monitoring and scheduling signals is reduced.

[0183] For example, the change in the number of newly added causal edges in the updated causal hypergraph before and after the execution of this feedback correction is used as the information gain of this round of regulation. Information gain refers to the amount of new and valuable causal structure information about the state of the distributed cluster obtained through this monitoring and regulation. Its role is to quantify the actual effect of this regulation and provide an objective basis for subsequent adaptive adjustments. Besides the change in the number of newly added causal edges, information gain can also be measured in other ways, such as obtaining the number of causal edges in the updated causal hypergraph with significantly increased causal strength, or obtaining the number of causal edges in the causal hypergraph with changed causal type labels. These changes also reflect the information increment brought about by the regulation. "Before and after the execution of the feedback correction" refers to the period before and after the system adjusts the access parameters of the multimodal state event stream according to monitoring scheduling signals and channel start / stop instructions. By comparing the changes in the causal hypergraph at these two points in time, the actual impact of this regulation on causal discovery can be accurately assessed. The updated causal hypergraph is a graph structure reflecting the current causal dependencies of the distributed cluster after incremental causal testing and topological measurement. It includes nodes (representing various modal events) and edges (representing causal relationships), as well as attributes such as causal type labels and confidence levels. The change in the number of newly added causal edges is considered as information gain. The rationale for this is that the newly added causal edges directly represent the discovery of previously unknown causal relationships by the system. These newly discovered causal relationships are the direct result of this regulation (e.g., by adjusting the sampling frequency or enabling new channels), and can intuitively reflect the improvement of the regulation's ability to discover causal relationships.

[0184] The information gain is compared with the expected gain threshold. When the information gain is lower than the expected gain threshold, the response sensitivity of the sampling frequency adjustment in the next round of monitoring and scheduling signals is reduced. The expected gain threshold is a pre-set benchmark value used to measure whether the current regulation has achieved the expected effect. It can be set based on historical data, expert experience, or system performance requirements. For example, it can be set to find at least X new causal edges in a regulation cycle, or the average increase in causal strength reaches Y. Response sensitivity refers to the degree or intensity of the monitoring and scheduling system's response to changes in the causal hypergraph. Reducing response sensitivity means that the system will adopt a more conservative or delayed sampling frequency adjustment strategy when facing subsequent changes in the causal hypergraph. For example, the step size of the sampling frequency adjustment may be reduced, or the time interval between two adjustments may be extended to avoid frequent and potentially ineffective adjustments. Sampling frequency adjustment refers to dynamically increasing or decreasing the sampling frequency of multimodal state event streams according to changes in cluster state to balance the real-time performance of data acquisition and resource consumption.

[0185] Furthermore, the information gain is correlated with the topology change metric for verification. When the topology change metric indicates that the causal hypergraph is in a change range while the information gain remains below the expected gain threshold, the weight of the topology change metric in the next round of monitoring and scheduling signal generation is reduced. The topology change metric is a quantitative assessment of the stability or dynamics of the causal hypergraph structure, measured, for example, by indicators such as the number of connected components and loop structures. It reflects the degree of change in the overall structure of the causal network. Correlation verification involves comprehensively comparing the information gain and the topology change metric to verify the effectiveness and accuracy of the topology change metric, aiming to identify potential misjudgments. The change range refers to the range of the degree of topology change in the causal hypergraph, such as a stationary range, a range of concern, or a range of significant change. These ranges are typically defined by stationary and significant change thresholds. Weight refers to the proportion or influence of a particular indicator in the decision-making process. Reducing the weight of the topology change metric means that when generating monitoring and scheduling signals, the system reduces its reliance on the topology change metric, instead considering other factors or adopting a more cautious strategy. For example, this can be achieved by adjusting the weighted average coefficient or by reducing its priority in the decision tree.

[0186] Through the above technical solution, after adjusting the monitoring parameters, a closed-loop feedback verification mechanism for the control effect is introduced. By obtaining the change in the number of newly added causal edges in the causal hypergraph before and after the feedback correction is executed as information gain, the actual effect of this control can be objectively quantified. When the information gain is lower than the expected gain threshold, the system can adaptively reduce the response sensitivity of the sampling frequency adjustment in the next round of monitoring and scheduling signals, thereby avoiding the waste of monitoring resources caused by frequent and inefficient adjustments and optimizing the allocation efficiency of monitoring resources. By correlating and verifying the information gain with the topology change metric, possible misjudgments in the topology change metric can be identified. When the topology change metric indicates that the causal hypergraph is in the change range but the actual information gain is consistently insufficient, the system can reduce the weight of the topology change metric in the generation of subsequent monitoring and scheduling signals, thereby reducing the misleading effect of incorrect topology judgments on monitoring and scheduling. This allows the monitoring and scheduling strategy to more accurately reflect the actual state changes of the distributed cluster, realizing the dynamic self-optimization of the monitoring and scheduling strategy and ensuring that the configuration of monitoring resources is synchronously matched with the actual state changes of the cluster.

[0187] This application further proposes a method to optimize the accuracy of incremental causality testing. After correcting the access parameters of the multimodal state event stream based on the monitoring and scheduling signal and the channel start / stop command, the method further includes: obtaining the causal edges marked as false alarms in the causal hypergraph in multiple consecutive rounds after the feedback correction, and extracting the modality source label combination features of the false alarm causal edges. Based on the modality source label combination features, the modality combination type that frequently generates false alarms is identified. In the next round of incremental causality testing, the significance threshold for causality testing is increased for the node pairs corresponding to the modality combination type that frequently generates false alarms.

[0188] For example, after feedback correction is executed, the system continuously acquires and tracks causal edges marked as false positives in the causal hypergraph across multiple consecutive rounds, and extracts the modal source label combination features of these falsely labeled causal edges. A "causal edge marked as a false positive" refers to a causal connection identified as a causal relationship in incremental causality testing but judged as a false positive in subsequent verification (e.g., through manual confirmation, comparison with other more reliable data sources, or system behavioral feedback). To ensure the stability rather than randomness of the identified false positives, this method emphasizes observation and labeling across "multiple consecutive rounds." For example, a threshold is set, such as only causal edges marked as false positives for 3 or 5 consecutive rounds are included in the statistics. For these falsely labeled causal edges, the system extracts their "modal source label combination features." This feature refers to the combination of modal types of the nodes constituting the two ends of the causal edge. For example, if a causal edge connects a continuous numerical indicator event node and a discrete text log event node, its modal source label combination feature is "numerical indicator-log event." This feature can be extracted by querying the modality source labels associated with the nodes in the causal hypergraph.

[0189] The system identifies "frequently false positive modal combination types" based on the extracted modality source label combination features. This is typically achieved through statistical analysis; for example, the system can maintain a counter to record the number of times each modality source label combination feature is marked as a false positive. When a certain modality combination type exceeds a frequency threshold within a set time window or cumulatively, it is identified as a frequently false positive type. For example, if the "numerical indicator-log event" combination has a significantly higher false positive rate than other combinations over the past N rounds, it is marked as a frequently false positive type. Identifying these frequently false positive modality combination types helps to accurately pinpoint error-prone links in causal testing.

[0190] In the next round of incremental causality testing, the system raises the significance threshold for causality tests on node pairs corresponding to modality combinations that frequently generate false positives. This means that for node pairs belonging to these specific modality combinations, stronger evidence is needed to determine a causal relationship during causality testing. For example, if the standard significance level (p-value) threshold for causality testing is 0.05, this threshold might be adjusted to 0.01 or lower for modality combinations that frequently generate false positives. This adjustment can be achieved by dynamically modifying the internal parameters of the causality testing algorithm; for example, by selecting different significance criteria based on the modality combination type of the node pair when performing linear or nonlinear causality tests.

[0191] The above technical solution addresses the persistent problem of false alarms in incremental causality testing. By statistically analyzing the causal edges of false alarms across multiple rounds after feedback correction and extracting their modality source label combination features, the system can accurately identify which specific modality combinations are more prone to generating false alarms. This adaptive adjustment based on historical false alarm patterns makes the causality testing process more targeted. In subsequent incremental causality testing, raising the significance threshold for node pairs corresponding to these frequently false alarm modality combinations can filter out false positive causal relationships that are prone to occur, thereby reducing the false alarm rate. This not only improves the accuracy and reliability of the causal hypergraph but also makes the monitoring scheduling signals and channel start / stop commands generated based on the causal hypergraph more accurate, avoiding resource waste or misjudgment caused by incorrect causal relationships, thus improving the overall efficiency of incremental monitoring of the distributed cluster status and the accuracy of fault location. This mechanism enables the monitoring system to learn and optimize itself, continuously improving its causal discovery capabilities in complex multimodal data environments.

[0192] This application further proposes that, after obtaining the topology change metric, the method further includes: when the topology change metric indicates the emergence of a new loop structure in the causal hypergraph, extracting the modality source labels of the nodes involved in the new loop structure. Based on the modality source labels, determining whether the new loop structure contains nodes from different modal sources. When the new loop structure contains nodes from different modal sources and the new loop structure persists in multiple consecutive rounds, adding a loop warning marker to the causal edges within the new loop structure in the updated causal hypergraph. This loop warning marker is used to indicate the cross-modal cyclic dependency relationship corresponding to the new loop structure.

[0193] For example, when the topology change metric indicates the presence of a new loop structure in the causal hypergraph, it means that a new circular dependency path has been detected by measuring the topology change of the causal hypergraph. This indication can be achieved by comparing the current round with the previous round of the causal topology persistence graph to identify the number of new loop structures or specific loop paths. For instance, graph traversal algorithms (such as depth-first search or breadth-first search) can be used to detect loops in the graph and compared with historical records to determine new loops. Alternatively, graph theory loop detection algorithms, such as Tarjan's algorithm or Kosaraju's algorithm, can be used to identify strongly connected components and thus determine whether a new loop structure exists. Once a new loop structure is detected, the system extracts the modal source label of the nodes involved in the new loop structure. This modal source label is an identifier assigned to each time-series signal component when uniformly representing the multimodal state event stream of the distributed cluster, indicating that the signal originates from continuous numerical index events, discrete text log events, or asynchronous call chain tracing events. In implementation, each node can store an attribute field, namely its modality source label, in the data structure of the causal hypergraph. After identifying the loop structure, the system traverses all nodes in the loop and obtains the modality source labels of these nodes.

[0194] Based on the modality source label, the system determines whether the newly added loop structure contains nodes with different modality sources. This determination aims to identify cross-modal circular dependencies, as these dependencies are often more complex and harder to pinpoint. Specifically, the system collects the modality source labels of all nodes in the loop and then checks if the set of these labels contains more than one modality type. For example, if the loop contains a numerical metric node and a log event node, it is considered to contain nodes with different modality sources.

[0195] When a newly added loop structure contains nodes from different modalities and persists across multiple consecutive rounds, the system attaches a loop warning flag to the causal edges within that new loop structure in the updated causal hypergraph. "Persistent across multiple consecutive rounds" means that the loop structure is still detected after several incremental causality checks and topological metrics. This can be achieved by maintaining a history and counter for loop structures, considering a loop structure as persistent only if it consistently appears within a consecutive round threshold. For example, a persistent counter can be assigned to each detected loop structure, incrementing each time the loop reappears in a subsequent round. If it does not appear, the counter is reset to zero. When the counter reaches a preset value (e.g., 2 or 3), it is considered persistent. The loop warning flag is an attribute attached to the causal edge, such as a boolean flag or an enumerated type field, indicating that the causal edge is part of a cross-modal cyclic dependency.

[0196] The above technical solution addresses the shortcomings of previous methods in incremental monitoring of distributed cluster status, which lacked targeted vulnerability labeling. This solution can identify persistent cross-modal cyclic dependencies composed of nodes from different modalities. These relationships are often key vulnerabilities that cause abnormal distributed cluster status and make fault location difficult. By attaching loop warning labels to causal edges within these key loop structures in the updated causal hypergraph, the system can provide clearer and more targeted guidance for subsequent root cause localization and status management. Furthermore, this solution fully reuses the topology change measurement results and modality source label information already generated in the original incremental monitoring process, avoiding redundant calculations and reducing additional computational overhead. Subsequent processing is triggered only when a new loop change occurs in the topology, saving computational resources for the monitoring system. By requiring the loop structure to persist for multiple consecutive cycles before issuing a warning, temporary and accidental false loop structures can be filtered out, improving the accuracy and reliability of warning results, avoiding unnecessary false alarms, and thus enhancing the practical value and operational efficiency of the monitoring system.

[0197] The following example will provide a more detailed explanation of the above technical solution: In a large-scale distributed e-commerce platform, user A reported a significant slowdown in checkout speed. The platform's operations team needed to quickly pinpoint the root cause of the slowdown and predict potential failure propagation paths. Traditional monitoring methods struggle to integrate the massive amounts of multimodal data generated by the platform, such as continuous numerical metrics like CPU utilization and memory usage for various services, discrete text log events like errors and warnings in service logs, and asynchronous call chain tracing events between services. Especially when dealing with causal relationships between cross-modal data, existing solutions often require pre-converting discrete or asynchronous events into numerical forms, making causal discovery complex and difficult to directly identify different types of cross-modal causal relationships within a unified framework. Furthermore, the evaluation of changes in the causal graph is limited to the addition and deletion of edges, failing to detect the formation and demise of causal loops, and monitoring strategy adjustments lack feedback mechanisms, making predictive adjustments based on trends in causal structure changes difficult.

[0198] To address the aforementioned issues, this method provides a unified representation of the multimodal state event streams generated by the distributed cluster. For example, the system collects continuous numerical indicator events such as OrderService_CPU and PaymentService_Latency, discrete text log events such as PaymentService_ErrorLog and InventoryService_WarningLog, and asynchronous call chain tracing events involved in the user settlement process, such as UserService→OrderService→PaymentService. For continuous numerical indicator events, the system performs time-series alignment to obtain numerical time-series components. For discrete text log events, such as PaymentService_ErrorLog, the system adaptively adjusts the kernel function bandwidth for continuous reconstruction based on the fluctuation amplitude of the current numerical time-series component (e.g., PaymentService_Latency) within the target time period. For instance, when PaymentService_Latency fluctuates drastically, the system reduces the kernel function bandwidth to capture the short-term time-series characteristics of the log event. When the fluctuation is stable, the system increases the kernel function bandwidth to integrate long-term effects, obtaining continuous log components. For asynchronous call chain tracing events, the system determines weighting coefficients based on the semantic features of discrete text log events (e.g., the severity level of PaymentService_ErrorLog). These weighting coefficients are then used to aggregate the call chain events by arrival time interval, resulting in continuous call chain components. These numerical time-series components, log continuous components, and call chain continuous components are merged on a unified time axis and assigned their respective modal source labels (e.g., OrderService_CPU is labeled as a "numerical indicator," and PaymentService_ErrorLog as a "text log"), forming a time-series signal carrying modal source labels. This unified representation avoids the independent data pre-transformation steps found in traditional schemes, laying the foundation for subsequent cross-modal causal analysis.

[0199] The system performs incremental causality tests and topological measurements on these time-series signals carrying modal source labels and the currently maintained causal hypergraph. For example, when the system detects a potential causal relationship between PaymentService_Latency (a numerical metric) and PaymentService_ErrorLog (a discrete text log), it adaptively selects a causality test method based on their modal source labels. If the test objects are OrderService_CPU and PaymentService_Latency (both numerical metrics), the system performs both linear and nonlinear causality tests. If both are significant, it is marked as a strong causal relationship. If only linearly significant, it is marked as a weak linear causal relationship. If only nonlinearly significant, it is marked as a nonlinear causal relationship. If the test objects are PaymentService_Latency (a numerical metric) and PaymentService_ErrorLog (a discrete text log), the system models the occurrence sequence of PaymentService_ErrorLog as a stochastic process and introduces PaymentService_Latency as an external explanatory variable. It determines the causal relationship by comparing the differences in model fit and marks it as an event-driven causal relationship. When a new causal relationship is discovered (e.g., an increase in PaymentService_ErrorLog leading to an increase in PaymentService_Latency), the system establishes a corresponding causal edge in the causal hypergraph and labels it with a causal type tag (e.g., "event-driven causal relationship"). The system adjusts the confidence accumulation rate of the causal edges based on the previous round of topology change metrics. For example, if the previous round of topology change metrics indicates that the causal hypergraph is in a stable state, a baseline confidence increment step size is used. If it is in a changing state, the confidence increment step size is increased to confirm or deny causal relationships more quickly. Based on the causal strength of each causal edge in the updated causal hypergraph, the system constructs a filtering sequence and measures the topology change of the causal hypergraph on this sequence to obtain a topology change metric. For example, by extracting the number of connected components and loop structures in the subgraph under different strength thresholds and comparing the differences with the previous round of causal topology persistence graph, the degree of change in causal structure is quantified. This method not only focuses on the addition and deletion of edges, but also can perceive the formation and demise of causal loops (such as ServiceA→ServiceB→ServiceC→ServiceA), making up for the shortcomings of existing schemes in measuring the stability of causal structures at the topological level.

[0200] Based on the topology change metrics, causal type markers, and modal source labels obtained in this round, the system generates monitoring and scheduling signals and channel start / stop instructions. For example, if the topology change metric indicates that the causal hypergraph is in a stable range (below the stable threshold), the system generates a first sampling frequency adjustment signal, indicating a reduction in the sampling frequency of the multimodal state event stream, and generates a first incremental reporting content selection signal, indicating that only the topology change metric identifier should be reported. If it is in a significant change range (not below the significant change threshold), a third sampling frequency adjustment signal is generated, indicating an increase in the sampling frequency, and reporting the complete causal subgraph and all state data of the affected areas. The system identifies cross-modal causal relationships based on the causal type markers. For example, if an event-driven causal relationship is found between PaymentService_ErrorLog (text log) and PaymentService_Latency (numerical metric), the system determines the modal channels (i.e., the text log acquisition channel and the numerical metric acquisition channel) to which the involved nodes belong based on the modal source label, and generates start / stop instructions for these modal channels. For example, if the text log acquisition channel is currently not enabled, an enable instruction is generated. This on-demand start / stop mechanism avoids continuous high-frequency acquisition of all modal channels, thus optimizing resource utilization.

[0201] Based on the generated monitoring and scheduling signals and channel start / stop commands, the system provides feedback correction to the access parameters of the multimodal state event stream and dynamically plans the neighborhood expansion range for the next round of incremental causality testing. Specifically, the system adjusts the signals according to the acquisition frequency to correct the sampling interval parameters of each modal channel. For example, if an instruction to increase the sampling frequency is received, the sampling interval of numerical indicators such as PaymentService_CPU and PaymentService_Latency is adjusted from 10 seconds to 1 second. For channel start / stop commands, the system executes the corresponding start or stop operation. For example, if an instruction to enable the PaymentService_ErrorLog acquisition channel is received, its data acquisition is resumed. If an instruction to stop a modal channel that no longer involves cross-modal causality is received, the system checks whether it is still a source node or target node of other causal edges. If not, a stop operation is executed and a state snapshot is saved for subsequent recovery. In addition, the system also dynamically determines the neighborhood expansion range for the next round of incremental causality testing based on the topology change metric of this round and the set of nodes involved in the newly added or changed causal edges in the updated causal hypergraph. For example, when the topology change metric is below the stability threshold, the neighborhood expansion layer is determined to be the first layer. When the change metric is not below the stability threshold, the system extracts the set of changed nodes (such as PaymentService_Latency and PaymentService_ErrorLog), and expands the neighborhood by one or more layers outward from these nodes based on the magnitude of the change metric exceeding the stability threshold; the greater the magnitude, the larger the expansion range. This feedback correction and dynamic planning mechanism enables the monitoring resource configuration to be synchronized with the actual changes in the cluster state, overcoming the limitations of existing open-loop control strategies.

[0202] After feedback correction, the system also acquires the change in the number of newly added causal edges in this round as the information gain for regulation and compares it with the expected gain threshold. If the information gain remains below the expected gain threshold, the system reduces the response sensitivity of the acquisition frequency adjustment in the next round of monitoring and scheduling signals to avoid excessively frequent adjustments. The system also correlates and verifies the information gain with the topology change metric. When the topology change metric indicates that the causal hypergraph is in a change range and the information gain remains low, the system reduces the weight of the topology change metric in the generation of the next round of monitoring and scheduling signals to avoid ineffective regulation due to false alarms or noise. In addition, the system continuously tracks causal edges marked as false alarms in the causal hypergraph, extracts their modality source label combination features, and identifies modality combination types that frequently generate false alarms. In the next round of incremental causality testing, for the node pairs corresponding to these modality combination types that frequently generate false alarms, the system increases the significance threshold for causality testing, thereby improving the accuracy and robustness of causality discovery. When the topology change metric indicates the appearance of a new loop structure in the causal hypergraph, the system extracts the modality source labels of the involved nodes to determine whether they contain nodes with different modality sources. If a loop is included and persists in multiple consecutive rounds, the system adds a loop warning marker to the causal edges within the loop in the causal hypergraph to indicate the corresponding cross-modal cyclic dependency, thereby providing early warning of potential system instability risks.

[0203] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0204] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for incremental monitoring of the status of a distributed cluster, characterized in that, The method includes: A unified representation is performed on the multimodal state event stream of the distributed cluster to obtain a time-series signal carrying a modality source label. The multimodal state event stream includes continuous numerical index events, discrete text log events, and asynchronous call chain tracing events. Incremental causality testing and topological measurement are performed on the time-series signal carrying modal source labels and the causal hypergraph to obtain an updated causal hypergraph, causal type label and topological change measurement. The incremental causality testing adaptively selects the testing method based on the modal source labels and adjusts the accumulation rate of causal edge confidence according to the previous round of topological change measurement. Based on the topology change metric, the causal type marker, and the modal source label, monitoring and scheduling signals and channel start / stop instructions are generated. The topology change metric drives the adjustment of the acquisition frequency and incremental reporting content. The causal type marker, combined with the modal source label, drives the on-demand start / stop of cross-modal acquisition channels. Based on the monitoring and scheduling signals and the channel start / stop instructions, the access parameters of the multimodal state event stream are corrected by feedback, and the neighborhood expansion range of the next round of incremental causality verification is dynamically planned.

2. The method according to claim 1, characterized in that, The unified representation of the multimodal state event stream of the distributed cluster to obtain a time-series signal carrying modality source labels includes: The continuous numerical index events are time-series aligned to obtain numerical time-series components. Based on the fluctuation range of the numerical time series components within the target time period, the kernel function bandwidth for continuous reconstruction of the discrete text log events is adaptively adjusted, and continuous reconstruction is performed with the adjusted kernel function bandwidth to obtain continuous log components. The weighting coefficients are determined based on the semantic features of the discrete text log events, and the arrival time intervals of the asynchronous call chain tracing events are weighted and aggregated based on the weighting coefficients to obtain the continuous components of the call chain. The numerical time-series component, the log continuous component, and the call chain continuous component are merged on a unified time axis and assigned corresponding modal source labels to obtain the time-series signal carrying the modal source labels.

3. The method according to claim 1, characterized in that, The incremental causality test and topological measurement are performed on the time-series signal carrying the modal source label and the causal hypergraph to obtain the updated causal hypergraph, causal type label, and topological change measurement, including: Based on the modality source label, for the node pairs in the local subgraph of the causal hypergraph that are affected by the time-series signal, an incremental causal test method is adaptively selected to perform causal test, so as to obtain the causal relationship between the node pairs and the causal type label. Based on the causal relationship and the causal type label, the structure and attributes of the local subgraph are updated, and the confidence accumulation rate of the causal edges in the local subgraph is adjusted according to the topological change metric of the previous round, so as to obtain the updated causal hypergraph. Based on the causal strength of each causal edge in the updated causal hypergraph, a filtering sequence is constructed, and the topological change of the updated causal hypergraph is measured on the filtering sequence to obtain the topological change metric.

4. The method according to claim 3, characterized in that, Based on the modality source label, for node pairs within the local subgraph of the causal hypergraph affected by the time-series signal, an incremental causal test method is adaptively selected to perform causal testing, obtaining the causal relationship between the node pairs and the causal type label, including: When the modal combination type of the node pair is numerical index and numerical index, a linear causality test is performed on the node pair to obtain a linear test statistic, and a nonlinear causality test is performed on the node pair to obtain a nonlinear test statistic. A strong causal relationship is defined as follows: when both the linear test statistic and the nonlinear test statistic satisfy the linear significance condition, the relationship is defined as follows: when only the linear test statistic satisfies the linear significance condition, the relationship is defined as follows: when only the nonlinear test statistic satisfies the nonlinear significance condition, the relationship is defined as follows: when only the nonlinear test statistic satisfies the nonlinear significance condition, the relationship is defined as follows: When the modal combination type of the node pair is a numerical indicator and a discrete text log event, a stochastic process model is performed on the occurrence sequence of the discrete text log events to obtain a baseline fitting model containing only historical event information; the numerical indicator is introduced as an external explanatory variable into the stochastic process model to obtain an extended fitting model containing numerical indicator information; by comparing the goodness-of-fit difference between the baseline fitting model and the extended fitting model, the causal relationship between the node pair is determined and marked as an event-driven causal relationship.

5. The method according to claim 4, characterized in that, The process of performing a linear causality test on the node pairs to obtain a linear test statistic, and performing a nonlinear causality test on the node pairs to obtain a nonlinear test statistic, includes: For the time series of the first node in the node pair, linear prediction modeling is performed using the time series of the second node to obtain the linear prediction residual sequence; Based on the linear prediction residual sequence and the time series of the second node, the nonlinear information residual in the linear prediction residual sequence explained by the time series of the second node is determined as the nonlinear test statistic. Based on the goodness of fit of the linear prediction model, the linear test statistic is determined.

6. The method according to claim 1, characterized in that, The generation of monitoring and scheduling signals and channel start / stop instructions based on the topology change metric, the causal type marker, and the modality source label includes: The topological change metric is compared with a stationary threshold and a significant change threshold to determine the change range of the current causal hypergraph. Based on the aforementioned variation range, a sampling frequency adjustment signal and an incremental reporting content selection signal are generated. The sampling frequency adjustment signal is used to control the sampling interval of the multimodal state event stream, and the incremental reporting content selection signal is used to determine whether to report an incremental causal subgraph or a complete causal subgraph. Based on the causal type label, cross-modal causal relationships are identified, and the modal channel to which the node involved in the cross-modal causal relationship belongs is determined by the modal source label, generating start / stop instructions for the modal channel.

7. The method according to claim 6, characterized in that, The process of identifying cross-modal causal relationships based on the causal type marker, determining the modal channel to which the nodes involved in the cross-modal causal relationship belong using the modal source label, and generating start / stop instructions for the modal channel includes: From the causal type tags, causal edges marked as event-driven causal relationships or cross-modal joint causal relationships are selected to obtain a set of cross-modal causal edges; Based on the modal source label, determine the modal channel type of the node connected to each causal edge in the cross-modal causal edge set, and obtain a list of modal channels to be enabled; The list of modal channels to be enabled is compared with the list of currently enabled modal channels to generate an enable command for the new modal channel and a deactivation command for the existing modal channels that no longer need to be monitored.

8. The method according to claim 1, characterized in that, The step of feeding back and correcting the access parameters of the multimodal state event stream based on the monitoring and scheduling signals and the channel start / stop commands, and dynamically planning the neighborhood expansion range for the next round of incremental causality testing, includes: Based on the acquisition frequency adjustment signal in the monitoring and scheduling signal, the sampling interval parameter of each modal channel in the multimodal state event stream is corrected; Based on the channel start / stop command, the corresponding modal channel in the multimodal state event stream is opened or stopped. Based on the topological change metric of this round and the set of nodes involved in the newly added or changed causal edges in the updated causal hypergraph, the number of neighborhood layers to be expanded in the next round of incremental causal testing is determined as the neighborhood expansion range.

9. The method according to claim 8, characterized in that, Based on the topological change metric of this round and the set of nodes involved in the newly added or changed causal edges in the updated causal hypergraph, the number of neighborhood layers to be expanded in the next round of incremental causality testing is determined as the neighborhood expansion range, including: When the topological change metric in this round is lower than the stability threshold, the neighborhood expansion layer number for the incremental causality test in the next round is determined as the first layer number. When the topological change metric in this round is not lower than the stability threshold, the nodes connected to the newly added causal edges and the causal edges with attribute changes in the updated causal hypergraph are extracted to form a set of changed nodes. Centered on each node in the set of changing nodes, the number of neighborhood layers to expand outward is determined based on the magnitude of the topology change metric exceeding the stability threshold, which serves as the neighborhood expansion range, wherein the neighborhood expansion range increases with the magnitude of the change.

10. The method according to claim 1, characterized in that, After correcting the access parameters of the multimodal state event stream based on the monitoring and scheduling signal and the channel start / stop command, the method further includes: After the feedback correction is executed, obtain the causal edges that are marked as false positives in the causal hypergraph in multiple consecutive rounds, and extract the modality source label combination features of the false positive causal edges; Based on the modality source label combination characteristics, identify the modality combination types that frequently generate false alarms; In the next round of incremental causality testing, the significance threshold for causality testing is increased for node pairs corresponding to the modality combination types that frequently generate false alarms.