Causal modeling method and device in computing power network, electronic equipment and storage medium

By constructing a cross-domain, multi-level dynamic causal graph, the problems of fragmented fault location and agent decision-making misjudgment in computing power networks are solved, achieving efficient and accurate fault tracing and agent-assisted decision-making.

CN122019133APending Publication Date: 2026-05-12INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing fault location technologies for computing power networks lack cross-level collaborative capabilities, rely on traditional monitoring tools and static rule bases, resulting in severe fragmentation of fault location, high error rate in agent decision-making, and a lack of a unified cross-domain, multi-level dynamic causal graph framework.

Method used

A cross-domain, multi-level dynamic causal graph is constructed by collecting computing power network data, extracting multi-dimensional features, and filtering based on topology and fault correlation. The resulting cross-regional, multi-level dynamic causal graph includes basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features. An initial causal graph architecture is constructed and causal subgraphs are built layer by layer to determine the causal relationships between nodes and update the causal graph to adapt to network changes.

Benefits of technology

It improves the accuracy of fault location, enhances the interpretability and reliability of agent decision-making, achieves high efficiency and accuracy in cross-domain fault tracing, and supports root cause analysis of faults by operation and maintenance agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019133A_ABST
    Figure CN122019133A_ABST
Patent Text Reader

Abstract

The invention provides a causal modeling method and device in a computing power network, electronic equipment and a storage medium, and the method comprises the steps: collecting cross-domain multi-level computing power network data in the computing power network; multi-dimensional features in the computing power network data are extracted, screening is carried out based on the association degree of the multi-dimensional features and faults, and multi-dimensional target features are obtained; and constructing a cross-domain multi-level dynamic causal graph based on a computing power network topology structure and the multi-dimensional target features. The invention provides a scheme for constructing a unified cross-domain multi-level dynamic causal graph adaptive to the computing power network, and based on the dynamic causal graph, core pain points such as difficulty in cross-domain fault tracing, weak feature association and inaccurate fault positioning in the computing power network can be solved, and the fault positioning accuracy in the complex computing power network is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing network technology, and in particular to a causal modeling method, apparatus, electronic device, and storage medium in computing networks. Background Technology

[0002] With the advancement of integrated computing networks, multiple layers exist across geographical regions, including physical, virtual, scheduling, application, and service layers. The objects and fault modes at each layer differ significantly, and cross-link correlations exist, posing a major challenge to the rapid location and handling of faults in computing networks. Existing fault location technologies for computing networks mainly rely on traditional monitoring tools, static rule bases, and fragmented indicator analysis, focusing on handling single-domain problems and lacking cross-dimensional collaborative capabilities, resulting in severe fragmentation in fault location.

[0003] Meanwhile, the rapid development of intelligent agent technology has brought new directions to the automated and intelligent operation and maintenance of computing networks. However, the interpretability and reliability of intelligent agents have hidden dangers, and there is an urgent need for a solution that can complement intelligent agents and provide them with auxiliary fault location decision-making. At present, there is no unified causal modeling framework for computing networks, resulting in a high error rate in the decision-making of operation and maintenance intelligent agents. Summary of the Invention

[0004] This invention provides a causal modeling method, apparatus, electronic device, and storage medium in computing power networks to address the deficiency in the prior art of lacking a unified cross-level dynamic causal graph adapted to computing power networks.

[0005] This invention provides a causal modeling method in a computing power network, comprising: Collect data from multi-domain, multi-level computing power networks within the computing power network; Multi-dimensional features are extracted from the computing power network data, and the multi-dimensional features are filtered based on their correlation with faults to obtain multi-dimensional target features; Based on the computing power network topology and the multi-dimensional target features, a cross-domain, multi-level dynamic causal graph is constructed.

[0006] According to a causal modeling method in a computing power network provided by the present invention, the step of constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features includes: Based on the computing power network topology, an initial causal graph architecture is constructed, and causal subgraphs are constructed hierarchically based on the initial causal graph architecture. Based on the multi-dimensional target features, the nodes in each of the causal subgraphs are determined; Based on the multi-dimensional target features, a first causal relationship is determined between nodes in the causal subgraph of the same layer, and a second causal relationship is determined between nodes in the causal subgraph of different layers. Based on the multi-dimensional target features, the first causal relationship, and the second causal relationship, the dynamic causal graph is constructed.

[0007] According to a causal modeling method in a computing power network provided by the present invention, the step of extracting multi-dimensional features from the computing power network data and filtering them based on the correlation between the multi-dimensional features and faults to obtain multi-dimensional target features includes: Extract multi-dimensional features from the computing power network data; wherein, the multi-dimensional features include basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features; Calculate the maximum information coefficient of each of the multi-dimensional features, and perform correlation filtering based on the maximum information coefficient to obtain the multi-dimensional target features.

[0008] According to the causal modeling method in a computing power network provided by the present invention, after constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: If a causal change triggering event is detected, the dynamic causal graph is updated based on the causal change triggering event.

[0009] According to the causal modeling method in a computing power network provided by the present invention, the step of collecting multi-level, cross-domain computing power network data includes: Based on the computing power network topology, collect the computing power network data in the computing power network; The computing network data is time-series calibrated.

[0010] According to the causal modeling method in a computing power network provided by the present invention, after constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: Construct the target hash index for each node in the dynamic causal graph; wherein the target hash index includes a region code, a location code, a hierarchy code, a category code, and an object identifier.

[0011] According to the causal modeling method in a computing power network provided by the present invention, after constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: Obtain fault information; Based on the fault information, root cause localization is performed in the dynamic cause-effect graph to obtain the fault root cause results.

[0012] The present invention also provides a causal modeling device in a computing power network, comprising: The acquisition module is configured to acquire multi-domain, multi-level computing network data in the computing network; The extraction module is configured to extract multi-dimensional features from the computing power network data and filter them based on the correlation between the multi-dimensional features and the faults to obtain multi-dimensional target features; The first construction module is configured to construct a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the causal modeling method in any of the above-described computing power networks.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the causal modeling method in the computing power network as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a causal modeling method in any of the above-described computing power networks.

[0016] This invention provides a causal modeling method, apparatus, electronic device, and storage medium for computing power networks. It collects cross-regional, multi-level computing power network data, providing foundational data for subsequent feature construction and causal graph modeling. Multi-dimensional features are extracted from the computing power network data, and filtering is performed based on the correlation between these features and faults to obtain multi-dimensional target features associated with faults. Based on the computing power network topology and multi-dimensional target features, a dynamic causal graph representing cross-regional, multi-level scenarios is constructed. By constructing a unified, cross-domain, multi-level dynamic causal graph adapted to computing power networks, core pain points such as difficulty in tracing cross-domain faults, weak feature correlations, and inaccurate fault location in computing power networks can be addressed, significantly improving the fault location accuracy in complex computing power networks. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1This is a flowchart illustrating the causal modeling method in the computing power network provided by the present invention.

[0019] Figure 2 This is a schematic diagram of the causal modeling device in the computing power network provided by the present invention.

[0020] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] Figure 1 This is a flowchart illustrating a causal modeling method in a computing power network according to an exemplary embodiment. Figure 1 As shown in an exemplary embodiment, the causal modeling method in a computing power network includes steps 110 to 130, which are described in detail below.

[0023] Step 110: Collect multi-level, cross-domain computing network data in the computing network.

[0024] In this embodiment of the invention, computing power network data that spans multiple regions and is located at multiple levels is collected to provide basic data for subsequent feature construction and causal graph modeling.

[0025] Step 120: Extract multi-dimensional features from the computing power network data, and filter them based on the correlation between the multi-dimensional features and the faults to obtain multi-dimensional target features.

[0026] In this embodiment of the invention, multi-dimensional features are extracted from the computing power network data, and the multi-dimensional features are filtered based on their correlation with the fault to obtain multi-dimensional target features associated with the fault.

[0027] Step 130: Based on the computing power network topology and the multi-dimensional target features, construct a cross-domain, multi-level dynamic causal graph.

[0028] In this embodiment of the invention, a dynamic causal graph representing cross-regional and multi-level characteristics is constructed based on the computing power network topology and multi-dimensional target features.

[0029] In this embodiment of the invention, computing power network data spanning multiple regions and levels is collected to provide foundational data for subsequent feature construction and causal graph modeling. Multi-dimensional features are extracted from the computing power network data, and filtering is performed based on the correlation between these features and faults to obtain multi-dimensional target features associated with faults. Based on the computing power network topology and multi-dimensional target features, a dynamic causal graph representing cross-regional and multi-level faults is constructed. By constructing a cross-domain, multi-level dynamic causal graph, the core pain points of cross-domain fault tracing, weak feature correlation, and inaccurate fault location in computing power networks are addressed, significantly improving the fault location accuracy in complex computing power networks. Furthermore, it can be combined with existing intelligent agent technology to assist in the root cause analysis of faults by the operation and maintenance intelligent agent, improving the interpretability and reliability of the agent's decisions.

[0030] In an exemplary embodiment of the present invention, the collection of multi-domain, multi-level computing network data in the computing network includes: Based on the computing power network topology, collect the computing power network data in the computing power network; The computing network data is time-series calibrated.

[0031] In this embodiment of the invention, firstly, based on the computing power network topology, multi-level computing power network data across domains is collected. The computing power network includes a physical layer, a virtual layer, a scheduling layer, a platform layer, and an application layer, with each layer having corresponding computing power network data.

[0032] Specifically, the physical layer mainly includes physical devices such as servers, routers, switches, firewalls, and storage, which carry virtual resources or run applications directly. Related computing power network data includes CPU / GPU utilization, memory / video memory utilization, storage size, network bandwidth, network latency, jitter, packet loss rate, BGP status, number of connections, and operation and maintenance logs.

[0033] The virtualization layer mainly includes virtual resources such as virtual machines / cloud hosts, containers, and virtual networks. Related computing power and network data involve the CPU / GPU, memory / video memory, and hard disk of virtual machines, the CPU / GPU / memory quotas and storage volumes of containers, as well as running status and log data.

[0034] The scheduling layer is mainly responsible for scheduling computing tasks and computing network resources. Its computing network data mainly involves scheduling queues, scheduling policies, scheduling logs, etc., such as container migration, scaling up and down, and cross-domain scheduling paths.

[0035] The platform layer primarily provides the basic environment for application operation, such as distributed computing (Spark), AI training and promotion (TensorFlow), message queues (Kafka), databases, and other middleware. Its computing power network data involves running status, resource consumption, running logs, etc.

[0036] The application layer mainly consists of various business applications running in the computing power network, such as cloud-native applications, microservice applications, big data applications, and AI applications. Its computing power network data includes running status, resource consumption, running logs, call volume, success rate, response time, and error rate.

[0037] Timing calibration addresses the inconsistency in the original granularity of data from various computing power networks. For example, the physical layer data granularity is at the millisecond level (100-500 milliseconds), while the application layer data granularity is at the second or minute level (5 seconds-1 minute). Here, low-granularity computing power network data is linearly interpolated and upsampled, while high-granularity computing power network data is downsampled using a sliding window average, unifying them to a target granularity of 1 second, thereby providing reliable data for causal discovery.

[0038] In an exemplary embodiment of the present invention, the step of extracting multi-dimensional features from the computing power network data and filtering them based on the correlation between the multi-dimensional features and faults to obtain multi-dimensional target features includes: Extract multi-dimensional features from the computing power network data; wherein, the multi-dimensional features include basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features; Calculate the maximum information coefficient of each of the multi-dimensional features, and perform correlation filtering based on the maximum information coefficient to obtain the multi-dimensional target features.

[0039] In this embodiment of the invention, a multi-dimensional feature system is constructed, including basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features. Various features from this feature system are extracted from the collected computing power network data.

[0040] Basic dimensional features are single-dimensional raw features that have been standardized, such as continuous or discrete values ​​like CPU utilization, bandwidth utilization, interface call latency, and business error rate. The computing power network data corresponding to the basic dimensional features are standardized to eliminate differences in units and ranges while preserving distribution characteristics.

[0041] Specifically, for numerical features that approximate a normal distribution, such as CPU utilization and interface response latency, Z-score is used for standardization. ,in, These are the original feature values, i.e., the specific computing power network data. This is the historical sample mean of this feature, which corresponds to the historical computing power network data collected within the historical time range. The standard deviation of this feature is the historical sample standard deviation, i.e., the standard deviation of historical computing power network data. The standardized data eliminates the difference in dimensions while preserving the relative fluctuation trend of the data.

[0042] For features with clearly defined value ranges, such as switch packet loss rate (0%-100%) and service success rate (0%-100%), deviation standardization (Min-Max) is used for standardization. ,in, The minimum value of the characteristic. This represents the maximum value of the characteristic.

[0043] For discrete non-numerical state features, such as device start / stop status, scheduling strategy type, and container health status, label encoding is used for standardization. For example, a 1 is used to identify device startup and a 0 is used to identify device shutdown.

[0044] For text-based features, such as unstructured data like fault logs and routing logs, the BERT pre-trained model is used for standardization. The BERT pre-trained model extracts the core semantics from the text-based computing network data and transforms them into quantifiable numerical vectors, providing a basis for constructing causal edges.

[0045] Topological association features mainly include object association relationships, cross-regional adjacency attributes, and path propagation length, which characterize the spatial association of objects. Temporal evolution features include sliding window trend slope, abrupt change amplitude, and fluctuation standard deviation, to capture dynamic changes.

[0046] Cross-dimensional coupling features refer to features obtained by integrating five dimensions: computing power, network, scheduling, service, and business. These features accurately quantify collaborative relationships, thus overcoming the limitations of traditional single-dimensional analysis. During the integration of these five dimensions, dynamic adjustments are achieved through weighted fusion and the introduction of scenario and time decay factors, effectively addressing the tidal fluctuations in computing power networks. The calculation formula for cross-dimensional coupling features is shown below: ; in, Standardize features for computing power (e.g., GPU utilization 0.65); Standardize network-level features (e.g., bandwidth utilization of 0.55); Standardize features for the scheduling dimension (e.g., scheduling matching degree 0.8); Standardize features for service dimensions (e.g., interface access latency of 1.6). Standardize features for business dimensions (e.g., transaction failure rate of 2.4%). As a scenario factor, it is dynamically adjusted based on the business type to highlight the weight of core business dimensions, such as setting the weight of financial transaction scenarios to 1.2 and AI training scenarios to 0.9; This is a time decay factor, which is adjusted based on the timeliness of the characteristics to reduce interference from historical data. For example, it is set to 1.0 for data from the last minute and 0.7 for data from 5 minutes ago. , , , , These are the weighting coefficients, and The initial weights can be obtained from historical sample data and continuously optimized.

[0047] When calculating cross-dimensional coupling features, the first step is to filter the strongly correlated indicators of each dimension by using the maximum information coefficient. For example, in the network dimension, only core indicators such as bandwidth usage and packet loss rate are retained. Then, the weighted average of multiple indicators within the same dimension is taken to generate a single-dimensional feature, such as the network dimension feature = 0.6 × bandwidth usage + 0.4 × packet loss rate. Next, the data is standardized by performing the corresponding standardization method according to the dimension type. Finally, the standardized data is substituted into the calculation formula to obtain the final coupling feature value.

[0048] The basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features extracted from the computing power network data will be filtered for correlation in order to remove irrelevant and redundant information and retain features with strong causal correlation related to faults, i.e., multi-dimensional target features, and thus form a core feature set.

[0049] Specifically, the correlation strength between each multi-dimensional feature and the fault is quantified using the Maximum Information Coefficient (MIC), eliminating multi-dimensional features with weak or no correlation to the fault, such as physical layer environmental temperature. MICs are calculated for four types of features: basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features, retaining multi-dimensional features with a MIC greater than 0.6. Then, Granger causality checks are used to verify the causal directionality among the retained multi-dimensional features, ensuring that the filtered features have a clear fault propagation logic. A causal feature correlation matrix is ​​constructed, recording the causal direction and strength between features (strength quantified by the F-statistic). For example, the F-value for bandwidth utilization causing service interface latency is 32.1, and the Granger causality p < 0.008. The F-statistic is an F-test used to measure causal relationships; the larger the value, the stronger the causal relationship.

[0050] In an exemplary embodiment of the present invention, the step of constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features includes: Based on the computing power network topology, an initial causal graph architecture is constructed, and causal subgraphs are constructed hierarchically based on the initial causal graph architecture. Based on the multi-dimensional target features, the nodes in each of the causal subgraphs are determined; Based on the multi-dimensional target features, a first causal relationship is determined between nodes in the causal subgraph of the same layer, and a second causal relationship is determined between nodes in the causal subgraph of different layers. Based on the multi-dimensional target features, the first causal relationship, and the second causal relationship, the dynamic causal graph is constructed.

[0051] In this embodiment of the invention, based on the hierarchical division of the physical layer, virtual layer, scheduling layer, platform layer, and application layer, the nodes in the dynamic cause-effect graph are clearly defined. Each node represents a monitoring indicator or entity status. For example, nodes in the physical layer include GPU utilization, memory utilization, and switch port packet loss rate.

[0052] An initial cause-effect graph architecture is established based on the computing network topology. This architecture includes physical connectivity, deployment dependencies, and service call chains. Physical connectivity refers to the connections between devices determined by the network topology, such as the connection between a server and a switch / router. Deployment dependencies refer to the attribution relationships between virtual machines / containers and physical machines, and the deployment relationships between services and virtual machines / containers, determined by the deployment topology. Service call chains refer to the call dependencies between services, such as the order query service depending on the authentication service, and the order service depending on the MySQL database service.

[0053] Based on the initial causal graph architecture, causal subgraphs are constructed layer by layer, with corresponding causal subgraphs for the physical layer, virtual layer, scheduling layer, platform layer, and application layer. Multi-dimensional target features at the same level are mapped to nodes, and internal edges within each level are constructed based on the first causal relationship of these multi-dimensional target features. Each causal subgraph independently carries the fault propagation logic for that level, while reserving interfaces for cross-level associations. For example, the physical layer focuses on the state and performance correlation of physical hardware, and the corresponding multi-dimensional target features are basic features of computing power and network. Therefore, the selected multi-dimensional target features of the physical layer are split into nodes according to hardware entity + indicator type. Each node is bound to a unique causal feature identifier and corresponding feature data, mapped according to the initial causal graph architecture, and internal edges are constructed based on the multi-dimensional target features of the physical layer.

[0054] Based on the fault propagation law in computing power networks, cross-level edges are generated based on cross-level feature causal chains. This means generating corresponding edges between multi-dimensional target features with causal relationships at different levels. Each edge is designed to include a quintuple of source identifier, target identifier, cross-domain attribute, time delay, and causal strength.

[0055] The cross-domain attribute marks the geographical combination of associated nodes, providing a basis for cross-regional fault location. Time delay is obtained based on the fault propagation lag time of time-series data. For example, after the network bandwidth of the physical layer switch fluctuates, the application layer service delay changes. If the average lag time between the two is 8 time steps, then the time delay can be determined to be 8.

[0056] Causality strength is calculated based on the MIC values ​​and F-statistics of the features from the upper and lower layers. The corresponding calculation formula is as follows: Causal strength = ; Where a and b are weighting coefficients, For two nodes at different levels The upper-level feature is the effect in a causal relationship. That is, in two nodes at different levels with a causal relationship, the node pointed to by the edge is the upper-level feature. For example, if bandwidth occupancy causes service interface latency, then the upper-level feature is service interface latency, and its MIC is calculated to be 0.99. The cross-level F It is 42.5. If both are 0.5, then the causal strength = 0.88 × 0.5 + 42.5 / 100 × 0.5 = 0.6525.

[0057] The cross-layer edges are integrated in the order of physical layer, virtual layer, scheduling layer, platform layer and application layer to form a complete fault propagation link. Among them, the virtual layer, scheduling layer, platform layer, etc. can be optional. For example, a CRM application deployed directly on a physical server only includes the physical layer and the application layer.

[0058] The causal subgraph and causal edges are optimized collaboratively by deleting intra-level edges with a causal strength of less than 0.4 and cross-level edges with a causal strength of less than 0.5, merging duplicate edges, and binding nodes in the causal subgraph to the target hash index. The associated edges are quickly located through the 16-bit unique identifier of the node, thereby improving the inference efficiency and accuracy of the dynamic causal graph.

[0059] In an exemplary embodiment of the present invention, after constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: Construct the target hash index for each node in the dynamic causal graph; wherein the target hash index includes a region code, a location code, a hierarchy code, a category code, and an object identifier.

[0060] In this embodiment of the invention, the target hash index refers to a hash index structure designed to improve inference and update efficiency. It mainly includes: a first-level index consisting of a region code + location code + hierarchy code, which allows for quick location of resources at the target region-data center node-hierarchy level during retrieval. A second-level index consisting of a region code + location code + hierarchy code + category code further narrows the retrieval scope to categories, such as computing power, network, etc. A third-level index, namely the target hash index, consisting of a region code + location code + hierarchy code + category code + object identifier. The target hash index has 16 bits, and the key value directly associates all features bound to the node.

[0061] Specifically, the generated target hash index is used to ensure the globally unique traceability of cross-regional and cross-level data. It consists of a 2-digit region code, a 4-digit location code, a 1-digit hierarchy code, a 1-digit category code, and an 8-digit object identifier. The region code identifies the region where the computing power resource is located, using the provincial code of that region. The location code is the encoding of objects such as data centers, server rooms, and computing power providers, mainly indicating the specific computing power affiliation, ranging from 0000 to 9999. The hierarchy code indicates the level within the computing power network, namely 1-Physical Layer, 2-Virtual Layer, 3-Scheduling Layer, 4-Platform Layer, and 5-Application Layer. The category code refers to a further division of computing power network objects, including computing power class, network class, scheduling class, application class, and business class, namely 1-Computing Power Class, 2-Network Class, 3-Scheduling Class, 4-Service Class, and 5-Business Class. Object identifiers are used to represent specific computing network objects. For example, a physical server is identified by 01000001. The design rules for object identifiers can be customized, such as the first two digits identifying the device model, like 01 identifying a server, 02 identifying a switch, etc. By uniformly encoding nodes, rapid data traceability can be achieved. For example, 3700111201000001 identifies the physical server with ID 000001 in the data center with identifier 11 in Shandong.

[0062] In an exemplary embodiment of the present invention, after constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: If a causal change triggering event is detected, the dynamic causal graph is updated based on the causal change triggering event.

[0063] In this embodiment of the invention, when causal change-triggered events such as changes in the computing power network topology, sudden changes in feature associations, and changes in business scenarios are detected in real time, the dynamic causal graph is updated incrementally in a timely manner.

[0064] Specifically, it quickly locates the changed nodes and related edges corresponding to the causal change triggering events, analyzes the related causal edges of the changed nodes, performs operations such as updating causal strength and adding / deleting related edges, and synchronously updates the target hash index and causal subgraph after the operations are completed to ensure that the index and causal subgraph are consistent. Figure 1To ensure the effectiveness of the update, local consistency checks and small-sample inference verification are employed. Local consistency checks verify whether the causal directions of the edges connecting changed nodes conform to the physical laws of the computing network and whether the causal strength is within a reasonable range. Small-sample inference verification uses the latest 10 fault samples to infer the updated local subgraph; if the root cause localization accuracy is within a preset range (e.g., >95%), the update is confirmed as effective; otherwise, a rollback mechanism is triggered.

[0065] This invention provides an embodiment of the invention that uses real-time topology perception of the computing power network to perform incremental updates of the causal graph in order to adapt to the highly dynamic characteristics of the computing power network.

[0066] In an exemplary embodiment of the present invention, after constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: Obtain fault information; Based on the fault information, root cause localization is performed in the dynamic cause-effect graph to obtain the fault root cause results.

[0067] In this embodiment of the invention, based on the regional, location, and hierarchical identifiers of the fault information, the abnormal node in the dynamic causal graph is quickly located by the target hash index, and all its incoming and outgoing edges are extracted with the abnormal node as the center to form an abnormal association subgraph. The number of nodes in the abnormal association subgraph is controlled to be within 50.

[0068] The process recursively traces the predecessor nodes of the abnormal nodes along the causal direction to construct a complete fault propagation chain. Valid paths are then filtered based on the causal strength and time delay of the edges. For example, a causal strength threshold is set, retaining only edges that meet the threshold condition. Time consistency verification is used for delay, ensuring that the occurrence time conforms to temporal logic. For instance, if a service interface delay occurs at 10:05, tracing back to bandwidth congestion at 10:02 is logically consistent, but if the bandwidth congestion occurs at 10:06, it is excluded. This ultimately forms a candidate causal chain from the physical layer root cause, intermediate layer propagation, to the business layer anomaly, thus yielding candidate root causes.

[0069] Based on the Bayesian probability model, combined with historical fault samples and real-time feature data, the confidence level of each candidate root cause is calculated. The candidate root causes are sorted in descending order according to the confidence level. Those with a confidence level greater than or equal to 85% are marked as high-confidence root causes and used as the result of the fault root cause, which can directly trigger operation and maintenance actions.

[0070] This invention provides an embodiment that avoids full graph traversal and ensures inference efficiency by performing root cause localization in the anomaly correlation subgraph. Simultaneously, it can visually display the dynamic causal graph of the computing power network across multiple domains and levels, allowing operations and maintenance personnel to intuitively grasp the fault evolution process, such as reverse root cause tracing and forward propagation verification.

[0071] The following describes the causal modeling apparatus in a computing power network provided by the present invention. The causal modeling apparatus in a computing power network described below can be referred to in correspondence with the causal modeling method in a computing power network described above. It should be noted that the apparatus provided in the embodiments below and the method provided in the embodiments above belong to the same concept, and the specific way in which each module and unit performs its operation has been described in detail in the method embodiments, and will not be repeated here.

[0072] In one exemplary embodiment of the present invention, please refer to Figure 2 , Figure 2 This is an exemplary embodiment of a causal modeling apparatus in a computing network, comprising the following modules.

[0073] The acquisition module 210 is configured to acquire multi-level, cross-domain computing network data in the computing network; Extraction module 220 is configured to extract multi-dimensional features from the computing power network data and filter them based on the correlation between the multi-dimensional features and the faults to obtain multi-dimensional target features; The first construction module 230 is configured to construct a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features.

[0074] In an exemplary embodiment of the present invention, the first construction module 230 includes: The first construction submodule is configured to construct an initial causal graph architecture based on the computing power network topology, and to construct causal subgraphs hierarchically based on the initial causal graph architecture; The first determining submodule is configured to determine the nodes in each of the causal subgraphs based on the multi-dimensional target features; The second determining submodule is configured to determine, based on the multi-dimensional target features, a first causal relationship between nodes in the causal subgraph of the same layer, and a second causal relationship between nodes in the causal subgraph of different layers; The second construction submodule is configured to construct the dynamic causal graph based on the multi-dimensional target features, the first causal relationship, and the second causal relationship.

[0075] In an exemplary embodiment of the present invention, the extraction module 220 includes: The extraction submodule is configured to extract multi-dimensional features from the computing power network data; wherein, the multi-dimensional features include basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features; The calculation submodule is configured to calculate the maximum information coefficient of each of the multi-dimensional features, and perform correlation filtering based on the maximum information coefficient to obtain the multi-dimensional target features.

[0076] In an exemplary embodiment of the present invention, the causal modeling module in the computing power network further includes: The update module is configured to update the dynamic causal graph based on a detected causal change trigger event.

[0077] In an exemplary embodiment of the present invention, the acquisition module 210 includes: The acquisition submodule is configured to acquire computing network data in the computing network according to the computing network topology. The timing calibration submodule is configured to perform timing calibration on the computing power network data.

[0078] In an exemplary embodiment of the present invention, the causal modeling apparatus in the computing power network further includes: The second construction module is configured to construct the target hash index of each node in the dynamic causal graph; wherein the target hash index includes a region code, a location code, a hierarchy code, a category code, and an object identifier.

[0079] In an exemplary embodiment of the present invention, the causal modeling apparatus in the computing power network further includes: The acquisition module is configured to acquire fault information; The root cause localization module is configured to perform root cause localization in the dynamic cause-effect graph based on the fault information to obtain the root cause result of the fault.

[0080] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 310, a communications interface 320, a memory 330, and a communication bus 340. The processor 310, communications interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions from the memory 330 to execute a causal modeling method in the computing power network. This method includes: collecting multi-level, cross-domain computing power network data. Multi-dimensional features are extracted from the computing power network data, and the multi-dimensional features are filtered based on their correlation with faults to obtain multi-dimensional target features; Based on the computing power network topology and the multi-dimensional target features, a cross-domain, multi-level dynamic causal graph is constructed.

[0081] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0082] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the causal modeling method in the computing power network provided by the above methods, the method including: collecting multi-level computing power network data across domains in the computing power network; Multi-dimensional features are extracted from the computing power network data, and the multi-dimensional features are filtered based on their correlation with faults to obtain multi-dimensional target features; Based on the computing power network topology and the multi-dimensional target features, a cross-domain, multi-level dynamic causal graph is constructed.

[0083] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the causal modeling method in the computing power network provided by the above methods, the method comprising: collecting computing power network data across multiple domains in the computing power network; Multi-dimensional features are extracted from the computing power network data, and the multi-dimensional features are filtered based on their correlation with faults to obtain multi-dimensional target features; Based on the computing power network topology and the multi-dimensional target features, a cross-domain, multi-level dynamic causal graph is constructed.

[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0085] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A causal modeling method in a computing power network, characterized in that, include: Collect data from multi-domain, multi-level computing power networks within the computing power network; Multi-dimensional features are extracted from the computing power network data, and the multi-dimensional features are filtered based on their correlation with faults to obtain multi-dimensional target features; Based on the computing power network topology and the multi-dimensional target features, a cross-domain, multi-level dynamic causal graph is constructed.

2. The causal modeling method in a computing network according to claim 1, characterized in that, The construction of a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features includes: Based on the computing power network topology, an initial causal graph architecture is constructed, and causal subgraphs are constructed hierarchically based on the initial causal graph architecture. Based on the multi-dimensional target features, the nodes in each of the causal subgraphs are determined; Based on the multi-dimensional target features, a first causal relationship is determined between nodes in the causal subgraph of the same layer, and a second causal relationship is determined between nodes in the causal subgraph of different layers. Based on the multi-dimensional target features, the first causal relationship, and the second causal relationship, the dynamic causal graph is constructed.

3. The causal modeling method in a computing network according to claim 1, characterized in that, The process involves extracting multi-dimensional features from the computing power network data and filtering them based on the correlation between these multi-dimensional features and faults to obtain multi-dimensional target features, including: Extract multi-dimensional features from the computing power network data; wherein, the multi-dimensional features include basic dimensional features, topological correlation features, temporal evolution features, and cross-dimensional coupling features; Calculate the maximum information coefficient of each of the multi-dimensional features, and perform correlation filtering based on the maximum information coefficient to obtain the multi-dimensional target features.

4. The causal modeling method in a computing network according to claim 1, characterized in that, After constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: If a causal change triggering event is detected, the dynamic causal graph is updated based on the causal change triggering event.

5. The causal modeling method in a computing network according to claim 1, characterized in that, The data collected from the cross-domain, multi-level computing power network includes: Based on the computing power network topology, collect the computing power network data in the computing power network; The computing network data is time-series calibrated.

6. The causal modeling method in a computing power network according to any one of claims 1 to 5, characterized in that, After constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: Construct the target hash index for each node in the dynamic causal graph; wherein the target hash index includes a region code, a location code, a hierarchy code, a category code, and an object identifier.

7. The causal modeling method in a computing power network according to any one of claims 1 to 5, characterized in that, After constructing a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features, the method further includes: Obtain fault information; Based on the fault information, root cause localization is performed in the dynamic cause-effect graph to obtain the fault root cause results.

8. A causal modeling device in a computing power network, characterized in that, include: The acquisition module is configured to acquire multi-domain, multi-level computing network data in the computing network; The extraction module is configured to extract multi-dimensional features from the computing power network data and filter them based on the correlation between the multi-dimensional features and the faults to obtain multi-dimensional target features; The first construction module is configured to construct a cross-domain, multi-level dynamic causal graph based on the computing power network topology and the multi-dimensional target features.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the causal modeling method in the computing power network as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the causal modeling method in the computing power network as described in any one of claims 1 to 7.