Digital twin modeling and resource dynamic scheduling system and method for cloud network integration
By using digital twin modeling and dynamic resource scheduling systems, the problems of inefficient fault detection and rigid resource allocation in cloud-network converged systems have been solved, enabling real-time health assessment and forward-looking operation and maintenance decisions, thereby improving operation and maintenance efficiency and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-07
AI Technical Summary
Existing cloud-network converged systems are inefficient in fault detection and localization, lack forward-looking operation and maintenance decisions, have rigid resource allocation, insufficient adaptability to intelligent operation and maintenance, and lack a unified digital mirror, resulting in low operation and maintenance efficiency and decision-making that relies on human experience.
A digital twin modeling and resource dynamic scheduling system is adopted. Multi-source heterogeneous data is collected through data acquisition devices, and the twin processing module performs multi-source side information analysis to generate real-time health status assessment and operation and maintenance decision instructions. Differentiated modeling and optimization are carried out by combining a hybrid intelligent model with a Bayesian learning network.
It enables accurate real-time health status assessment and operation and maintenance decision-making for cloud network systems, reduces the risk of false alarms and missed alarms, improves the scientific nature of operation and maintenance and the efficiency of emergency response, supports long-term trend analysis and prediction, and improves resource utilization.
Smart Images

Figure CN121807467A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, specifically to a digital twin modeling and dynamic resource scheduling system and method for cloud-network convergence. Background Technology
[0002] With the deep integration of cloud computing and communication networks, modern information infrastructure is evolving into a highly complex and dynamically changing cloud-network integrated system. While this evolution enhances business flexibility, it also presents unprecedented challenges to its operation and maintenance management. The complexity of the system is growing exponentially. The massive number of network devices, virtualized resources, microservice instances, and their intricate dependencies make traditional operation and maintenance models, which rely on manual experience, static thresholds, and isolated monitoring tools, increasingly inadequate. Current operation and maintenance practices generally suffer from the following prominent problems: First, fault detection and localization are lagging and inefficient. Existing systems mainly rely on passive, single-dimensional indicator thresholds for alarms, lacking deep integration and correlation analysis of multi-source heterogeneous data (such as performance indicators, logs, topology relationships, and configuration information). When complex cross-domain and cross-layer faults occur, alarm storms occur frequently, root causes are difficult to trace quickly, and mean time to repair (MTTR) is long, seriously affecting business continuity. Second, operation and maintenance decisions lack foresight and scientific rigor. Traditional methods are essentially "response-based," only addressing issues after they occur, lacking the ability to use historical data and system models for trend prediction and risk prevention. Meanwhile, rigid resource allocation strategies fail to dynamically and precisely adjust based on real-time business load and system health status, leading to both low resource utilization and performance bottlenecks. Furthermore, intelligent operations and maintenance (O&M) lacks adaptability. Although AI technologies such as machine learning have been introduced into the O&M field, existing solutions often employ fixed, universal models and parameters, making it difficult to adapt to the differentiated characteristics of different components and business scenarios in the cloud network environment. An anomaly detection model that performs well on core databases may not be applicable to edge network devices, resulting in high false alarm and false negative rates, significantly diminishing the effectiveness of intelligent systems. Finally, the system status lacks a unified and evolving digital mirror. O&M personnel lack a digital copy that accurately reflects the real-time status of the physical system and fully records its historical evolution, making simulation verification, solution pre-running, and knowledge accumulation difficult, and the decision-making process still heavily relies on personal experience. Summary of the Invention
[0003] To achieve the above objectives, the present invention provides the following technical solution: a digital twin modeling and dynamic resource scheduling system for cloud-network convergence, comprising: At least one data acquisition device is configured to acquire real-time operational status data of multiple network devices, virtualization resources and service instances in the cloud network module, receive monitoring signals and performance indicators reflected by the service traffic and hardware status running in the cloud network module, and generate a structured data set characterizing the operational health status and topology connection relationship of the system components. A twin data storage device, configured to store multi-source side information representing the historical and real-time status of the cloud network module; and A twin processing module is configured to receive the structured data set from the at least one data acquisition device, receive the multi-source side information from the twin data storage, determine hyperparameters for an intelligent operation and maintenance analysis model based on the multi-source side information, and use the hyperparameters to process the structured data set through the intelligent operation and maintenance analysis model to generate real-time health status assessment and operation and maintenance decision instructions for the cloud network module. The multi-source side information includes one or more of the following: The static or dynamic digital topology and configuration model of the cloud network module. The historical or real-time monitoring visualization images of the cloud network module, or The previously generated digital twin or system state snapshot of the cloud network module.
[0004] Furthermore, the multi-source edge information includes multiple historical monitoring images from the cloud network module and multiple previously generated digital twin versions; The twin processing module is configured to: extract feature sequences of performance degradation and abnormal events by performing time-series analysis and pattern recognition on the multiple historical monitoring images; and construct a unified twin with time continuity and state evolution capabilities by fusing the multiple previously generated digital twin versions to support long-term operation and maintenance trend analysis and prediction.
[0005] Furthermore, the twin processing module generates detected abnormal event data by performing computer vision-based abnormal pattern and event detection on the multiple historical monitoring images of the cloud network module. The detected abnormal event data includes location markers for abnormal regions identified in the multiple images, event type classifications, and corresponding detection confidence scores.
[0006] Furthermore, the twin processing module generates a topological risk map, including predictions of potential failure points, performance bottlenecks, and impact ranges, based on the system behavior logic and fault propagation relationships contained in the multiple previously generated digital twin versions, through causal reasoning and graph computation.
[0007] Furthermore, the twin processing module, based on the detected abnormal event data and the topology risk map, performs comprehensive analysis through evidence fusion and probabilistic graphical models to generate a comprehensive and quantitative operation and maintenance status assessment map. This map identifies the instantaneous risk level and fault probability estimate of each logical unit or physical component in the cloud network module.
[0008] Furthermore, the twin processing module determines the hyperparameters for differentiated modeling and optimization of the diagnostic and prediction algorithms for different risk level areas or components in the intelligent operation and maintenance analysis model based on the comprehensive operation and maintenance situation assessment map.
[0009] Furthermore, the intelligent operation and maintenance analysis model is a hybrid intelligent model based on Bayesian learning networks, and the hyperparameters determined based on the comprehensive operation and maintenance situation assessment map include: A first index set is determined, which includes a set of identifiers of system components that are assessed as high-risk or have shown abnormalities in the operation and maintenance status assessment map; Determine a second index set, which is the complement of the first index set, containing component identifiers that are assessed as low risk or functioning properly; A first set of hyperparameters is determined for the components in the first index set, the first set of hyperparameters being configured to assign higher anomaly weights to the monitoring data of these components by the model, enable more complex feature analysis, and set more sensitive warning thresholds; and A second set of hyperparameters is determined for the components in the second index set. This second set of hyperparameters is configured to enable the model to employ a conventional monitoring strategy for these components with a high noise tolerance. When used by the intelligent operation and maintenance analysis model, the first set of hyperparameters is associated with a higher probability of fault warning and a deeper root cause tracing capability than the second set of hyperparameters.
[0010] Furthermore, at least one of the data acquisition devices includes an IoT sensor array consisting of software probes, log acquisition agents, and hardware sensors deployed on network nodes, servers, and cloud platforms, enabling real-time multi-dimensional data acquisition of traffic, latency, resource utilization, and energy consumption.
[0011] Furthermore, at least one of the data acquisition devices includes a simulation test engine that supports an active probing protocol for simulating user service flows and measuring end-to-end service performance, generating synthetic performance probe data as an important component of the structured dataset.
[0012] This invention also provides a digital twin modeling and dynamic resource scheduling method for cloud-network convergence, including: The monitoring data stream corresponding to the operating status signals reflected by one or more network components or services in the cloud network environment monitored by the at least one data acquisition device is received by at least one data acquisition device. The monitoring data stream is aggregated, cleaned, and structured by the at least one data acquisition device to generate a set of operational data indicating the health status and performance level of the component or service; The set of operational data is received by the twin processing module; The twin processing module receives multi-source side information from the twin data storage, the multi-source side information including at least one digital representation of the cloud network module; The twin processing module determines the hyperparameters for the intelligent operation and maintenance analysis model based on the multi-source side information; and The twin processing module uses the hyperparameters and the intelligent operation and maintenance analysis model to process the operational data set, thereby generating a real-time health diagnosis report and automated operation and maintenance action instructions for the cloud network module. The multi-source side information includes one or more of the following: At least one digital topology and business logic model of the cloud network module. At least one historical or real-time operation and maintenance monitoring view image of the cloud network module, or At least one previously generated and optimized digital twin instance of the cloud network module.
[0013] The multi-source edge information includes multiple historical operation and maintenance panoramic view images of the cloud network module and multiple evolved versions of digital twins.
[0014] The twin processing module generates fine-grained historical anomaly event feature data by performing deep learning-based pattern recognition and event segmentation on the multiple historical operation and maintenance panoramic view images of the cloud network module. The historical anomaly event feature data includes labeling information and confidence scores for identifying historical fault modes, occurrence locations, and evolution processes in the multiple images.
[0015] The twin processing module generates a predictive system state view containing potential future abnormal events and performance degradation trends by using time-series prediction and graph neural network analysis, based on the system state transition and fault association knowledge recorded in the multiple evolved versions of digital twins.
[0016] The twin processing module performs spatiotemporal correlation analysis and probability fusion based on the fine-grained historical abnormal event feature data and the predictive system status view to generate a dynamic operation and maintenance risk assessment map covering the entire system. This map assigns a quantified real-time risk index to each managed entity.
[0017] The twin processing module determines the hyperparameters for differentiated modeling and strategy configuration of system units in the intelligent operation and maintenance analysis model for different risk assessment levels based on the operation and maintenance risk assessment map. The intelligent operation and maintenance analysis model is a decision-making model based on the fusion of Bayesian inference and machine learning, and the hyperparameters determined based on the operation and maintenance risk assessment map include: Determine a first set of indexes containing all system unit indexes identified as high-risk or faulty states; Determine a second index set that serves as the complement of the first index set and contains all low-risk or normal state unit indices; Determine a first set of hyperparameters for the cells within the first index set, which enables the model to perform high-frequency sampling, deep packet inspection, and multi-dimensional correlation analysis; and A second set of hyperparameters is determined for the cells within the second index set. This set of parameters enables the model to employ low-frequency inspection and summary statistical analysis. In the process of generating operation and maintenance decisions, units that apply the first set of hyperparameters will trigger a more prioritized and rigorous handling process and resource pre-allocation action than units that apply the second set of hyperparameters.
[0018] The beneficial effects are as follows: By introducing multi-source side information, including historical digital twins, topology models, and monitoring images, as prior knowledge, and dynamically determining the hyperparameters of the intelligent operation and maintenance analysis model based on this, the model acquires context awareness and adaptive capabilities. This allows the model to customize its sensitivity, focus, and judgment thresholds for abnormal data according to the system's specific historical state, logical structure, and visual performance. For example, for areas with historically frequent failures, the system assigns them higher monitoring weights and more stringent abnormal conditions through hyperparameters. This parameter adaptation process based on multi-source side information transforms isolated alarm events into accurate risk alerts calibrated with prior knowledge and accompanied by probability assessment. The generated operation and maintenance decision instructions (such as isolation and expansion) thus integrate real-time data and historical experience, and are built on quantitative evidence and model reasoning, significantly reducing the risk of misoperation due to false alarms or missed alarms, and improving the effectiveness of emergency response and the scientific nature of decision-making.
[0019] By integrating multiple historical digital twins to construct a unified twin with temporal continuity and state evolution capabilities, the system can support long-term operational trend analysis and prediction. This unified twin, acting as a dynamic knowledge base, enables the system to perform deep time-series analysis and pattern recognition, extracting performance degradation feature sequences and fault cycle patterns from historical monitoring images. Based on this, the system can not only provide early warnings of immediate faults but also predict performance inflection points or potential compatibility issues caused by accumulated pressure within the next few hours or days. This predictive capability allows for proactive operational work, shifting from passively handling existing faults to actively preventing potential risks.
[0020] This technology leverages computer vision-based image analysis, graph computing-based topological reasoning, and probabilistic graphical model-based evidence fusion to achieve deep fusion and intelligent correlation analysis of multimodal operation and maintenance data. This approach enables the correlation and reasoning between surface anomalies identified from images and logical dependencies in the topological graph, thereby tracing back to root cause chains across domains and layers. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the system framework of the present invention; Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Reference Figure 1 and Figure 2This invention discloses a digital twin modeling and dynamic resource scheduling system for cloud-network convergence, comprising: at least one data acquisition device configured to acquire real-time operational status data of multiple network devices, virtualized resources, and service instances in a cloud-network module; receive monitoring signals and performance indicators reflected by service traffic and hardware status running in the cloud-network module; and generate a structured data set characterizing the operational health status and topological connectivity of the system components; a twin data storage device configured to store multi-source side information representing the historical and real-time status of the cloud-network module; and a twin processing module. The processing module is configured to receive the structured data set from the at least one data acquisition device, receive the multi-source side information from the twin data storage, determine hyperparameters for the intelligent operation and maintenance analysis model based on the multi-source side information, and use the hyperparameters to process the structured data set through the intelligent operation and maintenance analysis model to generate real-time health status assessment and operation and maintenance decision instructions for the cloud network module. The multi-source side information includes one or more of the following: a static or dynamic digital topology and configuration model of the cloud network module, a historical or real-time monitoring visualization image of the cloud network module, or a previously generated digital twin or system status snapshot of the cloud network module.
[0024] In some embodiments, the data acquisition device includes a distributed lightweight data acquisition agent and a centralized data aggregation gateway. The lightweight data acquisition agent is embedded in the host operating system kernel, virtual machine monitor, or container runtime environment and is configured to acquire low-level fine-grained performance counters and event tracking data, including CPU instruction cycles, memory page faults, I / O wait queue depths, and network protocol stack buffer states, at the millisecond level. The centralized data aggregation gateway is configured to receive raw data streams from each agent through an asynchronous message queue, perform data normalization, unit unification, and outlier filtering preprocessing based on a sliding time window, and reorganize the processed data into a structured data set with a unified timestamp, device identifier, and multi-dimensional indicator labels according to a preset spatiotemporal data model.
[0025] The twin data storage is implemented as a hybrid storage engine supporting multi-version concurrency control and spatiotemporal data indexing. This storage engine includes: a relational database module for storing static digital topology models, device configuration metadata, and business logic relationships; a time-series database module for efficiently storing and compressing massive amounts of timestamped monitoring indicator sequence data; and a graph database module for dynamically storing and traversing a dynamic knowledge graph composed of component dependencies, data flow patterns, and fault propagation chains. The twin processing module can perform cross-database queries on the above three modules through a unified query interface to reconstruct a complete system state view at any historical moment or time period.
[0026] The twin processing module is a microservice architecture consisting of a data access layer, a model calculation layer, and a decision output layer. The data access layer is responsible for real-time alignment and feature engineering of the received structured data set and multi-source edge information to generate standardized feature vectors for use by the model calculation layer. The model calculation layer encapsulates the intelligent operation and maintenance analysis model and has a built-in hyperparameter manager, which dynamically loads or calculates the optimal hyperparameter combination based on multi-source edge information. The decision output layer receives the calculation results of the model and, based on a predefined strategy library, converts the results into specific executable operation and maintenance instructions. These instructions include, but are not limited to, automatically triggering alarms, executing elastic scaling scripts, issuing network policy changes, or generating repair plan work orders awaiting manual review.
[0027] In some embodiments, the multi-source edge information includes multiple historical monitoring images from the cloud network module and multiple previously generated digital twin versions; the twin processing module is configured to: extract feature sequences of performance degradation and abnormal events by performing time-series analysis and pattern recognition on the multiple historical monitoring images; and construct a unified twin with time continuity and state evolution capabilities by fusing the multiple previously generated digital twin versions to support long-term operation and maintenance trend analysis and prediction.
[0028] In some embodiments, temporal analysis and pattern recognition are performed on multiple historical surveillance images, specifically including: extracting high-level feature vectors from each historical surveillance image using a pre-trained convolutional neural network; subsequently, inputting the chronologically ordered sequence of feature vectors into a long short-term memory network or a temporal convolutional network to learn the dynamic evolution pattern of system performance indicators in the image space; the network is trained to identify two types of feature sequences: one is a "gradual pattern" representing slow performance degradation, characterized by a continuous shift in the direction of the feature vector in the latent space; the other is a "mutation pattern" representing sudden anomalies, characterized by significant jumps in the feature vector at adjacent time points; the extracted feature sequences are labeled with their pattern type, start and end times, and severity scores.
[0029] In some embodiments, a unified twin is constructed by fusing multiple previously generated digital twin versions. Specifically, this includes: defining a graph-based representation of twin version differences, where each version is a graph containing nodes, edges, and node / edge attributes; employing a graph difference algorithm to calculate the sequence of addition, deletion, and modification operations in the graph structure between adjacent versions; analyzing multiple consecutive version difference sequences to summarize common patterns of system evolution, such as node expansion, link addition / deletion, and configuration drift; the constructed unified twin not only contains the complete graph of the current latest version but also includes an evolution graph that records the lifecycle of each component and relationship, all key change histories, and their context, enabling the system to query the "common performance characteristics of a component before the past three expansions."
[0030] In some embodiments, long-term operation and maintenance trend analysis and prediction are supported as follows: a unified twin is used as input to drive a time series prediction model based on an attention mechanism; the model not only considers the historical values of the performance indicators themselves, but also takes the change events (such as version upgrades and configuration modifications) extracted from the "evolution map" as external covariate inputs; the model can output the predicted values and confidence intervals of the system's key performance indicators at multiple future time scales (such as the next 1 hour, 24 hours, and 7 days); at the same time, by combining the extracted gradual change patterns, the model can predict when the performance indicators will reach the preset alarm threshold, thereby achieving early warning based on trends, rather than delayed alarms based on instantaneous thresholds.
[0031] In some embodiments, the twin processing module generates detected abnormal event data by performing computer vision-based abnormal pattern and event detection on the plurality of historical monitoring images of the cloud network module. The detected abnormal event data includes location markers for abnormal regions identified in the plurality of images, event type classifications, and corresponding detection confidence scores.
[0032] In some embodiments, computer vision-based anomaly pattern and event detection employs an architecture combining a multi-scale feature pyramid network (FPN) and a region proposal network (RPN) to process monitoring images. The FPN extracts multi-scale features from monitoring dashboards, topology maps, or traffic heatmaps at different levels. The RPN slides across the feature map, automatically generating candidate boxes that may contain anomaly regions. For each candidate box, a region of interest alignment technique is used to extract fixed-size features, which are then fed into a classification-regression branch network. This branch network performs two tasks in parallel: first, it performs fine-grained event classification on the regions within the candidate boxes, distinguishing specific event types such as "traffic storm," "connection exhaustion," and "CPU bottleneck"; second, it performs accurate regression on the location bounding boxes of the anomaly regions to achieve pixel-level anomaly region labeling on the image.
[0033] In some embodiments, the detected abnormal event data is associated with a global event context knowledge base; the knowledge base stores typical image patterns, root causes, handling measures, and final effects of various abnormal events in history; when a new abnormal event is detected, the system calculates the similarity between its image feature vector and the features of historical events in the knowledge base; if the similarity exceeds a threshold, it automatically associates the matching historical event records, and outputs the root cause inferences and handling suggestions in the historical records as additional information along with the current detection result, forming auxiliary diagnostic information of "detecting a similar historical event [X]", thereby accelerating the judgment process of operation and maintenance personnel.
[0034] In some embodiments, the calculation of the detection confidence score is based not only on the Softmax probability output by the classification network, but also on an uncertainty estimation method based on test-time data augmentation and Monte Carlo Dropout. Specifically, the same input monitoring image is randomly cropped and color-dithered multiple times, and multiple forward inferences are performed using a model with Dropout enabled. The distribution of event classifications in the multiple inference results is statistically analyzed. The final confidence score consists of two parts: the average classification probability and the information entropy of the distribution (used to measure model uncertainty). Results with high average probability and low information entropy are considered high-confidence detections, while results with high information entropy are marked as uncertain, triggering more detailed manual review or initiating additional, non-image-dimensional diagnostic processes.
[0035] In some embodiments, the twin processing module generates a topological risk map, including predictions of potential failure points, performance bottlenecks, and impact ranges, based on the system behavior logic and fault propagation relationships contained in the plurality of previously generated digital twin versions, through causal reasoning and graph computation.
[0036] In some embodiments, a causal graph is learned from a sequence of historical events using a causal discovery algorithm. The algorithm takes the historical state time series and event logs of each component of the system as input and uses algorithms such as PC or constraint-based structural learning models to infer the direction and strength of potential causal relationships between component state variables. The learned causal graph clarifies causal chains such as the failure of component A leading to a performance degradation of component B, providing a data-driven and quantifiable logical model for understanding fault propagation, which goes beyond dependency judgment based solely on static topology.
[0037] In some embodiments, causal reasoning and graph computation generate a topological risk graph that is a graph with a three-layer structure: the bottom layer is the physical / logical topology layer, representing devices, links, service instances and their connections; the middle layer is the "causal influence layer," which is superimposed on the topology layer in the form of a directed weighted graph, with edge weights representing the causal strength learned from the causal discovery algorithm or the co-occurrence probability statistically derived from historical failures; the top layer is the real-time status layer, which dynamically labels the current health status (normal, sub-healthy, faulty) of each component and real-time abnormal events as attributes on the corresponding nodes and edges; by running graph traversal algorithms on these three layers of graph, the influence radius of a fault event can be calculated in real time, and the most likely root cause node of the fault can be located.
[0038] In some embodiments, the generation of the topological risk graph is dynamic and incremental. The system maintains a baseline version of the graph. Whenever new monitoring data arrives or a new abnormal event is detected, the system does not regenerate the entire graph, but instead starts an incremental update engine. This engine first maps the new data / event to the corresponding node or edge of the graph and updates its state attributes. Subsequently, it triggers a message-passing-based graph reasoning process, propagating the impact of state changes (such as an increase in risk value) to downstream nodes along the edges of the causal influence layer. At the same time, the system checks for new, frequently occurring state co-occurrence patterns that can be used to fine-tune the edge weights of the causal influence layer. This incremental mechanism ensures that the graph can reflect the latest risk situation of the system with low latency.
[0039] Furthermore, the twin processing module, based on the detected abnormal event data and the topology risk map, performs comprehensive analysis through evidence fusion and probabilistic graphical models to generate a comprehensive and quantitative operation and maintenance status assessment map. This map identifies the instantaneous risk level and fault probability estimate of each logical unit or physical component in the cloud network module.
[0040] In some embodiments, a dynamic Bayesian network (DBN) is used as the core probabilistic graphical model for comprehensive analysis through evidence fusion and probabilistic graphical models. The nodes of this DBN represent the health status, performance indicators, and detected abnormal events of key system components. The conditional probability tables between nodes are partially initialized with domain knowledge and partially learned from historical data. The detected abnormal event data serve as "soft evidence," and the influence relationships derived from the topological risk graph serve as constraints on the network structure. During inference, various types of evidence observed in real time are input into the DBN, and precise inference is performed using algorithms such as connection trees or approximate inference is performed using sampling methods to calculate the marginal probability of each component being in a "fault" or "high-risk" state. This probability is the estimate of the probability of the fault occurring.
[0041] In some embodiments, a comprehensive and quantitative operational status assessment map is generated, which is represented as a multi-dimensional risk scoring matrix. The rows of the matrix represent all managed logical units or physical components, and the columns represent different risk dimensions, including: real-time health status risk (based on current indicators), correlation propagation risk (based on topology risk map), historical recurrence risk (based on the component's historical failure frequency), and business criticality risk (preset weights). The value of each cell is obtained by normalizing and weighting the original data of the corresponding dimension. Finally, the comprehensive risk level of each component is obtained by weighted aggregation of all risk dimension scores in the row (such as weighted geometric mean), and supports dynamic adjustment of dimension weights according to different operational perspectives (such as network perspective and application perspective) to generate status views with different focuses.
[0042] In some embodiments, the operational status assessment map has self-explanatory capabilities; for each high-risk component identified in the map, the system automatically generates a concise explanation report; the report content is obtained by tracing the reasoning path and evidence source in the probabilistic graphical model.
[0043] In some embodiments, the twin processing module determines hyperparameters for differentiated modeling and optimization of diagnostic and prediction algorithms for different risk level regions or components in the intelligent operation and maintenance analysis model based on the comprehensive operation and maintenance situation assessment map.
[0044] In some embodiments, hyperparameters for differentiated modeling and optimization are determined based on a comprehensive operational status assessment map, specifically implemented as a hyperparameter strategy mapping table. This mapping table defines the mapping relationship from the combination of "risk level" and "component type" to a set of specific model hyperparameters. For example, for a "high-risk" component of the "network router" type, the mapped hyperparameters may include: setting the learning rate of the anomaly detection model to 0.01 to accelerate the learning of anomaly patterns, setting the pollution parameter of the isolated forest algorithm to a low level to improve the anomaly detection sensitivity, and setting the input window length of the time series prediction model to a longer level to capture long-period patterns. The system looks up the corresponding hyperparameter set for each component in real time according to the assessment map and dynamically loads it into the model instance that processes the data stream of that component.
[0045] Differentiated modeling is not only reflected in hyperparameters but also extends to the selection of model structure. The system maintains a lightweight model library containing various algorithms of varying complexity, such as logistic regression, random forest, and lightweight neural networks. Based on the real-time risk level and historical behavioral characteristics of components, the system executes an online model selection process. For newly emerging components or those with rapidly increasing risk, a simple and fast-training model may be selected for initial monitoring. For components that are stable in the long term but on the critical path, a more complex and interpretable model may be selected. The selection process is based on a meta-learner trained offline, which recommends the most suitable model type and its initial hyperparameters based on the component's characteristics (such as indicator dimensions and data stationarity) and current risk level.
[0046] Hyperparameter tuning is a continuous online learning and feedback process. The system deploys a lightweight performance monitor for each model instance with differentiated configurations to continuously evaluate its prediction accuracy, alert timeliness, and false alarm rate. When the performance metrics of a model instance on a specific component (such as consecutive false alarms) drop below a threshold, an automatic tuning process is triggered. This process uses the component's recent historical data in an isolated environment to perform small-scale Bayesian optimization or grid search on the current hyperparameters to find a better parameter combination. After the optimized parameters are verified to be effective through a simplified A / B test, they are seamlessly applied to model instances in the production environment, achieving a closed loop of model adaptive capability.
[0047] In some embodiments, the intelligent operation and maintenance analysis model is a hybrid intelligent model based on Bayesian learning networks, and the determination of hyperparameters based on the comprehensive operation and maintenance situation assessment map includes: determining a first index set, which includes a set of identifiers of system components assessed as high-risk or abnormal in the operation and maintenance situation assessment map; determining a second index set, which is the complement of the first index set and contains identifiers of components assessed as low-risk or operating normally; determining a first set of hyperparameters for the components in the first index set, which is configured to assign higher anomaly weights to the monitoring data of these components, enable more complex feature analysis, and set more sensitive early warning thresholds; and determining a second set of hyperparameters for the components in the second index set, which is configured to enable the model to adopt conventional monitoring strategies and higher noise tolerance for these components, wherein, when used by the intelligent operation and maintenance analysis model, the first set of hyperparameters is associated with a higher probability of fault early warning and a deeper root cause tracing capability than the second set of hyperparameters.
[0048] The hybrid intelligent model based on Bayesian learning networks consists of a shared feature extraction layer and multiple parallel Bayesian inference heads. The feature extraction layer transforms the input structured data into high-dimensional features. For components in the first index set (high-risk set), their data is routed to a deep inference unit, which uses a Bayesian neural network. Its weights follow a distribution rather than fixed values, and it can output prediction results and the uncertainty of the model itself, making it suitable for complex scenarios with small samples or high data noise. For components in the second index set (low-risk set), their data is routed to a lightweight inference unit, which uses a Bayesian linear model or a Gaussian process. This lightweight inference unit is computationally efficient and suitable for scenarios with relatively simple data patterns. Two sets of hyperparameters control the concentration of the prior distribution of the two inference heads and the model complexity, respectively.
[0049] The specific configuration of the first set of hyperparameters includes: setting a high sparsity facilitation parameter for a Gaussian-scale mixture prior to encourage the hybrid intelligent model to learn a few more discriminative key features from the complex data of high-risk components; thereby generating a stronger likelihood ratio change for small abnormal fluctuations and achieving more sensitive early warning; in addition, under the variational inference framework, setting fewer training iterations but more frequent hybrid intelligent model update cycles for the deep inference unit to quickly adapt to the rapid changes in the state of high-risk components.
[0050] The hybrid intelligent model achieves deeper root cause tracing capabilities by utilizing the first set of hyperparameters, specifically through a tracing method based on a Bayesian attention mechanism. In the Bayesian neural network of the deep inference unit, an attention gate following a Bernoulli distribution is introduced for each dimension of the input features. During inference, the model not only outputs the prediction results but also obtains the posterior distribution of the attention gates through sampling. Analyzing this distribution can reveal which input feature dimensions contribute the most to the current anomaly prediction. Combined with the system topology, it can trace back to which specific indicator of which upstream component is abnormal, ultimately triggering this high-risk alarm. Thus, the conclusion that component A is high-risk is refined into a root cause chain where the [indicator Y] of component B is abnormal, leading to the deterioration of [indicator X] of component A.
[0051] In some embodiments, at least one of the data acquisition devices includes an Internet of Things (IoT) sensor array consisting of software probes, log acquisition agents, and hardware sensors deployed on network nodes, servers, and cloud platforms, to achieve multi-dimensional real-time data acquisition of traffic, latency, resource utilization, and energy consumption.
[0052] In some embodiments, at least one of the data acquisition devices includes a simulation test engine that supports an active probing protocol for simulating user service flows and measuring end-to-end service performance, generating synthetic performance probe data as an important component of the structured data set.
[0053] This invention also provides a digital twin modeling and dynamic resource scheduling method for cloud-network convergence, including: The monitoring data stream corresponding to the operating status signals reflected by one or more network components or services in the cloud network environment monitored by the at least one data acquisition device is received by at least one data acquisition device. The monitoring data stream is aggregated, cleaned, and structured by the at least one data acquisition device to generate a set of operational data indicating the health status and performance level of the component or service; The set of operational data is received by the twin processing module; The twin processing module receives multi-source side information from the twin data storage, the multi-source side information including at least one digital representation of the cloud network module; The twin processing module determines the hyperparameters for the intelligent operation and maintenance analysis model based on the multi-source side information; and The twin processing module uses the hyperparameters and the intelligent operation and maintenance analysis model to process the operational data set, thereby generating a real-time health diagnosis report and automated operation and maintenance action instructions for the cloud network module. The multi-source side information includes one or more of the following: At least one digital topology and business logic model of the cloud network module. At least one historical or real-time operation and maintenance monitoring view image of the cloud network module, or At least one previously generated and optimized digital twin instance of the cloud network module.
[0054] The multi-source edge information includes multiple historical operation and maintenance panoramic view images of the cloud network module and multiple evolved versions of digital twins.
[0055] The twin processing module generates fine-grained historical anomaly event feature data by performing deep learning-based pattern recognition and event segmentation on the multiple historical operation and maintenance panoramic view images of the cloud network module. The historical anomaly event feature data includes labeling information and confidence scores for identifying historical fault modes, occurrence locations, and evolution processes in the multiple images.
[0056] The twin processing module generates a predictive system state view containing potential future abnormal events and performance degradation trends by using time-series prediction and graph neural network analysis, based on the system state transition and fault association knowledge recorded in the multiple evolved versions of digital twins.
[0057] The twin processing module performs spatiotemporal correlation analysis and probability fusion based on the fine-grained historical abnormal event feature data and the predictive system status view to generate a dynamic operation and maintenance risk assessment map covering the entire system. This map assigns a quantified real-time risk index to each managed entity.
[0058] The twin processing module determines the hyperparameters for differentiated modeling and strategy configuration of system units in the intelligent operation and maintenance analysis model for different risk assessment levels based on the operation and maintenance risk assessment map. The intelligent operation and maintenance analysis model is a decision-making model based on the fusion of Bayesian inference and machine learning, and the hyperparameters determined based on the operation and maintenance risk assessment map include: Determine a first set of indexes containing all system unit indexes identified as high-risk or faulty states; Determine a second index set that serves as the complement of the first index set and contains all low-risk or normal state unit indices; Determine a first set of hyperparameters for the cells within the first index set, which enables the model to perform high-frequency sampling, deep packet inspection, and multi-dimensional correlation analysis; and A second set of hyperparameters is determined for the cells within the second index set. This set of parameters enables the model to employ low-frequency inspection and summary statistical analysis. In the process of generating operation and maintenance decisions, units that apply the first set of hyperparameters will trigger a more prioritized and rigorous handling process and resource pre-allocation action than units that apply the second set of hyperparameters.
[0059] The beneficial effects are as follows: By introducing multi-source side information, including historical digital twins, topology models, and monitoring images, as prior knowledge, and dynamically determining the hyperparameters of the intelligent operation and maintenance analysis model based on this, the model acquires context awareness and adaptive capabilities. This allows the model to customize its sensitivity, focus, and judgment thresholds for abnormal data according to the system's specific historical state, logical structure, and visual performance. For example, for areas with historically frequent failures, the system assigns them higher monitoring weights and more stringent abnormal conditions through hyperparameters. This parameter adaptation process based on multi-source side information transforms isolated alarm events into accurate risk alerts calibrated with prior knowledge and accompanied by probability assessment. The generated operation and maintenance decision instructions (such as isolation and expansion) thus integrate real-time data and historical experience, and are built on quantitative evidence and model reasoning, significantly reducing the risk of misoperation due to false alarms or missed alarms, and improving the effectiveness of emergency response and the scientific nature of decision-making.
[0060] By integrating multiple historical digital twins to construct a unified twin with temporal continuity and state evolution capabilities, the system can support long-term operational trend analysis and prediction. This unified twin, acting as a dynamic knowledge base, enables the system to perform deep time-series analysis and pattern recognition, extracting performance degradation feature sequences and fault cycle patterns from historical monitoring images. Based on this, the system can not only provide early warnings of immediate faults but also predict performance inflection points or potential compatibility issues caused by accumulated pressure within the next few hours or days. This predictive capability allows for proactive operational work, shifting from passively handling existing faults to actively preventing potential risks.
[0061] This technology leverages computer vision-based image analysis, graph computing-based topological reasoning, and probabilistic graphical model-based evidence fusion to achieve deep fusion and intelligent correlation analysis of multimodal operation and maintenance data. This approach enables the correlation and reasoning between surface anomalies identified from images and logical dependencies in the topological graph, thereby tracing back to root cause chains across domains and layers.
Claims
1. A digital twin modeling and dynamic resource scheduling system for cloud-network convergence, characterized in that, include: At least one data acquisition device is configured to acquire real-time operational status data of multiple network devices, virtualization resources and service instances in the cloud network module, receive monitoring signals and performance indicators reflected by the service traffic and hardware status running in the cloud network module, and generate a structured data set characterizing the operational health status and topology connection relationship of the system components. A twin data storage device, configured to store multi-source side information representing the historical and real-time status of the cloud network module; as well as A twin processing module is configured to receive the structured data set from the at least one data acquisition device, receive the multi-source side information from the twin data storage, determine hyperparameters for an intelligent operation and maintenance analysis model based on the multi-source side information, and use the hyperparameters to process the structured data set through the intelligent operation and maintenance analysis model to generate real-time health status assessment and operation and maintenance decision instructions for the cloud network module. The multi-source side information includes one or more of the following: The static or dynamic digital topology and configuration model of the cloud network module. The historical or real-time monitoring visualization images of the cloud network module, or The previously generated digital twin or system state snapshot of the cloud network module.
2. The digital twin modeling and dynamic resource scheduling system for cloud-network convergence according to claim 1, characterized in that, The multi-source edge information includes multiple historical monitoring images from the cloud network module and multiple previously generated digital twin versions; The twin processing module is configured to: extract feature sequences of performance degradation and abnormal events by performing time-series analysis and pattern recognition on the multiple historical monitoring images; and construct a unified twin with time continuity and state evolution capabilities by fusing the multiple previously generated digital twin versions to support long-term operation and maintenance trend analysis and prediction.
3. The digital twin modeling and dynamic resource scheduling system for cloud-network convergence according to claim 2, characterized in that, The twin processing module generates detected abnormal event data by performing computer vision-based abnormal pattern and event detection on the multiple historical monitoring images of the cloud network module. The detected abnormal event data includes location markers for abnormal regions identified in the multiple images, event type classifications, and corresponding detection confidence scores.
4. The digital twin modeling and dynamic resource scheduling system for cloud-network convergence according to claim 3, characterized in that, The twin processing module generates a topological risk map, including predictions of potential failure points, performance bottlenecks, and impact ranges, based on the system behavior logic and fault propagation relationships contained in the multiple previously generated digital twin versions, through causal reasoning and graph computation.
5. The digital twin modeling and dynamic resource scheduling system for cloud-network convergence according to claim 4, characterized in that, The twin processing module, based on the detected abnormal event data and the topology risk map, performs comprehensive analysis through evidence fusion and probabilistic graphical models to generate a comprehensive and quantitative operation and maintenance status assessment map. This map identifies the real-time risk level and fault probability estimate of each logical unit or physical component in the cloud network module.
6. The digital twin modeling and dynamic resource scheduling system for cloud-network convergence according to claim 5, characterized in that, The twin processing module determines the hyperparameters for differentiated modeling and optimization of the diagnostic and prediction algorithms for different risk level areas or components in the intelligent operation and maintenance analysis model based on the comprehensive operation and maintenance situation assessment map.
7. The digital twin modeling and dynamic resource scheduling system for cloud-network convergence according to claim 6, characterized in that, The intelligent operation and maintenance analysis model is a hybrid intelligent model based on Bayesian learning networks, and the hyperparameters determined based on the comprehensive operation and maintenance situation assessment map include: A first index set is determined, which includes a set of identifiers of system components that are assessed as high-risk or have shown abnormalities in the operation and maintenance status assessment map; Determine a second index set, which is the complement of the first index set, containing component identifiers that are assessed as low risk or functioning properly; A first set of hyperparameters is determined for the components in the first index set, the first set of hyperparameters being configured to assign higher anomaly weights to the monitoring data of these components by the model, enable more complex feature analysis, and set more sensitive warning thresholds; and A second set of hyperparameters is determined for the components in the second index set. This second set of hyperparameters is configured to enable the model to employ a conventional monitoring strategy for these components with a high noise tolerance. When used by the intelligent operation and maintenance analysis model, the first set of hyperparameters is associated with a higher probability of fault warning and a deeper root cause tracing capability than the second set of hyperparameters.
8. A digital twin modeling and dynamic resource scheduling method for cloud-network convergence, characterized in that, include: The monitoring data stream corresponding to the operating status signals reflected by one or more network components or services in the cloud network environment monitored by the at least one data acquisition device is received by at least one data acquisition device. The monitoring data stream is aggregated, cleaned, and structured by the at least one data acquisition device to generate a set of operational data indicating the health status and performance level of the component or service; The set of operational data is received by the twin processing module; The twin processing module receives multi-source side information from the twin data storage, the multi-source side information including at least one digital representation of the cloud network module; The twin processing module determines the hyperparameters for the intelligent operation and maintenance analysis model based on the multi-source edge information; as well as The twin processing module uses the hyperparameters and the intelligent operation and maintenance analysis model to process the operational data set, thereby generating a real-time health diagnosis report and automated operation and maintenance action instructions for the cloud network module. The multi-source side information includes one or more of the following: At least one digital topology and business logic model of the cloud network module. At least one historical or real-time operation and maintenance monitoring view image of the cloud network module, or At least one previously generated and optimized digital twin instance of the cloud network module.
9. The digital twin modeling and dynamic resource scheduling method for cloud-network convergence according to claim 8, characterized in that, The multi-source edge information includes multiple historical operation and maintenance panoramic view images of the cloud network module and multiple evolved versions of digital twins; The twin processing module generates fine-grained historical anomaly event feature data by performing deep learning-based pattern recognition and event segmentation on the multiple historical operation and maintenance panoramic view images of the cloud network module. The historical anomaly event feature data includes labeling information and confidence scores for identifying historical fault modes, occurrence locations and evolution processes in the multiple images.
10. The digital twin modeling and dynamic resource scheduling method for cloud-network convergence according to claim 8 or 9, characterized in that, Based on the system state transition and fault correlation knowledge recorded in the multiple evolved versions of digital twins, the twin processing module generates a predictive system state view containing potential future abnormal events and performance degradation trends through time series prediction and graph neural network analysis. The twin processing module performs spatiotemporal correlation analysis and probability fusion based on the fine-grained historical abnormal event feature data and the predictive system status view to generate a dynamic operation and maintenance risk assessment map covering the entire system. This map assigns a quantified real-time risk index to each managed entity. The twin processing module determines the hyperparameters for differentiated modeling and strategy configuration of system units in the intelligent operation and maintenance analysis model that are oriented towards different risk assessment levels, based on the operation and maintenance risk assessment map. The intelligent operation and maintenance analysis model is a decision-making model based on the fusion of Bayesian inference and machine learning, and the hyperparameters determined based on the operation and maintenance risk assessment map include: Determine a first set of indexes containing all system unit indexes identified as high-risk or faulty states; Determine a second index set that serves as the complement of the first index set and contains all low-risk or normal state unit indices; Determine a first set of hyperparameters for the cells within the first index set, which enables the model to perform high-frequency sampling, deep packet inspection, and multi-dimensional correlation analysis; and A second set of hyperparameters is determined for the cells within the second index set. This set of parameters enables the model to employ low-frequency inspection and summary statistical analysis. In the process of generating operation and maintenance decisions, units that apply the first set of hyperparameters will trigger a more prioritized and rigorous handling process and resource pre-allocation action than units that apply the second set of hyperparameters.