Network performance optimization system and method based on data analysis
By using a data analysis-based network performance optimization system, which leverages multimodal data acquisition and a graph attention time-series prediction model, the system addresses the issues of insufficient business intent perception and adaptive capabilities in network optimization. It achieves efficient network performance prediction and proactive resource orchestration, thereby enhancing the network's adaptive and self-optimizing capabilities.
Patent Information
- Application Number
- CN202511910510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-02-27
AI Technical Summary
Existing network optimization solutions suffer from problems such as insufficient business intent perception, lack of data correlation analysis, limited predictive capabilities, and delayed optimization decisions and a lack of adaptive capabilities.
The network performance optimization system based on data analysis includes a multimodal network data acquisition module, a multimodal network knowledge graph construction and management module, an adaptive business intent parsing module, an intent-driven graph attention temporal prediction model module, and a multi-objective intelligent orchestration and feedback module. Through multimodal data acquisition, entity recognition, semantic association, business intent parsing, graph attention temporal prediction, and reinforcement learning, it achieves network performance optimization.
It enables automatic quantitative analysis of high-level business requirements, constructs a unified multimodal knowledge graph, performs accurate network performance prediction and proactive resource orchestration, and has adaptive and self-optimizing capabilities, thereby improving network performance assurance capabilities.
Smart Images

Figure CN121585553A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of network management, data analysis, artificial intelligence and network automation, and in particular relates to using multimodal knowledge graphs, deep learning (graph neural networks, temporal convolutional networks) and reinforcement learning techniques to achieve network performance prediction and proactive resource orchestration and configuration optimization based on business intent. Background Technology
[0002] The current network environment is becoming increasingly complex. Technologies such as cloud computing, 5G, and the Internet of Things have led to an unprecedented expansion of network scale and diversification of service types, placing higher demands on network performance and availability. Traditional network management and optimization methods face numerous challenges:
[0003] Lack of business intent awareness: Most systems only focus on low-level network metrics and cannot understand high-level business needs (such as "smooth high-definition video conferencing"), making it difficult to effectively correlate network performance with business SLAs (Service Level Agreements).
[0004] Data silos and heterogeneity: Network data is scattered across different devices, systems, and platforms, with heterogeneous formats, making it difficult to perform unified analysis and correlation.
[0005] Insufficient predictive capability: Traditional predictive models are mostly based on a single indicator or simple time series analysis, which makes it difficult to capture the complex nonlinear relationships between multidimensional indicators, and in particular ignores the impact of network topology and service dependencies on performance.
[0006] Passive optimization: Most systems only alert and intervene after performance problems occur, failing to achieve early warning and proactive prevention.
[0007] High cost of manual intervention: Diagnosing and optimizing complex network problems still requires a lot of human experience and manual operation, which is inefficient and prone to errors.
[0008] Lack of adaptive and self-optimizing capabilities: Optimization strategies are often static, making it difficult to adapt to dynamic changes in the network, and lack continuous learning and improvement mechanisms.
[0009] Existing technologies such as threshold-based alarm systems, simple capacity planning tools, and some machine learning-based anomaly detection and prediction systems, while making progress, generally lack a deep understanding of business intent, the ability to perform complex cross-domain and multi-level correlation analysis, and the decision-making ability to achieve intelligent proactive orchestration. Summary of the Invention
[0010] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0011] The purpose of this invention is to provide a network performance optimization system based on data analysis, which aims to solve the problems of insufficient business intent perception, lack of data correlation analysis, limited predictive ability, and delayed optimization decision-making and lack of adaptive ability in existing network optimization schemes.
[0012] To address the aforementioned technical problems, this invention provides the following technical solution: a network performance optimization system based on data analysis, comprising the following components: a. A multimodal network data acquisition module: used to collect multimodal heterogeneous data from network devices, application systems, user terminals, service management platforms, and operation and maintenance knowledge bases; b. A multimodal network knowledge graph (MM-NKG) construction and management module: used to perform entity recognition, relation extraction, and semantic association on the collected data, constructing and dynamically updating a knowledge graph that integrates network topology, performance, services, and expert experience; c. An adaptive service intent parsing module: used to parse high-level service requirements into a set of quantifiable network performance goals and constraints, and associate them with service entities in the MM-NKG; d. An intent-driven graph attention temporal prediction model (IGATP) module: used to predict future trends of network performance indicators and identify potential intent violation risks by integrating the service intent parsing results with the MM-NKG as the framework; e. Multi-objective intelligent orchestration and feedback module: Used to generate and execute network resource and configuration orchestration schemes based on the prediction results of the IGATP module, network resource information in MM-NKG, and the goals and constraints set by the service intent parsing module, and continuously optimize the orchestration strategy through reinforcement learning; f. Optimization execution and feedback interface: Used to interact with actual network devices, SDN controllers, or cloud platform APIs, execute orchestration schemes, and feed back the execution results and network performance changes to the system.
[0013] As a preferred embodiment of the network performance optimization system based on data analysis described in this invention, the multimodal network knowledge graph (MM-NKG) includes at least the following entity types: devices, interfaces, links, services, applications, performance metrics, and expert experience; and at least the following relationship types: connections, bearers, dependencies, configurations, generation, and influences.
[0014] As a preferred embodiment of the network performance optimization system based on data analysis described in this invention, the adaptive service intent parsing module includes: a natural language understanding unit for parsing service requirements in natural language form; an intent-to-metric mapping unit for mapping service requirements into quantifiable network performance metrics and thresholds; and a service entity association unit for associating the parsing results with entities in MM-NKG.
[0015] As a preferred embodiment of the network performance optimization system based on data analysis described in this invention, the Intent-Driven Graph Attention Temporal Prediction Model (IGATP) module uses MM-NKG as the graph structure input, takes network performance indicators and business intent targets as node features, and employs a multi-head attention mechanism to capture dynamic associations between nodes, ultimately predicting future performance trends and intent violation probabilities. Specifically, it includes the following models: a. Graph Attention Encoder: used to aggregate neighbor node information through a multi-head graph attention mechanism based on the MM-NKG graph structure to generate node context embeddings; b. Intent Fusion Layer: used to fuse intent targets with context embeddings to generate intent-enhanced node embeddings; c. Temporal Convolutional Network (TCN): used to extract temporal features from the intent-enhanced node embedding sequence to generate temporal context; d. Attention Decoder: used to fuse intent-enhanced node embeddings, temporal context, and expert experience associated with intent in MM-NKG to predict future network performance indicators and calculate intent violation probabilities.
[0016] As a preferred embodiment of the network performance optimization system based on data analysis described in this invention, the loss function of the IGATP model includes a mean squared error term, a binary cross-entropy term, and a regularization term that encourages the model to maintain a structure and semantics consistent with MM-NKG.
[0017] As a preferred embodiment of the network performance optimization system based on data analysis described in this invention, the multi-objective intelligent orchestration and feedback module includes: a policy space generator, used to generate candidate optimized orchestration actions based on predicted risks; a multi-objective optimizer, used to balance multiple objectives such as intent achievement, resource utilization, and cost to select the optimal solution; and a reinforcement learning policy optimization unit, used to continuously optimize the orchestration strategy through reinforcement learning.
[0018] As a preferred embodiment of the network performance optimization system based on data analysis described in this invention, the reward function of the reinforcement learning strategy optimization unit comprehensively considers the degree of intent achievement, resource consumption cost, operation cost, and side effect penalty.
[0019] As a preferred embodiment of the network performance optimization system based on data analysis described in this invention, the optimization execution and feedback interface supports Restful API, Netconf / YANG, OpenFlow protocol or vendor SDK to achieve automated orchestration of physical network devices, SDN controllers, cloud platforms and virtualization resources.
[0020] To address the aforementioned technical problems, this invention also provides the following technical solution: a network performance optimization method based on data analysis, employing the aforementioned network performance optimization system based on data analysis, comprising the following steps: a. Multimodal network data acquisition: acquiring multimodal heterogeneous data from multiple sources, including network devices, application systems, user terminals, service management platforms, and operation and maintenance knowledge bases; b. Multimodal network knowledge graph (MM-NKG) construction and management: performing entity recognition, relation extraction, and semantic association on the acquired data to construct and dynamically update the MM-NKG; c. Adaptive service intent parsing: parsing high-level service requirements into quantified network performance goals and constraints, and associating them with the MM-NKG; d. Intent-driven graph attention temporal prediction: using the MM-NKG as a framework, integrating service intents, and predicting network performance trends and intent violation risks through the IGATP model; e. Multi-objective intelligent orchestration and feedback: generating and executing network resource and configuration orchestration schemes based on prediction results, MM-NKG information, and intent goals, and continuously optimizing the orchestration strategy through reinforcement learning; f. Optimize execution and feedback: Execute the orchestration scheme through the interface and feed back the execution results and network performance changes to the system.
[0021] This invention provides a network performance optimization system and method based on data analysis, which has the following beneficial effects:
[0022] 1. Adaptive Business Intent Parsing: The system can automatically parse high-level, natural language-based business requirements (such as "ensure zero transaction lag during e-commerce promotions") into quantifiable network performance metrics (such as API response time P99 < 100ms, payment service availability > 99.99%) and constraints. This upgrades network optimization from "focusing on metrics" to "ensuring business intent," representing a crucial step in intelligent network autonomy.
[0023] 2. Multimodal Network Knowledge Graph (MM-NKG): Constructs a unified, dynamically updated knowledge graph that integrates heterogeneous data such as network devices, interfaces, links, configurations, performance metrics, services, applications, users, faults, and expert experience. MM-NKG semantically and structurally represents this data through entities and relationships, clearly expressing the complex dependencies and influences between various components, services, and performance metrics in the network, serving as the foundation for all subsequent analysis and decision-making.
[0024] 3. Intent-Driven Graph Attention Temporal Prediction Model (IGATP): This deep learning model uses MM-NKG as the graph structure input and incorporates network performance metrics, configuration information, and business intent objectives as node features. The model combines:
[0025] Graph Attention Networks (GAT): Capture the dynamic spatial relationships between entities in a network topology (e.g., the impact of link congestion on the performance of services that depend on it).
[0026] Intent Fusion: Deeply integrates the quantitative goals of business intent with node features, enabling the prediction model to "perceive" business objectives and predict not only future performance trends, but also "intent violation risks".
[0027] Temporal Convolutional Networks (TCNs): Efficiently capture long-distance temporal dependencies and periodic patterns in network performance data.
[0028] Expert experience fusion: Expert experience relevant to the current scenario (such as troubleshooting steps and optimization suggestions) from MM-NKG is embedded into the prediction process to enhance the model's interpretability and robustness. The IGATP model can predict precise values of network performance indicators over a future period and quantify the probability of business intent being violated.
[0029] 4. Multi-objective Intelligent Orchestration and Feedback: Based on the prediction results of the IGATP model (especially the probability of intent violation) and the objectives set by the business intent, the system obtains available resources and policies from the MM-NKG. Through a reinforcement learning-based multi-objective optimizer, it automatically generates and executes the optimal cross-domain, cross-layer network resource and configuration orchestration scheme. This module can balance multiple objectives such as performance, cost, and resource utilization. Through the closed-loop feedback of reinforcement learning, the system can continuously learn and optimize orchestration strategies, achieving true self-adaptation and self-optimization. Attached Figure Description
[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0031] Figure 1 The flowchart of the network performance optimization method based on data analysis provided by the present invention is shown. Detailed Implementation
[0032] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0033] The purpose of this invention is to provide a network performance optimization system and method based on data analysis, which aims to solve the problems of insufficient business intent perception, lack of data correlation analysis, limited predictive ability, and delayed optimization decision-making and lack of adaptive ability in existing network optimization schemes.
[0034] Specifically, this invention provides a network performance optimization system based on data analysis, comprising:
[0035] a. Multimodal network data acquisition module: used to collect multimodal heterogeneous data from network devices, application systems, user terminals, service management platforms and operation and maintenance knowledge bases, including but not limited to network performance indicators, device configuration information, service requirement parameters, user experience feedback and operation and maintenance expert experience;
[0036] To further clarify, the multimodal network data acquisition module specifically includes:
[0037] Performance data: Real-time and historical metrics such as CPU, memory, bandwidth utilization, packet loss rate, latency, jitter, throughput, API response time, and error rate are collected from routers, switches, firewalls, servers, virtualization platforms (VMware, OpenStack), container orchestration platforms (Kubernetes), and cloud platforms (AWS, Azure) using tools such as SNMP, NetFlow / IPFIX, Syslog, APM probes (such as Dynatrace, AppDynamics), and Prometheus.
[0038] Configuration data: Periodically collect network device configuration information (routing table, ACL, QoS policy, VLAN configuration), service configuration (microservice configuration, database parameters), and topology connection information via Netconf / YANG, SSH / Telnet, API interface, etc.
[0039] Business data: Obtain information such as business type, SLA level, user priority, traffic pattern, and peak business period forecast from the business management platform (BSS / OSS), CRM system, and billing system.
[0040] User experience data: Collect user-perceived data such as page load time, video stuttering rate, download speed, and user access success rate through Web / App tracking, CDN logs, and probes (such as Pingdom and Catchpoint).
[0041] Operations and Maintenance Knowledge Base: Extracts unstructured or semi-structured data such as fault modes, troubleshooting steps, optimization practices, equipment models, and suppliers from the fault ticket system, operations and maintenance logs, expert documents, and CMDB (Configuration Management Database).
[0042] Data preprocessing: Time synchronization, format conversion, data cleaning (outlier and missing value handling), data aggregation, and feature engineering are performed on the collected multimodal data.
[0043] b. Multimodal Network Knowledge Graph (MM-NKG) Construction and Management Module: This module is used to perform entity recognition, relation extraction, and semantic association on the data acquired by the collection module, constructing a multimodal knowledge graph that integrates network topology, configuration, performance, business, and expert experience, and supporting dynamic updates and queries of the graph;
[0044] It should be noted that the Multimodal Network Knowledge Graph (MM-NKG) includes the following entity types and relation types:
[0045] 1. Entity types: devices (routers, switches, servers), interfaces, links, VLANs, IP addresses, services (VoIP, video conferencing, e-commerce transactions), applications (databases, message queues), users, geographical locations, configuration parameters, performance metrics (bandwidth, latency, CPU utilization), fault codes, expert experience entries, etc.
[0046] 2. Relationship types: connection (device-interface, interface-link), bearer (link-service), dependent (service-application, application-database), configured in (device-configuration parameters), generated (device-performance indicators), affected (fault-performance, configuration-performance), belongs to (user-service), located in (device-geographical location), related to (expert experience-fault), etc.
[0047] To further clarify, MM-NKG is the core data foundation of this system. It is implemented using a graph database (such as Neo4j), enabling efficient storage and querying of entities and relationships.
[0048] Entity recognition and relation extraction:
[0049] Structured data: Entities (devices, interfaces, IPs, services, applications, users) and their attributes (CPU core count, memory size, SLA level) and predefined relationships (connections, bearers, dependencies) are directly extracted from CMDB, performance monitoring systems, etc.
[0050] Unstructured data: Utilizing Natural Language Processing (NLP) techniques, such as Named Entity Recognition (NER), Relation Extraction (RE), and Event Extraction, new entities (such as specific fault types and optimization suggestions) and relationships (such as "Fault A affects device B" and "Expert experience C addresses fault D") are extracted from operation and maintenance logs, fault tickets, and expert documents. For example, from the log "Input errors occurred on Router-X interface GigabitEthernet0 / 1," "Router-X" is identified as a device entity, "GigabitEthernet0 / 1" as an interface entity, and the relationship "Router-X owns GigabitEthernet0 / 1" is extracted. Simultaneously, "Input errors" can be used as a fault entity or performance indicator entity and associated with the interface.
[0051] Semantic association and fusion: Matching and merging identical entities from different data sources eliminates redundancy and constructs a unified entity ID. Through ontology modeling, the semantics of different entity types and relationship types are defined. For example, "Web service on server A" and "application B" can be associated as the same logical entity.
[0052] Graph Storage and Query: Leveraging the characteristics of graph databases, this feature enables efficient storage and querying of complex network entity relationships. For example, it allows for quick queries of "all physical links affecting e-commerce transactions" or "all applications that may be affected by a device failure."
[0053] Dynamic updates: MM-NKG is not static. It continuously ingests new configuration, performance, business data and operation and maintenance events, and updates the entity attributes and relationship status in the graph periodically or in real time (such as the link status changing from "UP" to "DOWN"), and even adds new entities and relationships.
[0054] c. Adaptive Service Intent Parsing Module: This module receives high-level service requirements or user expectations, and uses natural language processing and machine learning techniques to parse them into a set of quantifiable network performance goals (such as QoS and SLA metrics) and constraints, and associates them with service entities in MM-NKG.
[0055] It should be noted that the adaptive business intent parsing module includes:
[0056] 1. Natural Language Understanding (NLU) Unit: Used to parse business requirements in natural language form from user input (e.g., "ensure smooth high-definition video conferencing" and "ensure zero-lag transactions during e-commerce promotions"), and extract key entities, actions, and modifiers;
[0057] 2. Intent to Metric Mapping Unit: Used to map the information extracted by the NLU unit to quantifiable network performance metrics and threshold ranges. For example, "smooth" is mapped to "latency < 50ms, packet loss rate < 0.1%", and "zero stuttering" is mapped to "API response time P99 < 100ms, service availability > 99.99%".
[0058] 3. Constraint Extraction Unit: Used to identify non-functional constraints in business intent, such as cost limits, security policies, and resource isolation requirements;
[0059] 4. Business Entity Association Unit: Used to associate the parsed business intent with specific business, application, or user entities in MM-NKG.
[0060] To elaborate further, the adaptive business intent parsing module is the key to achieving "intent-driven" operation.
[0061] Natural Language Understanding (NLU) Unit: Users (business departments, end users) input business requirements through a web interface or API, such as "Ensure smooth access to the company's CRM system for users in the Beijing area" or "Ensure high availability of transaction services in response to the upcoming Double Eleven promotion." The NLU unit utilizes pre-trained language models (such as BERT, GPT) for intent recognition, slot filling, and extraction of key information.
[0062] Intent: To ensure smooth operation and high availability.
[0063] Entities: Beijing area, CRM system, Double Eleven promotion, transaction services.
[0064] Modifiers: Smooth, highly available.
[0065] Intent-to-metric mapping unit: Maps non-quantitative intents to quantifiable network performance metrics and threshold ranges using predefined rule bases, expert knowledge, or machine learning-based methods such as classification models.
[0066] Example:
[0067] The top management's intention: "To ensure smooth high-definition video conferencing."
[0068] Analysis results:
[0069] Business Entity: Video Conferencing Service (linked to specific servers and applications via MM-NKG)
[0070] Performance goals:
[0071] End-to-end latency < 50ms (P95)
[0072] Packet loss rate < 0.1% (P99)
[0073] Jitter < 20ms (P95)
[0074] Video stuttering rate < 0.01%
[0075] Constraint: Bandwidth guaranteed > 10Mbps / user
[0076] The top management's intention: "To ensure zero-lag transaction services during major e-commerce promotions."
[0077] Analysis results:
[0078] Business entities: e-commerce transaction services, payment services, inventory services
[0079] Performance goals:
[0080] API response time (payment service) P99 < 80ms
[0081] API response time (inventory service) P99 < 150ms
[0082] Transaction success rate > 99.99%
[0083] Service availability > 99.999%
[0084] Constraints: Service downgrades are not allowed, and cost increases cannot exceed 20%.
[0085] Constraint extraction unit: Identifies non-functional constraints such as cost limits, security levels, geographical restrictions, and resource isolation.
[0086] Business Entity Association Unit: This unit precisely associates the parsed intent with specific business, application, or user entities in the MM-NKG to obtain information such as their related network topology and dependencies.
[0087] d. Intent-Driven Graph Attention Temporal Prediction Model (IGATP) module: Used as the backbone of MM-NKG, it integrates the results of business intent parsing to predict the future trend of network performance indicators and identify potential intent violation risks;
[0088] It should be noted that the Intent-Driven Graph Attention Temporal Prediction Model (IGATP) module uses MM-NKG as the graph structure input, takes network performance metrics and business intent objectives as node features, and employs a multi-head attention mechanism to capture the dynamic relationships between nodes, ultimately predicting future performance trends and the probability of intent violation. Its specific model includes the following components and mathematical expressions:
[0089] 1. Graph structure input: Adjacency matrix A and node feature matrix F of MM-NKG t F tIncludes performance metrics, configuration information, and intent target parameters for each network entity (device, link, service) at time step t;
[0090] 2. Graph Attention Encoder: At each time step t, based on the MM-NKG graph structure, it aggregates neighbor node information through a multi-head graph attention mechanism and learns the dynamic spatiotemporal relationships between nodes to generate the node's context embedding H. t GAT ;
[0091] 3. Intent Fusion Layer: Used to integrate the quantified intent targets (such as target delay L) output by the adaptive business intent parsing module. target Target bandwidth B target ) and H t GAT Perform fusion to generate intent-enhanced node embeddings H t Intent ;
[0092] 4. Temporal Convolutional Networks (TCNs): Used for embedding sequences H of intent-enhanced nodes. Intent =[H t−Tin+1 Intent H t Intent Perform temporal feature extraction, capture long-distance temporal dependencies, and generate temporal context C. t TCN ;
[0093] 5. Attention Decoder: Used for fusing H t Intent C t TCN And expert experience related to intent in MM-NKG, predicting future T out Network performance metrics P^ at each time step t+1 , ..., P^ t+Tout And calculate the probability P of intentional violation. violate ;
[0094] 6. Loss function: Taking into account both prediction error and the risk of intention violation.
[0095] Specifically, the mathematical expression and parameter definitions of the above IGATP model are as follows:
[0096] 1. Graph structure input:
[0097] MM-NKG is represented by graph G=(V,E), where V is the set of entity nodes and E is the set of relation edges.
[0098] At time step t, the eigenvector of node v∈V is f. t v∈R D This includes performance metrics, configuration status, etc.
[0099] Intent target vector I v ∈R K , which represents the business intent target associated with node v (such as target latency, packet loss rate, etc.).
[0100] ①Graph Attention Encoder (GAT Layer):
[0101] For node v, the graph attention at the l-th layer is:
[0102] h v l+1 =ReLU(∑ k=1 H α v,neighbors(v) (k,l) ⋅W (k,l) h v l +∑ u∈N(v) ∑ k=1 H α v,u (k,l) ⋅W (k,l) h u l )
[0103] Among them, h v l It is the embedding representation of node v at layer l (initial h) v 0 =f t v N(v) are the neighbors of node v, and α v,u (k,l) W is the attention weight from node v to u calculated by the k-th attention head at layer l. (k,l) It is a learnable weight matrix, and H is the number of attention heads.
[0104] Attention weight calculation:
[0105] e v,u (k,l) =LeakyReLU(a (k,l) T [W (k,l) h v l ∥W (k,l) h u l ])
[0106] α v,u (k,l) =exp(e v,u (k,l) ) / ∑j∈N(v)∪{v} exp(e v,j (k,l) )
[0107] Final output H t GAT =[h1 LGAT ,…,h ∣V∣ LGAT ] is the graph attention encoding for all nodes at time step t.
[0108] ②Intent Fusion Layer: h v Intent =FC(Concat(h v GAT I v Concat is a vector concatenation layer, and FCFC is a fully connected layer used to fuse the intent target.
[0109] ③ Temporal Convolutional Network (TCN): For all nodes in T in Intent-enhanced embedding H at each time step Intent =[H t−Tin+1 Intent H t Intent The TCN layer uses causal convolution and residual connections for processing. t TCN =TCN(H Intent ), C t TCN This represents the global context features for capturing temporal dependencies.
[0110] ④ Attention Decoder:
[0111] Embed the current intent-enhancing node into H t Intent Timing Context C t TCN And the expert experience embedding related to intent in MM-NKG expert (Obtained via MM-NKG query) as input. DecInput t =Concat(H t Intent C t TCN E expert )
[0112] The decoder uses a feedforward network or LSTM layer to process DecInput. t Process it and output the future T. out Predictive performance P^ at each time step t+1 , ..., P^ t+Tout .
[0113] Simultaneously, the intention violation probability P is output through a sigmoid activation layer. violate ∈[0,1]:
[0114] P^ future =FC pred (DecInput t )
[0115] P violate =Sigmoid(FC violate (DecInput t ))
[0116] ⑤ Loss function:
[0117] L=λ1⋅MSE(P^ future ,P true )+λ2⋅BCE(P violate ,Y violate )+λ3⋅KL(MM-NKG,PredictedRelations)+Ω(W)
[0118] Where MSEMSE is the mean squared error (for prediction performance), BCEBCE is the binary cross-entropy (for intent violation), and Y violate λ1 is the label (0 or 1) indicating whether the actual intent is violated, KLKL is the KL divergence (used to encourage the model to maintain a structure and semantics consistent with MM-NKG during prediction), and Ω(W) is the L2 regularization term. λ1, λ2, and λ3 are adjustable weight parameters.
[0119] Furthermore, the IGATP model is the innovative aspect of this invention for achieving accurate prediction and risk assessment. It uses MM-NKG as a graph structure to integrate business intent into the prediction process.
[0120] Model details and data examples to aid understanding:
[0121] Suppose we have a simplified network consisting of device A (router), device B (server), link L (connecting A and B), and a service S (e-commerce transaction) running on B. Our goal is to "ensure that the API response time P99 of service S is <100ms".
[0122] a. Graph structure input: MM-NKG provides the graph structure. Nodes include A, B, L, and S.
[0123] Adjacency matrix A: Describes the connection relationships between nodes.
[0124] Node feature matrix F t At time step t, the feature vector f of node Bt B This may include: CPU utilization, memory utilization, current API response time P99 for service S, and current API throughput. The feature vector f of node L. t L This may include: bandwidth utilization, packet loss rate, and latency. The feature vector f of node S. t S Possible states include: current intent violation status (0 / 1).
[0125] Intent target vector I v For node S, its intended target I S =[100ms] (Target API response time P99). Other node intent targets are 0 or empty.
[0126] b. Graph Attention Encoder (GAT Layer):
[0127] The goal of the GAT layer is to capture the dynamic spatiotemporal relationships between nodes.
[0128] For example, to compute the embedding h of node B B GAT GAT will consider the characteristics of its neighboring nodes (such as device A, link L, service S).
[0129] Calculation process:
[0130] Initial node characteristics: h A 0 =f t A h B 0 =f t B h L 0 =f t L h S 0 =f t S .
[0131] Attention weight calculation: For example, calculating the attention weight of node B to node L: e B,L (k,l) =LeakyReLU(a (k,l) T [W (k,l) h B 0 ∥W (k,l) h L 0 Assume that h B 0=[0.8,0.7,90,1000] (CPU, Mem, P99,Throughput), h L 0 =[0.9,0.05,20] (Bandwidth, Loss, Latency) After linear transformation W (k,l) The results are then concatenated and processed using LeakyReLU to obtain a non-normalized attention score. α B,L (k,l) =exp(e B,L (k,l) ) / ∑ j∈N(B)∪{B} exp(e B,j (k,l) ).
[0132] Assuming α is calculated B,L =0.6 indicates that the current state of node B is greatly affected by the state of link L.
[0133] Neighbor aggregation: h B GAT =ReLU(α B,A ⋅Wh A 0 +α B,L ⋅Wh L 0 +α B,S ⋅Wh S 0 +α B,B ⋅Wh B 0 This process iterates multiple times, capturing broader neighbor information. The final output is the H of all nodes. t GAT .
[0134] c. Intent Fusion Layer:
[0135] Intent Target I v Node embedding with GAT output h v GAT The components are joined together and merged through a fully connected layer. For example, for business S, its h... S GAT Will with I S =[100ms] concatenation:
[0136] h S Intent =FC(Concat(h S GAT [100ms]))
[0137] This operation allows the model to directly "consider" the target SLA of business S when making subsequent predictions.
[0138] d. Temporal Convolutional Networks (TCN):
[0139] TCN processing T in H at each time step Intent Sequence. For example, if T in =60 (data from the past hour). TCN will use multi-layer causal convolution to extract the temporal evolution pattern of the intended enhancement features of each node in the past hour.
[0140] C t TCN =TCN([H t−59 Intent H t Intent ])
[0141] C t TCN It is a high-dimensional vector that summarizes the global temporal context over a past period.
[0142] e. Attention Decoder:
[0143] Decoder receives:
[0144] Current intent-enhanced node embedding H t Intent .
[0145] Timing Context C t TCN .
[0146] Expert experience queried and embedded from MM-NKG expert For example, if high CPU utilization and slow API response are detected, MM-NKG might associate it with expert-level embeddings related to "optimizing database indexes." (DecInput) t =Concat(H t Intent C t TCN E expert )
[0147] Prediction performance: P^ future =FC pred (DecInput t Assuming a prediction of the next 5 minutes (T) out =5) API response time of business S P99: P^ t+1 S,P99 =95ms, P^ t+2 S,P99 =105ms, P^ t+3S,P99 =110ms ...
[0148] Probability of intention violation: P violate =Sigmoid(FC violate (DecInput t Combined with intentional goal I S =[100ms] and predicted P^ future S,P99 The model predicts that within the next 5 minutes, there is an 80% probability that the API response time P99 of business S will violate the intended target of 100ms. That is, P... violate =0.8.
[0149] f. Loss function:
[0150] The goal of model training is to minimize the loss function L.
[0151] Assumption:
[0152] Real performance P true =[92,103,108,…]
[0153] Actual intention to violate Y violate =1 (if the predicted value does indeed exceed 100ms)
[0154] KL divergence is used to maintain the consistency of the network's structural information during model learning, preventing the model's predictions from deviating from the network's actual dependencies. Backpropagation is used to adjust model parameters, making the predictions closer to the true values while accurately identifying intent violations.
[0155] e. Multi-objective intelligent orchestration and feedback module: It is used to automatically generate and execute cross-domain and cross-level network resource and configuration orchestration schemes based on the prediction results of the IGATP module, network resource information in MM-NKG, and the goals and constraints set by the adaptive service intent parsing module, and continuously optimize the orchestration strategy through reinforcement learning;
[0156] It should be noted that the multi-objective intelligent orchestration and feedback module includes:
[0157] 1. Policy Space Generator: Based on the network resource capabilities and topology constraints in MM-NKG, as well as the predicted intent violation risks, it generates a set of candidate optimization orchestration actions, such as bandwidth adjustment, route switching, service expansion, priority adjustment, configuration parameter modification, etc.
[0158] 2. Multi-objective optimizer: Based on multiple objectives such as predicted intent violation probability, resource utilization, cost, and system load, a multi-objective optimization algorithm (such as NSGA-II or MOPSO) is used to select the optimal orchestration scheme from candidate strategies.
[0159] 3. Reinforcement Learning Strategy Optimization Unit: Through reinforcement learning (such as DQN, PPO) and interaction with the environment (simulated network or real network), the execution effect of the orchestration scheme is used as a reward signal to continuously learn and optimize the parameters of the strategy generator in order to adapt to the dynamic changes of the network and improve the optimization effect.
[0160] Furthermore, the reinforcement learning strategy optimization unit is performed in the following ways:
[0161] ①State definition: Current network performance indicators, MM-NKG status, service intent status, IGATP prediction results, historical orchestration actions and effects;
[0162] ② Action Space: A predefined set of network orchestration actions, such as "increase the bandwidth of link A by 10%", "switch the traffic of service B to route C", and "expand the capacity of service D instance by 1".
[0163] ③ Reward Function: This function comprehensively considers factors such as the degree to which business objectives are met (e.g., whether latency and packet loss rates satisfy the SLA), resource utilization, operating costs, and side effects (e.g., impact on other business operations). For example:
[0164] R t =w1⋅(1−P violate )−w2⋅C resource −w3⋅C op −w4⋅D side_effect
[0165] Among them, P violate C is the probability of predicting or actually violating the intended outcome. resource For resource consumption costs, C op For operating costs, D side_effect As a punishment for side effects, w i These are the weighting coefficients.
[0166] ④ Learning algorithm: Train a policy network using algorithms such as Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) to output the optimal orchestration action in the current state.
[0167] To further explain: the multi-objective intelligent orchestration and feedback module receives the prediction results of IGATP (especially the probability of intent violation and specific performance prediction values), and uses the network resources and topology information provided by MM-NKG to generate the optimal orchestration scheme through reinforcement learning.
[0168] Policy Space Generator: Based on the predicted intent violation risk (e.g., the API response time P99 of service S will exceed 100ms), the system queries MM-NKG to identify all network resources and potential optimization actions related to service S. Example: If MM-NKG shows that service S is running on server B, and server B is connected to router A via link L, candidate actions:
[0169] Scaling up: Add an instance of server B (assuming it's based on Kubernetes).
[0170] Traffic routing: Distribute some of the traffic from service S to other servers or regions.
[0171] QoS Adjustment: Increase the QoS priority of service S on link L.
[0172] Configuration optimization: Adjust the configuration parameters of a certain application on server B (such as the database connection pool size).
[0173] Bandwidth adjustment: Request to increase the bandwidth of link L.
[0174] Multi-objective optimizer: The reinforcement learning agent plays the role of decision-maker here. It attempts to weigh trade-offs among multiple objectives:
[0175] Maximize the achievement rate of intent (reduce P) violate )
[0176] Minimize resource costs (CPU, memory, bandwidth usage)
[0177] Minimize operational side effects (impact on other services, risk of service interruption).
[0178] Response speed (execution time of the orchestration scheme) Reward function example: Current state: Business S's API P99 predicts that it will reach 105ms in the next 5 minutes, P violate =0.8. Current resource cost C resource =100, no side effects. Agent selection action A: Expand server B instance by 1. After execution (or simulated execution in a simulation environment): New state: API P99 of business S actually drops to 90ms, P violate =0.1. Resource cost becomes C. resource =120 (Cost of adding an instance). Reward calculation (assuming w1=10, w2=0.5, w3=1, w4=5): R=10⋅(1−0.1)−0.5⋅(120−100)−1⋅5−5⋅0=10⋅0.9−0.5⋅20−5=9−10−5=−6 This reward function guides the agent to choose actions that significantly reduce the probability of intent violation, while controlling resource costs and side effects.
[0179] Reinforcement Learning Policy Optimization Unit: A policy network is trained using the DQN or PPO algorithm. This network takes the current network state (including MM-NKG information, IGATP prediction, and intent state) as input and outputs an optimal action probability distribution. The system continuously learns and optimizes its orchestration policy by interacting with real networks or high-fidelity simulators (exploration-exploitation). Example: In a simulated environment, if DQN finds that "immediately scaling up server instances when P99 prediction exceeds 95ms" yields a higher long-term reward than "scaling up after P99 actually exceeds 100ms," then this policy will be reinforced.
[0180] f. Optimize execution and feedback interface: Used to interact with actual network devices, SDN controllers, or cloud platform APIs, execute orchestration schemes, and feed back execution results and network performance changes to the system for iterative learning.
[0181] It should be noted that the optimized execution and feedback interface supports Restful API, Netconf / YANG, OpenFlow protocol and vendor SDK, enabling automated orchestration of SDN controllers, physical network devices, cloud platforms and virtualization resources.
[0182] It should be further explained that this module is responsible for converting the orchestrated actions generated by the reinforcement learning agent into actual network operations.
[0183] API / protocol integration: For example, if the policy is "scaling up server B instances", the system will call the Kubernetes API or cloud platform API to achieve this. If the policy is "adjusting the QoS priority of link L", the system will interact with router A through the Netconf / YANG protocol.
[0184] Execution verification: After the orchestration action is completed, the system will immediately monitor the affected performance indicators through the data acquisition module to verify the optimization effect.
[0185] Feedback learning: Optimized performance data and intent achievement are fed back as new states and rewards to the reinforcement learning agent for further training of the policy network. If optimization fails or produces side effects, negative rewards are given to encourage the agent to avoid such actions in the future. Simultaneously, this new data also updates the MM-NKG and IGATP models, forming a closed loop of continuous iterative optimization.
[0186] To better illustrate the technical solution of the present invention, please refer to... Figure 1 It also provides a network performance optimization method based on data analysis, including the following steps:
[0187] a. Multimodal network data acquisition: Acquire multimodal heterogeneous data from multiple sources, including network devices, application systems, user terminals, service management platforms, and operation and maintenance knowledge bases;
[0188] b. Construction and management of multimodal network knowledge graph (MM-NKG): Entity recognition, relation extraction and semantic association are performed on the collected data to construct and dynamically update the MM-NKG;
[0189] c. Adaptive service intent parsing: Parses high-level service requirements into quantified network performance goals and constraints, and associates them with MM-NKG;
[0190] d. Intent-driven graph attention temporal prediction: Using MM-NKG as the backbone, integrating business intent, and predicting network performance trends and intent violation risks through the IGATP model;
[0191] e. Multi-objective intelligent orchestration and feedback: Based on prediction results, MM-NKG information and intent objectives, automatically generate and execute cross-domain and cross-level network resource and configuration orchestration schemes, and continuously optimize orchestration strategies through reinforcement learning;
[0192] f. Optimize execution and feedback: Execute the orchestration scheme through the interface and feed back the execution results and network performance changes to the system for iterative learning.
[0193] This solution has the following beneficial effects:
[0194] 1) Introduce an adaptive service intent parsing module to transform high-level service requirements into actionable network performance targets;
[0195] 2) Construct a multimodal network knowledge graph (MM-NKG) that integrates multimodal heterogeneous data such as network topology, configuration, performance, services, and expert experience to establish complex relationships between entities;
[0196] 3) Based on MM-NKG and the Intent-driven Graph Attention Temporal Prediction Model (IGATP), accurate prediction of network performance is achieved, and potential risks are predicted in conjunction with the intent target;
[0197] 4) Multi-objective intelligent orchestration and feedback module: Based on prediction results and business intent, it automatically generates and executes cross-domain and cross-level network resource and configuration orchestration schemes, and continuously optimizes the orchestration strategy through reinforcement learning.
[0198] This invention creatively combines business intent, knowledge graphs, and deep learning to achieve intelligent network autonomy from "data-based" to "intent-based," significantly improving network service quality, resource utilization, and user experience.
[0199] To verify the technical effects of this invention, the following experimental report is presented:
[0200] 1. Experimental Objective
[0201] This experiment aims to verify the superiority of the "Intent-Driven Knowledge Graph Network Performance Prediction and Active Orchestration System" (hereinafter referred to as "the System of this Intent" or "ID-KOS") described in this invention compared to existing technologies. Specific verification objectives include:
[0202] Prediction accuracy: Verify the accuracy advantage of the IGATP model in the system of this invention over traditional time series prediction models (such as LSTM) in predicting network performance.
[0203] Proactive optimization and risk avoidance capabilities: This demonstrates the superiority of the present invention's system in proactively warning and automatically orchestrating business intentions before they are violated, compared to passive threshold alarm systems.
[0204] Overall Business Assurance Effectiveness: By simulating real business scenarios, the overall effectiveness of the system of this invention in ensuring business SLAs (Service Level Agreements), improving resource utilization, and reducing manual intervention is comprehensively evaluated.
[0205] 2. Test Environment and Setup
[0206] Hardware environment:
[0207] Servers: 3 Dell PowerEdge R740 servers (Intel Xeon Gold 6248R CPU, 256GB RAM, 2TB NVMe SSD).
[0208] Network equipment: 2 H3C S6800 switches and 1 Cisco ASR 1001-X router, configured as the core network environment.
[0209] Traffic generator: Uses Ixia IxNetwork devices to simulate real network traffic.
[0210] Software environment:
[0211] Virtualization / Containerization: Kubernetes (K8s) v1.28 cluster, used to deploy simulated business applications.
[0212] Database: Neo4j v5.13 is used to build a multimodal network knowledge graph (MM-NKG).
[0213] Data acquisition and monitoring: Prometheus + Grafana + APM probe (SkyWalking).
[0214] Model training and inference: Python 3.10, PyTorch 2.1, PyTorch Geometric, DGL.
[0215] Automated execution: Ansible + self-developed API gateway interacts with SDN controller.
[0216] Simulated Business Scenario: Deploy a simulated e-commerce microservice application in a Kubernetes cluster, including a user gateway, product service, order service, and payment service. Core Business Intent: Define "ensuring high availability and low latency of the payment service during peak sales periods," which is then parsed into the following quantitative indicators by the intent parsing module of this invention:
[0217] Payment service API response time P99 < 100ms
[0218] Payment success rate > 99.9%
[0219] 3. Comparison of Schemes
[0220] To fully demonstrate the superiority of the system of this invention, we set up the following three schemes for comparative testing:
[0221] Option 1: Traditional Threshold-based System (TTS)
[0222] Mechanism: Based on Prometheus Alertmanager, a static threshold (e.g., 100ms) is set for the payment service API response time P99. When the metric exceeds the threshold, the system issues an alarm, requiring manual intervention.
[0223] Representative technologies: traditional passive, rule-based network monitoring.
[0224] Option 2: Standard ML-based Prediction System (SMPS)
[0225] Mechanism: Collect single time-series data of the payment service API response time P99, and use a standard LSTM (Long Short-Term Memory) model to predict future trends. When the predicted value exceeds a threshold, an alert is issued in advance, but optimization decisions and execution still require manual intervention.
[0226] Representative technology: The current mainstream AIOps prediction and alerting solution based on a single metric.
[0227] Option 3: The Intent-Driven Knowledge Orchestration System (ID-KOS) of this invention
[0228] Mechanism: As described in the invention, an MM-NKG is constructed to parse business intent, and an IGATP model is used to perform multi-dimensional association prediction and intent violation risk prediction. Optimization strategies (such as service expansion and traffic scheduling) are automatically executed through a reinforcement learning-driven multi-objective orchestration module.
[0229] 4. Test Scenario and Evaluation Indicators
[0230] We designed two typical test scenarios to simulate real-world challenges.
[0231] Scenario 1: Predictable traffic peaks
[0232] Description: Simulating a major e-commerce promotional event, the request traffic for the payment service smoothly increases from 1000 QPS to 8000 QPS within 30 minutes, and then drops back down. This scenario is used to test the system's predictive and proactive scaling capabilities.
[0233] Scenario 2: Hidden Dependency Service Failure
[0234] Description: The payment service relies on a separate "coupon service." Under stable traffic (2000 QPS), a slow query occurred in the simulated coupon service's database, causing its API response time to gradually degrade from 20ms to 200ms. This scenario is used to test the system's ability to perform root cause analysis and intelligent decision-making for complex, interconnected failures.
[0235] Evaluation indicators:
[0236] Prediction accuracy metrics: Root Mean Square Error (RMSE) and Mean Absolute Error (MAE), used to measure the difference between predicted and actual values.
[0237] Proactively optimize metrics:
[0238] Proactive Alert Lead Time: The time difference (in seconds) between the system issuing a valid alert and the actual violation of the SLA. A negative value indicates an early warning.
[0239] Mean Time to Mitigate (MTTM): The time from when a problem occurs until the system completes optimization and performance returns to normal (in seconds).
[0240] Business Assurance Metrics:
[0241] SLA Compliance Rate: The percentage of time during which the payment service API response time (P99) is less than 100ms throughout the entire testing period.
[0242] Resource Overhead: The ratio of additional computing resources (in Core Hours) required by the system to cope with traffic surges or failures to the ideal minimum required resources.
[0243] Manual Interventions: The number of events requiring manual intervention during the entire testing process.
[0244] 5. Experimental Process and Results Analysis
[0245] 5.1 Comparison of Prediction Accuracy (for Scenario 1)
[0246] We had both Scheme 2 (SMPS) and Scheme 3 (ID-KOS) predict the payment service API response time P99 for the next 5 minutes and compare the results with the actual values. Scheme 1 (TTS) has no predictive capabilities and is not included in this comparison.
[0247] Table 1: Comparison of P99 Prediction Accuracy for API Response Time
[0248] plan Predictive Model Data input RMSE (ms) MAE (ms) Option 2 (SMPS) LSTM Single metric: Payment service API response time (P99) 12.54 9.81 Option 3 (ID-KOS) IGATP MM-NKG (contains multi-dimensional correlated data such as traffic, CPU, and performance of dependent services) 5.82 3.97
[0249] Results Analysis: As shown in Table 1, the IGATP model of the system (ID-KOS) of this invention significantly outperforms the standard LSTM model in prediction accuracy. Its RMSE is reduced by 53.6%, and MAE by 59.5%. This is mainly attributed to:
[0250] Multidimensional data input: The IGATP model utilizes MM-NKG, which not only considers the API response time itself, but also integrates multidimensional correlation features such as upstream traffic, its own CPU / memory utilization, and the performance of downstream dependent services (such as order services).
[0251] Graph Structure Correlation Analysis: The graph attention network in IGAT can learn the non-linear relationship between "traffic growth", "CPU consumption" and "API latency", thus more accurately predicting the subsequent latency increase trend in the early stage of traffic growth.
[0252] 5.2 Comparison of proactive optimization and risk avoidance capabilities
[0253] We documented the performance of each solution from problem identification to problem resolution in two scenarios.
[0254] Table 2: Comparison of Proactive Optimization Capabilities in Scenario 1 (Traffic Peak)
[0255] plan Risk warning lead time (s) Triggering mechanism Optimize actions Mean time to remission (MTTM) (s) Option 1 (TTS) N / A (Alarm delay +35s) Static threshold (P99 > 100ms) Manual assessment and manual capacity expansion 360 Option 2 (SMPS) -90 Predicted value > 100ms Alarms, manual judgment, manual capacity expansion 240 Option 3 (ID-KOS) -180 Probability of violation of intent > 70% Automatic orchestration: service expansion 45
[0256] Table 3: Comparison of Proactive Optimization Capabilities in Scenario 2 (Dependency Service Failure)
[0257] plan Risk warning lead time (s) Triggering mechanism / root cause analysis Optimize actions Mean time to remission (MTTM) (s) Option 1 (TTS) N / A (Alarm delay +60s) Static threshold (P99 > 100ms) cannot pinpoint the root cause. Manually check multiple services 900+ Option 2 (SMPS) N / A (Cannot be predicted effectively) The predictive model failed because the payment service itself showed no obvious precursors. Alarm, manual investigation 900+ Option 3 (ID-KOS) -120 MM-NKG Correlation Analysis: Increased Delays in Coupon Services are Strongly Correlated with Risk of Intent Violation in Payment Services Automatic orchestration: Service circuit breaking / degradation 60
[0258] Results analysis:
[0259] Scenario 1: The system of this invention (ID-KOS) predicts risks 180 seconds in advance and automatically completes capacity expansion, while MTTM takes only 45 seconds, resolving the problem almost imperceptibly to the user. In contrast, the TTS system only alerts 35 seconds after an SLA violation, resulting in lengthy manual processing. Although SMPS can provide early warnings, it still relies on manual decision-making, placing its efficiency in the middle range.
[0260] Scenario 2: This scenario fully demonstrates the core advantages of the system of this invention. The LSTM model of SMPS, due to its focus on only a single indicator, cannot predict the risk of the "payment service" from the deterioration of the "coupon service," thus its prediction fails. While the TTS system issues an alert, it cannot pinpoint the root cause, resulting in extremely long manual investigation times. In contrast, the ID-KOS system of this invention, through the "dependency" relationship of MM-NKG, quickly locates the root cause in the coupon service and issues a warning 120 seconds in advance. Simultaneously, it automatically executes service circuit breaking or degradation (e.g., temporarily disabling the coupon function to ensure the core payment process), with MTTM taking only 60 seconds, demonstrating powerful complex fault handling capabilities.
[0261] 5.3 Comparison of Comprehensive Business Support Effectiveness
[0262] Finally, we compiled the overall business performance of each solution throughout the entire trial period (including two scenarios).
[0263] Table 4: Comparison of Overall Business Support Effectiveness and Resource Consumption
[0264] plan SLA compliance rate (%) Resource expenditure (%) Number of manual interventions Option 1 (TTS) 99.15% 35% (to maintain a high water level buffer in order to cope with emergencies) 12 Option 2 (SMPS) 99.72% 28% (based on forecasts, but still requires manual buffering) 7 Option 3 (ID-KOS) 99.98% 18% (Accurate forecast, dynamically adjusted as needed) 1 (System-suggested complex strategies requiring manual confirmation)
[0265] Results analysis:
[0266] SLA Compliance Rate: The SLA compliance rate of this invention system (ID-KOS) reached 99.98%, almost perfectly guaranteeing business intent. This is thanks to its accurate prediction and fast, automated orchestration capabilities.
[0267] Resource overhead: ID-KOS has the lowest resource overhead, at only 18%. This is because it can schedule resources "just right"—expanding capacity in advance when needed and scaling down promptly after a risk has passed. This avoids the TTS system from maintaining a high level of resource redundancy for a long time due to unpredictability, and is also superior to the SMPS system in terms of resource waste caused by the lag in human decision-making.
[0268] Number of manual interventions: ID-KOS achieves a high degree of automation, reducing the number of manual interventions from 12 (TTS) and 7 (SMPS) to only 1. This greatly frees up maintenance manpower and reduces the risk of human error.
[0269] 6. Experimental Conclusions
[0270] Based on the above experiments and data comparison analysis, the following conclusions can be drawn:
[0271] The "Intent-Driven Knowledge Graph Network Performance Prediction and Active Orchestration System (ID-KOS)" proposed in this invention exhibits significant superior technical performance compared to traditional threshold alarm systems and standard machine learning prediction systems in the following aspects:
[0272] More accurate predictions: By integrating multi-dimensional relational information from knowledge graphs, the IGATP model's prediction accuracy is far higher than that of traditional models that rely solely on single time-series data.
[0273] More proactive response: The system can anticipate the risk of intent violation several minutes in advance and reduce the mean time to mitigation (MTTM) from minutes or even hours (manual handling) to seconds.
[0274] Smarter decision-making: It can handle complex dependency failure scenarios, quickly locate the root cause through MM-NKG and take the optimal orchestration strategy, which is something that traditional solutions cannot achieve.
[0275] Superior Results: Ultimately, it achieved a higher SLA compliance rate, better resource utilization, and the lowest cost of human intervention, fully demonstrating its enormous potential and practical value in building efficient, reliable, and intelligent autonomous networks. Specific Implementation
[0276] 1. Business Intent and Scenario Setting
[0277] Business Scenario: The online payment process of a large e-commerce platform. This process is handled by the core "Payment Service (Payment-Svc)," which calls the "Coupon Service (Coupon-Svc)" to apply discounts before making the payment.
[0278] Senior management's business objective: "To ensure a smooth payment process and a seamless user experience."
[0279] Problem Simulation: Under stable business traffic conditions, the database (Coupon-DB) upon which the "Coupon Service" relies begins to experience slow queries due to index failure. This causes the "Coupon Service" to respond slowly, which in turn affects the performance of the "Payment Service," ultimately threatening the core business objectives.
[0280] 2. Construction of Multimodal Network Knowledge Graph (MM-NKG)
[0281] The system first constructs an MM-NKG containing relevant entities. For simplicity, we only show the core entities and their relationships.
[0282] Entities (Nodes):
[0283] S1: Payment Service (Payment-Svc)
[0284] S2: Coupon Service (Coupon-Svc)
[0285] DB2: Coupon Database
[0286] H1: Host A (running S1)
[0287] H2: Host B (running S2)
[0288] L1: Network link (link between H1 and H2)
[0289] Edges:
[0290] S1 -[DEPENDS_ON]-> S2
[0291] S2 -[DEPENDS_ON]-> DB2
[0292] S1 -[HOSTED_ON]-> H1
[0293] S2 -[HOSTED_ON]-> H2
[0294] H1 -[CONNECTS_TO]-> L1
[0295] H2 -[CONNECTS_TO]-> L1
[0296] Adjacency Matrix A: Represents the graph structure of the MM-NKG. A 1 in the matrix indicates that there is a direct relationship between two nodes.
[0297] S1 S2 DB2 H1 H2 L1 S1 0 1 0 1 0 0 S2 1 0 1 0 1 0 DB2 0 1 0 0 0 0 H1 1 0 0 0 0 1 H2 0 1 0 0 0 1 L1 0 0 0 1 1 0
[0298] 3. Multimodal data acquisition and characterization
[0299] The system collects the performance metrics of each entity at consecutive time steps (t-2, t-1, t, each time step being 1 minute) and constructs feature vectors.
[0300] Feature definition:
[0301] P99_Latency (ms): P99 response time
[0302] CPU_Util (%): CPU utilization
[0303] Throughput (QPS): Queries per second
[0304] Bandwidth_Util (%): Bandwidth utilization
[0305] Slow_Query_Count: The number of slow queries.
[0306] Error_Rate (%): Error rate
[0307] Feature Matrix (F_t): At time step t, the feature vectors of each node are shown in the table below. Observe the changing trends of Slow_Query_Count in DB2 and P99_Latency in S2 and S1.
[0308] The node feature matrix F at time step t t
[0309] node P99_Latency CPU_Util Throughput Bandwidth_Util Slow_Query_Count Error_Rate S1 (Payment) 85 35 2000 N / A N / A 0.01 S2 (Coupon) 75 30 2000 N / A N / A 0.01 DB2 (Coupon) 65 45 2000 N / A 30 0 H1 (Host-A) N / A 40 N / A N / A N / A N / A H2 (Host-B) N / A 38 N / A N / A N / A N / A L1 (Link) 2 N / A N / A 25 N / A 0
[0310] Historical data trends:
[0311] At t-2: DB2.Slow_Query_Count=5, S2.P99_Latency=30ms, S1.P99_Latency=45ms.
[0312] At t-1: DB2.Slow_Query_Count=15, S2.P99_Latency=50ms, S1.P99_Latency=65ms.
[0313] At t (current): DB2.Slow_Query_Count=30, S2.P99_Latency=75ms, S1.P99_Latency=85ms.
[0314] 4. Adaptive Business Intent Parsing
[0315] The senior management's intention: "To ensure a smooth payment process and a seamless user experience."
[0316] Analysis results:
[0317] Related entity: S1:Payment-Svc
[0318] Performance target: P99_Latency < 100ms
[0319] Intent Target Vector (I): Generates an intent vector for each node that matches the feature dimension. For S1, the position corresponding to P99_Latency is 100, and the rest are NaN. The intent vectors of other nodes are all NaN.
[0320] I^S1 = [100, NaN, NaN, NaN, NaN, NaN]
[0321] 5. Detailed Explanation of the IGATP Model Prediction Process
[0322] The system uses A and F t (And historical F sequence) and I are used as inputs for prediction.
[0323] 5.1 Graph Attention Encoder (GAT Layer)
[0324] The core of GAT is to compute the attention weights between nodes in order to learn their mutual influence. We take the computation of the new embedding hS1^GAT of S1 at time t as an example.
[0325] Input: Feature vectors of S1's neighbor nodes S2 and H1.
[0326] Attention calculation (conceptual): The model uses a multi-head attention mechanism to calculate S1's attention to itself and its neighbors. After training, the model learns:
[0327] The performance of S1 (especially latency) is highly correlated with the performance of S2 (because S1 depends on S2).
[0328] The performance of S1 is also affected by the resource status of its host H1.
[0329] Assume the calculated attention weights α are as follows:
[0330] α(S1, S1) (Self-attention): 0.2
[0331] α(S1, S2): 0.6 (High attention to S2)
[0332] α(S1, H1): 0.2
[0333] Neighbor information aggregation: The new embedding hS1^GAT of S1 is a weighted sum of the features of its neighbors.
[0334] hS1^GAT=σ(α(S1,S1)·W·ft^S1+α(S1,S2)·W·ft^S2+α(S1,H1)·W·ft^H1)(where W is the learnable weight matrix and σ is the activation function). Key insight: Due to the high weight of α(S1, S2), the feature vector of S2 (especially its deteriorating 75ms latency) will be largely incorporated into the new representation of S1. GAT, through its graph structure, allows S1 to "sense" the state changes of its dependency S2. Similarly, S2 will also "sense" the state changes of DB2.
[0335] 5.2 Intent Fusion Layer
[0336] The embedded hS1^GAT of S1 after GAT encoding is concatenated and linearly transformed with its intention target vector I^S1. The operation hS1^Intent = FC(Concat(h_S1^GAT, I^S1)) encodes the information "the target is 100ms" into the features of S1, allowing the model to be aware of this constraint in subsequent predictions.
[0337] 5.3 Temporal Convolutional Networks (TCN)
[0338] TCN processes the h^Intent sequence of all nodes' past (t-2, t-1, t). TCN's causal convolution and long receptive field characteristics enable it to capture:
[0339] The growth trend of DB2.Slow_Query_Count.
[0340] S2.P99_Latency and S1.P99_Latency represent the corresponding delay ramp-up modes. The TCN outputs a context vector Ct^{TCN} that contains this temporal dynamic.
[0341] 5.4 Attention Decoder & Predicted Output
[0342] The decoder combines ht^Intent, Ct^{TCN}, and expert experience that may be extracted from MM-NKG (e.g., the embedding that "high slow query count is usually related to index failure") to generate the final prediction.
[0343] Performance prediction (Pfuture):
[0344] t+1 (next minute): S1.P99_Latency predicted value is 98ms.
[0345] t+2 (2 minutes from now): S1.P99_Latency is predicted to be 110ms.
[0346] t+3 (future 3 minutes): S1.P99_Latency predicted value is 125ms.
[0347] Probability of Intent Violation (Pviolate): Based on the above prediction, the model calculates the probability that S1's business intent (P99 <100ms) will be violated within the next 5 minutes.
[0348] Pviolate(S1) = 0.92 (92%)
[0349] 6. Multi-objective intelligent orchestration and feedback
[0350] 6.1 Triggering and Root Cause Analysis
[0351] Trigger: When P_violate(S1) > 0.8 (preset threshold), the system triggers the active orchestration process.
[0352] Root cause analysis revealed that by analyzing the internal attention weights and feature importance of the IGATP model, the features contributing most to S1 delay prediction were, in order: S2.P99_Latency, DB2.Slow_Query_Count, and the historical trend of S2.P99_Latency. The MM-NKG dependency chain S1 -> S2 -> DB2 clearly indicated the fault propagation path. The system generated an alert: "Payment service intent is about to be violated; the root cause is highly suspected to be an increase in slow queries in the coupon database (DB2)."
[0353] 6.2 Policy Space Generation and Reinforcement Learning Decision Making
[0354] Reinforcement learning (RL) agents receive the current state and root cause cues, and generate and evaluate candidate actions.
[0355] Candidate Actions (Action Space):
[0356] A1: Expanding the S1 (payment service) instance. (The traditional operations and maintenance approach's first reaction, but it's only a temporary solution.)
[0357] A2: Execute a service circuit breaker when S1 calls the interface of S2, and return the default value (no discount). (This ensures the core payment process, but sacrifices some functionality.)
[0358] A3: Notify the Database Administrator (DBA) team, along with evidence of the increased slow queries in DB2. (A radical solution, but time-consuming)
[0359] A4: Restart the S2 (coupon service) instance. (Invalid action)
[0360] RL Policy Network Output (Action Scores): A pre-trained RL model (such as DQN) estimates a Q-value (long-term reward expectation) for each action.
[0361] action Q value (estimated) reason A1 (Expansion S1) -5.2 Simulated environment training shows that scaling up S1 has little effect on alleviating latency caused by slow queries that depend on downstream dependencies, and it is a waste of resources. A2 (Circuit Breaker S2) +8.5 It can quickly restore S1 latency to normal levels, ensuring the core intent, with manageable side effects (no discounts available at the moment), making it the best immediate relief measure. A3 (Notify the DBA) +7.0 It is a necessary step to resolve the problem, but it is not an immediate action and is usually performed in combination with A2. A4 (Reboot S2) -10.0 Simulated environment training shows that restarting the service is ineffective in addressing the slow database query problem.
[0362] 6.3 Automatic Orchestration Execution and Feedback
[0363] Decision-making and execution: The system selects the action combination with the highest Q value and automatically executes A2 and A3.
[0364] Execute A2: Immediately configure the call from S1 to S2 to the circuit breaker state via the API of the service mesh (such as Istio), set the timeout to a very short time (such as 10ms), and trigger fast failure.
[0365] Execute A3: Automatically create a high-priority work order via the work order system API or alarm channel (such as Slack), assign it to the DBA team, and attach a detailed root cause analysis report and data charts.
[0366] Effect verification: The system continuously monitors after the action is executed.
[0367] At time step t+1, S1.P99_Latency drops rapidly to 40ms (because it no longer waits for the slow S2).
[0368] With business intent secured, Pviolate decreased to 0.05.
[0369] Feedback learning: This successful "diagnosis-decision-execution" process (state, action, reward, new state) was used as high-quality experience data to update and fine-tune the RL policy network, enabling it to make more accurate and faster decisions when facing similar scenarios in the future.
[0370] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A network performance optimization system based on data analysis, characterized in that, Includes the following components: a. Multimodal network data acquisition module: used to collect multimodal heterogeneous data from network devices, application systems, user terminals, business management platforms, and operation and maintenance knowledge bases; b. Multimodal Network Knowledge Graph (MM-NKG) Construction and Management Module: Used to perform entity recognition, relation extraction and semantic association on the collected data, and build and dynamically update a knowledge graph that integrates network topology, performance, business and expert experience; c. Adaptive Service Intent Parsing Module: Used to parse high-level service requirements into a set of quantifiable network performance goals and constraints, and associate them with service entities in MM-NKG; d. Intent-Driven Graph Attention Temporal Prediction Model (IGATP) module: Used as the backbone of MM-NKG, it integrates the results of business intent parsing to predict the future trend of network performance indicators and identify potential intent violation risks; e. Multi-objective intelligent orchestration and feedback module: It is used to generate and execute network resource and configuration orchestration schemes based on the prediction results of the IGATP module, network resource information in MM-NKG, and the objectives and constraints set by the service intent parsing module, and continuously optimize the orchestration strategy through reinforcement learning; f. Optimize execution and feedback interface: Used to interact with actual network devices, SDN controllers, or cloud platform APIs, execute orchestration schemes, and feed back execution results and network performance changes to the system.
2. The network performance optimization system based on data analysis according to claim 1, characterized in that, A multimodal network knowledge graph (MM-NKG) includes at least the following entity types: devices, interfaces, links, services, applications, performance metrics, and expert experience; and at least the following relationship types: connections, bearers, dependencies, configurations, generation, and impact.
3. The network performance optimization system based on data analysis according to claim 2, characterized in that, The adaptive business intent parsing module includes: a natural language understanding unit, used to parse business requirements in natural language form; an intent-to-metric mapping unit, used to map business requirements into quantifiable network performance metrics and thresholds; and a business entity association unit, used to associate the parsing results with entities in MM-NKG.
4. The network performance optimization system based on data analysis according to claim 3, characterized in that: The Intent-Driven Graph Attention Temporal Prediction Model (IGATP) module uses MM-NKG as the graph structure input, takes network performance metrics and business intent objectives as node features, and employs a multi-head attention mechanism to capture dynamic relationships between nodes. Ultimately, it predicts future performance trends and the probability of intent violation, specifically including the following model: a. Graph Attention Encoder: Used for MM-NKG-based graph structures, it aggregates neighbor node information through a multi-head graph attention mechanism to generate the context embedding of nodes; b. Intent Fusion Layer: Used to fuse the intent target with the context embedding to generate intent-enhanced node embeddings; c. Temporal Convolutional Network (TCN): Used to extract temporal features from intent-enhanced node embedding sequences and generate temporal context; d. Attention Decoder: Used to fuse intent-enhanced node embeddings, temporal context, and expert experience associated with intent in MM-NKG to predict future network performance metrics and calculate the probability of intent violation.
5. The network performance optimization system based on data analysis according to claim 4, characterized in that: The loss function of the IGATP model includes a mean squared error term, a binary cross-entropy term, and a regularization term that encourages the model to maintain a structure and semantics consistent with MM-NKG.
6. The network performance optimization system based on data analysis according to claim 5, characterized in that, The multi-objective intelligent orchestration and feedback module includes: A strategy space generator is used to generate candidate optimized orchestration actions based on predicted risks. A multi-objective optimizer is used to balance multiple objectives such as intent achievement, resource utilization, and cost to select the optimal solution. The reinforcement learning strategy optimization unit is used to continuously optimize orchestration strategies through reinforcement learning.
7. The network performance optimization system based on data analysis according to claim 6, characterized in that: The reward function of the reinforcement learning strategy optimization unit comprehensively considers the degree of intent achievement, resource consumption cost, operation cost, and side effect penalty.
8. The network performance optimization system based on data analysis according to claim 7, characterized in that: The optimized execution and feedback interface supports Restful API, Netconf / YANG, OpenFlow protocol or vendor SDK, enabling automated orchestration of physical network devices, SDN controllers, cloud platforms and virtualization resources.
9. A network performance optimization method based on data analysis, employing the network performance optimization system based on data analysis as described in any one of claims 1 to 8, characterized in that, Includes the following steps: a. Multimodal network data acquisition: Acquire multimodal heterogeneous data from multiple sources, including network devices, application systems, user terminals, service management platforms, and operation and maintenance knowledge bases; b. Construction and management of multimodal network knowledge graph (MM-NKG): Entity recognition, relation extraction and semantic association are performed on the collected data to construct and dynamically update the MM-NKG; c. Adaptive service intent parsing: Parses high-level service requirements into quantified network performance goals and constraints, and associates them with MM-NKG; d. Intent-driven graph attention temporal prediction: Using MM-NKG as the backbone, integrating business intent, and predicting network performance trends and intent violation risks through the IGATP model; e. Multi-objective intelligent orchestration and feedback: Based on prediction results, MM-NKG information and intent objectives, generate and execute network resource and configuration orchestration schemes, and continuously optimize orchestration strategies through reinforcement learning; f. Optimize execution and feedback: Execute the orchestration scheme through the interface and feed back the execution results and network performance changes to the system.