Network system, service processing method, device, apparatus, and storage medium
By establishing an intelligent system integrating perception, monitoring, and decision-making within the computing network, the problem of the disconnect between network perception and resource scheduling is solved, thereby improving resource allocation efficiency and user experience.
Patent Information
- Application Number
- CN202311317111.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-11
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2043-10-11
AI Technical Summary
In the computing power network resource scheduling framework, network awareness, service quality and resource scheduling are disconnected, making it difficult to guarantee the performance of heterogeneous resource scheduling under real-time, efficient and multi-source constraints, resulting in the inability to improve resource allocation efficiency and user experience.
Establish an intelligent network system that integrates perception, monitoring, and decision-making based on user service quality. Through computing network service perception module, computing network service monitoring module, and computing network service decision-making module, it comprehensively considers computing resources, network performance, and service quality, and integrates computing network and service quality information to improve resource utilization.
It achieves real-time, efficient, and multi-source-constrained heterogeneous resource scheduling performance, improving resource allocation efficiency and user experience.
Smart Images

Figure CN118827480B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computing power network, and particularly relates to a network system, a service processing method and device, equipment and a storage medium. BACKGROUND
[0002] In the related art, due to the large scale of the computing power network, the strong resource heterogeneity and the high dynamicity, the network perception, the service quality and the resource scheduling are in a state of fragmentation in the computing power network resource scheduling framework, and it is difficult to guarantee the heterogeneous resource scheduling performance under real-time, efficient and multi-source constraints, so as to realize the common improvement of resource configuration efficiency and user experience. SUMMARY
[0003] To solve the technical problems in the related art, the embodiments of the present application provide a network system, a service processing method and device, equipment and a storage medium.
[0004] To achieve the above purpose, the technical scheme of the embodiments of the present application is as follows:
[0005] In a first aspect, the embodiments of the present application provide a network system, which comprises a computing network service perception module, a computing network service monitoring module and a computing network service decision module; wherein
[0006] The computing network service perception module is configured to acquire each computing power node in the computing network service and computing network perception data; the computing network perception data comprises at least one of the following: computing power resource perception data; network performance perception data; service quality perception data;
[0007] The computing network service monitoring module is configured to monitor the service quality of the computing network service based on the computing network perception data, and obtain a monitoring result;
[0008] The computing network service decision module is configured to perform service decision optimization on the computing network service based on the monitoring result.
[0009] In the above scheme, the computing network service perception module further comprises a third sub-module.
[0010] The third sub-module is configured to periodically report the computing network perception data of the corresponding computing power node collected by each perception probe to the second sub-module.
[0011] In the above scheme, the computing network service monitoring module further comprises a seventh sub-module.
[0012] The seventh sub-module is configured to present alarm information after the fifth sub-module determines that the computing network service has a fault based on the analysis result.
[0013] In a second aspect, the embodiments of the present application further provide a service processing method, which comprises:
[0014] obtaining each computing power node in the computing network service and computing network sensing data; the computing network sensing data comprises at least one of the following: computing power resource sensing data; network performance sensing data; service quality sensing data;
[0015] monitoring service quality of the computing network service based on the computing network sensing data to obtain a monitoring result;
[0016] optimizing service decision of the computing network service based on the monitoring result.
[0017] In a third aspect, the embodiments of the present application further provide a service processing apparatus, which comprises:
[0018] an obtaining unit configured to obtain each computing power node in the computing network service and computing network sensing data; the computing network sensing data comprises at least one of the following: computing power resource sensing data; network performance sensing data; service quality sensing data;
[0019] a monitoring unit configured to monitor service quality of the computing network service based on the computing network sensing data to obtain a monitoring result;
[0020] a decision unit configured to optimize service decision of the computing network service based on the monitoring result.
[0021] In a fourth aspect, the embodiments of the present application further provide a service processing device, which comprises a processor and a memory for storing a computer program capable of running on the processor;
[0022] When the processor runs the computer program, it executes the steps of the service processing method according to the embodiments of the present application.
[0023] In a fifth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the service processing method according to the embodiments of the present application are implemented.
[0024] The network system provided by the embodiment of the application, the business processing method, the device, the equipment and the storage medium, the network system comprises: an algorithm network business perception module, an algorithm network business monitoring module and an algorithm network business decision module; wherein the algorithm network business perception module is used for acquiring each computing power node in the algorithm network business and algorithm network perception data; the algorithm network perception data comprises at least one of the following: algorithm resource perception data; network performance perception data; service quality perception data; the algorithm network business monitoring module is used for monitoring the service quality of the algorithm network business based on the algorithm network perception data to obtain a monitoring result; and the algorithm network business decision module is used for performing service decision optimization on the algorithm network business based on the monitoring result.
[0025] It can be seen that the embodiment of the application establishes an intelligent network system based on user service quality perception-monitoring-decision integration, by setting an algorithm network business perception module, an algorithm network business monitoring module and an algorithm network business decision module in the network system, considering algorithm resource, network performance and service quality as a whole, by fusing algorithm network and service quality information, comprehensively improving the resource utilization of the entire algorithm network, and providing high-quality algorithm network service for algorithm network users, that is, guaranteeing the heterogeneous resource scheduling performance under real-time, efficient and multi-source constraints, and finally realizing the common improvement of resource configuration efficiency and user experience. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 FIG. 1 is a structural schematic diagram of the network system of the embodiment of the application;
[0027] Figure 2 FIG. 2 is a flowchart of the business processing method of the embodiment of the application; Figure 1
[0028] Figure 3 FIG. 3 is an architecture schematic diagram of the intelligent perception-monitoring-decision integrated network system for algorithm network business of the embodiment of the application;
[0029] Figure 4 FIG. 4 is a deployment framework schematic diagram of a perception probe in an algorithm power network of the embodiment of the application;
[0030] Figure 5 FIG. 5 is a framework schematic diagram of a passive fusion algorithm network state perception technology of the embodiment of the application;
[0031] Figure 6 FIG. 6 is a framework schematic diagram of business exception monitoring and fault diagnosis analysis of the embodiment of the application;
[0032] Figure 7 FIG. 7 is a flowchart of the business processing method of the embodiment of the application; Figure 2
[0033] Figure 8 Fig. 1 is a schematic diagram of a service processing device according to an embodiment of the present application;
[0034] Figure 9 Fig. 2 is a schematic diagram of a hardware structure of a service processing device according to an embodiment of the present application. DETAILED DESCRIPTION
[0035] The present application will be further described below in conjunction with the accompanying drawings and embodiments.
[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application.
[0037] With the development of the fifth generation mobile communication technology (5G, 5th Generation Mobile Communication Technology) and artificial intelligence, a large number of emerging computing needs have emerged, and have brought about an explosive growth in the transmission data in the network. At the same time, users have increasingly high requirements for network quality, seeking more stable, high-speed and high-quality network responses, and have gradually put forward different quality of experience (QoE, Quality of Experience) requirements. Cloud computing and edge computing technology provides users with the ability to access various computing resources anytime, anywhere, and conveniently. Edge and cloud computing further solve the low latency requirements of services and alleviate the congestion problem caused by a large amount of data in the backbone network. Therefore, how to integrate the performance of the network and the state of the distributed computing resources to achieve flexible transmission of task data, flexible deployment of application programs and coordinated scheduling of distributed resources to guarantee the end-to-end quality of services and improve user experience becomes crucial.
[0038] Traditional network infrastructure centered on information transmission is transforming into a new heterogeneous hierarchical network architecture that integrates computing, storage and network resources represented by the computing power network. How to interconnect dynamically distributed computing and storage resources, and through the unified and coordinated scheduling of multi-dimensional resources such as network, storage and computing power, to enable massive applications to call ubiquitous distributed computing resources in real time on demand, to realize the global optimization of connectivity and computing power in the network, and to provide consistent user experience, has become a research hotspot in the industry in recent years.
[0039] In the related art, due to the large scale of computing power network, strong resource heterogeneity and high dynamicity, in the computing power network resource scheduling framework, network perception, service quality and resource scheduling are in a fragmented state, and it is difficult to guarantee the heterogeneous resource scheduling performance under real-time, efficient and multi-source constraints. Therefore, it is urgent to establish an intelligent network system integrating perception, monitoring and decision-making based on user service quality, and finally realize the common improvement of resource configuration efficiency and user experience.
[0040] Based on this, the embodiment of the application provides a network system, in various embodiments of the application, an intelligent network system integrating perception, monitoring and decision-making based on user service quality is established, by setting a computing network service perception module, a computing network service monitoring module and a computing network service decision module in the network system, the computing power resources, network performance and service quality are considered as a whole, the resource utilization of the entire computing network is improved by fusing the computing network and service quality information, and high-quality computing network services are provided for computing network users, that is, the heterogeneous resource scheduling performance under real-time, efficient and multi-source constraints is guaranteed, and finally the common improvement of resource configuration efficiency and user experience is realized.
[0041] The embodiment of the application provides a network system, Figure 1 The structure of the network system of the embodiment of the application is shown in Figure 1 As shown, the network system comprises a computing network service perception module 11, a computing network service monitoring module 12 and a computing network service decision module 13; wherein,
[0042] The computing network service perception module 11 is configured to acquire each computing power node in the computing network service and computing network perception data; the computing network perception data comprises at least one of the following: computing power resource perception data; network performance perception data; service quality perception data;
[0043] The computing network service monitoring module 12 is configured to monitor the service quality of the computing network service based on the computing network perception data, and obtain a monitoring result;
[0044] The computing network service decision module 13 is configured to perform service decision optimization on the computing network service based on the monitoring result.
[0045] In actual application, the computing network service perception module 11 acquires the perception data of computing power resources, network performance and service quality on demand in combination with active and passive perception data and the like, and provides unified data services for the upper computing network service monitoring module 12 and the computing network service decision module 13.
[0046] Based on this, in an embodiment, the computing network service perception module 11 comprises a first sub-module and a second sub-module; wherein,
[0047] The first sub-module is configured to deploy a perception probe on each computing power node, and the perception probe is configured to actively collect computing network perception data of the corresponding computing power node.
[0048] The second sub-module is configured to receive the computing network perception data of the corresponding computing power node reported by each perception probe.
[0049] Here, the perception probe is a distributed probe, i.e., a distributed probe deployed on each computing power node. The distributed probe can be used to realize fine perception of computing power resources, network performance, and service quality. The distributed probe is a proxy program that can be run on a computing power node in a lightweight manner.
[0050] Here, the computing network perception data in the embodiments of the present application includes at least one of the following: computing power resource perception data; network performance perception data; service quality perception data; wherein,
[0051] The computing power resource perception data includes at least one of the following: resource pool location; resource quota; resource load; wherein the resource quota includes at least one of the following: central processing unit (CPU) quota; memory quota; graphic processing unit (GPU) quota; storage quota; and the resource load includes at least one of the following: utilization rate of network resources; storage input / output (I / O) performance; network I / O performance;
[0052] The network performance perception data includes at least one of the following: physical link topology and capacity of network resources including intelligent private line and cloud private network; resource load; pre-configured transmission circuit and tunnel network quality of service (QoS) information; wherein the resource load includes latency, packet loss, and jitter;
[0053] The service quality awareness data includes at least one of the following: service resource information; service application performance; service quality key performance indicators (KPIs); service quality key quality indicators (KQIs); wherein the service resource information includes at least one of the following: application service interface address; resource requirement specification of an application instance; the service application performance includes at least one of the following: interface throughput performance; interface response performance; the service quality KPIs include at least one of the following: transmission control protocol (TCP) quality KPIs; domain name system (DNS) quality KPIs; hypertext transfer protocol (HTTP) quality KPIs; and the service quality KQIs include at least one of the following: uplink average rate; downlink average rate; average packet loss rate.
[0054] Here, the algorithm network service awareness module 11 further includes a third submodule;
[0055] The third submodule is configured to periodically report the algorithm network awareness data of the corresponding computing power node collected by each awareness probe to the second submodule.
[0056] It should be noted that the distributed probes in adjacent positions can discover and detect each other. In the embodiments of the present application, the distributed probe can be understood as a service quality awareness probe deployed on the service user side, the network side and the computing power side respectively. The distributed probe supports multiple detection protocols, such as Iperf, two-way active measurement protocol (TWAMP), trace route (TraceRoute), Ping, etc. By configuring active detection tasks on the real network path of the service, the real-time service running status is mastered, and the real-time awareness and detection index collection and reporting of the service quality, computing power resource and network performance in an end-to-end manner (such as cross-professional, cross-vendor and cross-geographical network segment) are realized.
[0057] Here, in the embodiments of the present application, the awareness probe, i.e. the distributed probe, periodically reports the algorithm network awareness data, which belongs to the active working mode, i.e. the algorithm network service awareness module 11 actively perceives the data. Of course, in actual application, in addition to actively perceiving the data, the algorithm network service awareness module 11 can also passively perceive the data, i.e. the passive working mode based on the data collection node to initiate a query and detection task.
[0058] Based on this, in an embodiment, the perception probe is further configured to receive a query task issued by the second submodule;
[0059] The query task is generated based on a query request of an upper service module of the second submodule, and the query request is carried in the query task.
[0060] In response to the query request, the algorithm network perception data of the corresponding algorithm node is collected.
[0061] Here, the second submodule can be a data collection node, also known as a data collection module. Specifically, the upper service module of the data collection node initiates a query request to the data collection node. After receiving the query request initiated by the upper service module, the data collection node generates a query task and issues the query task to the perception probe for execution. That is, after receiving the query task issued by the data collection node, the perception probe parses the query task to obtain the query request, and then collects the algorithm network perception data of the corresponding algorithm node in response to the query request.
[0062] As can be seen, the embodiments of the present application obtain the state data of each algorithm node and network link in the algorithm network (including at least one of the algorithm resource perception data, the network performance perception data, and the service quality perception data) through the lightweight algorithm network state perception mechanism of active and passive fusion, solve the problems of heterogeneous network, cross-vendor, and segmented measurement data inconsistency, guarantee data timeliness, cross-vendor compatibility, and end-to-end integrity, and realize end-to-end data perception and multi-dimensional data analysis capability for tenants.
[0063] In actual application, the algorithm network service monitoring module 12 analyzes and understands the service quality of the algorithm network service in real time based on the algorithm network perception data reported by the algorithm network service perception module 11, provides a data basis for the algorithm network service decision module 13, and improves the decision quality.
[0064] Based on this, in an embodiment, the algorithm network service monitoring module 12 includes a fourth submodule and a fifth submodule; wherein,
[0065] The fourth submodule is configured to analyze the performance indicators of the algorithm network service based on the algorithm network perception data to obtain an analysis result. The performance indicators include at least one of the following: the state of the algorithm resource; the state of the network resource; and the state of the service quality.
[0066] The fifth submodule is configured to determine whether the algorithm network service has a fault based on the analysis result.
[0067] The processing process of the fourth submodule is described below.
[0068] Here, the fourth submodule is specifically configured to:
[0069] determine the service quality of the computing network service by using a pre-established service quality evaluation model;
[0070] determine the health status of the computing power network and the target device and the running condition of the computing network service based on the computing network perception data; the target device is a device corresponding to the deployment of the computing network service;
[0071] obtain the analysis result based on the service quality of the computing network service, the health status of the computing power network and the target device, and the running condition of the computing network service.
[0072] Here, in order to determine the service quality of the to-be-deployed task, i.e., the computing network service in the target device, the characteristics of the elastic scheduling of the computing power network resources need to be combined to analyze from two aspects of device local information and network status. That is, a service quality evaluation model is established based on network status information and device local information, so as to evaluate the service quality of the computing network service through the established service quality evaluation model.
[0073] The processing process of the fifth submodule will be described below.
[0074] Here, the fifth submodule is specifically configured to:
[0075] compare the analysis result with a target threshold to obtain a comparison result;
[0076] in a case where the comparison result represents that the analysis result is less than the target threshold, determine that the computing network service has a fault;
[0077] in a case where the comparison result represents that the analysis result is greater than or equal to the target threshold, determine that the computing network service has no fault.
[0078] It should be noted that the target threshold is set according to actual needs, which is not limited here. After the analysis result is determined by the fourth submodule, the fifth submodule is used to compare the analysis result with the pre-set target threshold. When the analysis result is less than the pre-set target threshold, it indicates that the current node task allocation has a problem, i.e., the computing network service has a fault; when the analysis result is greater than or equal to the pre-set target threshold, it indicates that the current node task allocation has no problem, i.e., the computing network service has no fault.
[0079] In actual application, after detecting that the computing network service fails, the AI (Artificial Intelligence) diagnosis capability needs to be combined to quickly find and locate the fault between the user and the VPC (Virtual Private Cloud), and the SLA (Service-Level Agreement) of the computing network service is comprehensively guaranteed.
[0080] Based on this, in an embodiment, the computing network service monitoring module 12 further includes a sixth submodule.
[0081] The sixth submodule is configured to, after the fifth submodule determines that the computing network service fails based on the analysis result, diagnose and analyze the fault of the computing network service to identify the root cause of the fault and locate the fault.
[0082] Here, for the fault of the computing network service, the computing network service can be detected on the fly based on the VPN (Virtual Private Network) traffic, the packet loss or delay degradation problem of the service can be analyzed, the problems of each network element device on the service path can be aggregated and analyzed, the root cause of the service fault can be analyzed, and the cause of the fault can be quickly understood.
[0083] Here, the computing network service monitoring module 12 further includes a seventh submodule.
[0084] The seventh submodule is configured to, after the fifth submodule determines that the computing network service fails based on the analysis result, present the alarm information.
[0085] It should be noted that the seventh submodule is configured to visually display the alarm information, so that the user can quickly view the fault of the computing network service, and the alarm information can be pushed in the form of an instant message.
[0086] In actual application, when the computing network service is abnormal, the computing network resources usually need to be rescheduled, the task deployment needs to be optimized, and the control capability of the deterministic service quality for the computing network service needs to be implemented.
[0087] Based on this, in an embodiment, the computing network service decision module 13 includes an eighth submodule and a ninth submodule, and the eighth submodule is configured to, when the monitoring result indicates that the computing network service fails, trigger a rescheduling request.
[0088] The eighth submodule is configured to, when the monitoring result indicates that the computing network service fails, trigger a rescheduling request.
[0089] The ninth sub-module is configured to perform service decision optimization on the network calculation service by using the service optimization strategy in response to the rescheduling request.
[0090] Here, the ninth sub-module is specifically configured to:
[0091] In response to the rescheduling request, determine whether the simulation effect of the service optimization strategy meets a first set condition, and obtain a determination result.
[0092] In a case where the determination result indicates that the simulation effect of the service optimization strategy meets the first set condition, perform service decision optimization on the network calculation service by using the service optimization strategy.
[0093] It should be noted that, in order to guarantee the accuracy of policy execution, it is necessary to determine whether the simulation effect of the service optimization strategy meets the first set condition before the service optimization strategy is issued. Only when the simulation effect of the service optimization strategy meets the first set condition, the service optimization strategy is issued, so as to perform service decision optimization on the network calculation service by using the service optimization strategy, that is, to realize the optimization process of the network calculation service.
[0094] The generation process of the service optimization strategy will be described below.
[0095] In the embodiment of the present application, the network service decision module 13 further includes a tenth sub-module.
[0096] The tenth sub-module is configured to generate the service optimization strategy.
[0097] The tenth sub-module is specifically configured to:
[0098] Determine the weights of edges in the current computing power network topology and the weights of computing power nodes;
[0099] Construct a dynamic network based on the weights of edges in the current computing power network topology and the weights of computing power nodes.
[0100] Select a target computing power node and a target routing path from the dynamic network, the target routing path being a routing path from a source computing power node to the target computing power node, and the routing cost of the target routing path meeting a second set condition.
[0101] Generate the service optimization strategy based on the target computing power node and the target routing path.
[0102] Here, the modeling process of the dynamic network is to abstract the dynamic network into a graph topology using graph theory, and the dynamic nature of the multi-dimensional performance indicators of the computing power resources and network resources is represented by the point and edge weight information in graph theory. The points and edges in the topology graph are given corresponding weights according to the basic performance indicators KPI and task QoE constraints. In the embodiments of the present application, for the computing power routing problem, the transmission process and the computing process of the computing task are considered simultaneously, therefore, the edge weight and the node weight in the dynamic network are defined simultaneously, that is, the weight of each edge in the current computing power network topology, and the weight of each computing power node, so as to construct the dynamic network. In addition, in the embodiments of the present application, the computing network service optimization problem is transformed into identifying the most suitable computing power node (i.e., the target computing power node) for the computing task in the dynamic network, and selecting the best routing path (i.e., the target routing path) from the source node to the target node, so as to minimize the overall cost. The routing cost of the target routing path satisfies the second set condition, which can be understood as the lowest routing cost of the target routing path.
[0103] Based on the foregoing network system, the embodiments of the present application provide a service processing method, which is applied to a service processing device, Figure 2 The flow of the service processing method of the embodiments of the present application is shown in Figure 1 ; as Figure 2 shown, the method comprises:
[0104] Step 201: acquiring each computing power node in the computing network service and computing network perception data; the computing network perception data comprises at least one of the following: computing power resource perception data; network performance perception data; service quality perception data.
[0105] In an embodiment, the acquiring computing network perception data comprises:
[0106] deploying a perception probe arranged on each computing power node, the perception probe being used to actively collect the computing network perception data of the corresponding computing power node;
[0107] receiving the computing network perception data of the corresponding computing power node reported by each perception probe.
[0108] Here, the perception probe is used to receive a query task issued by a second sub-module;
[0109] The query task is generated based on a query request of an upper service module of the second sub-module, and the query request is carried in the query task.
[0110] In response to the query request, the computing network perception data of the corresponding computing power node is collected.
[0111] Step 202: monitoring the service quality of the computing network service based on the computing network perception data, and obtaining a monitoring result.
[0112] In an embodiment, the monitoring of the service quality of the computing network service based on the computing network perception data comprises:
[0113] The performance indicators of the computing network service are analyzed based on the computing network perception data to obtain an analysis result, wherein the performance indicators comprise at least one of the following: a state of computing resource; a state of network resource; a state of service quality.
[0114] Based on the analysis result, it is determined whether the computing network service has a fault.
[0115] The analysis of the performance indicators of the computing network service based on the computing network perception data to obtain an analysis result comprises:
[0116] The service quality of the computing network service is determined by using a pre-established service quality evaluation model.
[0117] Based on the computing network perception data, the health status of the computing network and the target device, and the running condition of the computing network service are determined, wherein the target device is a device corresponding to the deployment of the computing network service.
[0118] The analysis result is obtained based on the service quality of the computing network service, the health status of the computing network and the target device, and the running condition of the computing network service.
[0119] Based on the analysis result, it is determined whether the computing network service has a fault, which comprises:
[0120] The analysis result is compared with a target threshold to obtain a comparison result.
[0121] In a case where the comparison result represents that the analysis result is less than the target threshold, it is determined that the computing network service has a fault.
[0122] In a case where the comparison result represents that the analysis result is greater than or equal to the target threshold, it is determined that the computing network service does not have a fault.
[0123] In an embodiment, after the determination that the computing network service has a fault based on the analysis result, the method further comprises:
[0124] The fault of the computing network service is diagnosed and analyzed to identify the root cause of the fault and to locate the fault.
[0125] In another embodiment, after the determination that the computing network service has a fault based on the analysis result, the method further comprises: presenting alarm information.
[0126] Step 203: based on the monitoring result, performing service decision optimization on the computing network service.
[0127] In an embodiment, the service decision optimization on the computing network service based on the monitoring result comprises:
[0128] In the case that the monitoring result represents that the computing network service has a fault, triggering a rescheduling request;
[0129] In response to the rescheduling request, performing service decision optimization on the computing network service by using a service optimization strategy.
[0130] In an embodiment, the service decision optimization on the computing network service by using the service optimization strategy in response to the rescheduling request comprises:
[0131] In response to the rescheduling request, determining whether the simulation effect of the service optimization strategy meets a first set condition to obtain a determination result;
[0132] In the case that the determination result represents that the simulation effect of the service optimization strategy meets the first set condition, performing service decision optimization on the computing network service by using the service optimization strategy.
[0133] In an embodiment, the method further comprises: generating the service optimization strategy.
[0134] In an embodiment, the generating the service optimization strategy comprises:
[0135] Determining the weight of each edge in the current computing network topology and the weight of each computing node;
[0136] Based on the weight of each edge in the current computing network topology and the weight of each computing node, constructing a dynamic network;
[0137] Selecting a target computing node and a target routing path from the dynamic network, the target routing path being a routing path from a source computing node to the target computing node, and the routing cost of the target routing path meeting a second set condition;
[0138] Based on the target computing node and the target routing path, generating the service optimization strategy.
[0139] It should be noted that the specific processing process of the service processing device for processing the service has been described in detail above, and specific reference can be made to the detailed description process of the network system, which will not be repeated here.
[0140] By adopting the technical solutions of the embodiments of the present application, an intelligent network system based on user service quality is established, the network service perception module, the network service monitoring module and the network service decision module are set in the network system, the computing power resources, the network performance and the service quality are considered as a whole, the computing power network and the service quality information are fused, the resource utilization of the entire computing power network is comprehensively improved, and high-quality computing power network services are provided for the computing power network users, that is, the heterogeneous resource scheduling performance under real-time, efficient and multi-source constraints is ensured, and finally the resource configuration efficiency and user experience are improved together.
[0141] The embodiments of the present application obtain the state data of each computing power node and network link in the computing power network through the lightweight computing power network state perception mechanism of active and passive fusion, take the perceived service quality, network state and computing power resource usage as input, abstract the computing power network scheduling as a multi-objective optimization problem, and thus better support the optimization and scheduling of the computing power network service. Moreover, the computing power network service monitoring module in the network system can perform 7*24 real-time end-to-end monitoring for the computing power network service in the whole time period and the whole network, discover network hidden faults in time, discover problems before the customers, and realize rapid alarm and accurate positioning of faults.
[0142] The present application will be described below in combination with application embodiments.
[0143] The schemes of computing power network service perception, monitoring and decision in the related art will be described first.
[0144] Currently, the known technical schemes for computing power network service perception, monitoring and decision mainly include the following:
[0145] The computing power network involves a large number of devices of traditional network architecture, and these devices usually only support some traditional network state perception methods, such as Simple Network Management Protocol (SNMP), Traceroute, Ping, passive measurement of delay, Packet trace-based tracking of network state and the like. These methods face problems such as difficulty in unifying the measurement period, high cost of collection, fragmented collection range and the like in the large-scale network domain data collection scenario. Overall, the current computing power network lacks accurate perception data for the service side, the technical schemes that can be implemented for different network link sections are limited by their own computing power and network architecture, the perceived information is not unified in terms of dimension, data specification and the like, it is difficult to provide unified and effective application perception results, and it is unable to effectively support further computing power optimization and scheduling.
[0146] The existing scheme generates a detection plan according to the data center network topology, sends IP-in-IP packets to detect the packet loss rate of the path, and then applies a high-precision link fault reasoning algorithm designed for the inconsistency problem in the real world data to determine the fault link or fault device. It can be seen that the existing scheme does not consider the characteristics of the computing network, and has the characteristics of focusing on a certain specific fault, small expandable range, long fault diagnosis time, unable to detect in real time, and large performance overhead.
[0147] The computing network service scheduling strategy currently found that network tenants share resources, which can easily affect each other, leading to unpredictable service performance, and a framework FAB is designed to provide predictable performance services for data center users. The core network switch of FAB is responsible for sensing network information through inband network telemetry (INT) and sending it to the edge, and the edge relies on accurate network information to efficiently control the path and rate of tenant traffic. However, the existing method lacks a multi-factor decision system for computing network services, making it difficult to provide globally optimal routing results to support intelligent computing network service scheduling.
[0148] Therefore, the quality of service assurance system in the related art has the problems of poor timeliness of computing network data sensing, long service fault positioning time, and low intelligent level of service dynamic scheduling, which leads to the inability to realize end-to-end quality assurance for user service experience. The following explains several problems:
[0149] (1) Poor timeliness of data sensing
[0150] The quality of segmented physical link network quality collection tasks is fragmented across device manufacturers, the collection cycle is not unified, and the collection reporting delay is large; there is a lack of service quality sensing capability and means, and the service SLA is not visible, and the service fault cooperative diagnosis lacks basis. In view of the individualized and fine-grained sensing needs of computing power networks, the sensing results provided by passive sensing technology cannot meet the needs, and active probing path analysis methods under dynamic network topology need to be researched, end-to-end hop-by-hop state sensing technology for computing power networks needs to be developed, and an efficient sensing system for super-large-scale computing power networks needs to be established.
[0151] (2) Long service fault positioning time
[0152] The current computing network service fault handling process analyzes and locates end-to-end problems through network signaling tracking, alarm and log analysis packet capture, configuration verification and other technical means; each single domain on the network side has scattered data collection and packet capture tools, but the data is scattered, and segmented and cross-manufacturer correlation analysis and processing are difficult. The current fault analysis and positioning link relies mainly on manual work and tools, and strongly depends on expert experience (the fault analysis and positioning of video backhaul takes 7.5 hours)
[0153] (3) The intelligent level of service dynamic scheduling is not high
[0154] Current algorithm network service dynamic scheduling mainly focuses on single field scheduling, such as selecting and arranging dynamic computing power resources based on comprehensive resource availability and computing power resource cost information, and selecting network links based on network delay, packet loss, and jitter information. The mapping system of service user experience quality and basic algorithm and network resource indicators has not been established, and the end-to-end multi-dimensional service quality model cannot be constructed to seek the optimal solution.
[0155] Therefore, based on the limitations of the above related technologies, the present application proposes an intelligent perception-monitoring-decision integrated network system and service processing method for algorithm network services. Through the lightweight computing power network state perception mechanism of active and passive fusion, the state data of each computing power node and network link in the computing power network is obtained, solving the problems of heterogeneous network, cross-vendor, and segmented measurement data inconsistency, ensuring data timeliness, cross-vendor compatibility, and end-to-end integrity, realizing end-to-end data perception and multi-dimensional data analysis capabilities for tenants; based on the intelligent monitoring module, a mapping system of user service experience quality and basic algorithm and network resource indicators is established, the user application is monitored in real time through key operation and maintenance indicators, and combined with AI diagnosis capability, the fault between the user and VPC is quickly found and located, realizing the control capability of deterministic service quality for algorithm network services; based on the service SLA index constraint, computing power resource and network performance index dynamic change, combined with the optimization algorithm, the recommended computing power node and network path are solved and output, realizing the multi-technology element fusion capability supply under multi-dimensional constraints; finally, the intelligent algorithm path of software defined network (SDN, Software Defined Network) and the path flexible arrangement capability based on IPv6 forwarding plane segment routing (SRv6, Segment Routing IPv6) are used to realize efficient calling of deployed computing power services through network level scheduling control, and provide deterministic algorithm network service.
[0156] Figure 3As an architecture schematic diagram of an intelligent perception-monitoring-decision integrated network system for computing network services of an embodiment of the present application, the network system is used to solve the problems of poor timeliness of computing network data perception, long service fault positioning time, and low intelligent level of service dynamic scheduling in the related art, and applies fine-grained distributed probes, in-situ flow information telemetry (IFIT), SRv6, QoS evaluation model, computing power routing, and other technologies to provide users with safe and reliable end-to-end quality perception, monitoring analysis, and dynamic scheduling integrated services of computing network services. The network system includes a computing network perception data warehouse (corresponding to the computing network service perception module), an intelligent analysis unit (corresponding to the computing network service monitoring module), and an intelligent decision unit (corresponding to the computing network service decision module). The following describes the components of the network system.
[0157] 1. Computing network perception data warehouse
[0158] The computing network perception data warehouse is mainly targeted at service lifecycle management, and formulates a computing network perception system from the dimensions of resources, events, configurations, and performance for heterogeneous computing power, networks, and service resources. In combination with active and passive perception, the computing network perception data warehouse obtains perception data of computing power, networks, and service resources on demand, and becomes a computing network perception data warehouse, inventorying computing network resources and perceiving service states to provide unified data services for the intelligent analysis unit and the intelligent decision unit.
[0159] Here, to address the problem of lack of service quality perception means in the related art, the present application deploys service quality probes based on in-situ detection technology, configures detection tasks on the real network path of the service, and realizes end-to-end service quality perception by real-time monitoring of service running conditions. Figure 4 As an architecture schematic diagram of a perception probe deployment framework in a computing power network of an embodiment of the present application, as shown in Figure 4 The service quality collection means specifically deploys service quality perception probes on the service user side, the network side, and the computing power side, configures active detection tasks on the real network path of the service, and realizes real-time perception of service quality, computing power resources, and network performance and collection and reporting of detection indexes end-to-end (across disciplines, manufacturers, and regional network segments).
[0160] Figure 5 As a main and passive fusion computing network state perception technology framework of an embodiment of the present application, as shown in Figure 5As shown, the present application proposes a service active and passive perception framework based on distributed probes, which realizes fine perception of computing power resources, network performance and service quality. The framework includes distributed probes deployed on each computing power node, a data collection module deployed on the computing power network control node, and a set of protocols for detecting network quality and transmission perception data. The framework includes a distributed probe that periodically reports the active mode of computing power network state, and a passive mode initiated by the data collection module to query and detect tasks.
[0161] Among them, the distributed probe is a proxy program that can be run on the computing power node in a lightweight manner, which mainly completes the following functions: 1) Periodically collect the computing resources, storage resources and network status of the node; 2) Periodically report the statistical data of each node in the algorithm network to the data collection module; 3) Adjacent distributed probes can discover and detect each other; 4) Respond to the query request of the data collection node.
[0162] The data collection node is the upward interface of the computing power network state perception, and all algorithm network related states are unified through this module to provide data export. Its main functions include: 1) Accepting active perception data reported by each distributed probe; 2) Monitoring the state of each distributed probe and applying control; 3) Accepting query requests from the upper layer, generating query tasks and distributing them to the distributed probe for execution. In specific implementation, the data collection node provides a data interface that can access various big data stream processing platforms and an interface that receives upper layer query and control commands, realizing a set of probe task planning and generation controller.
[0163] Among them, the following data can be perceived through the distributed probe:
[0164] (1) Computing resource perception data, mainly including resource pool location, resource quota (CPU, memory, GPU, storage quota), resource load (utilization rate of computing, storage, network resources, storage I / O performance, network I / O performance), etc.
[0165] (2) Network performance perception data, mainly including physical link topology (service tunnel, transmission circuit) and capacity (cloud export bandwidth, bandwidth utilization) of intelligent private line, cloud private network and other network resources, resource load (delay, packet loss, jitter), pre-configured transmission circuit and tunnel Network QOS information (QOS configuration of allocated resources).
[0166] (3) Service quality perception data, mainly including service resource information (application service interface address, resource requirement specification of application instance), service application performance (interface throughput performance, interface response performance), service quality KPI (TCP quality KPI, DNS quality KPI, HTTP quality KPI), service quality KQI (uplink and downlink average rate, average packet loss rate).
[0167] 2. Intelligent analysis unit
[0168] The intelligent analysis unit is configured to analyze and understand the state of computing resources, network resources, and service quality in real time based on data provided by the algorithm network perception data warehouse, to provide an analysis basis for the intelligent decision unit and improve decision quality; by establishing a mapping system of service user experience quality and basic algorithm network resource indicators, real-time monitoring of user key operation and maintenance indicators, and cooperating with AI diagnosis capabilities, quickly discovering and locating faults between users and VPCs, and comprehensively guaranteeing algorithm network service SLA. The intelligent analysis unit includes service experience analysis, service fault analysis, and alarm presentation, which are described below.
[0169] (1) Service experience analysis
[0170] For user experience QoE, the service quality KQI and KPI indicators are converted into corresponding QOE scores according to a specific mapping model, the network fault handling response priority is quickly matched, and the algorithm network user experience is prioritized.
[0171] 1) Establish a QoS evaluation model
[0172] To determine the QoS of the task to be deployed in the target device, the characteristics of the computing network resource elastic scheduling are combined to analyze from two aspects of device local information and network status.
[0173] From the network perspective, the QoS of the current network in the corresponding device is defined according to whether the current related device network is available, the proportion of available bandwidth, packet loss rate, and other attributes; from the device local information perspective, the standard configuration of various computing indicators such as CPU, memory, storage, FPGA, and GPU is pre-selected, and the corresponding information is parameterized as the basis for measuring the performance of each heterogeneous device, and the application or computing task to be deployed is also parameterized with the same standard, and the weighted sum of the target device parameters is used to estimate the actual performance of the corresponding heterogeneous device. Based on the results obtained from the network and device local information, the weighted calculation obtains the expected running of the deployed application on the target device, i.e., the overall QoS level.
[0174] In the calculation process, the weight is introduced to effectively reduce the gap between the expected result and the actual running. For the weight, a gradient can be set according to the importance of each resource to the deployed task, and the task requirement can be selected specifically. By updating the local information and network status of each heterogeneous device in real time in the controller, the QoS of each task can be calculated in real time after deployment.
[0175] 2) Determine the abnormal state in combination with node feedback
[0176] After the related task is deployed on the computing power node, the abnormal condition is determined in combination with the network and device abnormal information obtained through monitoring and the expected QoS calculated by the controller. The determination f of the abnormality is calculated by the state expectation of the QoS, the health state Q n , Q d , the running condition k of the task, and the specific calculation method is determined based on the performance of the task and the computing power node. d and Q n According to the results of fault monitoring and the importance, a suitable gradient is selected. If the task is normally running, k = 1, otherwise k = 0.
[0177] f = Σk(QoS + aQ n + βQ d ) (1)
[0178] In actual application, the final goal of the abnormality determination is to ensure that the task is successfully deployed on the corresponding computing power node, and to re-schedule the computing network resources when the QoS is too low or the computing network appears abnormal, and to optimize the task deployment. When the result of f calculation is less than the pre-set threshold value, it is considered that the task allocation of the current node has a problem, at this time, one or several tasks occupying the most resources are optimized, for example, if multiple tasks are concurrently deployed on the same device, it will cause the QoS to quickly decrease; if any result is 0, the corresponding task is not successfully deployed, and the computing network resources need to be re-allocated.
[0179] (2) Business fault analysis
[0180] Figure 6 is a schematic diagram of the framework of the business abnormality monitoring and fault diagnosis analysis of the embodiments of the present application. For business faults, based on the IFIT technology and the distributed probe, the on-the-fly detection is performed based on the dedicated line VPN traffic, the business packet loss or delay degradation problem is analyzed, the problem aggregation is performed based on the network path, and the problem position is delimited. The position of the problem can be delimited by using the stream-by-hop manner.
[0181] Here, in the diagnosis and analysis link, after the business fault occurs, the network fault is automatically associated, based on the indexes such as the alarm / configuration / KPI in the network element device, the abnormality is determined by using the corresponding abnormality detection algorithm, the single network element abnormality is identified, the problem identification is completed through the multi-abnormal event clustering, and the root cause analysis is performed. Through the aggregation analysis of the problems of each network element on the business path, the root cause analysis of the business fault is completed. The root cause analysis can be obtained by at least one of the following ways: root cause reasoning; processing suggestion; fault portrait; similar fault.
[0182] (3) Alarm presentation
[0183] Real-time performance alarm display is realized at minute level, real-time monitoring is performed on resource pool, computing resource, storage resource, network resource, tenant and other levels, and topology visualization is performed. Among them, the real-time pushing of alarm information supports interfaces such as SNMP, WeChat, Email and Webservice.
[0184] The embodiment of the application is based on an intelligent analysis unit, quickly builds a service experience KQI model of each service scenario, forms an end-to-end service quality KPI / KQI system covering computing power, network and service application, realizes real-time visualization of service quality and minute-level positioning of service faults based on IFIT on-stream detection technology and real-time sensing data, and is used for service quality evaluation and service problem diagnosis.
[0185] 3. Intelligent decision unit
[0186] Based on the reported data sensing current running state of the algorithm and network resources, in combination with the fault delimiting results generated by the intelligent analysis unit, the algorithm nodes and network links are reselected from the perspective of user service end-to-end guarantee, and service optimization suggestions are given. Through supporting policy life cycle management, policy scheduling model design and policy issuing safety control, unified design and resource cooperation of opening policy and scheduling policy are realized. The intelligent decision unit includes service dynamic optimization, service policy simulation and service policy scheduling, which are described below.
[0187] (1) Service optimization policy (service dynamic optimization)
[0188] Based on different service SLA, network overall load, available computing resource pool distribution and other factors, the optimal cooperation strategy of algorithm, network and number is intelligently and dynamically calculated, the service scheduling policy is dynamically generated as needed, the application request is scheduled to the algorithm node along the optimal path, the efficiency of algorithm and network resources is improved, and the user experience is guaranteed. In the modeling process, the dynamic network is abstracted into a graph topology structure by using graph theory, the dynamic nature of the multi-dimensional performance indicators of algorithm and network resources is described by using the point and edge weight information in graph theory, and the points and edges in the topology graph are given corresponding weights according to the basic performance indicators KPI and task QoE constraints.
[0189] Among them, the dynamic network model can be expressed as G(t)=(V, E(t), W V (t), W E (t)) where V={v1, v2..., v n}(n=|V|) represents a node set, each node represents an actual network element device; E={e1, e2,..., e m(t)}(m=|E(t)|) represents a set of edges, each edge represents an actual physical connection link; time t is the time when the algorithm and network information is collected. a set of weights representing nodes, quantifying computing resource status; a set of weights representing edges, quantifying network link performance.
[0190] Since the computing resource routing problem should consider both the transmission process and the computing process of the computing task, the dynamic network model proposed in the present application defines both edge weights and node weights. Meanwhile, considering the QoE constraints of the computing task, W v (t) is related to the task. Generally, a computing task τ may have requirements for the available computing power, available memory of the computing resource, the key KPIs such as the delay, jitter, packet loss of the network link, and the like, based on the QoE constraints. E (t). Generally, a computing task τ may have requirements for the available computing power, available memory of the computing resource, the key KPIs such as the delay, jitter, packet loss of the network link, and the like, based on the QoE constraints.
[0191] wherein, for W E (t), the KPI indicators considered include the link delay D e (t), the jitter delay DV e (t), the packet loss rate L e (t), and the available bandwidth B e (t), and the weight of the edge E is calculated according to the following formula (2) in combination with the network link performance requirements of the task:
[0192]
[0193] wherein, for w v (t), the KPI indicators considered include the available computing power (FLOPS) and the available memory (GB), and the weight of the node i is calculated according to the following formula (3) in combination with the computing resource performance requirements of the task:
[0194]
[0195] wherein, represents the predicted processing time of the task by the computing node i according to the performance at the time t, S i (t) and C i (t) represent the available memory and the available computing power of the computing node i at the time t.
[0196] Therefore, in the present application, the computing network service optimization problem is converted into identifying the most suitable computing node v q for the computing task in the dynamic network G(t) and selecting the best routing path q from the source node u to the target node v so as to minimize the overall cost. The optimization problem can be modeled as: wherein, ψ(·) calculates the routing cost of selecting the computing node v q and the routing path The routing path is composed of The corresponding edge e i,j The weights of node j are summed to obtain e. i,j This represents an edge with origin i and endpoint j. Routing cost. It can be calculated using the following formula (4):
[0197]
[0198] (2) Business Strategy Simulation
[0199] The strategy simulation in this application supports routing simulation during network path opening and scheduling, and can perform communication process simulation and AI algorithm model training iteration. It also supports routing simulation calculation based on key factors (latency, bandwidth, etc.) to ensure the accuracy of strategy execution.
[0200] (3) Business strategy scheduling
[0201] This application transforms the generated strategies into a programmable computing and network domain system language, enabling network capability configuration and scheduling, as well as computing capability configuration and scheduling. The activation and scheduling strategies are instantiated and executed through this module, outputting the optimally routed business instance. The strategy calculation results are provided externally in the form of application programming interfaces (APIs). This mainly includes resource scheduling (CPU, memory, storage), network scheduling (latency, bandwidth, packet loss), application scheduling (automatic deployment of mirrored applications), and traffic scheduling (automatic switching and adjustment of service traffic).
[0202] In this embodiment, a computing power routing optimization model is considered in the dynamic optimization strategy for services, prioritizing computing power nodes and routing paths. By establishing a computing power network optimization model, the main objective of computing power routing is clarified: to identify the most suitable computing power nodes for computing tasks within the dynamic network and select the optimal routing path from the source node to the target node, thereby minimizing the overall cost. Based on the above modeling process, the computing power routing problem is transformed into the shortest path problem in classical path planning. The recommended routing paths obtained based on this model can better meet the comprehensive service quality requirements of computing network services.
[0203] Figure 7 This is a flowchart illustrating the business processing method of an embodiment of this application. Figure 2 ,like Figure 7 As shown, this business processing method includes the following steps:
[0204] Step 701: Computing network data perception.
[0205] Step 702: Obtain network sensing data and task parameter information.
[0206] Here, the algorithm network perception data includes at least one of the following: computing power resource perception data; network performance perception data; service quality perception data.
[0207] Step 703: Establish a QoS evaluation model according to specific algorithm network services, and calculate and analyze the result f using f =∑k(Qos+αQ n +βQ d ).
[0208] Step 704: Determine whether the analysis result f is less than the target threshold value. If it is less than the target threshold value, execute step 705, otherwise return to step 701.
[0209] Step 705: Service fault diagnosis and alarm presentation.
[0210] Step 706: Trigger rescheduling and calculate the weights of edges and points in the current algorithm network topology.
[0211] Step 707: Select a candidate source node set and a candidate target node set according to the task parameter requirements.
[0212] Step 708: Determine the minimum value of the output routing cost.
[0213] Step 709: Determine whether the optimization strategy simulation meets the expectation. If it meets the expectation, execute step 710, otherwise return to step 706.
[0214] Step 710: Issue the service optimization strategy.
[0215] Compared with the related technical solution, the present application has the following beneficial effects:
[0216] (1) In the computing power network resource scheduling framework of the related technology, network perception, service quality and resource scheduling are in a fragmented state, making it difficult to guarantee the performance of heterogeneous resource scheduling under real-time, efficient and multi-source constraints. Therefore, an intelligent system based on user service quality perception-monitoring-decision is established, and finally the resource configuration efficiency and user experience are improved together.
[0217] (2) At the algorithm network service perception level, tenant-level service quality perception is realized, solving the problems of heterogeneous network, cross-vendor, segmented measurement data inconsistency, etc., ensuring data timeliness, cross-vendor compatibility and end-to-end integrity.
[0218] (3) At the algorithm network service monitoring level, the control ability of the algorithm network service deterministic quality of service is realized. A mapping system of service user experience quality and basic algorithm network resource indicators is established, the user application is monitored in real time through key operation and maintenance indicators, and combined with AI diagnosis capability, the faults between users and VPCs are quickly found and located, and the algorithm network user experience is fully guaranteed.
[0219] (4) At the network service decision optimization level, realize the intelligent decision-making ability of network path facing multi-dimensional constraints. It is urgent to solve and output the recommended computing power nodes and network paths based on the dynamic changes of service SLA index constraints, computing power resources and network performance indicators, combined with optimization algorithms, to realize the multi-technology factor fusion ability supply under multi-dimensional constraints.
[0220] In order to realize the service processing method of the embodiments of the application, the embodiments of the application also provide a service processing device, Figure 8 The schematic diagram of the composition structure of the service processing device of the embodiments of the application is shown in Figure 8 The service processing device comprises:
[0221] The acquisition unit 81 is configured to acquire each computing power node in the computing network service and computing network perception data; the computing network perception data comprises at least one of the following: computing power resource perception data; network performance perception data; service quality perception data;
[0222] The monitoring unit 82 is configured to monitor the service quality of the computing network service based on the computing network perception data, and obtain a monitoring result;
[0223] The decision unit 83 is configured to perform service decision optimization on the computing network service based on the monitoring result.
[0224] In an embodiment, the acquisition unit 81 comprises a deployment subunit and a receiving subunit; wherein,
[0225] The deployment subunit is configured to deploy a perception probe arranged on each computing power node, and the perception probe is configured to actively collect computing network perception data of the corresponding computing power node;
[0226] The receiving subunit is configured to receive the computing network perception data of the corresponding computing power node reported by each perception probe.
[0227] In an embodiment, the perception probe is further configured to receive a query task issued by the receiving subunit;
[0228] The query task is generated based on a query request of an upper service module of the receiving subunit, and the query request is carried in the query task;
[0229] The receiving subunit is specifically configured to collect the computing network perception data of the corresponding computing power node in response to the query request.
[0230] In an embodiment, the acquisition unit 81 further comprises a reporting subunit;
[0231] The reporting subunit is configured to periodically report the computing network perception data of the respective computing power node collected by the perception probe to the receiving subunit.
[0232] In an embodiment, the monitoring unit 82 comprises an analyzing subunit and a determining subunit, wherein,
[0233] The analyzing subunit is configured to analyze the performance indicators of the computing network service based on the computing network perception data to obtain an analysis result. The performance indicators comprise at least one of the following: a state of computing power resources; a state of network resources; and a state of service quality.
[0234] The determining subunit is configured to determine whether the computing network service has a fault based on the analysis result.
[0235] In an embodiment, the analyzing subunit is specifically configured to:
[0236] determine the service quality of the computing network service by using a pre-established service quality evaluation model.
[0237] determine the health states of the computing power network and target devices and the running condition of the computing network service based on the computing network perception data. The target devices are devices corresponding to the deployment of the computing network service.
[0238] obtain the analysis result based on the service quality of the computing network service, the health states of the computing power network and target devices, and the running condition of the computing network service.
[0239] In an embodiment, the determining subunit is specifically configured to:
[0240] compare the analysis result with a target threshold to obtain a comparison result.
[0241] In a case where the comparison result indicates that the analysis result is less than the target threshold, it is determined that the computing network service has a fault.
[0242] In a case where the comparison result indicates that the analysis result is greater than or equal to the target threshold, it is determined that the computing network service does not have a fault.
[0243] In an embodiment, the monitoring unit 82 further comprises a processing subunit.
[0244] The processing subunit is configured to, after the determining subunit determines that the computing network service has a fault based on the analysis result, diagnose and analyze the fault of the computing network service to identify the root cause of the fault and locate the fault.
[0245] In an embodiment, the monitoring unit 82 further comprises a presenting subunit.
[0246] The presenting sub-unit is configured to present alarm information after the determining sub-unit determines that the computing network service is faulty based on the analysis result.
[0247] In an embodiment, the decision unit 83 comprises a triggering sub-unit and a decision sub-unit, wherein:
[0248] The triggering sub-unit is configured to trigger a rescheduling request when the monitoring result indicates that the computing network service is faulty.
[0249] The decision sub-unit is configured to perform service decision optimization on the computing network service by using a service optimization strategy in response to the rescheduling request.
[0250] In an embodiment, the decision sub-unit is specifically configured to:
[0251] In response to the rescheduling request, determine whether the simulation effect of the service optimization strategy meets a first set condition to obtain a determination result.
[0252] When the determination result indicates that the simulation effect of the service optimization strategy meets the first set condition, perform service decision optimization on the computing network service by using the service optimization strategy.
[0253] In an embodiment, the decision unit 83 further comprises a generating sub-unit.
[0254] The generating sub-unit is configured to generate the service optimization strategy.
[0255] In an embodiment, the generating sub-unit is specifically configured to:
[0256] Determine the weights of edges and the weights of nodes in a current computing network topology.
[0257] Construct a dynamic network based on the weights of edges and the weights of nodes in the current computing network topology.
[0258] Select a target node and a target routing path from the dynamic network, wherein the target routing path is a routing path from a source node to the target node, and the routing cost of the target routing path meets a second set condition.
[0259] Generate the service optimization strategy based on the target node and the target routing path.
[0260] In actual application, the obtaining unit 81 can be implemented by a communication interface in a service processing device; the monitoring unit 82 and the decision unit 83 can be implemented by a processor in the service processing device.
[0261] It should be noted that the service processing apparatus provided in the above embodiment is only exemplified by the division of the above program modules when performing service processing. In actual application, the above processing distribution can be completed by different program modules according to needs, that is, the internal structure of the apparatus is divided into different program modules to complete all or part of the above-described processing. In addition, the service processing apparatus and the service processing method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the service processing method embodiments, which will not be repeated here.
[0262] Based on the hardware implementation of the above program modules, and in order to realize the service processing method of the embodiment of the application, the embodiment of the application further provides a service processing device, Figure 9 The hardware component structure of the service processing device of the embodiment of the application is shown in FIG. 9, which comprises: Figure 9
[0263] The communication interface 91 can interact with other devices.
[0264] The processor 92 is connected with the communication interface 91 to realize information interaction with other devices, and is used to run a computer program to execute the service processing method provided above, and the computer program is stored on the memory 93.
[0265] Specifically, the communication interface 91 is used to obtain each computing power node in the computing network service and computing network sensing data; the computing network sensing data comprises at least one of the following: computing power resource sensing data; network performance sensing data; service quality sensing data.
[0266] The processor 92 is used to monitor the service quality of the computing network service based on the computing network sensing data to obtain a monitoring result, and to perform service decision optimization on the computing network service based on the monitoring result.
[0267] In an embodiment, the processor 92 is further used to deploy a sensing probe arranged on each computing power node, and the sensing probe is used to actively collect computing network sensing data of the corresponding computing power node.
[0268] The communication interface 91 is specifically used to receive the computing network sensing data of the corresponding computing power node reported by each sensing probe.
[0269] Here, the sensing probe is further used to receive a query task.
[0270] The query task is generated based on a query request, and the query request is carried in the query task.
[0271] The communication interface 91 is specifically configured to: in response to the query request, collect the algorithm network perception data of the corresponding computing power node.
[0272] In an embodiment, the communication interface 91 is further configured to: periodically report the algorithm network perception data of the corresponding computing power node collected by each perception probe to the receiving subunit.
[0273] In an embodiment, the processor 92 is specifically configured to:
[0274] analyze a performance index of the algorithm network service based on the algorithm network perception data to obtain an analysis result; the performance index includes at least one of the following: a state of a computing power resource; a state of a network resource; a state of a service quality;
[0275] determine whether the algorithm network service has a fault based on the analysis result.
[0276] In an embodiment, the processor 92 is specifically configured to:
[0277] determine a service quality of the algorithm network service by using a pre-established service quality evaluation model;
[0278] determine a health state of an algorithm power network and a target device and an operation of the algorithm network service based on the algorithm network perception data; the target device is a device corresponding to the algorithm network service;
[0279] obtain the analysis result based on the service quality of the algorithm network service, the health state of the algorithm power network and the target device, and the operation of the algorithm network service.
[0280] In an embodiment, the processor 92 is specifically configured to:
[0281] compare the analysis result with a target threshold to obtain a comparison result;
[0282] in a case where the comparison result represents that the analysis result is less than the target threshold, determine that the algorithm network service has a fault;
[0283] in a case where the comparison result represents that the analysis result is greater than or equal to the target threshold, determine that the algorithm network service does not have a fault.
[0284] In an embodiment, the processor 92 is further configured to:
[0285] after determining that the algorithm network service has a fault based on the analysis result, diagnose and analyze the fault of the algorithm network service to identify a root cause of the fault and locate the fault.
[0286] In an embodiment, the processor 92 is further configured to:
[0287] present alarm information after determining that the computing network service is faulty based on the analysis result.
[0288] In an embodiment, the processor 92 is specifically configured to:
[0289] trigger a rescheduling request in a case where the monitoring result indicates that the computing network service is faulty.
[0290] optimize service decision of the computing network service by using a service optimization strategy in response to the rescheduling request.
[0291] In an embodiment, the processor 92 is specifically configured to:
[0292] determine whether a simulation effect of the service optimization strategy meets a first set condition in response to the rescheduling request, to obtain a determination result.
[0293] optimize service decision of the computing network service by using the service optimization strategy in a case where the determination result indicates that the simulation effect of the service optimization strategy meets the first set condition.
[0294] In an embodiment, the processor 92 is further configured to generate the service optimization strategy.
[0295] In an embodiment, the processor 92 is specifically configured to:
[0296] determine weights of edges and weights of computing nodes in a current computing network topology.
[0297] construct a dynamic network based on the weights of the edges and the weights of the computing nodes in the current computing network topology.
[0298] select a target computing node and a target routing path from the dynamic network, the target routing path being a routing path from a source computing node to the target computing node, and a routing cost of the target routing path meeting a second set condition.
[0299] generate the service optimization strategy based on the target computing node and the target routing path.
[0300] It should be noted that the specific processing process of the communication interface 91 and the processor 92 can be understood with reference to the above service processing method.
[0301] Of course, in practice, the various components of the business processing device 90 are coupled together by a bus system 94. It will be appreciated that the bus system 94 is used to facilitate communication between the various components and is comprised of one or more buses, including a power bus, a control bus and a status signal bus, among others. For the sake of clarity, the various buses are illustrated in FIG. 1 as the bus system 94. Figure 9
[0302] The memory 93 stores various types of data used for operation of the business processing device 90. Examples of this data include any computer programs for operation on the business processing device 90.
[0303] The business processing method disclosed in the embodiments of the present application can be applied to or implemented by the processor 92. The processor 92 can be an integrated circuit chip having a signal processing capability. In the implementation process, the steps of the business processing method can be completed by hardware integrated logic circuit or software form of instructions in the processor 92. The processor 92 can be a general purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The processor 92 can implement or execute the business processing methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the business processing method disclosed in the embodiments of the present application, the hardware decoding processor can be directly embodied to execute the steps, or the hardware and software modules in the decoding processor can be combined to execute the steps. The software module can be located in the storage medium in the memory 93, and the processor 92 reads the information in the memory 93 and combines the hardware to complete the steps of the aforementioned business processing method.
[0304] In an exemplary embodiment, the service processing device 90 can be implemented by one or more Application Specific Integrated Circuits (ASICs), DSPs, Programmable Logic Devices (PLDs), Complex Programmable Logic Devices (CPLDs), Field-Programmable Gate Arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors (Microprocessors), or other electronic elements for performing the aforementioned service processing methods.
[0305] It can be understood that the memory 93 of the embodiments of the present application can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM). The magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory 93 described in the embodiments of the present application is intended to include, but not limited to, these and any other suitable types of memories.
[0306] In the exemplary embodiments, the embodiments of the present application also provide a storage medium, specifically a computer readable storage medium, such as the memory 93 storing a computer program executable by the processor 92 in the service processing device 90 to complete the steps of the service processing method described in the embodiments of the present application. The computer readable storage medium can be FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0307] It should be noted that "first", "second", "third", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0308] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0309] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A network system, characterized in that, The network system includes: a network service perception module, a network service monitoring module, and a network service decision-making module; wherein... The computing network service perception module is used to acquire each computing power node and computing network perception data in the computing network service; the computing network perception data includes at least one of the following: computing power resource perception data; network performance perception data; service quality perception data; The computing network service monitoring module is used to monitor the service quality of the computing network service based on the computing network perception data and obtain monitoring results. The computing network service decision module is used to optimize the computing network service based on the monitoring results and using service optimization strategies. The computing network service decision module includes a tenth submodule, used to generate the service optimization strategy; The tenth submodule is specifically used for: Determine the weights of each edge in the current computing power network topology, as well as the weights of each computing power node; Based on the weights of each edge and each computing node in the current computing power network topology, a dynamic network is constructed. Select a target computing power node and a target routing path from the dynamic network. The target routing path is a routing path from the source computing power node to the target computing power node, and the routing cost of the target routing path satisfies a second set condition. The service optimization strategy is generated based on the target computing power node and the target routing path.
2. The network system according to claim 1, characterized in that, The computing network service perception module includes: a first submodule and a second submodule; wherein... The first submodule is used to deploy sensing probes on each computing power node, and the sensing probes are used to actively collect computing network sensing data of the corresponding computing power node; The second submodule is used to receive the network sensing data of the corresponding computing power nodes reported by each sensing probe.
3. The network system according to claim 2, characterized in that, The sensing probe is also used to receive query tasks issued by the second submodule; The query task is generated based on the query request from the upper-layer business module of the second sub-module, and the query task carries the query request. In response to the query request, network perception data of the corresponding computing power node is collected.
4. The network system according to claim 1, characterized in that, The network service monitoring module includes a fourth sub-module and a fifth sub-module; wherein... The fourth submodule is used to analyze the performance indicators of the computing network service based on the computing network perception data, and obtain the analysis results; the performance indicators include at least one of the following: the status of computing resources; the status of network resources; and the status of service quality. The fifth submodule is used to determine whether the computing network service has malfunctioned based on the analysis results.
5. The network system according to claim 4, characterized in that, The fourth submodule is specifically used for: The service quality of the computing network service is determined using a pre-established service quality assessment model. Based on the computing network sensing data, the health status of the computing network and target devices, as well as the operation status of the computing network services, are determined; the target devices are the devices that deploy the computing network services. The analysis results are obtained based on the service quality of the computing network service, the health status of the computing network and the target device, and the operation status of the computing network service.
6. The network system according to claim 4, characterized in that, The fifth submodule is specifically used for: The analysis results are compared with the target threshold to obtain the comparison results; If the comparison result indicates that the analysis result is less than the target threshold, it is determined that the computing network service has failed. If the comparison result indicates that the analysis result is greater than or equal to the target threshold, it is determined that the computing network service has not experienced a failure.
7. The network system according to claim 4, characterized in that, The network service monitoring module also includes: a sixth sub-module; The sixth submodule is used to perform diagnostic analysis on the faults of the computing network service after the fifth submodule determines whether the computing network service has failed based on the analysis results, so as to identify the root cause of the fault and locate the fault.
8. The network system according to claim 1, characterized in that, The computing network service decision module further includes: an eighth sub-module and a ninth sub-module; wherein... The eighth submodule is used to trigger a rescheduling request when the monitoring results indicate that the computing network service has failed. The ninth submodule is used to respond to the rescheduling request and optimize the computing network service using a service optimization strategy.
9. The network system according to claim 8, characterized in that, The ninth submodule is specifically used for: In response to the rescheduling request, determine whether the simulation effect of the service optimization strategy meets the first set condition, and obtain the determination result; If the simulation effect of the business optimization strategy, as indicated by the judgment result, satisfies the first set condition, the business optimization strategy is used to optimize the computing network service.
10. A business processing method, characterized in that, The method includes: Acquire each computing node and computing network perception data in the computing network service; the computing network perception data includes at least one of the following: computing resource perception data; network performance perception data; service quality perception data; Based on the computing network sensing data, the service quality of the computing network services is monitored to obtain monitoring results; Based on the monitoring results, business decision optimization is performed on the computing network services using business optimization strategies. The method further includes: generating the business optimization strategy; The generation of the business optimization strategy includes: Determine the weights of each edge in the current computing power network topology, as well as the weights of each computing power node; Based on the weights of each edge and each computing node in the current computing power network topology, a dynamic network is constructed. Select a target computing power node and a target routing path from the dynamic network. The target routing path is a routing path from the source computing power node to the target computing power node, and the routing cost of the target routing path satisfies a second set condition. The service optimization strategy is generated based on the target computing power node and the target routing path.
11. A business processing apparatus, characterized in that, include: The acquisition unit is used to acquire each computing power node and computing network perception data in the computing network service. The network perception data includes at least one of the following: computing power resource perception data; Network performance perception data; service quality perception data; The monitoring unit is used to monitor the service quality of the computing network services based on the computing network perception data and obtain monitoring results. The decision-making unit is used to optimize the computing network services based on the monitoring results and using business optimization strategies. The decision-making unit includes: a generation subunit, used to generate the business optimization strategy; The generating subunit is specifically used for: Determine the weights of each edge in the current computing power network topology, as well as the weights of each computing power node; Based on the weights of each edge and each computing node in the current computing power network topology, a dynamic network is constructed. Select a target computing power node and a target routing path from the dynamic network. The target routing path is a routing path from the source computing power node to the target computing power node, and the routing cost of the target routing path satisfies a second set condition. The service optimization strategy is generated based on the target computing power node and the target routing path.
12. A business processing device, characterized in that, include: A processor and a memory for storing computer programs capable of running on the processor; When the processor is used to run the computer program, it performs the steps of the method of claim 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 10.
Citation Information
Patent Citations
Computing power reporting method, computing power obtaining method, computing power network element and computing power perception control network element
CN113840317A
Method and device for designing perception center in computing power network operating system
CN115766768A