Cloud management platform-based network anomaly determination method and cloud management platform

By predicting and collecting traffic on the communication paths of computing nodes through the cloud management platform, the problem of resource consumption by monitoring components was solved, efficient network anomaly detection was achieved, and the efficiency and experience of tenants' job tasks were improved.

WO2026066371A1PCT designated stage Publication Date: 2026-04-02HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing technologies, cloud providers monitor network anomalies by inserting monitoring components on compute nodes, which increases compute node resource consumption, affecting job efficiency and tenant experience.

Method used

A cloud management platform is used to predict and collect traffic on the communication paths between computing nodes. By detecting the match between predicted traffic and actual traffic, network anomalies can be identified, thereby reducing the resource consumption of computing nodes.

Benefits of technology

It improves the resource utilization of compute nodes, enhances the efficiency of executing tenant jobs, improves the tenant experience, and prevents monitoring components from encroaching on compute node resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025104804_02042026_PF_FP_ABST
    Figure CN2025104804_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a cloud management platform-based network anomaly determination method and a cloud management platform, which can improve tenant experience. The method of the present application comprises: a cloud management platform creates a plurality of computing nodes for a tenant to execute a job task of the tenant, wherein a plurality of communication paths are provided between the plurality of computing nodes, and each communication path can transmit traffic generated by two computing nodes connected thereto during execution of the job task; on this basis, the cloud management platform can perform traffic prediction on the plurality of communication paths by using information of the job task, so as to obtain predicted traffic transmitted over the plurality of communication paths; the cloud management platform can also collect traffic over the plurality of communication paths, so as to obtain actual traffic transmitted over the plurality of communication paths; and if the predicted traffic of the first communication path among the plurality of communication paths does not match the actual traffic of the first communication path, the cloud management platform can alert the tenant to an anomaly occurring on two computing nodes connected to the first communication path.
Need to check novelty before this filing date? Find Prior Art

Description

A network anomaly determination method based on a cloud management platform and the cloud management platform

[0001] The present application claims priority to the Chinese patent application No. 202411345892.9, filed on September 25, 2024, and entitled "A network anomaly determination method based on a cloud management platform and the cloud management platform", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of cloud technology, and in particular to a network anomaly determination method based on a cloud management platform and the cloud management platform. BACKGROUND

[0003] With the rapid development of cloud technology, more and more tenants choose to complete their job tasks in the cloud. Specifically, a tenant can deploy a computing node cluster in the cloud, and a communication network can be formed between multiple computing nodes in the computing node cluster. Since the multiple computing nodes often need to communicate when executing the job tasks of the tenant, the multiple computing nodes can complete communication through the network, thereby completing the job tasks.

[0004] In related technologies, in order to ensure that the computing node cluster of the tenant can successfully complete the job tasks of the tenant, the cloud vendor often needs to ensure that the network between the multiple computing nodes in the cluster is running normally. Therefore, the cloud vendor can insert a monitoring component into the multiple computing nodes to monitor whether the multiple computing nodes generate abnormal traffic in real time, thereby causing the communication volume between the computing nodes to become large or the communication speed to become slow, and various slow network conditions. When slow network conditions occur, the cloud vendor can feed back to the tenant in real time to timely solve these conditions through certain means, thereby ensuring that the network between the multiple computing nodes is running normally.

[0005] However, since the monitoring component runs on these computing nodes, it often occupies the resources of these computing nodes, which can cause the efficiency of these computing nodes in executing job tasks to decrease, thereby affecting the experience of the tenant. SUMMARY

[0006] Embodiments of the present application provide a network anomaly determination method based on a cloud management platform and the cloud management platform. In the process of determining network anomalies, the resources of the computing nodes can be saved, the efficiency of the computing nodes in executing the job tasks of the tenant can be improved, and the experience of the tenant can be improved.

[0007] A first aspect of embodiments of the present application provides a network anomaly determination method based on a cloud management platform. A cloud management platform for implementing the method can manage the infrastructure for providing cloud services to tenants. The method comprises:

[0008] When the tenant needs to process the job task, the cloud management platform provides a processing interface to the tenant, so that the tenant can send a processing request of the job task to the processing interface, since the processing request of the job task contains information of the job task and performance requirements of the job task, the cloud management platform can create a plurality of computing nodes in the infrastructure that meet the performance requirements of the job task, the plurality of computing nodes can be used to execute the job task of the tenant, the plurality of computing nodes have a plurality of communication paths between them, each communication path can connect two computing nodes in the plurality of computing nodes, so each communication path can be used to transmit the traffic generated by the two computing nodes connected by the communication path in the process of executing the job task.

[0009] Then, the cloud management platform can use the information of the job task to perform traffic prediction on the plurality of communication paths between the plurality of computing nodes, thereby obtaining the predicted traffic transmitted by the plurality of communication paths, wherein the information of the job task can include data used to execute the job task, models used to execute the job task, and policies followed to execute the job task, etc.

[0010] Then, the cloud management platform can notify the plurality of computing nodes to execute the job task, and in the process of executing the job task by the plurality of computing nodes, the cloud management platform can collect traffic on the plurality of communication paths, thereby obtaining the real traffic transmitted by the plurality of communication paths. Then, if the cloud management platform detects that the predicted traffic transmitted by a first communication path in the plurality of communication paths does not match the real traffic transmitted by the first communication path, the cloud management platform can determine that the first communication path is abnormal, and return a network exception notification to the client of the tenant, the notification can be used to indicate the two computing nodes connected by the first communication path. In this way, the tenant can determine the two computing nodes connected by the first communication path based on the notification that cause the network exception.

[0011] From the above method, it can be seen that the cloud management platform can monitor the traffic transmitted by the communication paths between the computing nodes to determine whether the network between the computing nodes is abnormal, and then determine whether the computing nodes are abnormal, without the need to insert monitoring components into the computing nodes to determine network abnormalities, which is beneficial to release the resources occupied by the computing nodes, thereby improving the resource utilization of the computing nodes, i.e. improving the efficiency of the computing nodes in executing the job tasks of the tenants, thereby improving the tenant experience.

[0012] In one possible implementation, the cloud management platform performs traffic prediction on the plurality of communication paths based on the information of the job task provided by the tenant, and obtains the predicted traffic transmitted by the plurality of communication paths, including: the cloud management platform simulates the process of the plurality of computing nodes executing the job task based on the information of the job task provided by the tenant, and obtains the predicted traffic transmitted by the plurality of communication paths. In the foregoing implementation, before the plurality of computing nodes are notified to execute the job task, the cloud management platform can call a simulation application, input the information of the job task provided by the tenant and the information of the plurality of computing nodes into the simulation application, so that the simulation application simulates the process of the plurality of computing nodes executing the job task, thereby accurately simulating the predicted traffic transmitted by the plurality of communication paths.

[0013] In a possible implementation, the plurality of communication paths includes the second communication path and a third communication path, and the second communication path and the third communication path share a sub-communication path, and the method further includes: during execution of the job task by the plurality of computing nodes, collecting, by the cloud management platform, real traffic transmitted by the sub-communication path to obtain real traffic transmitted by the sub-communication path; decomposing the real traffic transmitted by the sub-communication path to obtain real traffic transmitted by the second communication path and real traffic transmitted by the third communication path; and when the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path overlap in a time dimension, the network anomaly notification is further used to indicate two computing nodes connected by the second communication path and two computing nodes connected by the third communication path. In the foregoing implementation, when the second communication path and the third communication path in the plurality of communication paths share a sub-communication path, during execution of the job task by the plurality of computing nodes, the cloud management platform can further collect real traffic transmitted by the sub-communication path, thereby obtaining real traffic transmitted by the sub-communication path. Then, the cloud management platform can decompose the real traffic transmitted by the sub-communication path, thereby obtaining real traffic transmitted by the second communication path and real traffic transmitted by the third communication path. When the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path overlap in the time dimension, the cloud management platform can cause the network anomaly notification sent to the tenant to further indicate two computing nodes connected by the second communication path and two computing nodes connected by the third communication path. In this way, the tenant can determine, based on the notification, that the two computing nodes connected by the first communication path, the two computing nodes connected by the second communication path, and the two computing nodes connected by the third communication path cause the network anomaly. As can be seen, the cloud management platform can determine whether a single communication path (i.e., the first communication path) transmits abnormal traffic to determine whether the communication path has a slow network condition of excessive traffic, and can determine whether a plurality of communication paths (i.e., the second communication path and the third communication path) transmit traffic that overlaps in time to determine whether a sub-communication path shared by the plurality of communication paths has a slow network condition of traffic congestion, thereby achieving determination and analysis of various network anomaly conditions.

[0014] In a possible implementation, the job task includes a first job task and a second job task, two computing nodes connected by the second communication path are used to execute the first job task, and two computing nodes connected by the third communication path are used to execute the second job task. The network exception notification is further used to indicate the first job task and the second job task. In the foregoing implementation, the job task required to be completed by the tenant can include the first job task and the second job task, and when the two computing nodes connected by the second communication path are used to execute the first job task and the two computing nodes connected by the third communication path are used to execute the second job task, the network exception notification sent by the cloud management platform to the tenant is further used to indicate the first job task and the second job task, to inform the tenant that the two computing nodes connected by the second communication path cause a network exception in the execution process of the first job task, and the two computing nodes connected by the third communication path cause a network exception in the execution process of the second job task. It can be seen that the cloud management platform can also assist the tenant in analyzing and determining the network exception of multiple job tasks, thereby maintaining the network exception confirmation requirement of the tenant on different job tasks, and further improving the experience of the tenant.

[0015] In a possible implementation, the model is a neural network model to be trained, the data is training data, and the strategy includes any one or any combination of the following pipeline parallel strategy, data parallel strategy, model parallel strategy, and sequence parallel strategy.

[0016] In a possible implementation, the computing node includes a physical server, a virtual machine, a container, a micro virtual machine, or a bare metal server.

[0017] In a possible implementation, the plurality of computing nodes are deployed in the same site or different sites, and the site includes a region, an availability zone, a data center, a machine room, or a physical server group.

[0018] The second aspect of the embodiment of the present application provides a cloud management platform, the cloud management platform is used for managing an infrastructure, the infrastructure comprises a plurality of computing nodes created by the cloud management platform for a tenant, the plurality of computing nodes are used for executing a job task of the tenant, the plurality of computing nodes have a plurality of communication paths, each communication path connects two computing nodes in the plurality of computing nodes, and each communication path is used for transmitting traffic generated by the two computing nodes connected by the communication path in the process of executing the job task, the cloud management platform comprises: a prediction module, configured to perform traffic prediction on the plurality of communication paths based on information of the job task provided by the tenant, to obtain predicted traffic transmitted by the plurality of communication paths, the information comprising data required for executing the job task, a model required for executing the job task, and a strategy required for executing the job task; a collection module, configured to perform traffic collection on the plurality of communication paths by the cloud management platform in the process of executing the job task by the plurality of computing nodes, to obtain real traffic transmitted by the plurality of communication paths; and a notification module, configured to determine a first communication path in the plurality of communication paths, in which the predicted traffic does not match the real traffic, and send a network anomaly notification to the tenant, the network anomaly notification being used for indicating two computing nodes connected by the first communication path.

[0019] In a possible implementation manner, the prediction module is configured to simulate the process of executing the job task by the plurality of computing nodes based on the information of the job task provided by the tenant, to obtain the predicted traffic transmitted by the plurality of communication paths.

[0020] In a possible implementation manner, the plurality of communication paths comprise a second communication path and a third communication path, and the second communication path and the third communication path have a shared sub-communication path, and the cloud management platform further comprises a decomposition module, configured to: perform traffic collection on the sub-communication path in the process of executing the job task by the plurality of computing nodes, to obtain real traffic transmitted by the sub-communication path; decompose the real traffic transmitted by the sub-communication path, to obtain real traffic transmitted by the second communication path and real traffic transmitted by the third communication path; and when the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path overlap in a time dimension, the network anomaly notification is further used for indicating two computing nodes connected by the second communication path and two computing nodes connected by the third communication path.

[0021] In a possible implementation manner, the job task comprises a first job task and a second job task, the two computing nodes connected by the second communication path are used for executing the first job task, the two computing nodes connected by the third communication path are used for executing the second job task, and the network anomaly notification is further used for indicating the first job task and the second job task.

[0022] In a possible implementation, the model is a neural network model to be trained, the data is training data, and the strategy includes any one or any combination of the following pipeline parallel strategy, data parallel strategy, model parallel strategy, and sequence parallel strategy.

[0023] In a possible implementation, the computing node includes a physical server, a virtual machine, a container, a micro virtual machine, or a bare metal server.

[0024] In a possible implementation, the plurality of computing nodes are deployed in the same site or different sites, and the site includes a region, an availability zone, a data center, a machine room, or a physical server group.

[0025] A third aspect of the embodiments of the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory: the memory is configured to store instructions; and the processor is configured to execute the instructions to cause the computing device cluster to perform the method in the first aspect or any possible implementation manner of the first aspect.

[0026] A fourth aspect of the embodiments of the present application provides a computer storage medium, which stores one or more instructions, and the instructions, when executed by one or more computers, cause the one or more computers to implement the method in the first aspect or any possible implementation manner of the first aspect.

[0027] A fifth aspect of the embodiments of the present application provides a computer program product, which stores instructions, and the instructions, when executed by a computer, cause the computer to implement the method in the first aspect or any possible implementation manner of the first aspect.

[0028] In the embodiment of the present application, when a tenant needs to complete a job task, the tenant can send a processing request for the job task to a processing interface provided by the cloud management platform, so that the cloud management platform can receive the processing request for the job task through the processing interface. Since the processing request contains the information of the job task and the performance requirement of the job task, the cloud management platform can create a plurality of computing nodes meeting the performance requirement of the job task to execute the job task. Among them, the plurality of computing nodes have a plurality of communication paths, each communication path can connect two computing nodes in the plurality of computing nodes, so that each communication path can transmit the traffic generated by the two computing nodes connected by the communication path in the process of executing the job task. Based on this, the cloud management platform can use the information of the job task to perform traffic prediction on the plurality of communication paths, thereby obtaining the predicted traffic transmitted by the plurality of communication paths. In the process of executing the job task by the plurality of computing nodes, the cloud management platform can also collect the traffic of the plurality of communication paths, thereby obtaining the real traffic transmitted by the plurality of communication paths. If the predicted traffic of a first communication path in the plurality of communication paths does not match the real traffic of the first communication path, the cloud management platform can send a network exception notification to the tenant to remind the two computing nodes connected by the first communication path of the exception. In the foregoing process, the cloud management platform can monitor the traffic transmitted by the communication paths between the computing nodes to determine whether the network between the computing nodes is abnormal, and then determine whether the computing nodes are abnormal, without the need to insert a monitoring component into the computing nodes to determine the network exception, which is beneficial to release the resources occupied by the computing nodes, thereby improving the resource utilization of the computing nodes, i.e., improving the efficiency of the computing nodes in executing the job tasks of the tenants, thereby improving the tenant experience. BRIEF DESCRIPTION OF DRAWINGS

[0029] FIG. 1 is a structural schematic diagram of a cloud service system provided by an embodiment of the present application;

[0030] FIG. 2 is another structural schematic diagram of a cloud service system provided by an embodiment of the present application;

[0031] FIG. 3 is a flow schematic diagram of a network exception determination method based on a cloud management platform provided by an embodiment of the present application;

[0032] FIG. 4 is another structural schematic diagram of a cloud service system provided by an embodiment of the present application;

[0033] FIG. 5 is a schematic diagram of traffic decomposition provided by an embodiment of the present application;

[0034] FIG. 6 is another structural schematic diagram of a cloud service system provided by an embodiment of the present application;

[0035] FIG. 7 is a structural schematic diagram of a cloud management platform provided by an embodiment of the present application;

[0036] FIG. 8 is a structural schematic diagram of a computing device according to an embodiment of the present application;

[0037] FIG. 9 is a structural schematic diagram of a computing device cluster according to an embodiment of the present application;

[0038] FIG. 10 is a schematic diagram of a network connection between computing devices in a computer cluster according to an embodiment of the present application. DETAILED DESCRIPTION

[0039] The embodiments of the present application provide a network anomaly determination method based on a cloud management platform and the cloud management platform, which can save resource occupation of a computing node in the process of network anomaly determination, improve efficiency of the computing node in executing a job task of a tenant, and thus improve tenant experience.

[0040] The terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms thus used can be interchanged under appropriate circumstances, which is merely a distinguishing manner adopted in the description of the embodiments of the present application for the objects with the same attribute in the description. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device containing a series of units does not have to be limited to those units, but can include other units not clearly listed or inherent to the process, method, product or device.

[0041] With the rapid development of cloud technology, more and more tenants choose to complete their job tasks in the cloud. Specifically, a tenant can deploy a computing node cluster in the cloud, and a plurality of computing nodes in the computing node cluster can form a communication network. Since the plurality of computing nodes often need to communicate when executing a job task of a tenant, the plurality of computing nodes can complete communication through the network, thereby completing the job task.

[0042] In the related art, in order to ensure that the computing node cluster of a tenant can successfully complete the job task of the tenant, the cloud vendor often needs to ensure that the network between the plurality of computing nodes in the cluster is normally running, so the cloud vendor can insert a monitoring component into the plurality of computing nodes to monitor in real time whether abnormal traffic is generated due to abnormal operation of the plurality of computing nodes in the process of executing the job task of the tenant, thereby causing the communication volume between the computing nodes to become large or the communication speed to become slow, etc. When the slow network situation occurs, the cloud vendor can feed back to the tenant in real time to timely solve the situation by certain means, thereby ensuring that the network between the plurality of computing nodes is normally running.

[0043] However, since the monitoring component runs on these computing nodes, it tends to occupy the resources of these computing nodes, which causes the efficiency of these computing nodes in performing job tasks to be reduced, thereby affecting the experience of tenants.

[0044] Further, since the monitoring component tends to intrude into the applications running in these computing nodes, it may affect the privacy of tenants, which also affects the experience of tenants.

[0045] To solve the above problems, an embodiment of the present application provides a network anomaly determination method based on a cloud management platform. The method can be implemented through a cloud service system. FIG. 1 is a structural schematic diagram of the cloud service system provided by an embodiment of the present application. As shown in FIG. 1, the cloud service system includes infrastructure that can provide cloud services and a cloud management platform that manages the infrastructure. The cloud management platform and the infrastructure are introduced respectively as follows:

[0046] The cloud management platform can manage the infrastructure in the whole cloud service system (e.g., in the infrastructure, create multiple computing nodes for a tenant according to the instruction of the tenant, the computing nodes are used to execute the job task of the tenant, thereby meeting the job requirements of the tenant, etc.), and the cloud management platform can also be open to each tenant outside the cloud service system and respond to their requests. For example, the cloud management platform can provide various interfaces such as a login interface and a processing interface for the client of the tenant (e.g., a terminal device used by the tenant or a browser on the terminal device, etc.) to access. Among them, the cloud management platform can authenticate the client of a certain tenant through the login interface, and allow the client of the tenant to log in to the cloud management platform after successful authentication. For another example, the cloud management platform can also allow the client of the tenant to send a processing request of the job task of the tenant to the cloud management platform through the processing interface. Since the processing request of the job task contains the information of the job task and the performance requirements of the job task, the cloud management platform can create multiple computing nodes (used to execute the job task of the tenant) in the infrastructure that meet the performance requirements of the job task. The multiple computing nodes are connected through multiple communication paths, each communication path can connect two computing nodes (i.e., a source computing node and a destination computing node) in the multiple computing nodes, so each communication path can be used to transmit the traffic generated by the two computing nodes connected by the communication path in the process of executing the job task. Then, the cloud management platform can use the information of the job task to perform traffic prediction on the multiple communication paths, thereby obtaining the predicted traffic transmitted by the multiple communication paths. Subsequently, in the process of executing the job task by the multiple computing nodes, the cloud management platform can collect the traffic of the multiple communication paths, thereby obtaining the real traffic transmitted by the multiple communication paths. Finally, in the multiple communication paths, if it is detected that the predicted traffic transmitted by a certain communication path does not match the real traffic transmitted by the communication path, the cloud management platform can determine that the communication path is abnormal, and return a network exception notification to the client of the tenant, the notification can be used to indicate the two computing nodes connected by the communication path. In this way, the tenant can determine that network exception occurs between the multiple computing nodes based on the notification, and the network exception is caused by the two computing nodes indicated by the notification.

[0047] The infrastructure includes a plurality of computing nodes that provide cloud services for the tenant service, the computing nodes can run applications of the tenant, and the applications can be used to execute a job task of the tenant to meet a job requirement of the tenant. It should be noted that the specifications of the plurality of computing nodes all meet the performance requirements of the job task (for example, the specifications of the computing resources, storage resources, and network resources required to execute the job task, and the like), and the plurality of computing nodes have a certain communication relationship (the communication relationship can be specified by the tenant (included in the information of the job task), or can be determined by the cloud management platform), so the plurality of computing nodes can form a plurality of communication pairs, each communication pair can include two computing nodes and a communication path between the two computing nodes, and the communication path is used to transmit the traffic generated by the two computing nodes in the process of executing the job task. As can be seen, the plurality of computing nodes have a plurality of communication paths, and the plurality of communication paths are used to transmit the traffic generated by the plurality of computing nodes in the process of executing the job task.

[0048] It is worth noting that for two computing nodes in the plurality of computing nodes that have a communication relationship, the two computing nodes can be communicatively connected through at least one switch, that is, the communication path between the two computing nodes includes at least one switch. As can be seen, the plurality of communication paths established between the plurality of computing nodes can be divided into two parts of communication paths, the first part of communication paths are completely independent communication paths, and there is no shared sub-communication path between the communication paths, and the second part of communication paths are partially shared communication paths, and there is a shared sub-communication path between the communication paths. For example, as shown in FIG. 2 (FIG. 2 is another structure schematic diagram of a cloud service system provided by an embodiment of the present application), the cloud management platform creates computing node 1 to computing node 14 for the tenant, and the 14 computing nodes form a large number of communication paths, each communication path connects two computing nodes, and passes through at least one switch. For example, the communication path between computing node 1 and computing node 2 passes through switch 1, and the communication path between computing node 3 and computing node 4 passes through switch 1, so the two communication paths are independent of each other. For example, the communication path between computing node 2 and computing node 7 passes through switch 1, switch 5, and switch 2 in sequence, and the communication path between computing node 9 and computing node 10 passes through switch 3, switch 5, and switch 2 in sequence, so the two communication paths are partially shared communication paths, and the two communication paths both include a sub-communication path between switch 5 and switch 2. The remaining communication paths are also similar to the two cases, which will not be described here.

[0049] It is also worth noting that for any one of the plurality of communication paths, the traffic transmitted by the communication path refers to the traffic transmitted between the two computing nodes connected by the communication path, for example, the traffic transmitted by the communication path between computing node 1 and computing node 2 refers to the traffic transmitted between computing node 1 and computing node 2. Of course, the traffic transmitted by some sub-communication paths shared between certain communication paths is usually superimposed by the traffic transmitted by these communication paths, for example, the traffic transmitted by the sub-communication path between switch 5 and switch 2 can be superimposed by the traffic transmitted by the communication path between computing node 2 and computing node 7, the traffic transmitted by the communication path between computing node 9 and computing node 10, and the traffic transmitted by the rest of the communication paths (also including the sub-communication path between switch 5 and switch 2).

[0050] Then, for any one of the plurality of communication paths, the cloud management platform can predict the predicted traffic transmitted by the communication path and collect the real traffic transmitted by the communication path, if the predicted traffic transmitted by the communication path does not match the real traffic transmitted by the communication path, the cloud management platform can determine that the real traffic transmitted by the communication path is abnormal, that is, the two computing nodes connected by the communication path are abnormal, causing the communication path to have a large traffic, which is a slow network situation. For several of the plurality of communication paths, the cloud management platform can collect the real traffic transmitted by the sub-communication paths shared between the several communication paths, and decompose the real traffic into the real traffic transmitted by the several communication paths respectively, if the real traffic transmitted by the several communication paths overlaps in the time dimension, it indicates that the two computing nodes connected by the several communication paths respectively are abnormal, causing the sub-communication path to have traffic congestion, which is a slow network situation. Based on this, the cloud management platform can inform the tenant of these situations in the form of network anomaly notification, thereby playing a role in reminding and preventing.

[0051] Further, for the plurality of compute nodes serving the tenant, the compute nodes can be presented in various forms, for example, the compute nodes can be physical servers (containing certain specifications of compute resources, storage resources, network resources, etc.) selected by the cloud management platform in the infrastructure, for another example, the compute nodes can also be bare metal servers (containing certain specifications of compute resources, storage resources, network resources, etc.) selected by the cloud management platform in the infrastructure. For another example, the compute nodes can also be virtual machines (VMs) created by the cloud management platform in the physical servers or bare metal servers through virtualization technology, for another example, the compute nodes can also be containers (docker) created by the cloud management platform in the physical servers or bare metal servers through virtualization technology, for another example, the compute nodes can also be micro VMs created by the cloud management platform in the physical servers or bare metal servers through virtualization technology, and the like.

[0052] Further, for the plurality of compute nodes serving the tenant, the compute nodes can be deployed in the same site or different sites, and the site can be presented in various forms, for example, the site can be a region in the infrastructure, for another example, the site can be an availability zone in the infrastructure, for another example, the site can be a data center (DC) in the infrastructure, for another example, the site can be a room in the infrastructure, for another example, the site can be a rack (also referred to as a physical server group) in the infrastructure, and the like.

[0053] Based on the cloud service system, when a tenant needs to complete a job task, the tenant can send a processing request for the job task to a processing interface provided by the cloud management platform, so that the cloud management platform can receive the processing request for the job task through the processing interface. Since the processing request contains the information of the job task and the performance requirement of the job task, the cloud management platform can create a plurality of computing nodes meeting the performance requirement of the job task to execute the job task. Among them, the plurality of computing nodes have a plurality of communication paths, each communication path can connect two computing nodes in the plurality of computing nodes, so that each communication path can transmit the traffic generated by the two computing nodes connected by the communication path in the process of executing the job task. Based on this, the cloud management platform can use the information of the job task to perform traffic prediction on the plurality of communication paths, thereby obtaining the predicted traffic transmitted by the plurality of communication paths. In the process of executing the job task by the plurality of computing nodes, the cloud management platform can also perform traffic collection on the plurality of communication paths, thereby obtaining the real traffic transmitted by the plurality of communication paths. If the predicted traffic of a communication path does not match the real traffic of the communication path, the cloud management platform can send a network exception notification to the tenant to remind the two computing nodes connected by the communication path of the exception. As can be seen, the cloud management platform can monitor the traffic transmitted by the communication paths between the computing nodes to determine whether the network between the computing nodes is abnormal, and then determine whether the computing nodes are abnormal, without the need to insert a monitoring component into the computing nodes to determine the network exception, which is beneficial to release the resources occupied by the computing nodes, thereby improving the resource utilization of the computing nodes, i.e., improving the efficiency of the computing nodes in executing the job tasks of the tenants, thereby improving the tenant experience. In order to further understand the working process of the cloud management platform, the process is further introduced below in combination with FIG. 3. FIG. 3 is a flowchart of a network exception determination method based on a cloud management platform according to an embodiment of the present application. As shown in FIG. 3, the method can be implemented by the cloud service system shown in FIG. 1. The cloud service system can include a plurality of computing nodes created for a tenant, the plurality of computing nodes can be used to execute the job tasks of the tenant, the plurality of computing nodes have a plurality of communication paths, each communication path connects two computing nodes in the plurality of computing nodes, and each communication path can be used to transmit the traffic generated by the two computing nodes connected by the communication path in the process of executing the job task. The method includes:

[0054] 301. The cloud management platform performs traffic prediction on the plurality of communication paths based on the information of the job task provided by the tenant to obtain the predicted traffic transmitted by the plurality of communication paths. The information includes data used to execute the job task, a model used to execute the job task, and a strategy followed to execute the job task.

[0055] In this embodiment, when a tenant needs to process a job task, the cloud management platform provides a processing interface (e.g., a job task processing column of a tenant interface) to the client of the tenant, so that the tenant can send a processing request of the job task of the tenant to the processing interface through the client, so that the processing request of the job task contains the information of the job task and the performance requirement of the job task, so that the cloud management platform can create a plurality of computing nodes in the infrastructure to meet the performance requirement of the job task, the plurality of computing nodes can be used to execute the job task of the tenant, the plurality of computing nodes have a plurality of communication paths, each communication path can connect two computing nodes in the plurality of computing nodes, so that each communication path can be used to transmit the traffic generated by the two computing nodes connected by the communication path in the process of executing the job task.

[0056] It should be noted that the performance requirement of the job task can include information such as the specification of the computing resource required to execute the job task, the specification of the storage resource, and the specification of the network resource, wherein the computing resource can include various types of processors such as central processing unit (CPU), graphics processing unit (GPU), neural processing unit (NPU), tensor processing unit (TPU), and data processing unit (DPU), the network resource can include a network card and the like, and the storage resource can include memory and disk and the like. The information of the job task can include information such as one or more data required to execute the job task, a model required to execute the job task, and a strategy required to execute the job task, for example, when the job task is a training task of a neural network model, the data is training data, the model is a neural network model to be trained, and the strategy includes any one or any combination of pipeline parallel strategy, data parallel strategy, model parallel strategy, and sequence parallel strategy. Of course, the job task can also be other types of tasks, which are not described one by one here.

[0057] Before notifying the plurality of computing nodes to execute the job task, the cloud management platform can use the information of the job task to perform traffic prediction on the plurality of communication paths between the plurality of computing nodes, thereby obtaining and recording the predicted traffic transmitted by the plurality of communication paths, that is, the ideal traffic that should be transmitted by the plurality of communication paths.

[0058] Specifically, the cloud management platform can obtain the predicted traffic transmitted by the plurality of communication paths in the following ways:

[0059] Before notifying the plurality of computing nodes to execute the job task, the cloud management platform can call a simulation application, input information of the job task provided by the tenant (including information required for executing the job task, a model required for executing the job task, and a strategy required for executing the job task, etc.) and information of the plurality of computing nodes (including a communication relationship between the plurality of computing nodes (for example, two computing nodes in the plurality of computing nodes form a communication pair, and there is a communication path between the two computing nodes), a network topology between the plurality of computing nodes (for example, a number and a layer of switches arranged between the plurality of computing nodes, the switches are used to connect the plurality of computing nodes, thereby forming a plurality of communication paths between the plurality of computing nodes), and a networking manner between the plurality of computing nodes, etc.) into the simulation application, so that the simulation application simulates a process of executing the job task by the plurality of computing nodes, thereby simulating predicted traffic transmitted by the plurality of communication paths.

[0060] For example, as shown in FIG. 4 (FIG. 4 is another structure schematic diagram of a cloud service system provided by an embodiment of the present application), when the tenant needs to execute a job task 1, the tenant can log in to the cloud management platform, and the cloud management platform can provide a tenant interface for the tenant, the tenant interface including a job task processing bar, the tenant can input a processing request for the job task 1 into the job task processing bar, the processing request including information of the job task 1 and performance requirements, so that the cloud management platform can select physical servers 1 to 14 that meet the performance requirements from a physical server cluster of the infrastructure, and there are 2-layer switches between the 14 physical servers, so that a large number of communication paths can be formed between the 14 physical servers, each communication path connecting two computing nodes and passing through at least one switch.

[0061] Then, before notifying the 14 physical servers to execute the job task 1, the cloud management platform can input information of the job task 1 and information of the 12 physical servers into the simulation application, so that the simulation application simulates a process of executing the job task 1 by the 14 physical servers, thereby obtaining predicted traffic transmitted by a plurality of communication paths between the 14 physical servers, for example, predicted traffic transmitted by a communication path between the physical server 1 and the physical server 2, predicted traffic transmitted by a communication path between the physical server 1 and the physical server 3,..., predicted traffic transmitted by a communication path between the physical server 12 and the physical server 14, and predicted traffic transmitted by a communication path between the physical server 13 and the physical server 14.

[0062] 302、In the process of executing the job task by the plurality of computing nodes, the cloud management platform collects traffic of the plurality of communication paths, and obtains real traffic transmitted by the plurality of communication paths.

[0063] 303、The cloud management platform determines, among the plurality of communication paths, a first communication path in which the predicted traffic does not match the real traffic, and sends a network anomaly notification to the tenant, the network anomaly notification indicating two computing nodes connected by the first communication path.

[0064] The cloud management platform can then notify the plurality of computing nodes to perform the job task, and collect traffic of the plurality of communication paths during the performance of the job task by the plurality of computing nodes, to obtain real traffic transmitted by the plurality of communication paths. For any one of the plurality of communication paths, referred to as a first communication path hereinafter, if it is detected that the predicted traffic transmitted by the first communication path does not match the real traffic transmitted by the first communication path (e.g., the predicted traffic transmitted by the first communication path is less than the real traffic transmitted by the first communication path, etc.), the cloud management platform can determine that the first communication path is abnormal, and return a network anomaly notification to the client of the tenant, the notification indicating two computing nodes connected by the first communication path. In this way, the tenant can determine, based on the notification, that a network anomaly occurs between the plurality of computing nodes serving the tenant, and is caused by the two computing nodes connected by the first communication path indicated by the notification.

[0065] The cloud management platform can also perform the type operation on the remaining communication paths other than the first communication path, which is not described herein.

[0066] Still referring to the above example, the cloud management platform can notify the 14 physical servers to perform the job task 1, and collect real traffic transmitted by the plurality of communication paths among the 14 physical servers during the performance of the job task 1 by the 14 physical servers, such as real traffic transmitted by a communication path between the physical server 1 and the physical server 2, real traffic transmitted by a communication path between the physical server 1 and the physical server 3,..., real traffic transmitted by a communication path between the physical server 12 and the physical server 14, and real traffic transmitted by a communication path between the physical server 13 and the physical server 14.

[0067] The cloud management platform can then compare the predicted traffic transmitted by the plurality of communication paths with the real traffic transmitted by the plurality of communication paths. Assuming that the real traffic transmitted by the communication path between the physical server 1 and the physical server 2 is greater than the predicted traffic transmitted by the communication path, the cloud management platform can send a network anomaly notification to the tenant to remind the tenant that the physical server 1 and the physical server 2 are abnormal during the performance of the job task 1, causing the communication path between the two physical servers to have a slow network condition of excessive traffic.

[0068] Specifically, the cloud management platform can further perform the following operations:

[0069] Suppose the multiple communication paths include two communication paths that exist partially shared, which are referred to as a second communication path and a third communication path respectively in the following, it can be understood that there is a shared sub-communication path between the second communication path and the third communication path. Then,

[0070] In the process of executing the job task by the multiple computing nodes, the cloud management platform can further perform traffic collection on the sub-communication path between the second communication path and the third communication path, so as to obtain the real traffic transmitted by the sub-communication path. Then, the cloud management platform can perform a series of decomposition operations on the real traffic transmitted by the sub-communication path, so as to obtain the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path respectively. If the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path overlap in the time dimension (that is, the sub-communication path needs to transmit the two parts of real traffic at the same time or at some times), the cloud management platform can make the network anomaly notification sent to the client of the tenant further indicate the two computing nodes connected by the second communication path and the two computing nodes connected by the third communication path.

[0071] In this way, the tenant can determine that there is a network anomaly between the multiple computing nodes serving itself based on the notification, which is not only caused by the two computing nodes connected by the first communication path, but also caused by the two computing nodes connected by the second communication path and the two computing nodes connected by the third communication path.

[0072] Still as the above example, the cloud management platform can instruct the 14 physical servers to perform the job task 1, and during the process of the 14 physical servers performing the job task 1, since the communication path between the physical server 2 and the physical server 7 and the communication path between the physical server 9 and the physical server 10 both contain the sub-communication path between the switch 5 and the switch 2 (assuming that the sub-communication path is only covered by the two communication paths), the cloud management platform can collect the real traffic transmitted by the sub-communication path and decompose the real traffic, so as to obtain the real traffic transmitted by the communication path between the physical server 2 and the physical server 7 and the real traffic transmitted by the communication path between the physical server 9 and the physical server 10, as shown in FIG. 5 (FIG. 5 is a schematic diagram of traffic decomposition provided by an embodiment of the present application). Since the real traffic transmitted by the communication path overlaps in the time dimension (such as the part of traffic circled in FIG. 5), the cloud management platform can send a network anomaly notification to the tenant, which can remind the tenant that the physical server 1 and the physical server 2 appear abnormal during the process of performing the job task 1, which causes the slow network situation of the communication path between the two physical servers being too large, and the physical server 2, the physical server 7, the physical server 9 and the physical server 10 appear abnormal during the process of performing the job task 1, which causes the slow network situation of the sub-communication path shared by the four physical servers being congested.

[0073] More specifically, the job task package required to be completed by the tenant can contain a first job task and a second job task. Then, the plurality of computing nodes serving the tenant can be divided into two parts of computing nodes, the first part of computing nodes being used to perform the first job task, and the second part of computing nodes being used to perform the second job task. Then, the two computing nodes connected by the second communication path and the two computing nodes connected by the third communication path can have multiple situations: (1) the two computing nodes connected by the second communication path and the two computing nodes connected by the third communication path are both used to perform the first job task. (2) the two computing nodes connected by the second communication path and the two computing nodes connected by the third communication path are both used to perform the second job task. (3) the two computing nodes connected by the second communication path are used to perform the first job task, and the two computing nodes connected by the third communication path are used to perform the second job task, in which case, the network anomaly notification sent by the cloud management platform to the client of the tenant is also used to indicate the first job task and the second job task, so as to notify the tenant that the network anomaly occurs in the execution process of the first job task, which is caused by the two computing nodes connected by the second communication path, and the network anomaly also occurs in the execution process of the second job task, which is caused by the two computing nodes connected by the third communication path.

[0074] For example, as shown in FIG. 6 (FIG. 6 is another structure schematic diagram of the cloud service system provided by the embodiment of the present application), assuming that a tenant needs to complete a job task 1 and a job task 2, the cloud management platform can select physical servers 1 to 14 for the tenant based on the performance requirements of the two job tasks, wherein the physical servers 1 to 7 are used to execute the job task 1, the physical servers 8 to 14 are used to execute the job task 2, and the communication path between the physical servers 2 and 7 and the communication path between the physical servers 9 and 10 both contain a sub-communication path between the switches 5 to 2.

[0075] Then, in the process of executing the job task 1 by the physical servers 1 to 6 and in the process of executing the job task 2 by the physical servers 7 to 12, the cloud management platform can collect the real traffic transmitted by the sub-communication path, and decompose the real traffic, so as to obtain the real traffic transmitted by the communication path between the physical servers 2 and 7 and the real traffic transmitted by the communication path between the physical servers 9 and 10. Since the real traffic transmitted by the communication path overlaps in the time dimension, the cloud management platform can send a network anomaly notification to the tenant, which can remind the tenant of the following contents: the physical servers 2 and 7 appear abnormal in the process of executing the job task 1, the physical servers 9 and 10 appear abnormal in the process of executing the job task 2, and the slow network condition of the sub-communication path shared by the four physical servers appears to be congested.

[0076] It should be understood that in the embodiment, the first communication path can be a completely independent communication path or a partially shared communication path. When the first communication path is a completely independent communication path, the cloud management platform can directly collect the real traffic transmitted by the first communication path, and when the first communication path is a partially shared communication path, the cloud management platform can also determine the sub-communication path shared by the first communication path and other communication paths, and then collect and decompose the real traffic transmitted by the sub-communication path, so as to obtain the real traffic transmitted by the first communication path, which will not be described herein again.

[0077] It should also be understood that in the embodiment, the first job task and the second job task can be different job tasks belonging to the same tenant, or can be job tasks belonging to different tenants, which is not limited herein.

[0078] In the embodiments of the present application, when a tenant needs to complete a job task, the tenant can send a processing request for the job task to a processing interface provided by the cloud management platform, so that the cloud management platform can receive the processing request for the job task through the processing interface. Since the processing request contains the information of the job task and the performance requirement of the job task, the cloud management platform can create a plurality of computing nodes meeting the performance requirement of the job task to execute the job task. Among them, the plurality of computing nodes have a plurality of communication paths, each communication path can connect two computing nodes in the plurality of computing nodes, so that each communication path can transmit the traffic generated by the two computing nodes connected by the communication path in the process of executing the job task. Based on this, the cloud management platform can use the information of the job task to perform traffic prediction on the plurality of communication paths, thereby obtaining the predicted traffic transmitted by the plurality of communication paths. In the process of executing the job task by the plurality of computing nodes, the cloud management platform can also collect the traffic of the plurality of communication paths, thereby obtaining the real traffic transmitted by the plurality of communication paths. If the predicted traffic of a first communication path in the plurality of communication paths does not match the real traffic of the first communication path, the cloud management platform can send a network exception notification to the tenant to remind that the two computing nodes connected by the first communication path are abnormal. In the foregoing process, the cloud management platform can monitor the traffic transmitted by the communication paths between the computing nodes to determine whether the network between the computing nodes is abnormal, and then determine whether the computing nodes are abnormal, without the need to insert a monitoring component into the computing nodes to determine the network exception, which is beneficial to release the resources occupied by the computing nodes, thereby improving the resource utilization of the computing nodes, i.e., improving the efficiency of the computing nodes in executing the job tasks of the tenants, thereby improving the tenant experience.

[0079] Further, in the embodiments of the present application, since the cloud management platform no longer needs to insert a monitoring component into the computing nodes to determine the network exception, such a way does not invade the applications deployed by the tenant on the computing nodes, which is beneficial to protect the privacy and data security of the tenant, thereby further improving the tenant experience.

[0080] Further, in the embodiments of the present application, the cloud management platform can determine and analyze a plurality of network exception conditions based on whether the traffic transmitted by a single communication path is abnormal to determine whether the communication path has a slow network condition of too much traffic, and based on whether the traffic transmitted by a plurality of communication paths overlaps in time sequence to determine whether a sub-communication path shared by the plurality of communication paths has a slow network condition of traffic congestion.

[0081] The above is a detailed description of the network anomaly determination method based on the cloud management platform provided by the embodiments of the present application. The cloud management platform provided by the embodiments of the present application will be introduced below. FIG. 7 is a structural schematic diagram of a cloud management platform provided by an embodiment of the present application. As shown in FIG. 7, the cloud management platform is used to manage infrastructure, and the infrastructure includes a plurality of computing nodes created by the cloud management platform for tenants. The plurality of computing nodes are used to execute job tasks of the tenants. The plurality of computing nodes have a plurality of communication paths therebetween. Each communication path connects two computing nodes in the plurality of computing nodes. Each communication path is used to transmit traffic generated by the two computing nodes connected thereby in the process of executing the job tasks. The cloud management platform includes:

[0082] A prediction module 701 is configured to perform traffic prediction on the plurality of communication paths based on information of the job tasks provided by the tenants, to obtain predicted traffic transmitted by the plurality of communication paths. The information includes data required to execute the job tasks, models required to execute the job tasks, and strategies required to be followed to execute the job tasks. For example, the prediction module 701 is configured to implement step 301 in the embodiment shown in FIG. 3.

[0083] A collection module 702 is configured to perform traffic collection on the plurality of communication paths by the cloud management platform in the process of executing the job tasks by the plurality of computing nodes, to obtain real traffic transmitted by the plurality of communication paths. For example, the collection module 702 is configured to implement step 302 in the embodiment shown in FIG. 3.

[0084] A notification module 703 is configured to determine, among the plurality of communication paths, a first communication path in which the predicted traffic does not match the real traffic, and send a network anomaly notification to the tenants. The network anomaly notification is used to indicate two computing nodes connected by the first communication path. For example, the notification module 703 is configured to implement step 303 in the embodiment shown in FIG. 3.

[0085] In a possible implementation manner, the prediction module 701 is configured to simulate the process of executing the job tasks by the plurality of computing nodes based on the information of the job tasks provided by the tenants, to obtain the predicted traffic transmitted by the plurality of communication paths.

[0086] In a possible implementation, the plurality of communication paths includes the second communication path and a third communication path, and the second communication path and the third communication path share a sub-communication path, and the cloud management platform further includes a decomposition module configured to: collect, during execution of the job task by the plurality of computing nodes, real traffic transmitted by the sub-communication path to obtain real traffic transmitted by the sub-communication path; decompose the real traffic transmitted by the sub-communication path to obtain real traffic transmitted by the second communication path and real traffic transmitted by the third communication path; and when the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path overlap in a time dimension, the network anomaly notification is further configured to indicate two computing nodes connected by the second communication path and two computing nodes connected by the third communication path.

[0087] In a possible implementation, the job task includes a first job task and a second job task, the two computing nodes connected by the second communication path are configured to execute the first job task, and the two computing nodes connected by the third communication path are configured to execute the second job task, and the network anomaly notification is further configured to indicate the first job task and the second job task.

[0088] In a possible implementation, the model is a neural network model to be trained, and the data is training data, and the strategy includes any one or any combination of a pipeline parallel strategy, a data parallel strategy, a model parallel strategy, and a sequence parallel strategy.

[0089] In a possible implementation, the computing node includes a physical server, a virtual machine, a container, a micro virtual machine, or a bare metal server.

[0090] In a possible implementation, the plurality of computing nodes are deployed in a same site or different sites, and the site includes a region, an availability zone, a data center, a machine room, or a physical server group.

[0091] It should be noted that the information interaction and implementation process between the modules / units of the apparatus described above are based on the same concept as the method embodiments of the present application, and the technical effects brought by the method embodiments of the present application are the same. The specific content can be referred to the description of the method embodiments described above, and will not be repeated here.

[0092] Referring to FIG. 8, FIG. 8 is a structural schematic diagram of a computing device provided by an embodiment of the present application. As shown in FIG. 8, the computing device 800 (which can be used to present the cloud management platform described above) includes a processor 801, a memory 802, a communication interface 803, and a bus 804, wherein the processor 801, the memory 802, and the communication interface 803 are coupled through the bus (not shown in the figure). The memory 802 stores instructions, and when the instructions stored in the memory 802 are executed, the computing device 800 performs the method performed by the cloud management platform in the method embodiments described above.

[0093] The computing device 800 can be one or more integrated circuits configured to implement the above method, for example, one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. For another example, when the units in the apparatus can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can invoke programs. For another example, these units can be integrated together to implement a system-on-a-chip (SOC).

[0094] The processor 801 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor can be a microprocessor, or any conventional processor.

[0095] The memory 802 can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0096] The executable program code stored in the memory 802 is executed by the processor 801 to realize the functions of the aforementioned prediction module, acquisition module, and notification module, etc. to realize the network anomaly determination method based on the cloud management platform. That is, the memory 802 has instructions for executing the network anomaly determination method based on the cloud management platform.

[0097] The communication interface 803 uses a transceiver module such as, but not limited to, a network interface card, a transceiver, etc. to realize the communication between the computing device 800 and other devices or communication networks.

[0098] The bus 804 can include, in addition to a data bus, a power bus, a control bus, and a state signal bus, etc. The bus can be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0099] Referring to FIG. 9, FIG. 9 is a structural schematic diagram of a computing device cluster provided by an embodiment of the present application. As shown in FIG. 9, the computing device cluster 900 includes at least one computing device 800.

[0100] As shown in FIG. 9, the computing device cluster 900 includes at least one computing device 800. The memory 802 in one or more computing devices 800 in the computing device cluster 900 can store the same instructions for executing the network anomaly determination method based on the cloud management platform described above.

[0101] In some possible implementation manners, the memory 802 in one or more computing devices 800 in the computing device cluster 900 can also respectively store partial instructions for executing the network anomaly determination method based on the cloud management platform described above. In other words, the combination of one or more computing devices 800 can collectively execute the instructions for executing the network anomaly determination method based on the cloud management platform described above.

[0102] It should be noted that the memory 802 in different computing devices 800 in the computing device cluster 900 can store different instructions, respectively, for executing part of the functions of the cloud management platform described above. That is, the instructions stored in the memory 802 in different computing devices 800 can implement the functions of one or more of the prediction module, the collection module, and the notification module, etc.

[0103] In some possible implementation manners, one or more computing devices 800 in the computing device cluster 900 can be connected through a network. The network can be a wide area network or a local area network, etc.

[0104] Referring to FIG. 10, FIG. 10 is a schematic diagram of the connection of the computer devices in the computer cluster provided in the embodiments of the present application through a network. As shown in FIG. 10, the two computer devices 800A and 800B are connected through a network. Specifically, the communication interface in each computer device is connected to the network.

[0105] In a possible implementation, the memory in the computer device 800A stores instructions for executing the functions of the prediction module and other modules. Meanwhile, the memory in the computer device 800B stores instructions for executing the functions of the collection module, the notification module and other modules.

[0106] It should be understood that the functions of the computer device 800A shown in FIG. 10 can also be completed by multiple computer devices. Similarly, the functions of the computer device 800B can also be completed by multiple computer devices.

[0107] The embodiments of the present application also relate to a computer storage medium, which stores a program for signal processing, and when the program runs on a computer, the computer executes the steps performed by the cloud management platform in the embodiment shown in FIG. 3.

[0108] The embodiments of the present application also relate to a computer program product, which stores instructions, and when the instructions are executed by a computer, the computer executes the steps performed by the cloud management platform in the embodiment shown in FIG. 3.

[0109] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.

[0110] In the several embodiments of the present application, it should be understood that the disclosed system, device and method can be implemented by other ways. For example, the device embodiments described above are only schematic, and the division of the units is only a logical function division, and there can be another division way in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0111] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiments of the present application.

[0112] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0113] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), magnetic disk or optical disk, and various media that can store program codes.

Claims

1. A network anomaly determination method based on a cloud management platform, characterized in that, The cloud management platform is used for managing an infrastructure, the infrastructure including a plurality of computing nodes created by the cloud management platform for a tenant, the plurality of computing nodes being used for executing a job task of the tenant, the plurality of computing nodes having a plurality of communication paths therebetween, each communication path connecting two computing nodes in the plurality of computing nodes, the each communication path being used for transmitting traffic generated by the two computing nodes connected thereby in the process of executing the job task, and the method comprising: The cloud management platform performs traffic prediction on the plurality of communication paths based on information of the job task provided by the tenant, the information including data to be used for executing the job task, a model to be used for executing the job task, and a policy to be followed for executing the job task, to obtain predicted traffic transmitted by the plurality of communication paths; In the process of executing the job task by the plurality of computing nodes, the cloud management platform performs traffic collection on the plurality of communication paths to obtain real traffic transmitted by the plurality of communication paths; The cloud management platform determines a first communication path in which the predicted traffic does not match the real traffic in the plurality of communication paths, and sends a network anomaly notification to the tenant, the network anomaly notification being used for indicating two computing nodes connected by the first communication path.

2. The method of claim 1, wherein, The cloud management platform performs traffic prediction on the plurality of communication paths based on information of the job task provided by the tenant, the information including data to be used for executing the job task, a model to be used for executing the job task, and a policy to be followed for executing the job task, to obtain predicted traffic transmitted by the plurality of communication paths; The cloud management platform performs traffic prediction on the plurality of communication paths based on information of the job task provided by the tenant, the information including data to be used for executing the job task, a model to be used for executing the job task, and a policy to be followed for executing the job task, to obtain predicted traffic transmitted by the plurality of communication paths; 3. The method according to claim 1 or 2, characterized in that, The plurality of communication paths include a second communication path and a third communication path, and the second communication path and the third communication path have a shared sub-communication path therebetween, and the method further comprises: In the process of executing the job task by the plurality of computing nodes, the cloud management platform performs traffic collection on the sub-communication path to obtain real traffic transmitted by the sub-communication path; The real traffic transmitted by the sub-communication path is decomposed to obtain real traffic transmitted by the second communication path and real traffic transmitted by the third communication path; When the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path overlap in a time dimension, the network anomaly notification is further used for indicating two computing nodes connected by the second communication path and two computing nodes connected by the third communication path.

4. The method of claim 3, wherein, The job task includes a first job task and a second job task, the two computing nodes connected by the second communication path being used for executing the first job task, and the two computing nodes connected by the third communication path being used for executing the second job task, and the network anomaly notification is further used for indicating the first job task and the second job task.

5. The method according to any one of claims 1 to 4, characterized in that, The model is a neural network model to be trained, the data is training data, and the strategy includes any one or any combination of the following pipeline parallel strategy, data parallel strategy, model parallel strategy, and sequence parallel strategy.

6. The method according to any one of claims 1 to 5, characterized in that, The computing nodes include physical servers, virtual machines, containers, micro virtual machines, or bare metal servers.

7. The method according to any one of claims 1 to 6, characterized in that, The plurality of computing nodes are deployed in the same site or different sites, and the site includes a region, an availability zone, a data center, a machine room, or a physical server group.

8. A cloud management platform, characterized by, The cloud management platform is used to manage infrastructure, which includes a plurality of computing nodes created by the cloud management platform for a tenant, the plurality of computing nodes are used to execute job tasks of the tenant, and the plurality of computing nodes have a plurality of communication paths therebetween, each communication path connecting two computing nodes in the plurality of computing nodes, each communication path being used to transmit traffic generated by the two computing nodes connected thereby in the process of executing the job tasks, and the cloud management platform includes: a prediction module configured to perform traffic prediction on the plurality of communication paths based on information of the job tasks provided by the tenant, to obtain predicted traffic transmitted by the plurality of communication paths, the information including data to be used in executing the job tasks, models to be used in executing the job tasks, and a strategy to be followed in executing the job tasks; a collection module configured to perform traffic collection on the plurality of communication paths in the process of executing the job tasks by the plurality of computing nodes, to obtain real traffic transmitted by the plurality of communication paths; a notification module configured to determine, among the plurality of communication paths, a first communication path in which predicted traffic does not match real traffic, and to send a network anomaly notification to the tenant, the network anomaly notification being used to indicate two computing nodes connected by the first communication path.

9. The cloud management platform of claim 8, wherein, The prediction module is configured to simulate the process of executing the job tasks by the plurality of computing nodes based on the information of the job tasks provided by the tenant, to obtain predicted traffic transmitted by the plurality of communication paths.

10. The cloud management platform of claim 8 or 9, wherein, The plurality of communication paths include a second communication path and a third communication path, and the second communication path and the third communication path have a shared sub-communication path therebetween, and the cloud management platform further includes a decomposition module configured to: perform traffic collection on the sub-communication path in the process of executing the job tasks by the plurality of computing nodes, to obtain real traffic transmitted by the sub-communication path; decompose the real traffic transmitted by the sub-communication path, to obtain real traffic transmitted by the second communication path and real traffic transmitted by the third communication path; when the real traffic transmitted by the second communication path and the real traffic transmitted by the third communication path overlap in the time dimension, the network anomaly notification is further used to indicate two computing nodes connected by the second communication path and two computing nodes connected by the third communication path.

11. The cloud management platform of claim 10, wherein, The job tasks include a first job task and a second job task, two computing nodes connected by the second communication path are used to execute the first job task, and two computing nodes connected by the third communication path are used to execute the second job task, and the network anomaly notification is further used to indicate the first job task and the second job task.

12. The cloud management platform of any of claims 8 to 11, wherein, The model is a neural network model to be trained, the data is training data, and the strategy includes any one or any combination of a pipeline parallel strategy, a data parallel strategy, a model parallel strategy, and a sequence parallel strategy.

13. The cloud management platform of any of claims 8 to 12, wherein, The computing node includes a physical server, a virtual machine, a container, a micro virtual machine, or a bare metal server.

14. The cloud management platform of any of claims 8 to 13, wherein, The plurality of computing nodes are deployed in the same site or different sites, and the site includes a region, an availability zone, a data center, a machine room, or a physical server group.

15. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, and each computing device includes a processor and a memory: The memory is configured to store instructions; The processor is configured to execute the instructions to cause the computing device cluster to perform the method of any one of claims 1-7.

16. A computer storage medium, comprising, The computer storage medium stores one or more instructions, which, when executed by one or more computers, cause the one or more computers to implement the method of any one of claims 1-7.

17. A computer program product, characterised in that, The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for managing traffic in cloud computing system

    CN108965025A

  • Link anomaly detection method and device

    CN111314121A

  • Traffic protection method and device in cloud environment

    CN114244576A

  • Traffic data processing method and device, storage medium and electronic equipment

    CN117675389A

  • Nuclear power industry-oriented network abnormal behavior detection and analysis method

    CN118487872A