Communication system, resource management method and apparatus
Through real-time monitoring and adjustment of the intelligent analysis and management platform, the network congestion problem caused by sudden and frequent changes in traffic from AI business applications was solved, improving the operational efficiency and resource utilization of the data center.
Patent Information
- Application Number
- CN202411248877.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-09-05
AI Technical Summary
Traditional network operation and maintenance strategies are unable to effectively cope with the problem of sudden high and frequent traffic fluctuations in AI business applications, which often leads to data center network congestion.
Through intelligent analysis and intelligent management platforms, the bandwidth of business applications and the performance parameters of computing nodes can be monitored and adjusted in real time, and network resource allocation and computing load can be dynamically adjusted to cope with sudden traffic and frequent changes in AI services.
It improves the efficiency and quality of data center network operation and maintenance, maximizes resource utilization, and ensures the stable operation of business applications.
Smart Images

Figure CN119135609B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One or more embodiments of the present specification relate to the technical field of communication, and in particular to a communication system, a resource management method and device. BACKGROUND
[0002] A data center (DC) is a facility for centralized storage, processing and distribution of data, and a data center interconnect (DCI) network is a network facility connecting multiple data centers, aiming to realize high-speed, reliable and secure communication between data centers.
[0003] In recent years, AI (Artificial Intelligence) technology has developed rapidly, and various AI large model applications and software have emerged in an endless stream. The rapid growth of network traffic in data centers has put tremendous pressure on data center network equipment. The characteristics of AI business applications are high burst traffic and frequent changes and adjustments. Traditional network operation and maintenance strategies cannot effectively cope with this, resulting in network congestion. SUMMARY
[0004] To improve the efficiency and quality of operation and maintenance of data centers and ensure stable operation of business applications, the embodiments of the present specification provide a communication system, a resource management method and device.
[0005] In a first aspect, the embodiments of the present specification provide a communication system, comprising a data center, a data center network and a management and operation platform, the data center is deployed with a plurality of business applications, each business application is running on a computing node of the data center, the management and operation platform comprises an intelligent analysis platform and an intelligent management and control platform;
[0006] The intelligent analysis platform is configured to obtain bandwidth information of each business flow in the data center network and performance parameters of each computing node included in the data center, and determine a business application to which each business flow belongs based on a pre-constructed application networking relationship, and determine bandwidth parameters of the business application according to bandwidth information of business flows belonging to the same business application, wherein the application networking relationship comprises a business flow transmission path of each business application.
[0007] The intelligent management and control platform is configured to, for each business application, adjust a management and control strategy for running the business application according to the bandwidth parameters of the business application and performance parameters of the computing node running the business application, wherein the management and control strategy comprises allocating network bandwidth for the business application in the case that the bandwidth parameters reach a set bandwidth, and adjusting the running load of the computing node in the case that the performance parameters reach a performance threshold.
[0008] In some embodiments, the data center network comprises a cloud cluster switch, the cloud cluster switch comprises a plurality of network nodes, flow information of each service flow in the cloud cluster switch is collected by each network node, the flow information comprises traffic identification of the service flow and the bandwidth information, and the intelligent analysis platform receives the flow information of each service flow sent by the cloud cluster switch.
[0009] In some embodiments, the plurality of network nodes of the cloud cluster switch comprises an NCP node and an NCC node, the NCP node is configured to collect traffic identification of service flows sent by the computing node, establish a first flow table based on the traffic identification of the service flows, and send the first flow table to the NCC node;
[0010] The NCC node is configured to count traffic of each service flow according to the received first flow table to obtain bandwidth information of each service flow.
[0011] In some embodiments, the NCP node is specifically configured to, in the case that storage space is sufficient, establish a first flow table according to the received traffic identification of the service flows and send the first flow table to the NCC node, and in the case that storage space is insufficient, establish a first flow table according to the received traffic identification of service flows of a preset message type and send the first flow table to the NCC node.
[0012] The NCC node is specifically configured to, in the case that storage space is sufficient, establish a second flow table according to the traffic identification of the service flows in the first flow table, in the case that storage space is insufficient, establish a second flow table according to the traffic identification of service flows of a preset message type in the first flow table, count traffic of each service flow in the second flow table, and obtain bandwidth information of each service flow.
[0013] In some embodiments, the intelligent analysis platform is specifically configured to,
[0014] obtain a path port number corresponding to each service flow, match the path port number with a transport layer port number corresponding to each service application in the application networking relationship, and determine a service application to which each service flow belongs;
[0015] sum bandwidth information of service flows belonging to the same service application to obtain a bandwidth parameter corresponding to each service application.
[0016] In some embodiments, the management and operation platform further comprises a computing management platform, the computing management platform is configured to collect performance parameters of each computing node and send the performance parameters of each computing node to the intelligent analysis platform.
[0017] In a case that the computing node is a GPU computing node, the performance parameter of the GPU computing node includes a floating point operation performance and a GPU memory bandwidth.
[0018] In some embodiments, the intelligent analysis platform is further configured to acquire packet loss information of each service flow in the data center network, and determine a packet loss rate of a service application according to the packet loss information of service flows belonging to the service application.
[0019] The intelligent management and control platform is further configured to adjust a management and control strategy for running the service application according to the bandwidth parameter, the packet loss rate, and a performance parameter of a computing node running the service application, wherein the management and control strategy includes adjusting a packet forwarding strategy of the service application in a case that the packet loss rate reaches a preset threshold.
[0020] In some embodiments, the management and operation platform further includes an intelligent deployment platform, which is configured to acquire all service applications deployed on the data center and a transport layer port number of each service application, wherein the transport layer port number includes path port numbers of service flows included in the service application, and to generate the application networking relationship according to the path port numbers of service flows of each service application, taking a service application as a node and a transmission path of a service flow as an edge.
[0021] In a second aspect, the embodiments of the present specification provide a resource management method applied to a communication system, the communication system including a data center and a data center network, the data center deploying a plurality of service applications, the service applications running on computing nodes of the data center, and the method including:
[0022] acquiring bandwidth information of each service flow in the data center network and a performance parameter of each computing node included in the data center;
[0023] determining a service application to which each service flow belongs based on a pre-constructed application networking relationship, and determining a bandwidth parameter of the service application according to bandwidth information of service flows belonging to the service application, wherein the application networking relationship includes a service flow transmission path of each service application;
[0024] for each service application, adjusting a management and control strategy for running the service application according to the bandwidth parameter of the service application and a performance parameter of a computing node running the service application, wherein the management and control strategy includes allocating network bandwidth for the service application in a case that the bandwidth parameter reaches a set bandwidth, and adjusting a running load of the computing node in a case that the performance parameter reaches a performance threshold.
[0025] In some embodiments, the data center network comprises a cloud cluster switch comprising a plurality of network nodes, and the obtaining the bandwidth information of each traffic flow in the data center network comprises:
[0026] The flow information of each traffic flow in the cloud cluster switch is collected by each network node, and the flow information comprises a traffic identifier of the traffic flow and the bandwidth information.
[0027] In some embodiments, the plurality of network nodes of the cloud cluster switch comprises an NCP node and an NCC node, and the collecting the flow information of each traffic flow in the cloud cluster switch by each network node comprises:
[0028] The NCP node collects the traffic identifier of the traffic flow sent by the computing node, and establishes a first flow table based on the traffic identifier of the traffic flow, and uploads the first flow table to the NCC node;
[0029] The NCC node counts the traffic of each traffic flow according to the received first flow table to obtain the bandwidth information of each traffic flow.
[0030] In some embodiments, the NCP node collects the traffic identifier of the traffic flow sent by the computing node, and establishes a first flow table based on the traffic identifier of the traffic flow, and uploads the first flow table to the NCC node, comprising:
[0031] The NCP node establishes a first flow table according to the received traffic identifier of the traffic flow and uploads the first flow table to the NCC node in the case of sufficient storage space, and establishes a first flow table according to the received traffic identifier of the traffic flow of a preset message type and uploads the first flow table to the NCC node in the case of insufficient storage space;
[0032] The NCC node counts the traffic of each traffic flow according to the received first flow table to obtain the bandwidth information of each traffic flow, comprising:
[0033] The NCC node establishes a second flow table according to the traffic identifier of the traffic flow in the first flow table in the case of sufficient storage space, and establishes a second flow table according to the traffic identifier of the traffic flow of a preset message type in the first flow table in the case of insufficient storage space, and counts the traffic of each traffic flow in the second flow table to obtain the bandwidth information of each traffic flow.
[0034] In some embodiments, the determining the business application to which each traffic flow belongs based on the pre-constructed application networking relationship, and determining the bandwidth parameter of the business application according to the bandwidth information of the traffic flows belonging to the same business application, comprises:
[0035] obtain a path port number corresponding to each service flow, and match the path port number with a transport layer port number corresponding to each service application in the application networking relationship to determine a service application to which each service flow belongs;
[0036] sum bandwidth information of service flows belonging to the same service application to obtain a bandwidth parameter corresponding to each service application.
[0037] In some embodiments, the data center includes a first computing node, and in a case where the first computing node is a GPU computing node, the process of obtaining the performance parameter of the GPU computing node includes:
[0038] obtaining a floating point operation performance and a GPU memory bandwidth of the GPU computing node, and the performance parameter of the GPU computing node includes the floating point operation performance and the GPU memory bandwidth.
[0039] In some embodiments, the adjusting a management and control policy for running the service application according to the bandwidth parameter of the service application and a performance parameter of a computing node running the service application includes:
[0040] determining a packet loss rate of the service application according to packet loss information of service flows belonging to the same service application;
[0041] adjusting a management and control policy for running the service application according to the bandwidth parameter of the service application, the packet loss rate, and a performance parameter of a computing node running the service application, wherein the management and control policy includes adjusting a packet forwarding strategy of the service application in a case where the packet loss rate reaches a preset threshold.
[0042] In some embodiments, the process of pre-constructing the application networking relationship includes:
[0043] obtaining all service applications deployed on the data center and a transport layer port number of each service application, wherein the transport layer port number includes a path port number of each service flow included in the service application;
[0044] taking a service application as a node and a transport path of a service flow as an edge, and generating the application networking relationship according to a path port number of a service flow of each service application.
[0045] In a third aspect, the embodiments of the present specification provide a resource management apparatus applied to a communication system, the communication system including a data center and a data center network, the data center deploying a plurality of service applications, the service applications running on computing nodes of the data center, and the apparatus including:
[0046] an information acquisition module, configured to acquire bandwidth information of each service flow in the data center network and performance parameters of each computing node included in the data center;
[0047] a bandwidth determination module, configured to determine a service application to which each service flow belongs based on a pre-constructed application networking relationship, and determine a bandwidth parameter of the service application according to bandwidth information of service flows belonging to the same service application, wherein the application networking relationship includes a service flow transmission path of each service application;
[0048] a policy control module, configured to, for each service application, adjust a control policy for running the service application according to the bandwidth parameter of the service application and the performance parameter of the computing node running the service application, wherein the control policy includes: allocating network bandwidth for the service application in a case where the bandwidth parameter reaches a set bandwidth, and adjusting a running load of the computing node in a case where the performance parameter reaches a performance threshold.
[0049] In some embodiments, the data center network includes a cloud cluster switch including a plurality of network nodes, and the information acquisition module is specifically configured to acquire flow information of each service flow in the cloud cluster switch through each network node, the flow information including traffic identification and the bandwidth information of the service flow.
[0050] In some embodiments, the plurality of network nodes of the cloud cluster switch include NCP nodes and NCC nodes, and the information acquisition module is specifically configured to,
[0051] the NCP nodes acquire traffic identification of service flows sent by the computing nodes, and establish a first flow table based on the traffic identification of the service flows, and upload the first flow table to the NCC nodes;
[0052] the NCC nodes count traffic of each service flow according to the received first flow table to obtain the bandwidth information of each service flow.
[0053] In some embodiments, the information acquisition module is specifically configured to,
[0054] in a case where the storage space of the NCP nodes is sufficient, establish a first flow table according to the received traffic identification of the service flow, and upload the first flow table to the NCC nodes, and in a case where the storage space is insufficient, establish a first flow table according to the received traffic identification of the service flow of a preset message type, and upload the first flow table to the NCC nodes;
[0055] In a case that the NCC node has sufficient storage space, a second flow table is established according to the traffic identification of the service flow in the first flow table, and in a case that the NCC node does not have sufficient storage space, a second flow table is established according to the traffic identification of the service flow of a preset message type in the first flow table, and the traffic of each service flow in the second flow table is counted to obtain the bandwidth information of each service flow.
[0056] In some embodiments, the bandwidth determination module is specifically configured to,
[0057] The path port number corresponding to each service flow is acquired, and the path port number is matched with the transport layer port number corresponding to each service application in the application networking relationship to determine the service application to which each service flow belongs.
[0058] The bandwidth information of the service flows belonging to the same service application is summed up to obtain the bandwidth parameter corresponding to each service application.
[0059] In some embodiments, the data center includes a first computing node, and in a case that the first computing node is a GPU computing node, the information acquisition module is specifically configured to acquire the floating point operation performance and GPU memory bandwidth of the GPU computing node, and the performance parameter of the GPU computing node includes the floating point operation performance and the GPU memory bandwidth.
[0060] In some embodiments, the policy control module is specifically configured to,
[0061] The packet loss rate of the service application is determined according to the packet loss information of the service flows belonging to the same service application.
[0062] The control policy for running the service application is adjusted according to the bandwidth parameter, the packet loss rate, and the performance parameter of the computing node running the service application, wherein the control policy includes adjusting the message forwarding strategy of the service application in a case that the packet loss rate reaches a preset threshold.
[0063] In some embodiments, the apparatus further includes a relationship construction module, and the relationship construction module is configured to,
[0064] All service applications deployed on the data center and the transport layer port number of each service application are acquired, wherein the transport layer port number includes the path port number of each service flow included in the service application.
[0065] The service application is taken as a node, the transmission path of the service flow is taken as an edge, and the application networking relationship is generated according to the path port number of the service flow of each service application.
[0066] The communication system of the embodiment of the present specification determines the service application to which each service flow belongs according to the pre-constructed application networking relationship, determines the bandwidth parameter of the service application according to the bandwidth information of the service flow belonging to the same service application, and adjusts the management and control strategy of the service application according to the bandwidth parameter of the service application and the performance parameter of the computing node running the service application. Real-time monitoring of the service application and the computing node in the communication system is realized, the network bandwidth congestion condition of the service application and the busy degree of the computing node are perceived, so that the running strategy of the service application can be adjusted accordingly, the AI service data center such as the intelligent computing center which has high burst traffic and needs frequent changes and adjustments can be effectively responded to, the network operation and maintenance efficiency and quality are improved, and the resource utilization rate and processing capacity of the data center can be maximized. BRIEF DESCRIPTION OF DRAWINGS
[0067] Figure 1 FIG. 1 is a logical architecture diagram of a communication system in an exemplary embodiment of the present specification.
[0068] Figure 2 FIG. 2 is a system architecture diagram of a communication system in an exemplary embodiment of the present specification.
[0069] Figure 3 FIG. 3 is a schematic diagram of an application networking relationship in an exemplary embodiment of the present specification.
[0070] Figure 4 FIG. 4 is a flow chart of a resource management method in an exemplary embodiment of the present specification.
[0071] Figure 5 FIG. 5 is a structural block diagram of a resource management device in an exemplary embodiment of the present specification. DETAILED DESCRIPTION
[0072] The technical solutions of the present specification will be described in detail below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present specification. In addition, the technical features involved in the different embodiments of the present specification described below can be combined with each other as long as they do not conflict with each other.
[0073] A data center (DC) is a facility that centrally stores, processes, and distributes data. A data center typically includes servers, storage devices, and other equipment to provide various computing and storage services. A data center interconnect (DCI) network is a network facility that connects multiple data centers. The DCI network aims to enable high-speed, reliable, and secure communication between data centers to meet the high-bandwidth and low-latency requirements of data and applications. A data center network includes network transmission technologies, fiber connections, routers, switches, and other devices.
[0074] In recent years, AI (Artificial Intelligence) technology has developed rapidly, and various AI large model applications and software have emerged in an endless stream. The computing capacity and network traffic of data centers have grown rapidly, putting a huge pressure on data center network equipment. Traditional network operation and management cannot meet the requirements, causing network congestion. For example, the cloud service business provided by traditional data centers is relatively fixed, and the traffic of each business is relatively clear and controllable. Traffic conflicts generally occur during expected business peak periods, so network operation can guarantee normal network operation by adjusting and changing the network according to the business peak. However, AI business applications have extremely high real-time requirements, and the amount of computation and communication varies geometrically according to user input, resulting in high burst traffic and frequent changes and adjustments. Traditional network operation strategies cannot effectively cope with this, causing network congestion.
[0075] An intelligent computing center is a data center that focuses on data intelligent analysis, computation, and processing. The intelligent computing center has stronger computing capacity to meet complex requirements such as high-performance computing and AI computing. Currently, the business of the intelligent computing center generally includes two parts: one is AI business services provided to users, which complete real-time computation according to user requests and have exponentially different computation and communication amounts according to user requirements; the other is AI business learning and training, which requires even larger computation and communication amounts.
[0076] Figure 1 The logical architecture of the intelligent computing center system is shown. The following describes the logical architecture of the communication system in this specification, taking the intelligent computing center system shown in Figure 1
[0077] As shown in Figure 1 , the bottom layer of the communication system is the hardware layer, which is also the data center. The hardware layer includes hardware that provides computing and storage capabilities, such as servers that provide traditional computing, servers that provide AI computing, and storage platforms. The hardware layer also includes machine rooms that provide power and cooling for the hardware. All of these are standardized hardware infrastructure.
[0078] The upper layer of the hardware layer is the network layer, i.e., a data center network. The network layer refers to a DCI network facility providing communication of the intelligent computing center. For example, the network layer adopts a network system based on a cloud cluster switch. The complex DCI network is deployed and managed through the cloud cluster switch. The entire network is like one or several devices to the outside, greatly simplifying the network complexity.
[0079] The upper layer of the network layer is a management and operation platform. The management and operation platform is used to provide operation and management of each part of the intelligent computing center. The management and operation platform of the traditional intelligent computing center includes a computing management platform, a network management platform, and a storage management platform. The computing management platform generally creates multiple virtual machine instances on a physical server by using virtualization to achieve flexible allocation and management of computing resources. The network management platform refers to a platform for managing the DIC network, such as the DCI network based on the cloud cluster switch described above. The storage management platform provides a high-performance, high-reliability, and high-extensibility storage solution for the intelligent computing center.
[0080] The uppermost layer of the communication system is a service operation platform. The service operation platform provides various service applications. The service applications of the intelligent computing center can include traditional cloud service applications and can also include AI application services. The service applications run on the hardware layer devices.
[0081] In the embodiments of the present specification, in addition to the computing management platform, the network management platform, and the storage management platform described above, the management and operation platform is further deployed with an intelligent management and control platform, an intelligent deployment platform, and an intelligent analysis platform. The intelligent management and control / deployment / analysis platform is used to implement the resource management method described below in the present specification to achieve real-time traffic monitoring, intelligent analysis, and network policy adjustment of all service applications of the intelligent computing center. The congestion degree of the service applications is determined by sensing the traffic of the service applications, and the service resources are intelligently allocated to the service applications to ensure normal development of the service and maximize the use of the capabilities of the data center.
[0082] In some embodiments, the present specification provides a resource management method. The method can be applied to a communication system, which can be, for example, the intelligent computing center shown in Figure 1 The present specification does not limit the communication system. Taking the intelligent computing center shown in Figure 1 For example, the resource management method of the embodiments of the present specification can be executed by the management and operation platform to achieve real-time traffic collection, intelligent analysis, resource adjustment, and the like of the service applications of the intelligent computing center by using the method of the present specification.
[0083] Figure 2 A network architecture diagram of a communication system in some embodiments of the present specification is shown. The following will be described in combination with Figure 2 .
[0084] Referring to Figure 2 As shown, the network layer of the communication system includes one or more cloud cluster switches, such as Figure 2 Two cloud cluster switches, cloud cluster switch 1 and cloud cluster switch 2, are shown in FIG. 1. Taking any one of the cloud cluster switches as an example, the cloud cluster switch includes a plurality of network nodes, such as a network cloud controller (NCC) node, a network cloud fabric (NCF) node, and a network cloud packet-forwarder (NCP) node, and the number of each network node can be multiple. Each network node can be a separate physical device, such as a switch, a white-box switch, and the like. The plurality of network nodes are physically connected through a network management transport (MGT) network and are collectively a cloud cluster switch, which presents as one network device externally.
[0085] In the cloud cluster switch, the connections between the NCC node, the NCF node, the NCP node, and the MGT are physical connections, and a cluster management network is established between the nodes through the MGT. The cluster management network can use the LIPC protocol. The LIPC protocol is in the same position as the TCP / IP protocol stack, is encapsulated through the Socket layer, and provides a uniform external interface to the service application. In some embodiments, the functions provided by the LIPC protocol for the nodes in the cloud cluster switch include, but are not limited to, addressing functions transparent to the protocol stack, topology management functions, location-transparent unicast, location-non-transparent unicast, connection-based transport layer reliable unicast, multicast group management functions, location-transparent multicast, connection-based transport layer reliable multicast, non-connection-based unicast datagram transmission, and the like.
[0086] In combination with Figure 2 As shown, in the embodiments of the present specification, the traffic collection in the cloud cluster switch is implemented based on the chip capability, for example Figure 2As shown, the traffic collector 400 can be arranged in each cloud cluster switch, and the traffic collector 400 is used to collect the flow information of each service flow in the cloud cluster switch, and the collected flow information of the service flow is sent to the intelligent analysis platform in the form of a flow table. The intelligent analysis platform analyzes the running condition of each service application according to the collected flow information. The running condition of the service application may, for example, include two aspects: one is whether the running bandwidth of the service application reaches the upper limit of the allocated bandwidth, and the other is whether the performance of the server device running the service application reaches the upper limit. The intelligent analysis platform can send the analysis result to the intelligent management and control platform. The intelligent management and control platform adjusts different management and control strategies based on different analysis results, so as to ensure that each service application can run smoothly. Specifically, the traffic collector 400 can be arranged in each physical device of the cloud cluster switch, such as the NCP node and the NCC node, so as to collect the corresponding flow information of each service flow.
[0087] The asset management method of the embodiments of the present specification can be applied to a communication system, which includes a data center and a data center network. The data center can include one or more computing nodes, and one computing node can refer to one physical server or multiple physical servers. For example, the data center can include one or more computing nodes, and one computing node can refer to one physical server or multiple physical servers. For example, the data center can include one or more computing nodes, and one computing node can refer to one physical server or multiple physical servers. Figure 2 As shown in the intelligent computing center system, the data center is the hardware layer, which can include multiple computing nodes, such as cloud server computing nodes, GPU (graphics processing unit) computing nodes, etc.
[0088] The data center is deployed with multiple service applications, which run on the computing nodes. One computing node can run one service application or multiple service applications, and one service application can run on multiple computing nodes. The service application may, for example, include traditional cloud service applications and AI applications, etc., which are not limited in the present specification.
[0089] The data center network is the network layer as described above, for example Figure 2 In the example, the data center network refers to the cloud cluster switch. The service applications running on the data center perform message forwarding through the data center network to realize service communication between the data centers. The service flow refers to the data flow of the message transmission of the service application. The transmission path of one service flow includes the out port and the in port of the traffic, which defines the source end and the destination end of the service flow. One service application can include one service flow path or multiple service flow paths.
[0090] It can be understood that in the data center network, the flow information of each service flow is collected by the foregoing traffic collector 400. In order to realize the corresponding relationship between the service flow and the service application, in the embodiments of the present specification, the application networking relationship can be constructed in advance by using the intelligent deployment platform, and the application networking relationship is used to associate the service application with the collected service flow, so as to facilitate the intelligent analysis platform to analyze the running situation of each service application. The construction process of the application networking relationship is described below.
[0091] In combination with Figure 1 As shown in the figure, first, the intelligent deployment platform can obtain all service applications deployed on the entire communication system from the service running platform, that is, the related information of each service application. The related information may, for example, include the computing node resources, storage resources, message identifier, priority and the like occupied by the service application. The message identifier may, for example, include the IP address of the service application, the Vxlan identifier, the transport layer port number and the like. The IP address is used to identify the address of the service application, the Vxlan identifier is used to identify and distinguish the identifiers of different virtual networks, and the transport layer port number is a logical address used by the transport layer protocol (such as TCP, UDP) to distinguish different service applications. In the embodiments of the present specification, the transport layer port number may, for example, include the path port number of each service flow included in the service application, such as the source port number and the destination port number of each service flow.
[0092] For example, in one example, the service application information collected by the intelligent deployment platform is shown in Table 1 as follows:
[0093] Table 1
[0094]
[0095]
[0096] For example, taking the first row in Table 1 as an example, the service application is "cloud service 1", the computing node resource occupied by the service application is "virtual machine (VM) 1", the storage resource occupied by the service application is "node 1", the corresponding message identifier includes "IP address 1, Vxlan identifier 1, transport layer port number 1", and the priority parameter is "100". In the example of Table 1, the priority represents the running priority of the service application, and the lower the priority parameter, the higher the priority.
[0097] According to the foregoing, it can be known that the related information of the service application obtained above includes the path port number of each service flow in the service application. The path port number can identify the transmission path of a service flow. Therefore, in the embodiments of the present specification, the application networking relationship used to represent the transmission path of each service application can be constructed by taking the service application as a node and taking the transmission path of the service flow as an edge.
[0098] For example Figure 3A schematic diagram of an application networking relationship is shown in some embodiments of the present specification, in Figure 3 In the example application networking relationship, each node represents a service application, and each edge represents a transmission path of a service flow. Through the application networking relationship shown, the service application to which each service flow belongs and the transmission path of the service flow can be known. Figure 3
[0099] In some embodiments, the intelligent deployment platform can perform network deployment for each service application according to the application networking relationship described above, and design a corresponding network forwarding strategy, for example, assign appropriate network bandwidth to the service application, to achieve network deployment of the service application. For the network deployment process of the service application, those skilled in the art can understand and fully implement it in combination with related technologies, and the present specification will not repeat it.
[0100] As shown in Figure 4 In some embodiments, the resource management method of the present specification includes:
[0101] S410, obtaining bandwidth information of each service flow in the data center network, and performance parameters of each computing node included in the data center.
[0102] According to the foregoing, in the embodiments of the present specification, the monitoring of the running condition of the service application mainly includes two aspects: one is whether the running bandwidth of the service application reaches the upper limit, and the other is whether the performance of the computing node running the service application reaches the bottleneck. Therefore, in some embodiments, during the running of the service application, the bandwidth of the service application and the performance of the computing node can be detected respectively.
[0103] Specifically, in combination with Figure 2 As shown, the traffic collector 400 deployed in the data center network can collect each service flow in the network, and statistically obtain the bandwidth information of the service flow, and upload the bandwidth information of the service flow to the intelligent analysis platform for analysis and processing.
[0104] In Figure 2 In the example, the traffic collector 400 can be a processing chip deployed in the NCP node and the NCC node for collecting traffic information, and in the process of forwarding network messages, the processing chip can collect the flow information of each service flow in the network.
[0105] In the embodiments of the present specification, the flow information of the service flow can include traffic identification of the service flow, and the traffic identification can include different information according to different message types. For example, taking TCP / UDP / Vxlan messages as an example, the traffic identification can include five-tuple information, that is, source IP, destination IP, source port number, destination port number and protocol number. For example, taking RoCEv2 messages as an example, the traffic identification can include source IP, destination IP and destination queue pair (Destination Queue Pair, used to identify a RoCEv2 flow).
[0106] In some embodiments, the flow information of the service flow can further include flow start time, flow end time, message entry port / exit port, TCP SN, and the like. The flow start time and the flow end time can be used to identify the duration of a service flow. For TCP messages, the time when the SYN message is received is the flow start time, and the time when the FIN message is received is the flow end time. For UDP / RoCEv2 messages, the time when the first message is received is the flow start time, and the time when the last message is received is the flow end time. For Vxlan messages, the inner-layer messages of Vxlan are collected, and the field meanings can be determined by the types of the inner-layer messages. The entry port / exit port of the message is the path port number described in the present specification, which is used to record the transmission path of the service flow, and realizes the visibility of the transmission path. The TCP SN represents the next TCP message sequence number, which is used to detect the packet loss of the TCP message. Of course, those skilled in the art can understand that the flow information of the collected service flow can further include more or less information types, and the present embodiments do not limit this.
[0107] Hereinafter, the data center network is taken as an example of the cloud cluster switch shown in Figure 2 to describe in detail the process of collecting the bandwidth information of the service flow.
[0108] In combination with Figure 2 shown, the NCP node can receive the service flow sent by the computing node, so that the traffic collector 400 in the NCP node can collect the traffic identification of each service flow. Hereinafter, the message type of the service flow is taken as a TCP message, and the traffic identification is five-tuple information, which includes source IP, destination IP, source port, destination port and message protocol.
[0109] For the NCP node, the forwarding of the service flow is performed by the forwarding chip of the NCP node. In the embodiments of the present specification, when the service flow is collected in the NCP node, a flow table needs to be established according to the memory storage space of the forwarding chip. The specific process is as follows:
[0110] When a service flow arrives at the NCP node, the current memory storage space of the NCP node needs to be judged. If the storage space is sufficient, it indicates that the current storage space can accommodate all service flow information, and thus a flow table can be established for the received service flow. If the storage space is not sufficient, it indicates that the current storage space is tight and cannot accommodate all service flow information, and thus a flow table can be established only for a preset message type message. The preset message type is, for example, a SYN / FIN message, which represents a start and end message. In the case of insufficient storage space, a flow table can be established only for the SYN / FIN message.
[0111] In the embodiments of the present specification, the process of judging whether the memory storage space of the NCP node is sufficient can be that a critical threshold is set in advance for the memory of the NCP node. If the memory usage of the NCP node reaches the critical threshold, it is determined that the storage space is not sufficient, and otherwise it is determined that the storage space is sufficient. The specific value of the critical threshold can be selected according to requirements, which is not limited in the present specification.
[0112] In the embodiments of the present specification, the flow table established by the NCP node is defined as a first flow table. After the NCP node creates the first flow table according to the foregoing process, the first flow table can be uploaded to the NCC node. In some embodiments, in the case that the storage space of the NCP node is sufficient, the first flow table can be uploaded to the NCC node after waiting for the first flow table to age or after reaching a reporting period. In the case that the storage space of the NCP node is not sufficient, the first flow table can be uploaded to the NCC node immediately after the first flow table is established based on the SYN / FIN message.
[0113] For example, in one example, the first flow table established by the NCP node can be as shown in Table Two below:
[0114] Table Two
[0115] Source IP Source Port Destination IP Destination Port Protocol 1.1.1.1 8756 3.3.3.1 80 TCP 1.1.1.1 20024 3.3.3.1 80 TCP 1.1.1.1 21124 3.3.3.1 80 TCP 2.2.2.1 9265 4.4.4.1 21 TCP 2.2.2.1 15623 4.4.4.1 21 TCP 2.2.2.1 23542 4.4.4.1 21 TCP
[0116] In the example of Table Two, each row represents a service flow. The first flow table includes the traffic identifier of the service flow, that is, the five-tuple information of the service flow, which is the source IP, the source port, the destination IP, the destination port, and the message protocol.
[0117] In combination with Figure 2As shown, after receiving the first flow table sent by the NCP node, the NCC node performs a flow table creation process for the service flows, which is similar to the NCP node. Specifically, the NCC node judges the current memory storage space. If the storage space is sufficient, it means that the current storage space can accommodate all the service flow information, and thus flow tables can be established for all the service flows in the received first flow table. If the storage space is not sufficient, it means that the current storage space is tight and cannot accommodate all the service flow information, and thus flow tables can be established only for the preset message type messages, such as SYN / FIN messages, which represent start and end messages. In the case of insufficient storage space, flow tables can be established only for SYN / FIN messages.
[0118] In the embodiments of the present specification, the process of judging whether the memory storage space of the NCC node is sufficient can be that a critical threshold is set in advance for the memory of the NCC node. If the memory usage of the NCC node reaches the critical threshold, it is determined that the storage space is not sufficient, and otherwise it is determined that the storage space is sufficient. The specific value of the critical threshold can be selected according to the needs, which is not limited in the present specification.
[0119] In the embodiments of the present specification, the flow table established by the NCC node is defined as a second flow table. After the NCC node creates the second flow table according to the foregoing process, the second flow table can be sent to the analyzer module in the NCC node. In some embodiments, in the case where the storage space of the NCC node is sufficient, the second flow table can be sent to the analyzer module after waiting for the second flow table to age or reach a reporting period. In the case where the storage space of the NCC node is not sufficient, the second flow table can be sent to the analyzer module in the NCC node immediately after the second flow table is established based on the SYN / FIN messages.
[0120] The analyzer module is a module for traffic statistics in the NCC node. By analyzing the traffic of each service flow in the second flow table, the bandwidth information corresponding to each service flow can be obtained. In some embodiments of the present specification, the bandwidth information of a service flow can be represented by the amount of message data and the number of data packets transmitted per unit time, and the unit time can be, for example, 1 second. For example, in one example, the statistical results obtained by the analyzer module of the NCC node for the traffic statistics of the service flows of the second flow table can be shown in Table Three as follows:
[0121] Table Three
[0122] Source IP Source Port Destination IP Destination Port Protocol Bandwidth Information 1.1.1.1 8756 3.3.3.1 80 TCP 88KB / 150Pkts 1.1.1.1 20024 3.3.3.1 80 TCP 188KB / 250Pkts 1.1.1.1 21124 3.3.3.1 80 TCP 488KB / 360Pkts 2.2.2.1 9265 4.4.4.1 21 TCP 1MB / 1100Pkts 2.2.2.1 15623 4.4.4.1 21 TCP 3MB / 3200Pkts 2.2.2.1 23542 4.4.4.1 21 TCP 2MB / 2018Pkts
[0123] In the example of Table III, each row represents a service flow, and the bandwidth information represents the bandwidth information of the service flow. For example, the bandwidth information of the first service flow is "88KB / 150Pkts", which means that the data amount of the service flow transmitted per unit time is 88KB, and the number of data packets transmitted per unit time is 150.
[0124] Through the above process, the flow information of each service flow in the data center network can be collected, including the traffic identifier and the bandwidth information. Next, the NCC node can upload the flow information of each service flow in the second flow table to the intelligent analysis platform, and the intelligent analysis platform can analyze and process the service flow data.
[0125] At the same time, the computing management platform can obtain the performance parameters of each computing node in the data center, and then the computing management platform can send the performance parameters of each computing node to the intelligent analysis platform. In combination with the foregoing Figure 1 It is shown that the computing management platform refers to a functional platform for managing computing resources of the data center. During the operation of the data center, the computing management platform can monitor the performance parameters of each computing node in real time, so that the computing management platform can obtain the performance parameters of each computing node and send the performance parameters of each computing node to the intelligent analysis platform.
[0126] It can be understood that the performance of the computing node refers to the ability of the computing node for data processing and calculation, and the performance parameter can reflect the busy degree of the current service operation of the computing node. For example, in an example scenario, a plurality of service applications are running in parallel on a computing node, and the plurality of service applications are simultaneously in a service high-traffic stage, and the computing resources consumed by the computing node increase, and the corresponding performance parameter also increases.
[0127] The performance parameters of the computing node can be different according to the type of the computing node. For example, taking a CPU computing node as an example, the CPU computing node mainly consumes CPU computing power, so for the CPU computing node, the corresponding performance parameter can be the CPU usage rate, that is, the higher the CPU usage rate, the closer the current performance of the CPU computing node to the performance bottleneck or upper limit. For example, taking a GPU computing node as an example, the CPU computing node can effectively cope with AI applications, AI model training and other services that require a large amount of concurrent calculation. For the CPU computing node, the two most important performance indicators are floating point operation performance and GPU memory bandwidth, so the corresponding performance parameters can include floating point operation performance and GPU memory bandwidth.
[0128] Of course, those skilled in the art can understand that the computing node can also include other types, and the corresponding performance parameters can also include any parameter that can represent the computing node capability, which will not be described herein.
[0129] S420, determine the service application to which each service flow belongs based on the pre-constructed application networking relationship, and determine the bandwidth parameter of the service application according to the bandwidth information of the service flows belonging to the same service application.
[0130] According to the foregoing, the intelligent analysis platform can receive the flow information of each service flow sent by the data center network on the one hand, and can receive the performance parameters of each computing node sent by the computing management platform on the other hand. The intelligent analysis platform can perform data analysis, perceive the running conditions of the service application and the computing node, and determine whether the service application is congested and whether the computing node is busy.
[0131] In combination with the foregoing Table 3 example, the intelligent analysis platform is sent the flow information of each service flow on the data center network. Different service flows can belong to the same service application. The intelligent analysis platform needs to process the service flows, determine the service flows belonging to the same service application, and determine the bandwidth parameter of the service application according to the service flows belonging to the same service application.
[0132] In the embodiments of the present specification, the intelligent analysis platform can utilize the foregoing pre-constructed application networking relationship. The application networking relationship includes all service applications and the transmission layer port numbers corresponding to the service applications. Based on the matching between the transmission layer port numbers and the traffic identifiers of the service flows, the service application to which each service flow belongs can be identified.
[0133] Specifically, the path port numbers of each service flow can be matched with the transmission layer port numbers of each service application. The path port numbers of the service flow can include, for example, the source port and the destination port shown in Table 3. By matching the path port numbers of each service flow with the transmission layer port numbers of each service application, the service application to which each service flow belongs can be determined. Then, for each service application, the bandwidth parameter corresponding to the service application can be obtained by summing the bandwidth information of all service flows included in the service application. For example, in one example, referring to Table 3, it is assumed that the first three service flows in Table 3 belong to the same service application A. The bandwidth information of the first three service flows can be summed to obtain the bandwidth parameter corresponding to the service application A.
[0134] Through the foregoing process, the intelligent analysis platform can process the service flows in combination with the application networking relationship to obtain the bandwidth parameter of each service application.
[0135] S430, for each service application, adjust the management and control strategy for operating the service application according to the bandwidth parameter of the service application and the performance parameter of the computing node operating the service application.
[0136] Taking a business application A as an example, it is assumed that the business application A runs on a computing node a. According to the foregoing method process of the present specification, the intelligent analysis platform can determine the bandwidth parameter of the business application A and the performance parameter of the computing node a running the business application A.
[0137] Then, the intelligent analysis platform can send the bandwidth parameter and the performance parameter corresponding to the business application A to the intelligent management and control platform, and the intelligent management and control platform adjusts the management and control strategy for the business application A according to the bandwidth of the business application and the busy degree of the computing node.
[0138] It can be understood that when the intelligent deployment platform deploys the business application in advance, it can allocate bandwidth to the business application A according to the business situation of the business application A, and can set a set bandwidth in advance according to the allocated bandwidth. The set bandwidth represents a critical value of the upper limit of the bandwidth of the business application A. The set bandwidth can be equal to the allocated bandwidth of the business application A, or can be less than the allocated bandwidth.
[0139] In some embodiments, the bandwidth parameter of the business application A can be compared with the set bandwidth. If the bandwidth parameter is greater than or equal to the set bandwidth, it indicates that the current network transmission of the business application A reaches the maximum processing capacity, and congestion is likely to occur. At this time, the intelligent management and control platform needs to allocate network bandwidth to the business application A to increase the allocated bandwidth of the business application A. For example, in one example, the intelligent management and control platform can use cloud cluster to expand the bandwidth, for example, a new network device can be added to the cloud cluster composed of the original network devices, and the port of the new member device can be added to the existing aggregation group to increase the network bandwidth.
[0140] In the embodiments of the present specification, a corresponding performance threshold value can be set for each computing node in advance. The performance threshold value represents a critical value of the current operation performance of the computing node reaching a busy state. If the performance parameter of the computing node is greater than or equal to the performance threshold value, it indicates that the current computing node is in a busy state.
[0141] For example, in the foregoing example, for the business application A, the performance parameter of the computing node a running the business application A can be compared with the performance threshold value. If the performance parameter of the computing node a is greater than or equal to the performance threshold value, it indicates that the current computing node a is in a busy state. At this time, the intelligent management and control platform needs to reduce the running load of the computing node a so that the computing node a runs in a normal state. For example, in one example, the intelligent management and control platform can adjust the number of business applications running on the computing node a, and reduce the number of business applications running on the computing node a through load balancing, so as to reduce the running load of the computing node a, so that the computing node a can run in a normal state.
[0142] In the above example, if the bandwidth parameter of the service application A is greater than or equal to the set bandwidth, and the performance parameter of the computing node a is greater than or equal to the performance threshold, the network bandwidth can be increased and the computing node load can be reduced simultaneously according to the above management and control strategy. Those skilled in the art can understand this, and the present specification will not be repeated.
[0143] In some embodiments, the flow information of the service flow collected in the data center network includes packet loss information. For example, in the foregoing example, when the traffic collector 400 collects the service flow, the TCP SN information of the service flow can be collected, which indicates the next TCP packet sequence number, so that the TCP SN information can be used as the packet loss information of the service flow.
[0144] The data center network uploads the collected flow information of the service flow to the intelligent analysis platform. The intelligent analysis platform determines the bandwidth parameter of the service application according to the bandwidth information of the service flow belonging to the same service application. At the same time, the intelligent analysis platform can also determine the packet loss rate of the service application according to the packet loss information of the service flow belonging to the same service application. For example, the packet loss rate corresponding to the service application can be determined according to the TCP SN information of each service flow included in the service application.
[0145] In the examples of the present specification, the intelligent analysis platform can send the bandwidth parameter, the packet loss information corresponding to the service application, and the performance parameter of the computing node running the service application to the intelligent management and control platform. The intelligent management and control platform adjusts the management and control strategy of the service application according to the bandwidth parameter and the performance parameter of the computing node, which is the same as the foregoing, and will not be repeated.
[0146] The intelligent management and control platform can pre-set a preset threshold corresponding to the service application, which represents the critical value of the packet loss rate of the service application. If the packet loss rate of a certain service application is greater than or equal to the preset threshold, it indicates that the service application is unstable, and the packet forwarding strategy corresponding to the service application has a problem, so that the packet forwarding strategy of the service application can be adjusted. The way to adjust the packet forwarding strategy includes but is not limited to bandwidth guarantee, priority queue, QoS (Quality of Service), network speed limit, etc., which is not limited in the present specification. After adjusting the packet forwarding strategy of the service application, the intelligent management platform can issue the packet forwarding strategy to each network device to guarantee sufficient network service for the service application.
[0147] In some embodiments, the flow information of the service flow can further include delay, jitter and other information. The intelligent analysis platform determines the running state of the service application according to the delay, jitter and other information, and adjusts the corresponding management and control strategy. On the basis of the foregoing embodiments, those skilled in the art can understand and fully implement this, and the present specification will not be repeated.
[0148] In some embodiments, after the application networking relationship is pre-built by the intelligent deployment platform, the application networking relationship can be displayed on the management and operation platform, for example, the display interface can be as shown in Figure 3 In this embodiment, the communication system can use the aforementioned resource management method to monitor whether congestion occurs in each transmission path in real time. For example, when the bandwidth parameter of a certain service application reaches the set bandwidth, it indicates that the transmission path of the service application is congested. At this time, the transmission path of the service flow included in the service application in the application networking relationship shown in Figure 3 can be identified to prompt the management personnel that network congestion occurs.
[0149] For example, in one example, the transmission path under normal circumstances can be as shown in Figure 3 , that is, the transmission path is blue. When congestion occurs in a certain transmission path, the transmission path can be displayed in red on the application networking relationship, or further combined with a flashing state, etc. The management personnel can observe which service application has a congested transmission path through the application networking relationship.
[0150] In some embodiments, the intelligent management and control platform can be used to count the congestion of the service application and the busy condition of the computing node. If the service application frequently appears congestion and / or the computing node frequently appears busy in the communication system, it indicates that the processing capacity of the entire communication system has reached the upper limit. At this time, it is necessary to consider expanding the entire communication system or adjusting the operation strategy, for example, giving up part of the low-priority service application to ensure the service quality of the high-priority service application.
[0151] From the above, in the embodiments of the present specification, through the intelligent operation and management strategy, the intelligent management and control platform, the intelligent analysis platform, and the intelligent deployment platform are cooperated to realize real-time monitoring of the service application and the computing node in the communication system, perceive the network bandwidth congestion condition of the service application and the busy degree of the computing node, so that the operation strategy of the service application can be adjusted accordingly, the AI service data center such as the intelligent computing center which has high burst traffic and needs to be frequently changed and adjusted can be effectively coped with, the network operation efficiency and quality are improved, and the resource utilization rate and processing capacity of the data center can be maximized.
[0152] In some embodiments, the present specification provides a resource management device applied to a communication system, the communication system including a data center and a data center network, the data center deploying a plurality of service applications, the service applications running on computing nodes of the data center, as shown in Figure 5 The device includes:
[0153] The information acquisition module 10 is configured to acquire bandwidth information of each service flow in the data center network and performance parameters of each computing node included in the data center.
[0154] The bandwidth determination module 20 is configured to determine a service application to which each service flow belongs based on a pre-constructed application networking relationship, and determine bandwidth parameters of the service application according to bandwidth information of service flows belonging to the same service application, wherein the application networking relationship includes a service flow transmission path of each service application.
[0155] The policy control module 30 is configured to, for each service application, adjust a control policy for running the service application according to the bandwidth parameters of the service application and the performance parameters of the computing node running the service application, wherein the control policy includes allocating network bandwidth for the service application in a case where the bandwidth parameters reach a set bandwidth, and adjusting a running load of the computing node in a case where the performance parameters reach a performance threshold.
[0156] In some embodiments, the data center network includes a cloud cluster switch including a plurality of network nodes, and the information acquisition module 10 is specifically configured to acquire flow information of each service flow in the cloud cluster switch through each network node, wherein the flow information includes traffic identification and the bandwidth information of the service flow.
[0157] In some embodiments, the plurality of network nodes of the cloud cluster switch include NCP nodes and NCC nodes, and the information acquisition module 10 is specifically configured to,
[0158] The NCP node acquires traffic identification of the service flow sent by the computing node, and establishes a first flow table based on the traffic identification of the service flow, and uploads the first flow table to the NCC node.
[0159] The NCC node counts traffic of each service flow according to the received first flow table to obtain the bandwidth information of each service flow.
[0160] In some embodiments, the information acquisition module 10 is specifically configured to,
[0161] In a case where the storage space of the NCP node is sufficient, the first flow table is established according to the received traffic identification of the service flow, and the first flow table is uploaded to the NCC node, and in a case where the storage space is insufficient, the first flow table is established according to the received traffic identification of the service flow of a preset message type, and the first flow table is uploaded to the NCC node.
[0162] In a case where the NCC node has sufficient storage space, a second flow table is established according to the traffic identifier of the service flow in the first flow table, and in a case where the NCC node does not have sufficient storage space, a second flow table is established according to the traffic identifier of the service flow of a preset message type in the first flow table, and the traffic of each service flow in the second flow table is counted to obtain the bandwidth information of each service flow.
[0163] In some embodiments, the bandwidth determination module 20 is specifically configured to,
[0164] The path port number corresponding to each service flow is obtained, and the path port number is matched with the transport layer port number corresponding to each service application in the application networking relationship to determine the service application to which each service flow belongs.
[0165] The bandwidth information of the service flows belonging to the same service application is summed up to obtain the bandwidth parameter corresponding to each service application.
[0166] In some embodiments, the data center includes a first computing node, and in a case where the first computing node is a GPU computing node, the information acquisition module is specifically configured to acquire the floating point operation performance and GPU memory bandwidth of the GPU computing node, and the performance parameter of the GPU computing node includes the floating point operation performance and the GPU memory bandwidth.
[0167] In some embodiments, the policy control module 30 is specifically configured to,
[0168] According to the packet loss information of the service flows belonging to the same service application, the packet loss rate of the service application is determined.
[0169] According to the bandwidth parameter, the packet loss rate, and the performance parameter of the computing node running the service application, the control strategy for running the service application is adjusted, wherein the control strategy includes adjusting the message forwarding strategy of the service application in a case where the packet loss rate reaches a preset threshold.
[0170] In some embodiments, the apparatus further includes a relationship construction module, and the relationship construction module is configured to,
[0171] All service applications deployed on the data center and the transport layer port number of each service application are acquired, wherein the transport layer port number includes the path port number of each service flow included in the service application.
[0172] The service application is taken as a node, the transmission path of the service flow is taken as an edge, and the application networking relationship is generated according to the path port number of the service flow of each service application.
[0173] Obviously, the above embodiments are merely example for clearly illustrating but not limitation to the embodiments. Based on the above description, other different forms of changes or variations can be made by those skilled in the art. Here, all the embodiments need not and can not be enumerated. The obvious changes or variations derived from the above description are still within the protection scope of the disclosure.
Claims
1. A communication system, characterized by The communication system comprises a data center, a data center network and a management and operation platform, the data center is deployed with a plurality of service applications, each service application runs on a computing node of the data center, and the management and operation platform comprises an intelligent analysis platform and an intelligent management and control platform; The intelligent analysis platform is configured to acquire bandwidth information of each service flow in the data center network and performance parameters of each computing node included in the data center, determine a service application to which each service flow belongs based on a pre-constructed application networking relationship, and determine bandwidth parameters of the service application according to bandwidth information of service flows belonging to the same service application, wherein the application networking relationship comprises a service flow transmission path of each service application; The intelligent management and control platform is configured to, for each service application, adjust a management and control strategy for running the service application according to the bandwidth parameters of the service application and performance parameters of the computing node running the service application, wherein the management and control strategy comprises allocating network bandwidth for the service application in a case where the bandwidth parameters reach a set bandwidth and adjusting a running load of the computing node in a case where the performance parameters reach a performance threshold.
2. The communication system of claim 1, wherein The data center network comprises a cloud cluster switch, the cloud cluster switch comprises a plurality of network nodes, flow information of each service flow in the cloud cluster switch is collected through each network node, the flow information comprises a traffic identifier and the bandwidth information of the service flow, and the intelligent analysis platform receives the flow information of each service flow sent by the cloud cluster switch.
3. The communication system of claim 2, wherein, The plurality of network nodes of the cloud cluster switch comprise an NCP node and an NCC node, The NCP node is configured to collect the traffic identifier of the service flow sent by the computing node, establish a first flow table based on the traffic identifier of the service flow, and send the first flow table to the NCC node; The NCC node is configured to count the traffic of each service flow according to the received first flow table to obtain the bandwidth information of each service flow.
4. The communication system of claim 3, wherein The NCP node is specifically configured to, in a case where storage space is sufficient, establish a first flow table according to the received traffic identifier of the service flow and send the first flow table to the NCC node, and in a case where storage space is insufficient, establish a first flow table according to the traffic identifier of the service flow of a preset message type and send the first flow table to the NCC node; The NCC node is specifically configured to, in a case where storage space is sufficient, establish a second flow table according to the traffic identifier of the service flow in the first flow table, in a case where storage space is insufficient, establish a second flow table according to the traffic identifier of the service flow of a preset message type in the first flow table, count the traffic of each service flow in the second flow table, and obtain the bandwidth information of each service flow.
5. The communication system of claim 1, wherein, The intelligent analysis platform is specifically configured to, obtain the path port number corresponding to each service flow, and match the path port number with the transport layer port number corresponding to each service application in the application networking relationship to determine the service application to which each service flow belongs; sum the bandwidth information of the service flows belonging to the same service application to obtain the bandwidth parameter corresponding to each service application.
6. The communication system of claim 1, wherein, The management and operation platform further comprises a computing management platform, configured to collect the performance parameter of each computing node and send the performance parameter of each computing node to the intelligent analysis platform. In the case where the computing node is a GPU computing node, the performance parameter of the GPU computing node comprises floating point operation performance and GPU memory bandwidth.
7. The communication system according to any one of claims 1 to 6, wherein The intelligent analysis platform is further configured to obtain the packet loss information of each service flow in the data center network, and determine the packet loss rate of the service application according to the packet loss information of the service flows belonging to the same service application. The intelligent management and control platform is further configured to adjust the management and control strategy for running the service application according to the bandwidth parameter, the packet loss rate of the service application and the performance parameter of the computing node running the service application, wherein the management and control strategy comprises adjusting the message forwarding strategy of the service application in the case where the packet loss rate reaches a preset threshold.
8. The communication system according to any one of claims 1 to 6, wherein The management and operation platform further comprises an intelligent deployment platform, configured to obtain all service applications deployed on the data center and the transport layer port number of each service application, wherein the transport layer port number comprises the path port number of each service flow included in the service application, and the application networking relationship is generated according to the path port number of the service flow of each service application, taking the service application as a node and the transmission path of the service flow as an edge.
9. A resource management method characterized by, The method is applied to a communication system, the communication system comprising a data center and a data center network, the data center deploying a plurality of service applications, the service applications running on computing nodes of the data center, and the method comprising: obtain the bandwidth information of each service flow in the data center network and the performance parameter of each computing node included in the data center; determine the service application to which each service flow belongs based on a pre-constructed application networking relationship, and determine the bandwidth parameter of the service application according to the bandwidth information of the service flows belonging to the same service application, wherein the application networking relationship comprises the service flow transmission path of each service application; for each service application, adjust the management and control strategy for running the service application according to the bandwidth parameter of the service application and the performance parameter of the computing node running the service application, wherein the management and control strategy comprises allocating network bandwidth for the service application in the case where the bandwidth parameter reaches a set bandwidth, and adjusting the running load of the computing node in the case where the performance parameter reaches a performance threshold.
10. The method of claim 9, wherein, The process of pre-constructing the application networking relationship comprises: obtaining all service applications deployed on the data center and a transport layer port number of each service application, wherein the transport layer port number comprises a path port number of each service flow included in the service application; generating the application networking relationship according to the path port number of the service flow of each service application, taking the service application as a node and taking the transmission path of the service flow as an edge.
11. The method according to claim 9 or 10, characterized in that, The bandwidth determination module is specifically used for: obtaining a path port number corresponding to each service flow, and matching the path port number with the transport layer port number corresponding to each service application in the application networking relationship to determine the service application to which each service flow belongs; summing bandwidth information of the service flows belonging to the same service application to obtain the bandwidth parameter corresponding to each service application.
12. A resource management apparatus characterized by comprising: The application is applied to a communication system, and the communication system comprises a data center and a data center network. The data center is deployed with a plurality of service applications. The service applications run on computing nodes of the data center. The device comprises: an information obtaining module, configured to obtain bandwidth information of each service flow in the data center network and performance parameters of each computing node included in the data center; a bandwidth determination module, configured to determine the service application to which each service flow belongs based on a pre-constructed application networking relationship and determine a bandwidth parameter of the service application according to bandwidth information of the service flows belonging to the same service application, wherein the application networking relationship comprises a service flow transmission path of each service application; a policy control module, configured to adjust a control policy for running each service application according to the bandwidth parameter of the service application and the performance parameter of the computing node running the service application, wherein the control policy comprises allocating network bandwidth for the service application in the case that the bandwidth parameter reaches a set bandwidth and adjusting the running load of the computing node in the case that the performance parameter reaches a performance threshold.
13. The apparatus of claim 12, wherein, The relationship construction module is further configured to: obtain all service applications deployed on the data center and a transport layer port number of each service application, wherein the transport layer port number comprises a path port number of each service flow included in the service application; generate the application networking relationship according to the path port number of the service flow of each service application, taking the service application as a node and taking the transmission path of the service flow as an edge.
14. The apparatus of claim 12 or 13, wherein, The bandwidth determination module is specifically used for: obtaining a path port number corresponding to each service flow, and matching the path port number with the transport layer port number corresponding to each service application in the application networking relationship to determine the service application to which each service flow belongs; summing bandwidth information of the service flows belonging to the same service application to obtain the bandwidth parameter corresponding to each service application.
Citation Information
Patent Citations
Router for sensing bearing state and service flow bandwidth distribution method thereof
CN102739507A
Network traffic control method and network equipment thereof
CN107196877A