Server cluster processing method and system, and electronic device

By acquiring node traffic data from the server cluster, determining the training period, and performing clustering, the problem of inaccurate node differentiation was solved, achieving accurate node differentiation and reducing the difficulty of monitoring.

WO2026026884A1PCT designated stage Publication Date: 2026-02-05CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/111637
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-02
Filing Date
2025-07-30
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

In a server cluster, the accuracy of distinguishing nodes performing different training tasks is low, making it difficult to monitor the training status.

Method used

By acquiring traffic data from multiple nodes in the server cluster, the training cycle of the nodes is determined based on the traffic data, and clustering is performed to obtain a node set, ensuring that nodes in the same set perform the same neural network model training task.

Benefits of technology

It enables accurate differentiation of nodes in the server cluster, reduces the difficulty of monitoring training tasks, and improves the accuracy of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025111637_05022026_PF_FP_ABST
    Figure CN2025111637_05022026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present disclosure are a server cluster processing method and system, and an electronic device. The method relates to the field of data processing in the field of artificial intelligence, and comprises: acquiring traffic data of a plurality of nodes in a server cluster, wherein the server cluster is used for executing training tasks for a plurality of neural network models, and the traffic data is used for representing data transmitted among the plurality of nodes during network communication while the server cluster executes the training tasks; on the basis of the traffic data of any node, determining a training period corresponding to any node, wherein the training period is used for representing a period during which any node executes the corresponding training task; and on the basis of the training periods corresponding to the plurality of nodes, clustering the plurality of nodes to obtain a plurality of node sets, wherein nodes in the same node set are used for executing training tasks for the same neural network model. The present disclosure solves the technical problem in the related art whereby the accuracy of differentiating nodes in a server cluster that execute different training tasks is low.
Need to check novelty before this filing date? Find Prior Art

Description

Processing method, system and electronic device of server cluster TECHNICAL FIELD

[0001] The present disclosure relates to the field of data processing in the field of artificial intelligence, in particular to a processing method, system and electronic device of server cluster. BACKGROUND

[0002] In the scenario of neural network model training, a server cluster deployed in advance is usually used to perform distributed training of a neural network model based on a training task. That is, there may be a case that multiple neural network models use the same server cluster for training at the same time, and in order to ensure that the same server cluster can successfully and efficiently complete the training of multiple neural network models, the training situation of the server cluster needs to be monitored.

[0003] At present, the monitoring of the training situation of the server cluster usually obtains training data of each neural network model, and analyzes the training data and each node (i.e. a single server) in the server cluster to achieve the purpose of monitoring the training situation. However, since multiple training tasks may exist in the server cluster at the same time, and the multiple training tasks are executed by multiple nodes, the relationship between the multiple nodes may be complex, which leads to low accuracy in distinguishing nodes executing different training tasks, and further leads to great difficulty in monitoring the execution of training tasks by the server cluster.

[0004] At present, there is no effective solution to the above problems. SUMMARY

[0005] The embodiments of the present disclosure provide a processing method, system and electronic device of server cluster to at least solve the technical problem of low accuracy in distinguishing nodes executing different training tasks in the server cluster in the related art.

[0006] According to an aspect of an embodiment of the present disclosure, a processing method of a server cluster is provided, including: obtaining traffic data of multiple nodes in the server cluster, wherein the server cluster is configured to execute training tasks of multiple neural network models, and the traffic data is used to represent data transmitted by network communication between the multiple nodes in the process of the server cluster executing the training tasks; determining a training period corresponding to any one node based on the traffic data of the any one node, wherein the training period is used to represent a period of the any one node executing a corresponding training task; clustering the multiple nodes based on the training periods corresponding to the multiple nodes to obtain multiple node sets, wherein the nodes in the same node set are configured to execute training tasks of the same neural network model.

[0007] According to a further aspect of the embodiments of the present disclosure, a processing method of a server cluster is also provided. The method comprises: obtaining traffic data of a plurality of nodes in the server cluster by calling a first interface, wherein the first interface comprises a first parameter, a parameter value of the first parameter comprises the traffic data of the plurality of nodes, the server cluster is configured to perform training tasks of a plurality of neural network models, and the traffic data is used to represent data transmitted in network communication between the plurality of nodes during execution of the training tasks by the server cluster; determining a training period corresponding to any one node based on the traffic data of the any one node, wherein the training period is used to represent a period for the any one node to perform a corresponding training task; clustering the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in a same node set are configured to perform training tasks of a same neural network model; and outputting the plurality of node sets by calling a second interface, wherein the second interface comprises a second parameter, and a parameter value of the second parameter comprises the plurality of node sets.

[0008] According to a further aspect of the embodiments of the present disclosure, a processing system of a server cluster is also provided. The system comprises: a server cluster comprising a plurality of nodes, the server cluster being configured to perform training tasks of a plurality of neural network models; and a monitoring device connected to the plurality of nodes, configured to determine a training period corresponding to any one node based on traffic data of the any one node, and cluster the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein the traffic data is used to represent data transmitted in network communication between the plurality of nodes during execution of the training tasks by the server cluster, the training period is used to represent a period for the any one node to perform a corresponding training task, and nodes in a same node set are configured to perform training tasks of a same neural network model.

[0009] According to a further aspect of the embodiments of the present disclosure, a processing apparatus of a server cluster is also provided. The apparatus comprises: an obtaining module configured to obtain traffic data of a plurality of nodes in the server cluster, wherein the server cluster is configured to perform training tasks of a plurality of neural network models, and the traffic data is used to represent data transmitted in network communication between the plurality of nodes during execution of the training tasks by the server cluster; a determining module configured to determine a training period corresponding to any one node based on the traffic data of the any one node, wherein the training period is used to represent a period for the any one node to perform a corresponding training task; and a clustering module configured to cluster the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in a same node set are configured to perform training tasks of a same neural network model.

[0010] According to another aspect of the embodiments of the present disclosure, a processing apparatus of a server cluster is also provided, comprising: an obtaining module configured to obtain traffic data of a plurality of nodes in the server cluster by calling a first interface, wherein the first interface comprises a first parameter, a parameter value of the first parameter comprises the traffic data of the plurality of nodes, the server cluster is configured to perform training tasks of a plurality of neural network models, and the traffic data is used to represent data transmitted in network communication between the plurality of nodes during execution of the training tasks by the server cluster; a determining module configured to determine a training period corresponding to any one node based on the traffic data of the any one node, wherein the training period is used to represent a period for the any one node to perform a corresponding training task; a clustering module configured to cluster the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in a same node set are used to perform training tasks of a same neural network model; and an output module configured to output the plurality of node sets by calling a second interface, wherein the second interface comprises a second parameter, and a parameter value of the second parameter comprises the plurality of node sets.

[0011] According to another aspect of the embodiments of the present disclosure, a computer terminal is also provided, comprising: a memory storing an executable program; and a processor configured to run the program, wherein the program performs the method in the various embodiments of the present disclosure when running.

[0012] According to another aspect of the embodiments of the present disclosure, a computer readable storage medium is also provided, comprising a stored executable program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to perform the method in the various embodiments of the present disclosure when the executable program runs.

[0013] According to another aspect of the embodiments of the present disclosure, a computer program product is also provided, comprising a computer program, which, when executed by a processor, implements the method in the various embodiments of the present disclosure.

[0014] According to another aspect of the embodiments of the present disclosure, a computer program product is also provided, comprising a non-volatile computer readable storage medium storing a computer program, which, when executed by a processor, implements the method in the various embodiments of the present disclosure.

[0015] According to another aspect of the embodiments of the present disclosure, a computer program is also provided, which, when executed by a processor, implements the method in the various embodiments of the present disclosure.

[0016] In the embodiments of the present disclosure, the traffic data of multiple nodes in the server cluster is acquired; based on the traffic data of any one node, the training period corresponding to any one node is determined; and the multiple nodes are clustered based on the training periods corresponding to the multiple nodes, to obtain multiple node sets. It is easy to note that the traffic data of the nodes is general traffic data generated by network communication between the nodes, that is, the traffic data of the nodes is relatively easy to acquire, and in addition, the training data of the training task does not need to be acquired by analyzing the easily acquired traffic data. In addition, by clustering the nodes corresponding to the traffic data, different nodes processing the same training task of the neural network model can be divided into the same node set, and different node sets correspond to different training tasks, so that the nodes executing different training tasks can be clearly distinguished, achieving the technical effect of accurately distinguishing the nodes executing different training tasks, reducing the difficulty of monitoring the execution of the training task of the server cluster, and solving the technical problem of low accuracy of distinguishing the nodes executing different training tasks in the server cluster in the related art.

[0017] It is easy to note that the general description above and the detailed description below are only for illustrating and explaining the present disclosure, and do not constitute a limitation on the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0018] The drawings described herein are used to provide further understanding of the present disclosure, and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure, and do not constitute an improper limitation on the present disclosure. In the drawings:

[0019] Fig. 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a processing method of a server cluster according to an embodiment of the present disclosure;

[0020] Fig. 2 is a structure block diagram of a computing environment according to an embodiment of the present disclosure;

[0021] Fig. 3 is a structure block diagram of a service mesh according to an embodiment of the present disclosure;

[0022] Fig. 4 is a flowchart of a processing method of a server cluster according to an embodiment 1 of the present disclosure;

[0023] Fig. 5 is a schematic diagram of clustering multiple nodes in a server cluster to obtain multiple node sets according to an embodiment 1 of the present disclosure;

[0024] Fig. 6 is a flowchart of a processing method of a server cluster according to an embodiment 2 of the present disclosure;

[0025] Fig. 7 is a schematic diagram of a processing system of a server cluster according to an embodiment 3 of the present disclosure;

[0026] FIG. 8 is a schematic diagram of a processing device of a server cluster according to an embodiment 4 of the present disclosure;

[0027] FIG. 9 is a schematic diagram of a processing device of a server cluster according to an embodiment 5 of the present disclosure;

[0028] FIG. 10 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] In order to enable persons skilled in the art to better understand the present disclosure scheme, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in combination with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by persons skilled in the art without creative labor should be within the scope of protection of the present disclosure.

[0030] It should be noted that the terms "first", "second" and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] First, some nouns or terms that appear in the description of the embodiments of the present disclosure are applicable to the following explanations:

[0032] Multiple neural network models of large models: refer to models trained using a large amount of data and complex algorithms, which can be used for prediction, classification, clustering and other tasks.

[0033] Distributed training: refers to the process of decomposing training tasks into multiple sub-tasks and training them on different computers, and then merging the results.

[0034] Monitoring view: refers to a view used to monitor the running state of the system, which can display various indicators and alarm information.

[0035] Mixing curve: refers to a curve used to represent the relationship between multiple data sets, which can be used to discover similarities and differences between data sets.

[0036] Clustering: refers to the process of grouping objects in a dataset, which can be used to discover the inherent structure and patterns in the dataset.

[0037] Traffic Periodicity: refers to the periodic variation of network traffic, which can be used to predict the trend of network traffic and improve the allocation of network resources.

[0038] Fourier Transform: refers to the process of converting a signal into a frequency domain representation, which can be used to analyze the frequency components of a signal and extract signal features.

[0039] Fundamental Frequency: refers to the lowest periodic component in a periodic signal, which can be used to extract the characteristics of the signal and identify the type of signal.

[0040] Time Complexity: refers to the relationship between the time an algorithm takes to execute and the size of the input, which can be used to evaluate the efficiency of the algorithm and improve the performance of the algorithm.

[0041] DBSCAN (Density-Based Spatial Clustering of Applications with Noise): a density-based clustering algorithm that can be used to discover dense regions and cluster objects in a dataset.

[0042] Example 1

[0043] According to an embodiment of the present disclosure, a processing method of a server cluster is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.

[0044] The method embodiments provided by the embodiments of the present disclosure can be executed in a mobile terminal, a computer terminal or similar computing device. FIG. 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a processing method of a server cluster according to an embodiment of the present disclosure. As shown in FIG. 1, the computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that the structure shown in FIG. 1 is only schematic, and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or fewer components than those shown in FIG. 1, or have a different configuration from that shown in FIG. 1.

[0045] It should be noted that the one or more processors 102 and / or other data processing circuits described above can be referred to as "data processing circuits" herein. The data processing circuit can be embodied in whole or in part as software, hardware, firmware or any other combination. In addition, the data processing circuit can be a single independent processing module, or all or part of any one of the other elements combined into the computer terminal 10 (or mobile device). As referred to in the embodiments of the present disclosure, the data processing circuit serves as a processor to control (for example, selection of a variable resistance terminal path connected to an interface).

[0046] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the methods in the embodiments of the present disclosure. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the methods in the above embodiments. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely disposed with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0047] The transmission device 106 is configured to receive or transmit data via a network. Examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can connect to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet wirelessly.

[0048] The display can be a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0049] The hardware structure diagram shown in FIG. 1 can not only be used as an exemplary block diagram of the computer terminal 10 (or mobile device) described above, but also be used as an exemplary block diagram of the server. In an alternative embodiment, FIG. 2 shows an embodiment of using the computer terminal 10 (or mobile device) shown in FIG. 1 as a computing node in a computing environment 201 in a block diagram. FIG. 2 is a structural block diagram of a computing environment according to an embodiment of the present disclosure. As shown in FIG. 2, the computing environment 201 includes a plurality of computing nodes (e.g., servers) running on a distributed network (shown as 210-1, 210-2, …, in the figure). The computing nodes all include local processing and memory resources, and an end user 202 can remotely run applications or store data in the computing environment 201. The applications can be provided as a plurality of services 220-1, 220-2, 220-3, and 220-4 in the computing environment 201, representing services “A”, “D”, “E”, and “H”, respectively.

[0050] The end user 202 can provide and access the services through a web browser or other software applications on a client, and in some embodiments, the provision and / or requests of the end user 202 can be provided to an entry gateway 230. The entry gateway 230 can include a corresponding proxy to process the provision and / or requests for the services (one or more services provided in the computing environment 201).

[0051] Services are provided or deployed in accordance with various virtualization technologies supported by the computing environment 201. In some embodiments, services can be provided in accordance with virtual machine (VM)-based virtualization, container-based virtualization, and / or the like. VM-based virtualization can be emulating a real computer by initializing a virtual machine, executing programs and applications without directly touching any actual hardware resources. While the virtual machine is virtualized, in accordance with container-based virtualization, a container can be launched to virtualize an entire operating system (OS) so that multiple workloads can run on a single OS instance.

[0052] In one embodiment of container-based virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, as shown in FIG. 2, the service 220-2 can be equipped with one or more Pods 240-1, 240-2, …, 240-N (collectively, Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, …, 242-M (collectively, containers). The one or more containers in a Pod handle requests related to one or more respective functions of the service, and the proxy 245 generally controls network functions related to the service, such as routing, load balancing, and the like. Other services can also be equipped with similar Pods.

[0053] In operation, executing a user request from the end user 202 can require invoking one or more services in the computing environment 201, and executing one or more functions of a service can require invoking one or more functions of another service. As shown in FIG. 2, the service “A” 220-1 receives a user request from the end user 202 from the ingress gateway 230, the service “A” 220-1 can invoke the service “D” 220-2, and the service “D” 220-2 can request the service “E” 220-3 to execute one or more functions.

[0054] The computing environment described above can be a cloud computing environment, and the allocation of resources is managed by a cloud service provider, allowing the development of functions without considering the implementation, adjustment, or expansion of servers. The computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be split into a set of functions that can automatically scale independently, rather than expanding a single hardware device to handle potential loads.

[0055] In another alternative embodiment, FIG. 3 illustrates, in a block diagram, an embodiment of using the computer terminal 10 (or mobile device) shown in FIG. 1 above as a service mesh. FIG. 3 is a structural block diagram of a service mesh according to an embodiment of the present disclosure, as shown in FIG. 3, the service mesh 300 is mainly used to facilitate secure and reliable communication between multiple microservices, which refers to decomposing an application into multiple smaller services or instances and running them on different clusters / machines.

[0056] As shown in FIG. 3, the microservices can include an application service instance A and an application service instance B, which form a functional application layer of the service mesh 300. In an implementation, the application service instance A is running in the form of a container / process 308 on a machine / workload container group 314 (Pod), and the application service instance B is running in the form of a container / process 310 on a machine / workload container group 316 (Pod).

[0057] In an implementation, the application service instance A can be a commodity query service, and the application service instance B can be a commodity ordering service.

[0058] As shown in FIG. 3, the application service instance A and a mesh proxy 303 coexist in the machine workload container group 314, and the application service instance B and a mesh proxy 305 coexist in the machine workload container 316. The mesh proxy 303 and the mesh proxy 305 form a data plane layer of the service mesh 300. Among them, the mesh proxy 303 and the mesh proxy 305 run in the form of a container / process 304, a container / process 306, respectively, can receive a request 312, be set to perform a commodity query service, and the mesh proxy 303 and the application service instance A can communicate bidirectionally, the mesh proxy 305 and the application service instance B can communicate bidirectionally. In addition, the mesh proxy 303 and the mesh proxy 305 can also communicate bidirectionally.

[0059] In an implementation, the network traffic of the application service instance A is all routed to the appropriate destination through the mesh proxy 303, and the network traffic of the application service instance B is all routed to the appropriate destination through the mesh proxy 305. It should be noted that the network traffic mentioned herein includes but is not limited to Hyper Text Transfer Protocol (HTTP), Representational State Transfer (REST), Google Remote Procedure Call (gRPC), and open source in-memory data structure storage system (Redis) and the like.

[0060] In an embodiment, the function of extending the data plane layer can be implemented by writing a custom filter for the proxies (Envoy) in the service mesh 300. The service mesh proxy configuration can be to make the service mesh correctly proxy service traffic, implement service interworking and service governance. The mesh proxy 303 and the mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.

[0061] As shown in FIG. 3, the service mesh 300 also includes a control plane layer. The control plane layer can be a group of services running in a dedicated namespace, hosted by the hosting control plane component 301 in the machine / Pod 302. As shown in FIG. 3, the hosting control plane component 301 communicates with the mesh proxy 303 and the mesh proxy 305 in a bidirectional manner. The hosting control plane component 301 is configured to perform some control management functions. For example, the hosting control plane component 301 receives telemetry data transmitted by the mesh proxy 303 and the mesh proxy 305, and can further aggregate the telemetry data. The services, the hosting control plane component 301 can also provide a user-oriented application programming interface (API) to facilitate manipulation of network behavior, and provide configuration data to the mesh proxy 303 and the mesh proxy 305, etc.

[0062] In the above running environment, the disclosure provides a processing method of a server cluster as shown in FIG. 4. FIG. 4 is a flowchart of a processing method of a server cluster according to Embodiment 1 of the disclosure, as shown in FIG. 4, including: a server cluster 20 and a specific server 10, wherein the server cluster 20 and the specific server 10 are connected through a network. The server cluster 10 performs: reporting traffic data of a plurality of nodes, and the specific server 20 performs: obtaining traffic data of a plurality of nodes in the server cluster, determining a training period corresponding to any one node based on the traffic data of the any one node, clustering a plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets. As shown in FIG. 4, the method includes the following steps:

[0063] In step S402, traffic data of multiple nodes in a server cluster is acquired, wherein the server cluster is configured to perform training tasks of multiple neural network models, and the traffic data is used to represent data transmitted by network communication between the multiple nodes in the server cluster during the execution of the training tasks.

[0064] The server cluster described above can train multiple neural network models based on training tasks of the multiple neural network models. The server cluster can be composed of multiple local servers, multiple edge servers, or multiple cloud servers, and the specific type of the server cluster is not limited in the embodiment. The multiple neural network models described above can be neural network models with a large number of parameters that need to be trained in a distributed manner. The multiple neural network models are used for various tasks such as prediction, classification, and clustering. The application scenarios of the multiple neural network models can include, but are not limited to, natural language processing, computer vision processing, reinforcement learning, and speech recognition and synthesis. The types of tasks performed by the multiple neural network models are different in different application scenarios.

[0065] The training task described above can be set in advance by a user, and based on the training task, the multiple neural network models can be trained into multiple neural network models that can handle specific tasks required by the user. The training task can include, but is not limited to, data management, memory and computing resource improvement, mixed precision training, model fine-tuning, hyperparameter optimization, model pruning and quantization, model distillation, model robustness testing, model evaluation, and environment and dependency management. During the training of the multiple neural network models by the server cluster based on the training task, the multiple nodes generate traffic data corresponding to the training task when performing network communication. The traffic data corresponding to different training tasks is also different. In the embodiment, the traffic data can include, but is not limited to, model parameter synchronization, gradient information, weight and bias, loss function value, performance index, control command, log and monitoring data, data shard, model checkpoint, resource scheduling information, configuration and parameter update, and error and exception information.

[0066] In an optional embodiment, in the process of currently training the neural network model, there are problems of difficulty in monitoring the training situation of the server cluster and inability to accurately distinguish different training tasks. In order to solve the above technical problems, in the method proposed in the embodiment of the present disclosure, when monitoring the training situation of the server cluster, first, the general traffic data of multiple nodes in the server cluster and the data volume of the traffic data can be obtained, wherein the server cluster is configured to perform multiple training tasks of neural network models, and the traffic data is used to represent the data transmitted during network communication in the process of the server cluster performing the training task. For example, when monitoring the server cluster composed of multiple cloud servers, in the natural language processing scene, the training situation of training multiple neural network models based on memory and computing resource improvement (i.e. training task), first, the traffic data generated by network communication between multiple nodes in the server cluster processing the training task and the data volume of the traffic data can be obtained, but not limited to this. For another example, when monitoring the server cluster composed of multiple edge servers, in the computer vision processing scene, the training situation of training multiple neural network models based on model robustness test (i.e. training task), first, the traffic data generated by network communication between multiple nodes in the server cluster processing the training task and the data volume of the traffic data can be obtained, but not limited to this. It should be noted that different training tasks correspond to different traffic data and different data volumes of traffic data.

[0067] In step S404, the training period corresponding to any one node is determined based on the traffic data of any one node, wherein the training period is used to represent the period of any one node performing the corresponding training task.

[0068] In an optional embodiment, in the case of obtaining the traffic data of any one node, the training period corresponding to any one node can be determined based on the traffic data corresponding to any one node. For example, first, the traffic data can be preprocessed, such as cleaning and normalization processing, to obtain preprocessed traffic data, then features related to the training period, such as traffic peak value and traffic fluctuation, can be extracted from the preprocessed traffic data, then the extracted features related to the training period can be input into the trained cycle processing model, that is, the training period corresponding to any one node can be obtained. For another example, the traffic data corresponding to any one node can be subjected to Fourier transform to obtain a frequency spectrum, and finally the training period corresponding to any one node can be obtained based on the frequency spectrum, but not limited to this.

[0069] In step S406, multiple nodes are clustered based on the training periods corresponding to the multiple nodes to obtain multiple node sets, wherein the nodes in the same node set are configured to perform the training task of the same neural network model.

[0070] In an optional embodiment, after obtaining the plurality of training periods corresponding to the plurality of nodes, the plurality of nodes can be clustered based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets. For example, after obtaining the plurality of training periods, the plurality of nodes of the same training period can be clustered to obtain a plurality of node sets, wherein the plurality of nodes in the same node set have the same training period, and the plurality of nodes in different node sets have different training periods. For another example, after obtaining the plurality of training periods corresponding to the plurality of nodes, the training periods can be classified based on different period length intervals, and then the nodes corresponding to the classified training periods can be clustered to obtain a plurality of node sets, wherein the plurality of nodes in the same node set correspond to the training periods in the same period length interval. For example, after obtaining the plurality of training periods corresponding to the plurality of nodes, the plurality of training periods can be classified based on the period length intervals of [0, 1], (1, 3], (3, 5], etc., and then the nodes corresponding to the classified training periods can be clustered, i.e., a plurality of node sets can be obtained.

[0071] In the embodiments of the present disclosure, the traffic data of the plurality of nodes in the server cluster is obtained, the training period corresponding to any one node is determined based on the traffic data of the node, and the plurality of nodes are clustered based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets. It is easy to note that the traffic data of the nodes is the general traffic data generated by the network communication between the nodes, i.e., the traffic data of the nodes is relatively easy to obtain, and in addition, the training data of the training task does not need to be obtained by analyzing the easily obtained traffic data. In addition, by clustering the nodes corresponding to the traffic data, different nodes processing the same training task of the neural network model can be divided into the same node set, and different node sets correspond to different training tasks, so that the nodes executing different training tasks can be clearly distinguished, thereby achieving the technical effect of accurately distinguishing the nodes executing different training tasks, reducing the difficulty of monitoring the execution of the training task of the server cluster, and thereby solving the technical problem of low accuracy of distinguishing the nodes executing different training tasks in the server cluster in the related art.

[0072] In the above embodiments of the present disclosure, the training period corresponding to any one node is determined based on the traffic data of the node, which includes Fourier transforming the traffic data of any one node to obtain a frequency spectrum, extracting a fundamental frequency from the frequency spectrum, and determining the training period based on the fundamental frequency.

[0073] The above-mentioned fundamental frequency can be the periodic component with the lowest frequency in a periodic signal, which can be used to extract the characteristics of the signal and identify the type of the signal.

[0074] In an optional embodiment, after obtaining the traffic data corresponding to any node, Fourier transform can be performed on the traffic data corresponding to any node to obtain a frequency spectrum corresponding to any node, then the maximum peak value can be determined from the frequency spectrum, and the frequency corresponding to the maximum peak value is determined as the fundamental frequency, and finally the training period can be obtained based on the fundamental frequency.

[0075] In the above embodiments of the present disclosure, the fundamental frequency is extracted from the frequency spectrum, and the training period is determined based on the fundamental frequency, which includes extracting the fundamental frequency from the frequency spectrum, performing inverse transform on the fundamental frequency to obtain an initial period, and performing an integer operation on the initial period to obtain the training period.

[0076] In an optional embodiment, after obtaining the frequency spectrum corresponding to any node, first, the fundamental frequency can be extracted from the frequency spectrum, then inverse transform can be performed on the fundamental frequency to obtain an initial period, and finally an integer operation can be performed on the initial period to obtain the training period. For example, in the case where the initial period is 1.5 times the inverse of the fundamental frequency, 1.5 can be rounded to obtain a training period that is 2 times the inverse of the fundamental frequency, but it is not limited thereto. It should be noted that the numerical value in this embodiment is only an example, and the specific numerical value is not limited in this embodiment.

[0077] In the above embodiments of the present disclosure, the plurality of nodes are clustered based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, which includes clustering nodes corresponding to the same training period into the same node set, and clustering nodes corresponding to different training periods into different node sets.

[0078] In an optional embodiment, after obtaining a plurality of training periods corresponding to a plurality of nodes, nodes corresponding to a plurality of training periods with the same period length can be aggregated according to the period length to obtain a plurality of node sets, wherein the training periods of the plurality of nodes in the same node set are the same, and the training periods of the plurality of nodes in different node sets are different.

[0079] In the above embodiments of the present disclosure, the traffic data of the plurality of nodes in the server cluster is obtained, which includes collecting the traffic data of the plurality of nodes by deploying a traffic collection device on the plurality of nodes to obtain the traffic data of the plurality of nodes, or determining a target port connected to the data forwarding device by the plurality of nodes, and collecting data transmitted by the target port to obtain the traffic data of the plurality of nodes.

[0080] The traffic collection device described above can be a network card deployed on the nodes of the server cluster, and the traffic information of the network card can be collected by a program also deployed on the nodes of the server cluster, i.e., the traffic data can be obtained. The data forwarding device described above can be a switch.

[0081] In an optional embodiment, the traffic information of the corresponding plurality of network cards deployed on the plurality of nodes can be collected by deploying a plurality of programs on the plurality of nodes, that is, the traffic data of the plurality of nodes can be obtained.

[0082] In another optional embodiment, the target port connected by the plurality of nodes and the switch can also be determined, and then the data transmitted by the target port can be collected by the program deployed on the node, that is, the corresponding traffic data of the plurality of nodes can be obtained.

[0083] In the above embodiments of the present disclosure, the method further comprises: determining the traffic data of at least one node in the same node set; and comparing the traffic data of the at least one node to determine the execution state of the training task corresponding to the same node set.

[0084] The execution state can include, but is not limited to, execution completion, execution in progress, and execution exception.

[0085] In an optional embodiment, in the case where the traffic data of the at least one node is obtained, the traffic data of the at least one node can be compared based on the traffic data of the normally executed completed training task and the data amount of the traffic data, in the case where the data amount of the traffic data of the at least one node is equal to the data amount of the normally executed completed traffic data, and the content of the traffic data of the at least one node is the same as the data content of the normally executed completed traffic data, it can be determined that the execution state of the training task corresponding to the same node set is execution completion. In the case where the data amount of the traffic data of the at least one node is less than the data amount of the normally executed completed traffic data, or the content of the traffic data of the at least one node is not the same as the data content of the normally executed completed traffic data, it can be determined that the execution state of the training task corresponding to the same node set is execution in progress or execution exception. In the case where the data amount of the traffic data of the at least one node is greater than the data amount of the normally executed completed traffic data, or the content of the traffic data of the at least one node is not the same as the data content of the normally executed completed traffic data, it can be determined that the execution state of the training task corresponding to the same node set is execution exception, but not limited thereto.

[0086] In the above embodiments of the present disclosure, the method further comprises: determining the communication performance between any one node and other nodes in the same node set; and in the case where the communication performance between the any one node and the other nodes is less than a preset performance, adjusting the any one node, so that the communication performance between any two nodes in the adjusted node set is greater than or equal to the preset performance, wherein the preset performance has a correlation with the training task corresponding to the same node set.

[0087] The communication performance described above is used to represent the efficiency and effect of data exchange and cooperation between nodes. The preset performance described above can be set in advance by a user to determine whether to adjust any node. The preset performance has a correlation relationship with the training task corresponding to the same node set.

[0088] In an optional embodiment, the communication performance between any node in the same node set and other nodes can also be determined. If the communication performance between any node and other nodes is less than the preset performance, it indicates that the distance between the node and other nodes is far, which will affect the communication performance. At this time, the position of any node can be adjusted to make the communication performance between any two nodes in the adjusted node set greater than or equal to the preset performance.

[0089] It should be noted that in the adjusted node set, the distance between any two nodes is close, for example, after adjusting the position of any node, any two nodes will be in the same cabinet, so the hop count to the switch is less, which will improve the communication performance between any two nodes.

[0090] In the above embodiments of the present disclosure, the method further comprises: obtaining historical traffic data of at least one node in the same node set; and comparing the historical traffic data with the traffic data of the at least one node to obtain a detection result of the at least one node, wherein the detection result is used to represent whether the at least one node fails.

[0091] The historical traffic data described above can be traffic data generated by network communication of multiple nodes in a historical period during which the server cluster trains multiple neural network models based on historical training tasks.

[0092] In another optional embodiment, historical traffic data of at least one node in the same node set can also be obtained, and then the historical traffic data can be compared with the traffic data of the at least one node to obtain a detection result of the at least one node, wherein the detection result is used to represent whether the at least one node fails. For example, in the case that the historical traffic data is less than or greater than the traffic data of the at least one node, it can be determined that the detection result is that the at least one node fails, and in the case that the historical traffic data is equal to the traffic data of the at least one node, it can be determined that the detection result is that the at least one node does not fail, but not limited thereto.

[0093] Fig. 5 is a schematic diagram of an optional method for clustering multiple nodes in a server cluster to obtain multiple node sets according to Embodiment 1 of the present disclosure. The method mainly utilizes the periodic similarity of the traffic between nodes in the distributed training process, and achieves the purpose of distinguishing different training tasks by aggregating nodes with similar periods together. As shown in Fig. 5, first, the time-traffic curves of all nodes in the cluster can be obtained (e.g., time series curves 1-N in Fig. 5). Second, the Fourier transform can be performed on multiple time series curves to obtain multiple frequency domain curves (e.g., frequency spectra 1-N in Fig. 5). Third, the fundamental frequency can be extracted from the frequency domain curves, which is the frequency of the traffic data. The inverse of the frequency is the training period (e.g., periods 1-N in Fig. 5). Finally, multiple nodes can be aggregated according to the rounded periods, and the clustering of different traffic characteristics can be obtained (e.g., clusters 1-K in Fig. 5).

[0094] The effect of the method is that the clustered curves show high similarity and synchronization, effectively solving the aliasing problem of multiple task monitoring curves in the cluster. The method has the advantage of being less restricted and can complete clustering only with the traffic index, which is a general, commonly collected, and low-sensitivity index, and can be collected from the network card side and the switch side. Therefore, the scheme has strong adaptability. In addition, the method also has the advantage of lower time complexity.

[0095] The method is based on the curve clustering method of the traffic period, which effectively solves the curve aliasing problem of multiple tasks running in parallel in the same cluster after clustering. The method uses a general traffic index for classification, avoids sensitive issues, and has greater adaptability. The method reduces the time complexity of the algorithm through period clustering, which is superior to other conventional schemes.

[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0097] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the order of the described actions, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.

[0098] Those skilled in the art can clearly understand that the method according to the above-mentioned embodiments can be realized by means of software and necessary general hardware platforms, and of course can also be realized by hardware, through the description of the above embodiments. Based on such understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device) to execute the method described in the various embodiments of the present disclosure.

[0099] Embodiment 2

[0100] According to the embodiments of the present disclosure, a processing method of a server cluster is also provided. FIG. 6 is a flowchart of a processing method of a server cluster according to Embodiment 2 of the present disclosure. As shown in FIG. 6, the method includes the following steps:

[0101] In step S602, traffic data of a plurality of nodes in the server cluster is obtained by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the traffic data of the plurality of nodes, the server cluster is configured to perform a training task of a plurality of neural network models, and the traffic data is used to represent data transmitted during network communication between the plurality of nodes in the process of the server cluster performing the training task;

[0102] In step S604, a training period corresponding to any one node is determined based on the traffic data of the any one node, wherein the training period is used to represent a period in which the any one node performs a corresponding training task;

[0103] In step S606, the plurality of nodes are clustered based on the training periods corresponding to the plurality of nodes, to obtain a plurality of node sets, wherein the nodes in the same node set are used to perform a training task of the same neural network model;

[0104] In step S608, the plurality of node sets are output by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the plurality of node sets.

[0105] The first interface described above can be an interface through which other servers obtain traffic data from the server cluster. The second interface described above can be an interface through which other servers output the plurality of node sets to the server cluster.

[0106] In an optional embodiment, in a case where it is necessary to monitor the training of the server cluster and distinguish different training tasks, first, the traffic data of the plurality of nodes in the server cluster can be obtained by calling the first interface, second, the training period corresponding to any one node can be determined based on the traffic data of any one node, then the plurality of nodes can be clustered based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, and finally the plurality of node sets can be output by calling the second interface.

[0107] The first interface includes a first parameter, a parameter value of the first parameter includes the traffic data of the plurality of nodes, the server cluster is configured to perform the training tasks of the plurality of neural network models, the traffic data is configured to represent the data transmitted by the network communication between the plurality of nodes in the process of the server cluster performing the training tasks, the training period is configured to represent the period of any one node performing the corresponding training task, and the nodes in the same node set are configured to perform the training tasks of the same neural network model. The second interface includes a second parameter, and a parameter value of the second parameter includes the plurality of node sets.

[0108] Embodiment 3

[0109] According to the embodiments of the present disclosure, a processing system of a server cluster is also provided. FIG. 7 is a schematic diagram of a processing system of a server cluster according to Embodiment 3 of the present disclosure. As shown in FIG. 7, the system includes a server cluster 72 and a monitoring device 74, wherein the server cluster 72 includes a plurality of nodes 72-1. As shown in FIG. 7, the monitoring device 74 is connected to the plurality of nodes 72-1 in the server cluster 72.

[0110] The server cluster is configured to perform the training tasks of the plurality of neural network models, the monitoring device is configured to determine the training period corresponding to any one node based on the traffic data of any one node, and cluster the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein the traffic data is configured to represent the data transmitted by the network communication between the plurality of nodes in the process of the server cluster performing the training tasks, the training period is configured to represent the period of any one node performing the corresponding training task, and the nodes in the same node set are configured to perform the training tasks of the same neural network model.

[0111] Embodiment 4

[0112] According to the embodiments of the present disclosure, a processing device of a server cluster for implementing the above-mentioned processing method of the server cluster is also provided. FIG. 8 is a schematic diagram of a processing device of a server cluster according to Embodiment 4 of the present disclosure. As shown in FIG. 8, the device includes an obtaining module 82, a determining module 84, and a clustering module 86.

[0113] The obtaining module is configured to obtain traffic data of a plurality of nodes in a server cluster, wherein the server cluster is configured to perform training tasks of a plurality of neural network models, and the traffic data is used to represent data transmitted by network communication between the plurality of nodes during execution of the training tasks by the server cluster. The determining module is configured to determine, based on the traffic data of any one node, a training period corresponding to the any one node, wherein the training period is used to represent a period for the any one node to perform a corresponding training task. The clustering module is configured to cluster the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in a same node set are configured to perform a same training task of a same neural network model.

[0114] It should be noted that the obtaining module 82, the determining module 84, and the clustering module 86 correspond to steps S402 to S406 in Embodiment 1, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b,..., 102n), or the above modules can be run in the computer terminal 10 provided in Embodiment 1 as part of the device.

[0115] In the above embodiments of the present disclosure, the determining module includes a transformation unit and a determination unit.

[0116] The transformation unit is configured to perform Fourier transform on the traffic data of any one node to obtain a frequency spectrum, and the determination unit is configured to extract a fundamental frequency from the frequency spectrum and determine the training period based on the fundamental frequency.

[0117] In the above embodiments of the present disclosure, the determination unit includes an extraction subunit, a transformation subunit, and an integral extraction subunit.

[0118] The extraction subunit is configured to extract the fundamental frequency from the frequency spectrum, the transformation subunit is configured to perform inverse transformation on the fundamental frequency to obtain an initial period, and the integral extraction subunit is configured to perform an integral operation on the initial period to obtain the training period.

[0119] In the above embodiments of the present disclosure, the clustering module includes a clustering unit.

[0120] The clustering unit is configured to cluster nodes corresponding to a same training period into a same node set and cluster nodes corresponding to different training periods into different node sets.

[0121] In the above embodiments of the present disclosure, the obtaining module includes a first acquisition unit or a second acquisition unit.

[0122] The first acquisition unit is configured to acquire the traffic data of the plurality of nodes by deploying a traffic acquisition device on the plurality of nodes.

[0123] In the above embodiment of the present disclosure, the device further comprises a data determination module and a first comparison module.

[0124] The data determination module is configured to determine the traffic data of at least one node in the same node set; and the first comparison module is configured to compare the traffic data of the at least one node to determine the execution state of the training task corresponding to the same node set.

[0125] In the above embodiment of the present disclosure, the device further comprises a performance determination module and an adjustment module.

[0126] The performance determination module is configured to determine the communication performance between any one node and other nodes in the same node set; and the adjustment module is configured to adjust the any one node when the communication performance between the any one node and the other nodes is less than a preset performance, so that the communication performance between any two nodes in the adjusted node set is greater than or equal to the preset performance, wherein the preset performance has an association relationship with the training task corresponding to the same node set.

[0127] In the above embodiment of the present disclosure, the device further comprises a data acquisition module and a second comparison module.

[0128] The data acquisition module is configured to acquire historical traffic data of at least one node in the same node set; and the second comparison module is configured to compare the historical traffic data with the traffic data of the at least one node to obtain a detection result of the at least one node, wherein the detection result is used to represent whether the at least one node fails.

[0129] It should be noted that the preferred embodiments involved in the above embodiments of the present disclosure have the same application scenarios and implementation processes as the schemes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0130] Embodiment 5

[0131] According to the embodiments of the present disclosure, a processing device of a server cluster for implementing the processing method of the server cluster is further provided. FIG. 9 is a schematic diagram of a processing device of a server cluster according to Embodiment 5 of the present disclosure. As shown in FIG. 9, the device comprises an acquisition module 92, a determination module 94, a clustering module 96 and an output module 98.

[0132] The obtaining module is configured to obtain the traffic data of the plurality of nodes in the server cluster by calling a first interface, the first interface including a first parameter, a parameter value of the first parameter including the traffic data of the plurality of nodes, the server cluster being configured to perform a training task of a plurality of neural network models, the traffic data being used to represent data transmitted in network communication between the plurality of nodes during execution of the training task by the server cluster; the determining module is configured to determine, based on the traffic data of any one node, a training period corresponding to the any one node, the training period being used to represent a period for which the any one node executes a corresponding training task; the clustering module is configured to cluster the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, nodes in a same node set being configured to execute a same training task of a same neural network model; and the output module is configured to output the plurality of node sets by calling a second interface, the second interface including a second parameter, a parameter value of the second parameter including the plurality of node sets.

[0133] It should be noted that the obtaining module 92, the determining module 94, the clustering module 96, and the output module 98 correspond to steps S602 to S608 in Embodiment 2, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the solutions disclosed in Embodiment 1. It should be noted that the modules or units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b,..., 102n), or can be a part of the apparatus and can run in the computer terminal 10 provided in Embodiment 1.

[0134] It should be noted that the preferred embodiments involved in the above embodiments of the present disclosure have the same solutions, application scenarios, and implementation processes as those provided in Embodiment 1, but are not limited to the solutions provided in Embodiment 1.

[0135] Embodiment 6

[0136] The embodiments of the present disclosure can provide an electronic device, which can be any one of electronic devices in an electronic device group. Alternatively, in the present embodiment, the electronic device can also be replaced by a terminal device such as a mobile terminal.

[0137] Alternatively, in the present embodiment, the electronic device can be located in at least one network device of a plurality of network devices in a computer network.

[0138] In the present embodiment, the computer terminal can execute program codes in the method.

[0139] Optionally, FIG. 10 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG. 10, the electronic device A can include one or more (only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface connected with a radio frequency module, an audio module, and a display.

[0140] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and apparatuses in the embodiments of the present disclosure. The processor executes various functions and data processing by running the software programs and modules stored in the memory, i.e., implements the methods in the above embodiments. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal A through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0141] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining traffic data of a plurality of nodes in a server cluster; determining a training period corresponding to any one node based on the traffic data of the any one node; clustering the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets.

[0142] Optionally, the processor can further execute program codes of the following steps: performing Fourier transform on the traffic data of any one node to obtain a frequency spectrum; extracting a fundamental frequency from the frequency spectrum, and determining the training period based on the fundamental frequency.

[0143] Optionally, the processor can further execute program codes of the following steps: extracting the fundamental frequency from the frequency spectrum; performing inverse transform on the fundamental frequency to obtain an initial period; performing an integer operation on the initial period to obtain the training period.

[0144] Optionally, the processor can further execute program codes of the following steps: clustering nodes corresponding to the same training period into the same node set, and clustering nodes corresponding to different training periods into different node sets.

[0145] Optionally, the processor can further execute program codes of the following steps: collecting traffic data of the plurality of nodes by deploying a traffic collection device on the plurality of nodes to obtain the traffic data of the plurality of nodes; or determining target ports at which the plurality of nodes are connected with a data forwarding device, and collecting data transmitted by the target ports to obtain the traffic data of the plurality of nodes.

[0146] Optionally, the processor can further execute program codes of the following steps: determining traffic data of at least one node in the same node set; comparing the traffic data of the at least one node to determine the execution state of the training task corresponding to the same node set.

[0147] Optionally, the processor can further execute program codes of the following steps: determining the communication performance between any one node and other nodes in the same node set; and adjusting the any one node when the communication performance between the any one node and other nodes is less than a preset performance, so that the communication performance between any two nodes in the adjusted node set is greater than or equal to the preset performance, wherein the preset performance is associated with the training task corresponding to the same node set.

[0148] Optionally, the processor can further execute program codes of the following steps: obtaining historical traffic data of at least one node in the same node set; and comparing the historical traffic data with the traffic data of the at least one node to obtain a detection result of the at least one node, wherein the detection result is used to represent whether the at least one node fails.

[0149] By adopting the embodiments of the present disclosure, a manner for obtaining traffic data of multiple nodes in a server cluster, determining a training period corresponding to any one node based on the traffic data of the any one node, and clustering the multiple nodes based on the training periods corresponding to the multiple nodes to obtain multiple node sets is provided. It is easy to note that the traffic data of the nodes is general traffic data generated by network communication between the nodes, that is, the traffic data of the nodes is relatively easy to obtain, and in addition, by analyzing the easily obtained traffic data, the training data of the training task need not be obtained. In addition, by clustering the nodes corresponding to the traffic data, different nodes processing the training task of the same neural network model can be divided into the same node set, and different node sets correspond to different training tasks, so that the nodes executing different training tasks can be clearly distinguished, achieving the technical effect of accurately distinguishing the nodes executing different training tasks, reducing the difficulty of monitoring the execution of the training task of the server cluster, and further solving the technical problem of low accuracy of distinguishing the nodes executing different training tasks in the server cluster in the related art.

[0150] Those skilled in the art can understand that the structure as shown in the figure is only schematic, and the electronic device can also be a terminal device such as a smart phone, a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, and the like. The figure does not limit the structure of the electronic device. For example, the electronic device A can further include more or less components (such as a network interface, a display device, and the like) than those shown in the figure, or have a different configuration from that shown in the figure.

[0151] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device by a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, and the like.

[0152] Embodiment 7

[0153] The embodiments of the present disclosure further provide a computer readable storage medium. Optionally, in the present embodiment, the computer readable storage medium can be used to save the program code executed by the method provided in the above embodiments.

[0154] Optionally, in the present embodiment, the storage medium can be located in any one of the electronic devices in the group of electronic devices in the computer network, or in any one of the mobile terminals in the group of mobile terminals.

[0155] Optionally, in the present embodiment, the computer readable storage medium is configured to store program code for performing the following steps: obtaining traffic data of a plurality of nodes in a server cluster; determining a training period corresponding to any one node based on the traffic data of the any one node; clustering the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets.

[0156] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: performing Fourier transform on the traffic data of any one node to obtain a frequency spectrum; extracting a fundamental frequency from the frequency spectrum to obtain the training period.

[0157] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: extracting the fundamental frequency from the frequency spectrum; performing inverse transform on the fundamental frequency to obtain an initial period; performing an integer operation on the initial period, and determining the training period based on the fundamental frequency.

[0158] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: clustering the nodes corresponding to the same training period into the same node set, and clustering the nodes corresponding to different training periods into different node sets.

[0159] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: collecting the traffic data of the plurality of nodes by deploying the traffic collection device on the plurality of nodes to obtain the traffic data of the plurality of nodes; or determining the target port of the plurality of nodes connected to the data forwarding device, and collecting the data transmitted by the target port to obtain the traffic data of the plurality of nodes.

[0160] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: determining the traffic data of at least one node in the same node set; comparing the traffic data of the at least one node to determine the execution state of the training task corresponding to the same node set.

[0161] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: determining the communication performance between any one node and other nodes in the same node set; in the case that the communication performance between any one node and other nodes is less than a preset performance, adjusting any one node so that the communication performance between any two nodes in the adjusted node set is greater than or equal to the preset performance, wherein the preset performance has a correlation with the training task corresponding to the same node set.

[0162] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: obtaining the historical traffic data of at least one node in the same node set; comparing the historical traffic data with the traffic data of the at least one node to obtain a detection result of the at least one node, wherein the detection result is used to represent whether the at least one node fails.

[0163] Embodiment 8

[0164] The embodiments of the present disclosure further provide a computer program product. Optionally, in the embodiments, the computer program product can include a computer program, and the computer program, when executed by a processor, implements the method provided by the above-mentioned embodiments.

[0165] Embodiment 9

[0166] The embodiments of the present disclosure further provide a computer program product. Optionally, the computer program product can include a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium can be used to store a computer program, and the computer program, when executed by a processor, implements the method provided by the above-mentioned embodiments.

[0167] Embodiment 10

[0168] Embodiments of the present disclosure also provide a computer program. Optionally, in the embodiments, the computer program is executed by a processor to implement the method provided by the above-mentioned embodiments.

[0169] The sequence numbers of the above-described embodiments of the present disclosure are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0170] In the above-described embodiments of the present disclosure, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0171] In several embodiments provided by the present disclosure, it should be understood that the disclosed technology can be implemented in other ways. Of course, the embodiments described above are only schematic. For example, the division of units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.

[0172] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiments.

[0173] In addition, each functional unit in each embodiment of the present disclosure can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware or software functional units.

[0174] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present disclosure, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, etc.

[0175] The above only describes the preferred embodiments of the present disclosure, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present disclosure, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as the protection scope of the present disclosure.

Claims

1. A processing method of a server cluster, comprising: obtaining traffic data of a plurality of nodes in a server cluster, wherein the server cluster is configured to perform training tasks of a plurality of neural network models, and the traffic data is used to represent data transmitted between the plurality of nodes in network communication during the server cluster performing the training tasks; determining a training period corresponding to an arbitrary node based on traffic data of the arbitrary node, wherein the training period is used to represent a period in which the arbitrary node performs a corresponding training task; clustering the plurality of nodes based on training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in a same node set are configured to perform a same training task of a same neural network model.

2. The method of claim 1, wherein, The determination of the training period corresponding to the arbitrary node based on the traffic data of the arbitrary node comprises: performing Fourier transform on the traffic data of the arbitrary node to obtain a frequency spectrum; extracting a fundamental frequency from the frequency spectrum, and determining the training period based on the fundamental frequency.

3. The method of claim 2, wherein, The extraction of the fundamental frequency from the frequency spectrum and the determination of the training period based on the fundamental frequency comprise: extracting the fundamental frequency from the frequency spectrum; performing inverse transform on the fundamental frequency to obtain an initial period; performing an integer operation on the initial period to obtain the training period.

4. The method of claim 1, wherein, The clustering of the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain the plurality of node sets comprises: clustering nodes corresponding to a same training period into a same node set, and clustering nodes corresponding to different training periods into different node sets.

5. The method of claim 1, wherein, The obtaining of the traffic data of the plurality of nodes in the server cluster comprises: collecting the traffic data of the plurality of nodes by deploying a traffic collection device on the plurality of nodes to obtain the traffic data of the plurality of nodes; or determining a target port at which the plurality of nodes are connected to a data forwarding device, and collecting data transmitted by the target port to obtain the traffic data of the plurality of nodes.

6. The method of any one of claims 1 to 5, wherein, The method further comprises: determining traffic data of at least one node in a same node set; comparing the traffic data of the at least one node to determine an execution state of a training task corresponding to the same node set.

7. The method of claim 6, wherein, The comparison of the traffic data of the at least one node to determine the execution state of the training task corresponding to the same node set comprises: comparing the traffic data of the at least one node with a data amount and data content of normally executed and completed traffic data to obtain a comparison result, wherein the comparison result is used to reflect a comparison situation of the traffic data of the at least one node and the normally executed and completed traffic data in terms of the data amount and the data content; determining the execution state of the training task corresponding to the same node set based on the comparison result.

8. The method of claim 7, wherein, The determination of the execution state of the training task corresponding to the same node set based on the comparison result comprises: in response to the comparison result being that the data amount of the traffic data of the at least one node is equal to the data amount of the normally executed traffic data and the data content of the traffic data of the at least one node is the same as the data content of the normally executed traffic data, determining that the execution state of the training task corresponding to the same node set is an execution completed state; in response to the comparison result being that the data amount of the traffic data of the at least one node is less than the data amount of the normally executed traffic data or the data content of the traffic data of the at least one node is not the same as the data content of the normally executed traffic data, determining that the execution state of the training task corresponding to the same node set is an execution in progress state or an execution abnormal state; in response to the comparison result being that the data amount of the traffic data of the at least one node is greater than the data amount of the normally executed traffic data or the data content of the traffic data of the at least one node is not the same as the data content of the normally executed traffic data, determining that the execution state of the training task corresponding to the same node set is the execution abnormal state.

9. The method of any one of claims 1 to 8, wherein, The method further comprises: determining the communication performance between any one node in the same node set and other nodes; in the case that the communication performance between the any one node and the other nodes is less than a preset performance, adjusting the any one node, so that the communication performance between any two nodes in the adjusted node set is greater than or equal to the preset performance, wherein the preset performance has an association relationship with the training task corresponding to the same node set.

10. The method of any one of claims 1 to 8, wherein, The method further comprises: obtaining historical traffic data of at least one node in the same node set; comparing the historical traffic data with the traffic data of the at least one node to obtain a detection result of the at least one node, wherein the detection result is used to represent whether the at least one node has a fault.

11. A processing method of a server cluster, comprising: obtaining traffic data of a plurality of nodes in a server cluster by calling a first interface, wherein the first interface comprises a first parameter, a parameter value of the first parameter comprises the traffic data of the plurality of nodes, the server cluster is set to execute a training task of a plurality of neural network models, and the traffic data is used to represent data transmitted during network communication between the plurality of nodes in the process of the server cluster executing the training task; determining a training period corresponding to any one node based on the traffic data of the any one node, wherein the training period is used to represent a period of the any one node executing a corresponding training task; clustering the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in the same node set are set to execute a training task of the same neural network model; outputting the plurality of node sets by calling a second interface, wherein the second interface comprises a second parameter, and a parameter value of the second parameter comprises the plurality of node sets.

12. A processing system of a server cluster, comprising: a server cluster comprising a plurality of nodes, the server cluster being configured to perform training tasks of a plurality of neural network models; a monitoring device connected to the plurality of nodes, configured to determine a training period corresponding to any one node based on traffic data of the any one node, and cluster the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein the traffic data is used to represent data transmitted between the plurality of nodes for network communication during the server cluster performing the training tasks, and the training period is used to represent a period for the any one node to perform a corresponding training task, and nodes in a same node set are configured to perform a same training task of a same neural network model.

13. An electronic device, comprising: a memory storing an executable program; a processor configured to run the program, wherein the program, when executed, performs the following method: obtaining traffic data of a plurality of nodes in a server cluster, wherein the server cluster is configured to perform training tasks of a plurality of neural network models, and the traffic data is used to represent data transmitted between the plurality of nodes for network communication during the server cluster performing the training tasks; determining a training period corresponding to any one node based on traffic data of the any one node, wherein the training period is used to represent a period for the any one node to perform a corresponding training task; and clustering the plurality of nodes based on the training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in a same node set are configured to perform a same training task of a same neural network model.

14. The electronic device of claim 13, wherein, The program, when executed, further performs the following method: performing Fourier transform on the traffic data of the any one node to obtain a frequency spectrum; extracting a fundamental frequency from the frequency spectrum, and determining the training period based on the fundamental frequency.

15. The electronic device of claim 14, wherein, The program, when executed, further performs the following method: extracting the fundamental frequency from the frequency spectrum; performing inverse transform on the fundamental frequency to obtain an initial period; performing an integer operation on the initial period to obtain the training period.

16. The electronic device of claim 13, wherein, The program, when executed, further performs the following method: clustering nodes corresponding to a same training period into a same node set, and clustering nodes corresponding to different training periods into different node sets.

17. The electronic device of claim 13, wherein, The program, when executed, further performs the following method: collecting the traffic data of the plurality of nodes by deploying a traffic collection device on the plurality of nodes to obtain the traffic data of the plurality of nodes; or determining a target port at which the plurality of nodes are connected to a data forwarding device, and collecting data transmitted by the target port to obtain the traffic data of the plurality of nodes.

18. The electronic device of claim 13, wherein, The program, when executed, further performs the following method: determining traffic data of at least one node in a same node set; comparing the traffic data of the at least one node to determine an execution state of a training task corresponding to the same node set.

19. A computer readable storage medium comprising a stored executable program, wherein, The executable program controls a device where the storage medium is located to perform the following method when the executable program is running: obtaining traffic data of a plurality of nodes in a server cluster, wherein the server cluster is used to perform training tasks of a plurality of neural network models, and the traffic data is used to represent data transmitted by network communication between the plurality of nodes in a process in which the server cluster performs the training tasks; determining a training period corresponding to any one node based on traffic data of the any one node, wherein the training period is used to represent a period in which the any one node performs a corresponding training task; and clustering the plurality of nodes based on training periods corresponding to the plurality of nodes to obtain a plurality of node sets, wherein nodes in a same node set are used to perform training tasks of a same neural network model.

20. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Distributed training method and device for deep learning model, equipment and storage medium

    CN110969198A

  • Training task deployment method, system and device and storage medium

    CN117555666A

  • A device for diverging mobile scent with deodorizing function

    KR1020220011240A