Communication method and apparatus
Patent Information
- Application Number
- PCT/CN2026/079723
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-02-24
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026079723_01102026_PF_FP_ABST
Abstract
Description
Communication methods and devices
[0001] This application claims priority to Chinese patent application No. 202510391915.8, filed with the State Intellectual Property Office of China on March 28, 2025, entitled "Communication Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of communications, and more particularly to communication methods and apparatus. Background Technology
[0003] As computing clusters grow in size, their failure rate also increases, leading to a higher probability of training jobs failing within them. To promptly detect training job failures, fault monitoring can be implemented. Monitoring training jobs first requires identifying which computing nodes within the cluster the job is running on, and then using the job information from those nodes to perform fault monitoring.
[0004] Currently, determining the computing nodes corresponding to training jobs typically relies on intermediate nodes, such as artificial intelligence (AI) platforms or cluster management nodes. However, users may not have an AI platform or may not agree to the integration of the AI platform with their operational platform. In such cases, it becomes impossible to determine the computing nodes running the training jobs, thus hindering fault detection for the training jobs. Summary of the Invention
[0005] This application provides a communication method and apparatus that can determine the computing resources for training jobs through information interaction between computing nodes and operation and maintenance nodes, thereby avoiding the dependence of computing resource determination on AI platforms and cluster management nodes.
[0006] Firstly, a communication method is provided. This method can be executed by an operation and maintenance node, by a module applied to the operation and maintenance node (e.g., a processor, chip, or chip system), or by a logical node, logical module, or software capable of implementing all or part of the operation and maintenance node's functions. The method includes: receiving first information from multiple processing units, wherein the first information of the first processing unit includes identification information of a training job run by the first processing unit and identification information of the computing node where the first processing unit is located, and the first processing unit is any one of the multiple processing units; and determining the computing resources corresponding to a first training job based on the first information of the multiple processing units, wherein the training jobs run by the multiple processing units include the first training job.
[0007] In the above technical solution, the first information of the processing unit includes the identification information of the training job running by the processing unit and the identification information of the computing node where the processing unit is located. Therefore, the operation and maintenance node can use the identification information of the first training job to determine which processing unit among multiple processing units is running the first training job. Then, by combining this information with the identification information of the computing node where the processing unit corresponding to the first training job is located, the node corresponding to the first training job (i.e., the computing node running the first training job, which can also be called the computing resource) can be determined. Thus, this application can determine the computing resources of the training job through information interaction between the computing node and the operation and maintenance node, without relying on the AI platform or cluster management node. This avoids the dependence of determining computing resources on the AI platform and cluster management node, and improves the flexibility of job restoration.
[0008] Optionally, the computing resources corresponding to the first training job include a processing unit that runs the first training job. Further, the computing resources also include the computing node where the processing unit running the first training job is located.
[0009] In conjunction with the first aspect described above, in one possible design approach, determining the computing resources corresponding to the first training job based on the first information of multiple processing units includes: dividing the multiple processing units into at least one processing unit group according to the identification information of the training jobs run by the processing units, wherein the identification information of the training jobs run by processing units in the same processing unit group is the same; determining that the processing units running the first training job include the processing units in the first processing unit group. The first processing unit group is a processing unit group within at least one processing unit group, and the identification information of the training jobs run by the processing units in the first processing unit group is the identification information of the first training job.
[0010] In conjunction with the first aspect mentioned above, in one possible design approach, the first information of the first processing unit also includes the identifier of the main process of the first processing unit and the number of processing units corresponding to the training jobs run by the first processing unit.
[0011] In conjunction with the first aspect above, in one possible design approach, determining that the processing unit running the first training job includes the processing units in the first processing unit group includes: determining that the processing unit running the first training job includes the processing units in the first processing unit group when the number of processing units in the first processing unit group is equal to the number of processing units corresponding to the first training job, or when the identifier of the main process of the processing unit in the first processing unit group includes all natural numbers within a first range; wherein, the first range is [0, N-1], and N is the number of processing units corresponding to the first training job.
[0012] In the above technical solution, all processing units in the first processing unit group are used to run the first training job, and the number of processing units corresponding to the first training job is pre-indicated as the total number of processing units used to run the first training job. Therefore, if the number of processing units in the first processing unit group is equal to the number of processing units corresponding to the first training job, it indicates that the first processing unit group includes all processing units running the first training job. Alternatively, the number of identifiers of the main process of a processing unit can represent the number of processing units. Therefore, if the identifiers of the main processes of processing units in the first processing unit group include all natural numbers within the range [0, N-1], it indicates that the first processing unit group includes all processing units running the first training job. Thus, based on the above technical solution, the operation and maintenance node can determine all processing units corresponding to a training job, and then obtain complete information about the training job based on all processing units corresponding to a training job, thereby providing data support for fault analysis of the training job.
[0013] In conjunction with the first aspect mentioned above, in one possible design approach, the first information of the first processing unit also includes the timestamp corresponding to the first processing unit. The timestamp corresponding to the first processing unit is the timestamp when the main process of the first processing unit starts running the training job, or the set of timestamps when multiple processes of the first processing unit start running the training job.
[0014] In conjunction with the first aspect above, in one possible design approach, determining that the processing unit running the first training job includes the processing units in the first processing unit group includes: determining that the processing units in the first processing unit group are different from the processing units running the second training job, and / or, when the timestamps corresponding to the processing units in the first processing unit group are different from the timestamps corresponding to the processing units of the second training job, the processing unit running the first training job includes the processing units in the first processing unit group; wherein, the identification information of the second training job is the same as the identification information of the first training job, the identification information of the first training job is the Internet Protocol IP address of the computing node where the master process of the first training job is located, and the identification information of the second training job is the IP address of the computing node where the master process of the second training job is located.
[0015] In the above technical solution, when a training job is running, the timestamps of the processing unit and the main process of that processing unit that begin running the training job remain unchanged. Therefore, it can be assumed that if it is the same training job, the running processing unit should be the same, and the timestamps corresponding to the processing units should also be the same. Furthermore, by comparing and analyzing the relevant information of the processing units in the first processing unit group with the relevant information of the processing units corresponding to the existing second training job, it can be determined whether the training job running in the first processing unit group is an existing second training job. When it is different from the second training job, it is determined that the processing units running the first training job include the processing units in the first processing unit group, avoiding repeated analysis of existing jobs and having the advantage of saving resources.
[0016] In conjunction with the first aspect mentioned above, in one possible design, the first information of the first processing unit also includes the identification information of the first processing unit. This allows the operation and maintenance node to identify the processing unit in the computing node used to run the training job based on the identification information of the processing unit. Consequently, it can obtain information about the corresponding processing unit in the computing node used for the training job, while avoiding obtaining information about processing units in the computing node not used to run the training job, thus improving the accuracy of job reconstruction.
[0017] Secondly, a communication method is provided. This method can be executed by a computing node, a module applied to the computing node (e.g., a processor, chip, or chip system), or a logical node, logical module, or software capable of implementing all or part of the computing node's functions. The method includes: determining first information of at least one processing unit; the first information of the second processing unit includes identification information of the training job run by the second processing unit and identification information of the computing node where the second processing unit is located, wherein the second processing unit is any one of the at least one processing unit; and sending the first information of the at least one processing unit, wherein the first information of the at least one processing unit is used to determine the computing resources corresponding to the training job run by each of the at least one processing unit. The technical effects of the second aspect are similar to those of the first aspect and will not be elaborated further here.
[0018] In conjunction with the second aspect above, in one possible design approach, the computational resources corresponding to the training job include a processing unit for running the training job.
[0019] In conjunction with the second aspect above, in one possible design, the first information of the second processing unit also includes the identifier of the main process of the second processing unit and the number of processing units corresponding to the training job run by the second processing unit.
[0020] In conjunction with the second aspect above, in one possible design, the first information of the second processing unit also includes the timestamp corresponding to the second processing unit. The timestamp corresponding to the second processing unit is the timestamp when the main process of the second processing unit starts running the training job, or the set of timestamps when multiple processes of the second processing unit start running the training job.
[0021] In conjunction with the second aspect above, in one possible design approach, when the second processing unit corresponds to multiple pieces of first information, the first information of the second processing unit is the first information with the largest timestamp among the multiple pieces of first information corresponding to the second processing unit.
[0022] In conjunction with the second aspect mentioned above, in one possible design approach, the first information of the second processing unit also includes the identification information of the second processing unit.
[0023] The technical effects of any possible design in the second aspect can be referenced from the technical effects of the corresponding or the same design in the first aspect, and will not be elaborated here.
[0024] Thirdly, a communication device is provided for implementing various methods. The communication device includes modules, units, or means corresponding to the implementation of the methods, which can be implemented in hardware, software, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the functions.
[0025] In some possible designs, the communication device may include a processing module and a transceiver module. The processing module can be used to implement the processing functions in any of the above aspects and any possible implementations thereof. The transceiver module may include a receiving module and a transmitting module, respectively used to implement the receiving function and the transmitting function in any of the above aspects and any possible implementations thereof.
[0026] In some possible designs, the transceiver module can consist of transceiver circuits, transceivers, transceivers, or communication interfaces.
[0027] Fourthly, a communication device is provided, comprising: a processor and a memory; the memory being used to store computer instructions that, when executed by the processor, cause the communication device to perform the method described in either aspect.
[0028] Fifthly, a communication device is provided, comprising: a processor and a communication interface; the communication interface being used to communicate with a module outside the communication device; the processor being used to execute a computer program or instructions to cause the communication device to perform the method described in any one of these aspects.
[0029] A sixth aspect provides a communication device comprising: at least one processor; said processor being configured to execute a computer program or instructions stored in a memory to cause the communication device to perform the method described in any of the aspects. The memory may be coupled to the processor, or may be independent of the processor.
[0030] In a seventh aspect, a communication device (e.g., the communication device may be a chip or a chip system) is provided, the communication device including a processor for implementing the functions involved in any one of the first to second aspects.
[0031] In some possible designs, the communication device includes a memory for storing necessary program instructions and data.
[0032] In some possible designs, when the device is a chip system, it can be composed of chips or contain chips and other discrete components.
[0033] It is understood that the communication device provided in the third to seventh aspects may be the operation and maintenance node in the first aspect, or a module or unit (e.g., a chip, chip system, or circuit) in the operation and maintenance node that performs the methods / operations / steps / actions described in the first aspect, or a module or unit that can be used in conjunction with the operation and maintenance node, or a logic node, logic module, or software that can realize all or part of the functions of the operation and maintenance node; or, the communication device may be the computing node in the second aspect, or a module or unit (e.g., a chip, chip system, or circuit) in the computing node that performs the methods / operations / steps / actions described in the second aspect, or a module or unit that can be used in conjunction with the computing node, or a logic node, logic module, or software that can realize all or part of the functions of the computing node.
[0034] It is understandable that when the communication device provided by any of the third to seventh aspects is a chip, the sending action / function of the communication device can be understood as outputting information, and the receiving action / function of the communication device can be understood as inputting information.
[0035] Eighthly, a computer-readable storage medium is provided that stores a computer program or instructions that, when executed on a communication device, enable the communication device to perform the method described in any one of the first to second aspects.
[0036] A ninth aspect provides a computer program product containing instructions that, when run on a communication device, enables the communication device to perform the method described in any one of the first to second aspects.
[0037] In a tenth aspect, a communication system is provided, comprising an operation and maintenance node and a computing node. The operation and maintenance node is used to execute the method described in the first aspect and any possible design thereof, and the computing node is used to execute the method described in the second aspect and any possible design thereof.
[0038] The technical effects of any of the design methods in aspects three through ten can be found in the technical effects of different design methods in aspects one through two, and will not be repeated here. Attached Figure Description
[0039] Figure 1 is a schematic diagram of a template for creating training jobs;
[0040] Figure 2 is a schematic diagram of AI model training;
[0041] Figure 3 is a schematic diagram of a training task diagnosis;
[0042] Figure 4 is a schematic diagram of the architecture of a communication system provided in this application;
[0043] Figure 5 is a schematic diagram of the architecture of another communication system provided in this application;
[0044] Figure 6 is a flowchart illustrating a communication method provided in this application;
[0045] Figure 7 is a schematic diagram of a computing resource provided in this application;
[0046] Figures 8-10 are schematic flowcharts of another communication method provided in this application;
[0047] Figures 11 and 12 are schematic diagrams of the communication device provided in this application. Detailed Implementation
[0048] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. A and B can be singular or plural.
[0049] In the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0050] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0051] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0052] It is understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, various embodiments throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0053] It is understood that in this application, "...when" and "if" both refer to the corresponding processing that will be carried out under certain objective circumstances, and are not limited to a specific time, nor do they require a judgment action to be performed during implementation, nor do they imply any other limitations.
[0054] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Correspondingly, the apparatus given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.
[0055] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, unless otherwise specified or there is a logical conflict, the terminology and / or descriptions between different embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following descriptions of the embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0056] To facilitate understanding of the technical solutions of the embodiments of this application, a brief introduction to the relevant technologies of this application is given below.
[0057] Currently, artificial intelligence (AI) has entered the era of large-scale models. The demand for computing power from these large models is increasing daily, prompting AI computing power to shift from single-machine mode to cluster mode. As the scale of computing clusters continues to expand, the overall failure rate of the computing cluster increases, assuming the individual node failure rate remains constant. This leads to a higher probability of training jobs failing, resulting in training job freezes, interruptions, or performance degradation. To promptly detect training job failures, fault monitoring can be implemented for training jobs running within the computing cluster. When monitoring training jobs, it is first necessary to determine which computing nodes in the cluster the training job is running on, and then perform fault monitoring based on the job information on these nodes.
[0058] Currently, determining the computing nodes corresponding to training jobs typically relies on intermediate nodes, which may include AI platforms or cluster management nodes.
[0059] In one related technology, an AI platform can record information about the creation of training jobs. Specifically, an AI platform (such as a cloud platform for enterprise management of AI training and inference tasks, e.g., the Nvidia base command platform) can create training jobs based on a predefined template, as shown in Figure 1. The template includes multiple parameters, such as name, accelerated computing environment (ACE), instance type, number of allocated graphics processing units (GPUs), number of nodes (NODES), container image used, and data input configuration. The template standardizes the resource configuration of training jobs and simplifies the process of creating training jobs. The Nvidia base command platform can use the template shown in Figure 1 to define the job parameters of training jobs and also provides the NVIDIA GPU Cloud (NGC) command tool for external querying of job information. In this way, Nvidia can use the NGC command-line tool to retrieve training job information using an application programming interface (API) GET request. For example, it can send a GET request to obtain a valid authorization token, and then use the token returned by the first request to send another GET request to retrieve the training job information. This allows it to determine the compute node corresponding to the training job.
[0060] In another related technology, as shown in Figure 2, the cloud training system determines the computing nodes corresponding to the training job based on the AI platform and the cluster management node (also known as the guidance system). Specifically, developers develop and run training code based on the AI platform. The AI platform obtains and sends the initial information of the training job. The guidance system can determine which computing nodes the training job should be scheduled on, and then obtains and uploads the training job information based on the initial information. The cloud training system can execute the training job according to the training job information, and then returns the trained AI model to the guidance system.
[0061] In another related technology, as shown in Figure 3, the operations and maintenance (O&M) node obtains relevant information about the training job from the AI platform or cluster management node (for example, the scheduler in the cluster management node can obtain relevant information about the training job based on each computing node and send this information to the O&M node through a proxy platform. Correspondingly, the O&M node receives relevant training job information from the cluster management node) to determine the computing node corresponding to the training job and obtain the operational information of each computing node (e.g., alarms, metrics, log information). The O&M node can then perform training job diagnosis based on the operational information of the computing node corresponding to the training job, such as job fault diagnosis, job degradation diagnosis, and job deadlock diagnosis.
[0062] In the aforementioned technologies, if there is no AI platform in the computing cluster, the computing nodes for training jobs cannot be determined (i.e., job information cannot be obtained), thus making fault detection for training jobs impossible. If the owner of the computing cluster does not agree to the integration of the AI platform and the operations and maintenance (O&M) nodes, the O&M nodes also cannot determine the computing nodes for training jobs, thus making fault detection for training jobs impossible. Similarly, if the owner of the computing cluster does not agree to the integration of the cluster management node and the O&M nodes, the O&M nodes also cannot determine the computing nodes for training jobs, and therefore cannot perform fault monitoring for training jobs.
[0063] Therefore, how to determine the computational resources of a training job while avoiding the dependence of the determination of the computational resources corresponding to the training job on intermediate nodes is an urgent problem to be solved.
[0064] In view of this, this application provides a communication method in which a computing node determines and sends first information of at least one processing unit, and an operation and maintenance node receives first information of multiple processing units. Since the first information of a processing unit includes the identification information of the training job running by that processing unit and the identification information of the computing node where that processing unit is located, the operation and maintenance node can determine the processing unit running the first training job from among the multiple processing units through the identification information of the first training job. Furthermore, by combining this with the identification information of the computing node where the processing unit corresponding to the first training job is located, the operation and maintenance node can determine the computing node corresponding to the first training job (i.e., the computing node running the first training job, which can also be called a computing resource). Thus, this application can determine the computing resources of a training job through information interaction between the computing node and the operation and maintenance node, without relying on an AI platform or cluster management node, avoiding the dependence of computing resource determination on AI platforms and cluster management nodes, and improving the flexibility of job restoration.
[0065] Figure 4 illustrates a possible, non-limiting system diagram. As shown in Figure 4, the communication system includes an operation and maintenance node 401 and at least one computing node 402 (multiple nodes are shown in the figure). The operation and maintenance node 401 and the computing node 402 can communicate with each other, such as through direct or indirect communication, or through wireless or wired means, without limitation.
[0066] Computation node 402 can be classified into a computing cluster. This computing cluster can be divided into multiple sub-clusters, each with a corresponding cluster ID (clusterId). Each sub-cluster includes one or more computing nodes. For example, as shown in Figure 4, the computing cluster includes two sub-clusters: Sub-cluster 1 includes 9 computing nodes. Training job 1 runs on computing nodes 1 through 4 (nodes 4), and training job 2 runs on computing nodes 5 through 9 (nodes 9); Sub-cluster 2 includes 9 computing nodes. Training job 3 runs on computing nodes 10 through 13 (nodes 13), and training job 4 runs on computing nodes 14 through 18 (nodes 14 through 18).
[0067] In some embodiments, the computing node 402 includes at least one processing unit. Each computing node 402 corresponds to a globally unique serial number (SN), and each processing unit corresponds to an identification information. For example, the identification information can be the number of the processing unit. Taking a computing node comprising 10 processing units as an example, the numbers of the 10 processing units can be sequentially: 0, 1, 2, ..., 9.
[0068] The number of processing units in computing node 402 can be 8 or 16, or other numbers, and this application does not limit this.
[0069] The processing unit can be a neural network processing unit (NPU), a central processing unit (CPU), a tensor processing unit (TPU), or a GPU; this application does not impose any restrictions on this.
[0070] Each processing unit runs a main process (or main thread). The main process can also be called a rank, and this application does not restrict the name of the main process. Taking a processing unit running one main process as an example, one main process corresponds to a global rankId.
[0071] As one possible implementation, compute node 402 can run a training job. For example, compute node 402 includes 10 NPUs, of which 10 CPUs can be used to run training job 1.
[0072] As another possible implementation, compute node 402 can run multiple training jobs. For example, compute node 402 includes 10 NPUs, of which 5 NPUs can be used to run training job 1 and the other 5 NPUs can be used to run training job 2.
[0073] The number of processing units corresponding to a training job can be called the total number of processes (rankNum). This number can be understood as the number of processing units occupied by the training job. For example, if training job 1 occupies 16 processing units on a computing node, then 16 is the number of processing units corresponding to training job 1. As shown in Figure 4, one training job corresponds to one master process (or master thread), or one training job corresponds to multiple processes. Among the multiple processes corresponding to one training job, there is one master process. The Internet Protocol (IP) address of the computing node where the master process resides can be used as the identification information of the training job. The IP address of the computing node where the master process resides is also called the master IP.
[0074] For example, as shown in Figure 4, taking the case where each computing node contains 8 processes, training job 1 runs on the 8 processes of the 1st computing node, the 8 processes of the 2nd computing node, the 8 processes of the 3rd computing node, and the 4th computing node. If the 1st computing node runs a master process (that is, the 8 processes of the 1st computing node include the master process), the IP address (master IP) of the computing node 1 where the master process is located can be used as the identification information of training job 1.
[0075] In some embodiments, as shown in FIG5, the operation and maintenance node 401 has functions such as fault diagnosis, deadlock diagnosis, and degradation diagnosis of training jobs. The computing node 402 may include a proxy platform. The computing node 402 can send relevant information about the training job of at least one processing unit, as well as the operating information of the computing node (e.g., alarms, metrics, log information, etc.), to the operation and maintenance node 401 through the proxy platform. The operation and maintenance node 401 can reconstruct the job based on the relevant information of the training job of at least one processing unit to determine the computing resources corresponding to the training job. Furthermore, the operation and maintenance node can perform training job diagnosis based on the operating information of the computing resources corresponding to the training job, such as training job fault diagnosis, training job deadlock diagnosis, and training job degradation diagnosis.
[0076] Optionally, as shown in Figure 5, the communication system may also include an AI platform and a cluster management node (shown as dashed lines in the figure). The AI platform and cluster management node can schedule training jobs to designated computing nodes.
[0077] In some embodiments, the maintenance node 401 in the communication system may be a server in the maintenance cluster, and the computing node 402 may be a server in the computing cluster.
[0078] Optionally, the server may be a standalone physical server, a server cluster consisting of multiple physical servers, a distributed file system, or at least one of the following cloud servers providing basic cloud computing services: cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Of course, the above is merely an illustrative description of the server; the server may also be a relational database management system (e.g., MySQL, Oracle Database, or Microsoft SQL Server) and / or a database based on distributed file storage (e.g., MongoDB, HBase, or Cassandra), and this application makes no limitations in this regard.
[0079] It should be noted that the communication system described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0080] The following description, using the communication system shown in Figure 4 as an example, illustrates the communication method provided in this application. It should be noted that the message names, parameter names, or information names between the operation and maintenance node and the computing node in the following embodiments are merely examples; other names may exist in other embodiments, and the method provided in this application does not specifically limit these.
[0081] It is understood that in the embodiments of this application, the operation and maintenance node or the computing node may execute some or all of the steps in the embodiments of this application. These steps or operations are merely examples, and the embodiments of this application may also execute other operations or variations thereof. Furthermore, the various steps may be executed in different orders as presented in the embodiments of this application, and it is not necessarily necessary to execute all the operations in the embodiments of this application.
[0082] It is understood that this application uses computing nodes and operation and maintenance nodes as examples to illustrate the execution subjects of the interaction, but this application does not limit the execution subjects of the interaction. For example, the method executed by the computing node in this application can also be executed by a module applied to the computing node (e.g., a chip, chip system, or processor), or by a logical node, logical module, or software that can implement all or part of the functions of the computing node; similarly, the method executed by the operation and maintenance node in this application can also be executed by a module applied to the operation and maintenance node (e.g., a chip, chip system, or processor), or by a logical node, logical module, or software that can implement all or part of the functions of the operation and maintenance node.
[0083] Furthermore, in this application, "sending information" can be understood as one device sending information to another device, or it can also be understood as one logical module within a device sending information to another logical module. For example, "a computing node sending information" can be understood as a computing node sending information to another device (such as an operations and maintenance node), or it can be understood as logical module 1 (such as a processing module) in a computing node sending information to logical module 2 (such as a transceiver module) in the computing node.
[0084] In this application, "receiving information" can be understood as one device receiving information from another device, or it can also be understood as a logical module within a device receiving information from another logical module. For example, "the maintenance node receiving information" can be understood as the maintenance node receiving information from another device (such as a computing node), or it can be understood as logical module 1 (such as a processing module) in the maintenance node receiving information from logical module 2 (such as a transceiver module) in the maintenance node.
[0085] In this application, "sending information to... (e.g., an operations and maintenance node)" or the relevant illustrations in the accompanying drawings can be understood as the destination of the information being the operations and maintenance node. This can include sending information directly or indirectly to the operations and maintenance node. Similarly, "receiving information from... (e.g., a compute node)," "receiving information from... (e.g., a compute node)," or "receiving information sent by (e.g., a compute node)," or the relevant illustrations in the accompanying drawings, can be understood as the source of the information being the compute node. This can include receiving information directly or indirectly from the compute node. Information may undergo necessary processing between the source and destination, such as format changes, but the destination can understand the valid information from the source. Similar expressions in this application can be interpreted similarly, and will not be elaborated further here.
[0086] Referring to Figure 6, which is a flowchart of a communication method provided in an embodiment of this application, the method may include the following steps:
[0087] S601, The computing node determines the first information of at least one processing unit.
[0088] In the first information of at least one processing unit, the first information of each processing unit includes the identification information of the training job run by the processing unit and the identification information of the computing node where the processing unit is located.
[0089] For example, the identification information of the training job run by the processing unit can be the masterIP corresponding to the training job, where the masterIP is the IP address of the computing node where the master process of the training job run by the processing unit is located. Alternatively, the identification information of the training job run by the processing unit can also be a unique identifier of the training job, such as a unique code (identity document, ID) for the training job. This application does not restrict the identification information of the training job.
[0090] For example, the identification information of the computing node where the processing unit is located can be the computing node's serial number (SN), identifier, or IP address.
[0091] For example, in the presence of multiple computing nodes, each computing node can determine the first information of at least one processing unit it includes. Taking computing node 1, computing node 2, and computing node 3 as an example, in S601, computing node 1 determines the first information of at least one processing unit within computing node 1, computing node 2 determines the first information of at least one processing unit within computing node 2, and computing node 3 determines the first information of at least one processing unit within computing node 3. That is, each processing unit determines the first information of at least one processing unit it includes.
[0092] It is understandable that in the presence of multiple computing nodes, each computing node can perform the functions of the computing nodes described in this application. This will be stated uniformly here and will not be repeated hereafter.
[0093] S602, the compute node sends first information of at least one processing unit to the operation and maintenance node. Correspondingly, the operation and maintenance node receives first information of multiple processing units.
[0094] In the case of a single computing node, when the computing node includes multiple processing units, the computing node sends the first information of the multiple processing units to the operation and maintenance node. At this time, the first information of the multiple processing units is the first information of the multiple processing units sent by the computing node. When there are multiple computing nodes, the first information of the multiple processing units received by the operation and maintenance node includes the first information of at least one processing unit sent by each of the multiple computing nodes.
[0095] For example, taking a computing node comprising computing node 1, computing node 2, and computing node 3 as an example, in S602, computing node 1 sends first information of at least one processing unit in computing node 1, computing node 2 sends first information of at least one processing unit in computing node 2, and computing node 3 sends first information of at least one processing unit in computing node 3. Correspondingly, the maintenance node receives first information from at least one processing unit in computing node 1 from computing node 1, and also receives first information from at least one processing unit in computing node 2 from computing node 2, and also receives first information from at least one processing unit in computing node 3 from computing node 3. That is, the first information of multiple processing units received by the maintenance node includes the first information of at least one processing unit in computing node 1, the first information of at least one processing unit in computing node 2, and the first information of at least one processing unit in computing node 3.
[0096] In one possible implementation, the computing node can send the first information of at least one processing unit to the operation and maintenance node based on a preset period, and the operation and maintenance node receives the first information of at least one processing unit from the computing node based on the preset period; or, the computing node can immediately (in real time) send the first information of a processing unit to the operation and maintenance node each time it records the first information of a processing unit, and the operation and maintenance node receives the first information of at least one processing unit from the computing node in real time; or, the operation and maintenance node can send an instruction message to the computing node, instructing the computing node to report the first information of at least one processing unit, and the computing node receives the instruction message and, based on the instruction message, sends the first information of at least one processing unit to the computing node, and the operation and maintenance node receives the first information of at least one processing unit from the computing node.
[0097] S603. The operation and maintenance node determines the computing resources corresponding to the first training job based on the first information from multiple processing units.
[0098] The training jobs run by multiple processing units include a first training job, for example, training job 1. The training jobs run by multiple processing units may include training job 1, training job 2 and training job 3.
[0099] In some embodiments, the computing resources corresponding to the first training job include processing units that run the first training job. For example, taking training job 1 as the first training job, if training job 1 runs on processing units 1 to 8 of computing node 1, then the computing resources corresponding to training job 1 include processing units 1 to 8 that run training job 1.
[0100] Optionally, the processing unit running the first training job may implicitly indicate the computing node to which it belongs. For example, in conjunction with the above example, processing units 1 to 8 running training job 1 may implicitly indicate the computing node to which they belong, that is, the computing node to which processing units 1 to 8 belong is computing node 1.
[0101] Furthermore, the computing resources also include the computing nodes where the processing units running the first training job are located. For example, as shown in Figure 7, taking training job 1 as an example, if training job 1 runs on processing units 1 to 4 of computing node 1 and processing units 1 to 4 of computing node 2, then the computing resources corresponding to training job 1 include processing units 1 to 4 of computing node 1 and processing units 1 to 4 of computing node 2.
[0102] For example, if training job 1 runs on all processing units of computing node 1 and computing node 2, the computing resources corresponding to training job 1 include computing node 1 and computing node 2 running training job 1.
[0103] As one possible implementation, the implementation process of S603 may include: the operation and maintenance node dividing multiple processing units into at least one processing unit group based on the identification information of the training jobs run by the processing units, wherein the identification information of the training jobs run by processing units in the same processing unit group is the same. Then, the operation and maintenance node determines that the processing units running the first training job include the processing units in the first processing unit group, wherein the first processing unit group is a group of at least one processing unit group, and the identification information of the training jobs run by the processing units in the first processing unit group is the identification information of the first training job.
[0104] For example, the operation and maintenance node receives the first information of processing units 1 to 8 of computing node 1, the first information of processing units 1 to 8 of computing node 2, and the first information of processing units 1 to 8 of computing node 3. Based on the identification information of the training jobs running by the processing units, the operation and maintenance node groups the processing units 1 to 8 of computing node 1, the processing units 1 to 8 of computing node 2, and the processing units 1 to 8 of computing node 3. Referring to Table 1 below, the identification information of the training jobs in the first information of processing units 1 to 8 of computing node 1 and the first information of processing units 1 to 8 of computing node 2 is the identifier of training job 1, and the identification information of the training jobs in the first information of processing units 1 to 8 of computing node 3 is the identifier of training job 2. Therefore, the processing units 1 to 8 of computing node 1 and the processing units 1 to 8 of computing node 2 can be divided into processing unit group 1, and the processing units 1 to 8 of computing node 3 can be divided into processing unit group 2.
[0105] Table 1
[0106] Therefore, it can be seen that the processing units in processing unit group 1 are all used to run training job 1, and the processing units in processing unit group 2 are all used to run training job 2, thus realizing the determination of the computing resources for a training job.
[0107] It should be understood that after the operation and maintenance node groups multiple processing units by the identification information of the training job, it can determine that each processing unit group corresponds to a training job. In this way, the computing resources of a training job can be determined as the processing units in the processing unit group corresponding to the training job, or the computing resources corresponding to a training job can be determined as the computing nodes corresponding to the processing unit group corresponding to the training job (for example, if processing unit group 1 includes all processing units of computing node 1 and all processing units of computing node 2, then the computing resources corresponding to a training job can be determined as the computing nodes corresponding to processing unit group 1 and computing node 2).
[0108] In the above technical solution, the computing node determines and sends first information of at least one processing unit, and the operation and maintenance node receives first information of multiple processing units. Since the first information of a processing unit includes the identification information of the training job running by that processing unit and the identification information of the computing node where that processing unit resides, the operation and maintenance node can determine the processing unit running the first training job from among the multiple processing units through the identification information of the first training job. Furthermore, by combining this with the identification information of the computing node where the processing unit corresponding to the first training job resides, the operation and maintenance node can determine the computing node corresponding to the first training job (i.e., the computing node running the first training job, which can also be called the computing resource). Thus, this application can determine the computing resources of a training job through information interaction between the computing node and the operation and maintenance node, without relying on the AI platform or cluster management node, avoiding dependence on the AI platform and cluster management node for determining computing resources, and improving the flexibility of job restoration.
[0109] As one possible embodiment of this application, when all processing units of a computing node are running a training job, as shown in FIG8, the communication method provided by this application may include the following steps:
[0110] S801, Computing node 1 determines the first information of at least one processing unit in computing node 1 based on the data acquisition function, and computing node 2 determines the first information of at least one processing unit in computing node 2 based on the data acquisition function.
[0111] The first information includes the identification information of the training job run by the processing unit (the main process of the processing unit) and the identification information of the computing node where the processing unit is located. For a detailed explanation of the identification information of the training job run by the processing unit and the identification information of the computing node where the processing unit is located, please refer to S601 above; it will not be repeated here.
[0112] Optionally, the first information may also include the identifier of the main process of the processing unit and the number of processing units corresponding to the training jobs run by the processing unit.
[0113] For example, compute node 1 comprises 8 processing units (numbered 0-7 sequentially), where the main process of the second processing unit can be identified by its number 1. If training job 1 runs on 16 processing units, then the number of processing units corresponding to training job 1 is 16.
[0114] In one possible implementation, in response to a user's activation of the data acquisition functions of computing node 1 and computing node 2, or when the data acquisition functions of computing node 1 and computing node 2 are configured to be enabled by default, the main process within each processing unit of computing node 1 can record the identification information of the training job run by the processing unit, the identification of the main process of the processing unit, and the number of processing units corresponding to the training job run by the processing unit. Then, computing node 1 determines the information recorded by the processing unit and the identification of computing node 1 as the first information of that processing unit.
[0115] S802, Computing node 1 sends first information of at least one processing unit to the operation and maintenance node, and correspondingly, the operation and maintenance node receives first information of at least one processing unit from computing node 1; Computing node 2 sends first information of at least one processing unit, and correspondingly, the operation and maintenance node receives first information of at least one processing unit from computing node 2.
[0116] The first information received by the operation and maintenance node from multiple processing units includes the first information of at least one processing unit of computing node 1 and the first information of at least one processing unit of computing node 2.
[0117] Optionally, the specific implementation of the computing node reporting the first information of at least one processing unit in S802 can be referred to the embodiment shown in S602, and will not be repeated here.
[0118] S803. The operation and maintenance node divides multiple processing units into at least one processing unit group based on the identification information of the training jobs run by the processing unit.
[0119] Optionally, the specific implementation of S803 can be referred to the embodiment shown in S603, which will not be repeated here.
[0120] S804. The operation and maintenance node determines the processing unit that runs the first training job, including the processing units in the first processing unit group.
[0121] Wherein, the first processing unit group is a processing unit group in at least one processing unit group, and the identification information of the training job run by the processing unit in the first processing unit group is the identification information of the first training job.
[0122] Optionally, for the first training job, the operation and maintenance node can also determine whether the processing units in the first processing unit group corresponding to the first training job are complete, or determine whether the first information of the processing units in the first processing unit group corresponding to the first training job is complete. If complete, it can be determined that the processing units running the first training job include the processing units in the first processing unit group.
[0123] In one possible implementation, the implementation process of S804 may include: when the number of processing units in the first processing unit group is equal to the number of processing units corresponding to the first training job, determining that the processing units running the first training job include the processing units in the first processing unit group.
[0124] In one example, taking the first training job as training job 1, and the processing units in processing unit group 1 as running training job 1, if the number of processing units corresponding to training job 1 is equal to the number of processing units in processing unit group 1, that is, processing unit group 1 includes all processing units running training job 1, then it can be confirmed that the processing units in processing unit group 1 are complete, and thus it can be determined that the processing units running training job 1 include the processing units in processing unit group 1.
[0125] Alternatively, in another possible implementation, the implementation process of S704 may include: if the identifier of the main process of the processing unit in the first processing unit group includes all natural numbers within a first range, determining that the processing unit running the first training job includes the processing units in the first processing unit group. Here, the first range is [0, N-1], and N is the number of processing units corresponding to the first training job.
[0126] In one example, the first training job is training job 1. The processing units in processing unit group 1 are used to run training job 1. The identifiers of the main processes of the processing units in processing unit group 1 include 0, 1, 2, 3, 4, 5, 6, 7. Taking the first range as [0,7] as an example, in this example, the identifiers of the main processes of the processing units in processing unit group 1 include all natural numbers in the first range. That is to say, processing unit group 1 includes all processing units that run training job 1. Therefore, it can be confirmed that the processing units in processing unit group 1 are complete, and thus it can be determined that the processing units that run training job 1 include the processing units in processing unit group 1.
[0127] Alternatively, in another possible implementation, the implementation process of S804 may include: if the number of processing units in the first processing unit group is not equal to the number of processing units corresponding to the first training job, or if the identifier of the main process of the processing unit in the first processing unit group does not include all natural numbers within the first range, the operation and maintenance node determines that the processing units in the first processing unit group are incomplete and needs to continue waiting for the computing node running the first training job to report the first information until the number of processing units in the first processing unit group is equal to the number of processing units corresponding to the first training job, and / or the identifier of the main process of the processing unit in the first processing unit group includes all natural numbers within the first range.
[0128] Thus, in the above technical solution, the computing resources of the training job can be determined through information interaction between the computing node and the operation and maintenance node. Then, the operation and maintenance node can diagnose the first training job based on the logs, alarms and indicators collected from all computing resources corresponding to the first training job.
[0129] As another possible embodiment of this application, in the case where the processing units in the computing node may run different training jobs at different times, as shown in FIG9, the communication method provided by this application may include the following steps:
[0130] S901, Computing node 1 determines the first information of at least one processing unit in computing node 1 based on the data acquisition function, and computing node 2 determines the first information of at least one processing unit in computing node 2 based on the data acquisition function.
[0131] The first information includes the identification information of the training job run by the processing unit (the main process of the processing unit) and the identification information of the computing node where the processing unit is located. For a detailed explanation of the identification information of the training job run by the processing unit and the identification information of the computing node where the processing unit is located, please refer to S601 above; it will not be repeated here.
[0132] Optionally, the first information also includes the identifier of the main process of the processing unit and the number of processing units corresponding to the training jobs run by the processing unit. For a detailed explanation of the identifier of the main process of the processing unit and the number of processing units corresponding to the training jobs run by the processing unit, please refer to S801 above, and it will not be repeated here.
[0133] Optionally, the first information also includes the timestamp corresponding to the processing unit. The timestamp corresponding to the processing unit is either the timestamp at which the main process of the processing unit starts running the training job (startTimeStamp), or a set of timestamps at which multiple processes of the processing unit start running the training job.
[0134] Optionally, when a computing node reports the first information of multiple processing units, if the first information also includes the timestamps corresponding to the processing units, the computing node can encapsulate the timestamps corresponding to the multiple processing units into a timestamp group.
[0135] Optionally, the first information may also include identification information of the processing unit. For example, the processing unit may be a logical NPU or a physical NPU, and the identification information of the processing unit may be the identifier of the logical NPU or the identifier of the physical NPU.
[0136] S902, Computing node 1 sends first information of at least one processing unit to the operation and maintenance node, and correspondingly, the operation and maintenance node receives first information of at least one processing unit from computing node 1; Computing node 2 sends first information of at least one processing unit, and correspondingly, the operation and maintenance node receives first information of at least one processing unit from computing node 2.
[0137] The first information of the multiple processing units includes the first information of at least one processing unit of computing node 1 and the first information of at least one processing unit of computing node 2.
[0138] In some embodiments, if the processing unit 1 in computing node 1 records multiple pieces of first information (that is, the first information of the identification information of multiple processing units 1 in computing node 1, wherein the first information can be stored in computing node 1 in the form of a disk file), the first information of the processing unit is the first information with the largest timestamp among the multiple pieces of first information corresponding to the processing unit.
[0139] For example, taking computing node 1 as an example, if there are multiple first information in computing node 1 where the identification information of the processing unit is the identification of processing unit 1, then for processing unit 1, the first information reported by the computing node is the first information with the largest timestamp among the multiple first information.
[0140] It should be understood that the first piece of information with the largest timestamp among the multiple pieces of first information in processing unit 1 is the information recorded by the processing unit when running the latest training job. Based on the timestamp in the first piece of information, computing node 1 can report the first piece of information with the largest timestamp so that the operation and maintenance node can know the information of the latest job currently running by processing unit 1.
[0141] It should be noted that when processing unit 1 of computing node 1 runs multiple training jobs (e.g., training job 1, training job 2, and training job 3), processing unit 1 records the first information when running training job 1, training job 2, and training job 3, and the identification information of the processing unit in the first information is the identification of processing unit 1. Thus, for processing unit 1, computing node 1 will have multiple instances of the same first information with the same identification information for multiple processing units. Optionally, the specific implementation of the computing node reporting the first information of at least one processing unit in S902 can be referred to the embodiment shown in S602, and will not be repeated here.
[0142] S903. The operation and maintenance node divides multiple processing units into at least one processing unit group based on the identification information of the training jobs run by the processing unit.
[0143] Among them, the training jobs run by processing units in the same processing unit group have the same identification information.
[0144] Optionally, the specific implementation of S903 can be referred to the embodiment shown in S603, which will not be repeated here.
[0145] Optionally, after determining at least one processing unit group, for each group within the at least one processing unit group, the correspondence between the identifier of the main process of the processing unit and the timestamp corresponding to the processing unit can be recorded. Alternatively, the identifier information of the training job (masterIP) and the timestamp corresponding to any processing unit in the corresponding group can be used as the unique identifier of the training job. If the computing cluster is divided into multiple sub-clusters, since the masterIP may be duplicated in different sub-clusters, the identifier of the training job (masterIP), the identifier of the sub-cluster, and the timestamp corresponding to any processing unit in the corresponding group can be used as the unique identifier of the training job.
[0146] Optionally, when executing S903, the operation and maintenance node can also record the correspondence between the identifier of the main process of the processing unit, the identifier information of the processing unit, and the identifier information of the computing node where the processing unit is located. Thus, the processing unit group can include the identifier information of the processing unit, as well as the information of the main process and the computing node corresponding to the processing unit. This allows the operation and maintenance node to determine the processing unit where the first training job is located and the computing nodes corresponding to each processing unit; or, if multiple training jobs are running simultaneously in a processing unit within a computing node, the operation and maintenance node can distinguish between different processing units corresponding to a single computing node.
[0147] In one example, referring to Figure 7, taking the processing unit group corresponding to training job 1 as processing unit group 1 as an example, processing unit group 1 may include processing unit 1, processing unit 2, processing unit 3, and processing unit 4 in computing node 1, and processing unit 1, processing unit 2, processing unit 3, and processing unit 4 in computing node 2. The operation and maintenance node can record the correspondence between the identifier of processing unit 1 and computing node 1, the correspondence between processing unit 2 and computing node 1, the correspondence between processing unit 3 and computing node 1, the correspondence between processing unit 4 and computing node 1, the correspondence between the identifier of processing unit 1 and computing node 2, the correspondence between processing unit 2 and computing node 2, the correspondence between processing unit 3 and computing node 2, and the correspondence between processing unit 4 and computing node 2. In this way, the operation and maintenance node can determine the computing node corresponding to the processing unit in processing unit group 1. The masterIP corresponding to training job 1 can be the IP address of computing node 1, and the masterIP corresponding to training job 2 can be the IP address of computing node 2.
[0148] S904. The operation and maintenance node determines whether the processing unit in the first processing unit group is the same as the processing unit running the second training job, and whether the timestamp corresponding to the processing unit in the first processing unit group is the same as the timestamp corresponding to the processing unit of the second training job.
[0149] The identification information of the second training job is the same as that of the first training job (that is, the computing nodes running the second training job are the same as those running the first training job). The identification information of the first training job is the IP address of the computing node where the master process of the first training job is located, and the identification information of the second training job is the IP address of the computing node where the master process of the second training job is located.
[0150] In one possible implementation, if the processing unit in the first processing unit group is different from the processing unit running the second training job, and / or the timestamp corresponding to the processing unit in the first processing unit group is different from the timestamp corresponding to the processing unit of the second training job, it is determined that the training job currently running by the processing unit in the first processing unit group is different from the second training job previously run in the first processing unit group, that is, the first training job is a new job.
[0151] For example, if the processing unit in the first processing unit group is different from the processing unit running the second training job, it can be determined that the training job currently running by the processing unit in the first processing unit group is different from the second training job previously run in the first processing unit group; that is, the first training job is a new job. For instance, if the processing units running the second training job include processing units 1 to 4 in computing node 1, and the processing units in the first processing unit group are processing units 5 to 8 in computing node 1, then the job currently running in the first processing unit group is not the second training job.
[0152] Alternatively, if the timestamp corresponding to a processing unit in the first processing unit group is different from the timestamp corresponding to a processing unit in the second training job, it can be determined that the training job currently running in the first processing unit group is different from the second training job previously run in the first processing unit group; that is, the first training job is a new job. For example, if the timestamp corresponding to the processing unit running the second training job is time 1, and the timestamp corresponding to the processing unit in the first processing unit group is time 2, it proves that the currently running job in the first processing unit group and the second training job did not start running at the same time, thus indicating that the currently running first training job in the first processing unit group is not the second training job. As another example, if the timestamps corresponding to the processing units running the second training job include time 1 (corresponding to process 1), time 1 (corresponding to process 2), and time 1 (corresponding to process 3), and the timestamps corresponding to the processing units in the first processing unit group are time 2 (corresponding to process 1), time 2 (corresponding to process 2), and time 2 (corresponding to process 3), it proves that the currently running job in the first processing unit group and the second training job did not start running at the same time, thus indicating that the currently running first training job in the first processing unit group is not the second training job.
[0153] Alternatively, if the processing unit in the first processing unit group is different from the processing unit running the second training job, and the timestamp corresponding to the processing unit in the first processing unit group is different from the timestamp corresponding to the processing unit of the second training job, then it is determined that the training job currently running by the processing unit in the first processing unit group is different from the second training job that has been previously run in the first processing unit group, that is, the first training job is a new job.
[0154] Based on the above examples, it can be seen that if either of the following conditions is different, the first training job is a new job: the processing unit in the first processing unit group is different from the processing unit running the second training job (denoted as condition 1) or the timestamp corresponding to the processing unit in the first processing unit group is different from the timestamp corresponding to the processing unit of the second training job (denoted as condition 2).
[0155] In another possible implementation, if the processing unit in the first processing unit group is the same as the processing unit running the second training job, and the timestamp corresponding to the processing unit in the first processing unit group is the same as the timestamp corresponding to the processing unit of the second training job, it indicates that the processing unit running the second training job is the same as the processing unit running the first training job, and the start time of the first training job and the start time of the second training job are also the same. Thus, it can be determined that the first training job currently being run by the processing unit in the first processing unit group is the same as the second training job previously run in the first processing unit group; that is, the first training job is an existing job.
[0156] S905. If the processing unit in the first processing unit group of the operation and maintenance node is different from the processing unit running the second training job, and / or the timestamp corresponding to the processing unit in the first processing unit group is different from the timestamp corresponding to the processing unit of the second training job, then it is determined that the processing unit running the first training job includes the processing unit in the first processing unit group.
[0157] The identification information of the second training job is the same as that of the first training job. The identification information of the first training job is the IP address of the computing node where the master process of the first training job is located, and the identification information of the second training job is the IP address of the computing node where the master process of the second training job is located.
[0158] As one possible implementation, if the processing unit in the first processing unit group is different from the processing unit running the second training job, and / or the timestamp corresponding to the processing unit in the first processing unit group is different from the timestamp corresponding to the processing unit of the second training job, the operation and maintenance node can determine that the first training job is a new job, and then determine that the processing unit running the first training job (new job) includes the processing unit in the first processing unit group.
[0159] It should be understood that if the first training job is an existing job, the processing unit that runs the first training job has already been determined in previous cycles. Therefore, the processing unit that runs the first training job only needs to be determined when the first training job is a new job, which can save resources.
[0160] Of course, this application may also determine that the processing unit running the first training job includes the processing unit in the first processing unit group if the processing unit in the first processing unit group is the same as the processing unit running the second training job, and the timestamp corresponding to the processing unit in the first processing unit group is the same as the timestamp corresponding to the processing unit of the second training job.
[0161] Optionally, for the operation and maintenance node, if the processing unit running the first training job is determined to include a first processing unit group, and the processing units in the first processing unit group are complete, and the processing units in the first processing unit group are different from the processing units running the second training job, and / or the timestamps corresponding to the processing units in the first processing unit group are different from the timestamps corresponding to the processing units in the second training job, then the processing unit running the first training job is determined to include the processing units in the first processing unit group. The specific implementation process for determining whether the processing units in the first processing unit group are complete refers to the embodiment shown in S804, and will not be repeated here.
[0162] In the above technical solution, the operation and maintenance node can determine whether the first training job is a new job or an existing job, and for a new job, determine that the processing unit running the first training job includes the processing units in the first processing unit group. In this way, it has the effect of saving resources.
[0163] As another possible embodiment of this application, when the processing unit in the computing node runs multiple training jobs simultaneously, as shown in FIG10, the communication method provided by this application may include the following steps:
[0164] S1001, Computing node 1 determines the first information of at least one processing unit in computing node 1 based on the data acquisition function, and computing node 2 determines the first information of at least one processing unit in computing node 2 based on the data acquisition function.
[0165] The first information includes the identification information of the training job run by the processing unit (the main process of the processing unit) and the identification information of the computing node where the processing unit is located. For a detailed explanation of the identification information of the training job run by the processing unit and the identification information of the computing node where the processing unit is located, please refer to S601 above; it will not be repeated here.
[0166] The first information also includes the identification information of the processing unit. For example, the processing unit may be a physical NPU, and the identifier of the processing unit may be the identifier of the physical NPU.
[0167] Optionally, the first information also includes the identifier of the main process of the processing unit and the number of processing units corresponding to the training jobs run by the processing unit. For a detailed explanation of the identifier of the main process of the processing unit and the number of processing units corresponding to the training jobs run by the processing unit, please refer to S801 above, and it will not be repeated here.
[0168] It should be understood that the first information includes both the identification information of the processing unit and the identification information of the computing node where the processing unit is located. Thus, it can be assumed that there is a corresponding relationship between the identification information of the processing unit and the identification information of the computing node where the processing unit is located in the same first information, so that the computing resources of the first training job can be determined subsequently based on the identification information of the processing unit and the identification information of the computing node where the processing unit is located.
[0169] S1002, Computing node 1 sends first information of at least one processing unit to the operation and maintenance node, and correspondingly, the operation and maintenance node receives first information of at least one processing unit from computing node 1; Computing node 2 sends first information of at least one processing unit, and correspondingly, the operation and maintenance node receives first information of at least one processing unit from computing node 2.
[0170] The first information of the multiple processing units includes the first information of at least one processing unit of computing node 1 and the first information of at least one processing unit of computing node 2.
[0171] Optionally, the specific implementation of the computing node reporting the first information of at least one processing unit in S1002 can be referred to the embodiment shown in S602, and will not be repeated here.
[0172] S1003. The operation and maintenance node divides multiple processing units into at least one processing unit group based on the identification information of the training jobs run by the processing unit.
[0173] Optionally, when multiple training jobs are running simultaneously in the processing units of a computing node, the identification information of the multiple training jobs is different. For example, if training job 1 and training job 2 are running simultaneously in computing node 1 and computing node 2, the identification information of training job 1 (the masterIP corresponding to training job 1) can be the IP of computing node 1, and the identification information of training job 2 (the masterIP corresponding to training job 2) can be the IP of computing node 2.
[0174] Optionally, the specific implementation of S1003 can be referred to the embodiment shown in S603, which will not be repeated here.
[0175] Optionally, when executing S1003, the operation and maintenance node can also record the correspondence between the identifier of the main process of the processing unit, the identifier information of the processing unit, and the identifier information of the computing node where the processing unit is located. In this way, the processing unit group can include the identifier information of the processing unit, as well as the information of the main process and the computing node corresponding to the processing unit, so that the operation and maintenance node can know the processing unit where the first training job is located and the computing node corresponding to each processing unit.
[0176] In one example, referring to Figure 7, taking the processing unit group corresponding to training job 1 as processing unit group 1 as an example, processing unit group 1 may include processing unit 1, processing unit 2, processing unit 3, and processing unit 4 in computing node 1, and processing unit 1, processing unit 2, processing unit 3, and processing unit 4 in computing node 2. The operation and maintenance node can record the correspondence between the identifier of processing unit 1 and computing node 1, the correspondence between processing unit 2 and computing node 1, the correspondence between processing unit 3 and computing node 1, the correspondence between processing unit 4 and computing node 1, the correspondence between the identifier of processing unit 1 and computing node 2, the correspondence between processing unit 2 and computing node 2, the correspondence between processing unit 3 and computing node 2, and the correspondence between processing unit 4 and computing node 2. In this way, the operation and maintenance node can determine the computing node corresponding to the processing unit in processing unit group 1. The masterIP corresponding to training job 1 can be the IP address of computing node 1, and the masterIP corresponding to training job 2 can be the IP address of computing node 2.
[0177] S1004. The operation and maintenance node determines the processing unit that runs the first training job, including the processing units in the first processing unit group.
[0178] Optionally, the specific implementation process of determining the processing unit for running the first training job for the operation and maintenance node, including the processing unit in the first processing unit group, refers to the embodiment shown in S804, and will not be repeated here.
[0179] In the above technical solution, when the processing unit is running in a part of the processing unit in the computing node, the computing node corresponding to the processing unit in the processing unit group can be determined by the identification information of the processing unit in the first information and the identification information of the computing node where the processing unit is located.
[0180] The method provided in this application has been described above. In addition, this application also provides a communication device for implementing the functions described in the above method embodiments.
[0181] It is understood that, in order to achieve the aforementioned functions, the communication device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0182] This application embodiment can divide the communication device into functional modules according to the above method embodiment. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0183] Figure 11 shows a schematic diagram of a communication device 110. The communication device 110 includes a processing module 1101 and a transceiver module 1102. The communication device 110 can be used to implement the functions of the aforementioned operation and maintenance node or computing node.
[0184] In some embodiments, the communication device 110 may further include a storage module (not shown in FIG11) for storing program instructions and data.
[0185] In some embodiments, the transceiver module 1102, also referred to as a transceiver unit, is used to implement sending and / or receiving functions. The transceiver module 1102 may consist of a transceiver circuit, a transceiver, a transceiver unit, or a communication interface.
[0186] In some embodiments, the transceiver module 1102 may include a receiving module and a sending module, respectively used to execute the receiving and sending steps performed by the operation and maintenance node or the computing node in the above method embodiments, and / or other processes used to support the technology described herein; the processing module 1101 may be used to execute the processing steps performed by the operation and maintenance node or the computing node in the above method embodiments, and / or other processes used to support the technology described herein.
[0187] When the communication device 110 is used to implement the functions of the operation and maintenance node:
[0188] In one possible implementation: transceiver module 1102 is used to receive first information from multiple processing units, the first information of the first processing unit includes the identification information of the training job run by the first processing unit and the identification information of the computing node where the first processing unit is located, the first processing unit being any one of the multiple processing units; processing module 1101 is used to determine the computing resources corresponding to the first training job based on the first information of the multiple processing units, the training jobs run by the multiple processing units including the first training job.
[0189] In one possible implementation, the processing module 1101 is specifically configured to: divide multiple processing units into at least one processing unit group according to the identification information of the training jobs run by the processing units, wherein the identification information of the training jobs run by the processing units in the same processing unit group is the same; determine that the processing unit running the first training job includes the processing units in the first processing unit group, the first processing unit group is a processing unit group in at least one processing unit group, and the identification information of the training jobs run by the processing units in the first processing unit group is the identification information of the first training job.
[0190] In one possible implementation, the processing module 1101 is specifically used to: determine that the processing unit running the first training job includes the processing unit in the first processing unit group when the number of processing units in the first processing unit group is equal to the number of processing units corresponding to the first training job, or when the identifier of the main process of the processing unit in the first processing unit group includes all natural numbers in a first range; wherein the first range is [0, N-1], and N is the number of processing units corresponding to the first training job.
[0191] In one possible implementation, the processing module 1101 is specifically configured to: determine that the processing unit running the first training job includes the processing unit in the first processing unit group if the processing unit in the first processing unit group is different from the processing unit running the second training job, and / or if the timestamp corresponding to the processing unit in the first processing unit group is different from the timestamp corresponding to the processing unit of the second training job; wherein the identification information of the second training job is the same as the identification information of the first training job, the identification information of the first training job is the Internet Protocol IP address of the computing node where the master process of the first training job is located, and the identification information of the second training job is the IP address of the computing node where the master process of the second training job is located.
[0192] When the communication device 110 is used to implement the functions of a computing node:
[0193] In one possible implementation: processing module 1101 is used to determine first information of at least one processing unit; the first information of the second processing unit includes the identification information of the training job run by the second processing unit and the identification information of the computing node where the second processing unit is located, and the second processing unit is any one of the at least one processing units; transceiver module 1102 is used to send the first information of at least one processing unit, and the first information of at least one processing unit is used to determine the computing resources corresponding to the training job run by each of the at least one processing units.
[0194] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0195] In this application, the communication device 110 can be presented in an integrated manner by dividing it into various functional modules. Here, "module" can refer to an application-specific integrated circuit (ASIC), a circuit, a processor and memory that executes one or more software or firmware programs, integrated logic circuits, and / or other devices that can provide the above functions.
[0196] In some embodiments, when the communication device 110 in FIG11 is a chip or chip system, the function / implementation process of the transceiver module 1102 can be implemented through the input / output interface (or communication interface) of the chip or chip system, and the function / implementation process of the processing module 1101 can be implemented through the processor (or processing circuit) of the chip or chip system.
[0197] Since the communication device 110 provided in this embodiment can execute the above method, the technical effects it can achieve can be referred to the above method embodiment, and will not be repeated here.
[0198] As a possible product form, the terminal or computing node described in the embodiments of this application can be implemented using one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this application.
[0199] As another possible product form, the terminal or computing node in this application may adopt the composition structure shown in FIG12, or include the components shown in FIG12. FIG12 is a schematic diagram of the composition of a communication device 1200 provided in this application. The communication device 1200 may be a terminal or a chip or system-on-a-chip in a terminal; or, it may be a computing node or a module or chip or system-on-a-chip in a computing node.
[0200] As shown in Figure 12, the communication device 1200 includes at least one processor 1201 and at least one communication interface (Figure 12 is merely an example illustrating the inclusion of a communication interface 1204 and a processor 1201). Optionally, the communication device 1200 may also include a communication bus 1202 and a memory 1203.
[0201] Processor 1201 can be a general-purpose central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a PLD, or any combination thereof. Processor 1201 can also be other devices with processing capabilities, such as circuits, devices, or software modules, without limitation. As one possible implementation, processor 1201 may include one or more CPUs, such as CPU0 and CPU1 in Figure 12.
[0202] The communication bus 1202 is used to connect different components in the communication device 1200, enabling communication between them. The communication bus 1202 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in Figure 12, but this does not indicate that there is only one bus or one type of bus.
[0203] Communication interface 1204 is used for communicating with other devices or communication networks. For example, communication interface 1204 can be a module, circuit, transceiver, or any device capable of communication. Optionally, the communication interface 1204 can also be an input / output interface located within processor 1201, used to implement signal input and signal output for the processor.
[0204] The memory 1203 may be a device with storage function, used to store instructions and / or data. The instructions may be computer programs.
[0205] For example, the memory 1203 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and / or instructions; it may also be a random access memory (RAM) or other type of dynamic storage device capable of storing information and / or instructions; it may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, etc., without limitation.
[0206] It should be noted that the memory 1203 can exist independently of the processor 1201, or it can be integrated with the processor 1201. The memory 1203 can be located inside or outside the communication device 1200, without limitation. The processor 1201 can be used to execute the instructions stored in the memory 1203 to implement the methods provided in the following embodiments of this application.
[0207] As an optional implementation, the communication device 1200 may also include an output device 1205 and an input device 1206. The output device 1205 communicates with the processor 1201 and can display information in various ways. For example, the output device 1205 may be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 1206 communicates with the processor 1201 and can receive user input in various ways. For example, the input device 1206 may be a mouse, keyboard, touchscreen device, or sensing device, etc.
[0208] In some embodiments, those skilled in the art will recognize that the communication device 110 shown in FIG11 can take the form of the communication device 1200 shown in FIG12 in terms of hardware implementation.
[0209] As an example, the function / implementation process of the processing module 1101 in Figure 11 can be implemented by the processor 1201 in the communication device 1200 shown in Figure 12 calling computer execution instructions stored in the memory 1203. The function / implementation process of the transceiver module 1102 in Figure 11 can be implemented by the communication interface 1204 in the communication device 1200 shown in Figure 12.
[0210] It should be noted that the structure shown in Figure 12 does not constitute a specific limitation on the terminal or computing node. For example, in other embodiments of this application, the terminal or computing node may include more or fewer components than shown in the figure, or combine some components, or split some components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0211] In some embodiments, this application also provides a communication device, which includes a processor for implementing the methods in any of the above method embodiments.
[0212] As one possible implementation, the communication device also includes a memory. This memory stores necessary computer programs and data. The computer program may include instructions, which a processor can invoke to instruct the communication device to execute the methods described in any of the above method embodiments. Alternatively, the memory may not be present in the communication device.
[0213] As another possible implementation, the communication device also includes an interface circuit, which is a code / data read / write interface circuit, used to receive computer execution instructions (which are stored in memory and may be read directly from memory or may be transmitted through other devices) and transmit them to the processor.
[0214] As another possible implementation, the communication device also includes a communication interface for communicating with modules outside the communication device.
[0215] It is understood that the communication device can be a chip or a chip system. When the communication device is a chip system, it can be composed of chips or may include chips and other discrete devices. This application does not specifically limit this.
[0216] This application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a computer, implements the functions of any of the above-described method embodiments.
[0217] This application also provides a computer program product that, when executed by a computer, implements the functions of any of the above method embodiments.
[0218] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0219] It is understood that the systems, apparatuses, and methods described in this application can also be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0220] The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. The components shown as units may or may not be physical units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0221] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0222] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive (SSD)). In this embodiment, the computer may include the aforementioned apparatus.
[0223] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0224] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the scope of this application. Accordingly, this specification and drawings are merely illustrative descriptions of the application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Thus, if such modifications and modifications fall within the scope of the claims and their equivalents, this application is also intended to include such modifications and modifications.
Claims
1. A communication method, characterized in that, The method includes: The system receives first information from multiple processing units. The first information of the first processing unit includes the identification information of the training job being run by the first processing unit and the identification information of the computing node where the first processing unit is located. The first processing unit is any one of the multiple processing units. Based on the first information of the plurality of processing units, the computing resources corresponding to the first training job are determined, and the training jobs run by the plurality of processing units include the first training job.
2. The method according to claim 1, characterized in that, The computing resources corresponding to the first training job include the processing unit that runs the first training job.
3. The method according to claim 2, characterized in that, Based on the first information from the plurality of processing units, the computing resources corresponding to the first training job are determined, including: Based on the identification information of the training jobs run by the processing units, the plurality of processing units are divided into at least one processing unit group, wherein the identification information of the training jobs run by the processing units in the same processing unit group is the same. The processing unit that runs the first training job is determined to include the processing unit in the first processing unit group, the first processing unit group being the processing unit group in the at least one processing unit group, and the identification information of the training job run by the processing unit in the first processing unit group being the identification information of the first training job.
4. The method according to claim 3, characterized in that, The first information of the first processing unit also includes the identifier of the main process of the first processing unit and the number of processing units corresponding to the training job run by the first processing unit.
5. The method according to claim 4, characterized in that, The processing unit that runs the first training job includes the processing units in the first processing unit group, including: If the number of processing units in the first processing unit group is equal to the number of processing units corresponding to the first training job, or if the identifier of the main process of the processing unit in the first processing unit group includes all natural numbers within a first range, then the processing unit running the first training job is determined to include the processing units in the first processing unit group. Wherein, the first range is [0, N-1], and N is the number of processing units corresponding to the first training job.
6. The method according to any one of claims 3-5, characterized in that, The first information of the first processing unit also includes the timestamp corresponding to the first processing unit. The timestamp corresponding to the first processing unit is the timestamp when the main process of the first processing unit starts running the training job, or the set of timestamps when multiple processes of the first processing unit start running the training job.
7. The method according to claim 6, characterized in that, The processing unit that runs the first training job includes the processing units in the first processing unit group, including: If the processing unit in the first processing unit group is different from the processing unit running the second training job, and / or if the timestamp corresponding to the processing unit in the first processing unit group is different from the timestamp corresponding to the processing unit of the second training job, then the processing unit running the first training job is determined to include the processing unit in the first processing unit group. The identification information of the second training job is the same as that of the first training job. The identification information of the first training job is the Internet Protocol IP address of the computing node where the master process of the first training job is located, and the identification information of the second training job is the IP address of the computing node where the master process of the second training job is located.
8. The method according to any one of claims 1-7, characterized in that, The first information of the first processing unit also includes the identification information of the first processing unit.
9. A communication method, characterized in that, The method includes: First information of at least one processing unit is determined; the first information of the second processing unit includes the identification information of the training job run by the second processing unit and the identification information of the computing node where the second processing unit is located, and the second processing unit is any one of the at least one processing units; Send first information of the at least one processing unit, the first information of the at least one processing unit being used to determine the computing resources corresponding to the training job run by each of the at least one processing unit.
10. The method according to claim 9, characterized in that, The computing resources corresponding to the training job include the processing unit that runs the training job.
11. The method according to claim 9 or 10, characterized in that, The first information of the second processing unit also includes the identifier of the main process of the second processing unit and the number of processing units corresponding to the training job run by the second processing unit.
12. The method according to any one of claims 9-11, characterized in that, The first information of the second processing unit also includes the timestamp corresponding to the second processing unit. The timestamp corresponding to the second processing unit is the timestamp when the main process of the second processing unit starts running the training job, or the set of timestamps when multiple processes of the second processing unit start running the training job.
13. The method according to claim 12, characterized in that, When the second processing unit corresponds to multiple pieces of first information, the first information of the second processing unit is the first information with the largest timestamp among the multiple pieces of first information corresponding to the second processing unit.
14. The method according to any one of claims 9-13, characterized in that, The first information of the second processing unit also includes the identification information of the second processing unit.
15. A communication system, characterized in that, Including operation and maintenance nodes and computing nodes, The computing node is used to determine first information of at least one processing unit; the first information of the first processing unit includes the identification information of the training job run by the first processing unit and the identification information of the computing node where the first processing unit is located, and the first processing unit is any one of the at least one processing units. Send the first information of the at least one processing unit to the operation and maintenance node; The operation and maintenance node is used to receive the first information, determine the computing resources corresponding to the first training job based on the first information, and the training jobs run by the multiple processing units include the first training job.
16. The system according to claim 15, characterized in that, Based on the first information, the operation and maintenance unit determines the computing resources corresponding to the first training job, specifically including: Based on the identification information of the training jobs run by the processing units, the plurality of processing units are divided into at least one processing unit group, wherein the identification information of the training jobs run by the processing units in the same processing unit group is the same; it is determined that the processing unit running the first training job includes the processing units in the first processing unit group, the first processing unit group is the processing unit group in the at least one processing unit group, the identification information of the training job run by the processing units in the first processing unit group is the identification information of the first training job, and the computing resources corresponding to the first training job include the processing unit running the first training job.
17. A communication device, characterized in that, The communication device includes a module for performing the method as described in any one of claims 1-8, or includes a module for performing the method as described in any one of claims 9-14.
18. A communication device, characterized in that, The communication device includes a processor; the processor is configured to run a computer program or instructions to cause the communication device to perform the method as described in any one of claims 1-8, or to cause the communication device to perform the method as described in any one of claims 9-14.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions or programs that, when executed on a computer, cause the method described in any one of claims 1-8 to be performed, or cause the method described in any one of claims 9-14 to be performed.
20. A computer program product, characterized in that, The computer program product includes computer instructions; when some or all of the computer instructions are run on a computer, they cause the method of any one of claims 1-8 to be performed, or cause the method of any one of claims 9-14 to be performed.