Model training method and device, equipment, storage medium and program product

By using a distributed interface to transfer data and allocate computing tasks among multiple computing devices during deep learning model training, the problem of low model training efficiency in existing technologies is solved, and efficient processing of large dataset training is achieved.

CN115700607BActive Publication Date: 2026-04-21SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
Filing Date
2022-06-06
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

During the training of deep learning models, existing technologies struggle to effectively optimize and update model parameters to improve model accuracy.

Method used

Data is transferred between multiple computing devices/nodes through distributed interfaces, and computing tasks of individual operators are allocated to enable training of large datasets. Data broadcasting, reduction, scattering, collection, and reduction broadcast interfaces are used to improve data transmission efficiency.

Benefits of technology

It improves the overall efficiency of deep learning model training and supports training tasks on large datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115700607B_ABST
    Figure CN115700607B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a model training method and device, equipment, a storage medium and a program product, wherein the method comprises: a distributed interface receiving first to-be-transmitted data of a first process in at least two distributed processes; the distributed interface transmitting the first to-be-transmitted data or second to-be-transmitted data determined based on the to-be-transmitted data to a second process in the at least two distributed processes; the first process and the second process are at least one of the at least two distributed processes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to, but is not limited to, the field of artificial intelligence technology, and in particular to a model training method, apparatus, device, storage medium, and program product. Background Technology

[0002] In the process of training a deep learning model, in order to make the output of the deep learning model closer to the real label and improve the accuracy of the deep learning model, it is necessary to build an optimizer to optimize and update the values ​​of the learnable parameters in the deep learning model. Summary of the Invention

[0003] In view of the above, the present disclosure provides at least one model training method, apparatus, device, storage medium, and program product.

[0004] The technical solution of this disclosure embodiment is implemented as follows:

[0005] On one hand, this disclosure provides a model training method, characterized in that the method includes:

[0006] The distributed interface receives the first data to be transmitted from the first process of at least two distributed processes.

[0007] The distributed interface transmits the first data to be transmitted or the second data to be transmitted, determined based on the data to be transmitted, to the second process among the at least two distributed processes.

[0008] The first process and the second process are at least one of the at least two distributed processes.

[0009] In some embodiments, the first process is the main process, and the second process is at least one of the other processes among the at least two distributed processes besides the main process; or...

[0010] The second process is the main process, and the first process is at least one of the other processes among the at least two distributed processes besides the main process; or,

[0011] The first process is a first part of the at least two distributed processes, and the first process is a second part of the at least two distributed processes. The first part process and the second part process are unrelated.

[0012] In some embodiments, the distributed interface includes at least one of the following: a data broadcast interface, a data reduction interface, a data scattering interface, a data collection interface, a data reduction broadcast interface, a data collection broadcast interface, a communication synchronization interface, a process ID acquisition interface, and a process count acquisition interface.

[0013] In some embodiments, the distributed interface includes the data broadcast interface, and the first data to be transmitted includes first tensor data to be transmitted; the distributed interface transmits the first data to be transmitted or a second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes, including:

[0014] The data broadcast interface receives the first tensor data to be transmitted from the first process and the first process identifier of the first process.

[0015] The data broadcast interface transmits the first tensor data to be transmitted to the second process; the second process is all processes in the process group corresponding to the first process identifier.

[0016] In some embodiments, the data broadcast interface receives the first tensor data to be transmitted from the first process; including:

[0017] The data broadcast interface receives the first tensor data to be transmitted from the first process, the first process identifier of the first process, and the first process group identifier of the first process.

[0018] The data broadcast interface transmits the first tensor data to be transmitted to the second process; the second process is all processes in the process group determined based on the first process group identifier.

[0019] In some embodiments, the distributed interface includes a data reduction interface, the first data to be transmitted includes second tensor data to be transmitted, and the second data to be transmitted includes third tensor data to be transmitted; the distributed interface transmits the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes, including:

[0020] The data reduction interface receives the second tensor data to be transmitted and the second process identifier from the first process.

[0021] The data protocol interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier.

[0022] In some embodiments, the data reduction interface receives the second tensor data to be transmitted and the second process identifier of the first process, including: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, and the reduction algorithm of the first process;

[0023] The data reduction interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier, including: the data reduction interface performs reduction calculation on the second tensor data to be transmitted based on the reduction algorithm to obtain the third tensor data, and transmits the third tensor data to the second process determined based on the second process identifier.

[0024] In some embodiments, the data reduction interface receives the second tensor data to be transmitted and the second process identifier of the first process, including: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, and the second process group identifier of the process group to be reduced.

[0025] The data reduction interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier, including: the data reduction interface determines the third tensor data to be transmitted based on the second tensor data to be transmitted of each of the first processes in the process group to be reduced, and transmits the third tensor data to be transmitted to the second process determined based on the second process identifier.

[0026] In some embodiments, the data reduction interface receives the second tensor data to be transmitted and the second process identifier of the first process, including: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, the second process group identifier of the process group to be reduced, and the reduction algorithm of the first process.

[0027] The data reduction interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier, including: the data reduction interface performs reduction calculation on the second tensor data to be transmitted of each of the first processes in the process group to be reduced based on the reduction algorithm to obtain the third tensor data to be transmitted, and transmits the third tensor data to be transmitted to the second process determined based on the second process identifier.

[0028] In some embodiments, the distributed interface includes a data scattering interface, the first data to be transmitted includes a first list of tensors to be transmitted, and the second data to be transmitted includes a fourth list of tensor data to be transmitted; the distributed interface transmits the first data to be transmitted or a second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes, including:

[0029] The data scattering interface receives the first list of tensors to be transmitted and the third process identifier of the first process;

[0030] The data scattering interface determines at least two second processes in the process group determined based on the third process identifier, and allocates tensor data in the first tensor list to be transmitted to obtain the fourth tensor data to be transmitted corresponding to each second process.

[0031] The data scattering interface transmits the fourth tensor data to be transmitted for each of the second processes to each of the second processes.

[0032] In some embodiments, the data scattering interface receives a first list of tensors to be transmitted and a third process identifier of the first process, including: the data scattering interface receiving a first list of tensors to be transmitted, a third process identifier of the first process, and a third process group identifier of the process group to be scattered.

[0033] The data scattering interface determines at least two second processes in the process group determined based on the third process identifier, and allocates tensor data in the first tensor list to be transmitted to obtain a fourth tensor data to be transmitted corresponding to each second process. This includes: the data scattering interface determines at least two second processes in the process group determined based on the third process group identifier, and allocates tensor data in the first tensor list to be transmitted to obtain a fourth tensor data to be transmitted corresponding to each second process.

[0034] In some embodiments, the distributed interface includes a data collection interface, the first data to be transmitted includes a fifth tensor data to be transmitted, and the second data to be transmitted includes a second list of tensors to be transmitted; the distributed interface transmits the first data to be transmitted or a second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes, including:

[0035] The data collection interface receives the fifth tensor data to be transmitted from the first process and the fourth process identifier from the second process;

[0036] The data collection interface determines the second list of tensors to be transmitted based on the fourth tensor data to be transmitted, and transmits the second list of tensors to be transmitted to the second process determined based on the fourth process identifier.

[0037] In some embodiments, the data collection interface receives the fifth tensor data to be transmitted from the first process and the fourth process identifier of the second process, including: the data collection interface receives the fifth tensor data to be transmitted from the first process, the fourth process identifier of the second process, and the fourth process group identifier; the fourth process group identifier is used to determine the process group to be collected;

[0038] The data collection interface determines the second list of tensors to be transmitted based on the fourth tensor data to be transmitted, and transmits the second list of tensors to be transmitted to the second process determined based on the fourth process identifier, including: the data collection interface collects the fifth tensor data to be transmitted from the first process in the process group to be collected, and generates the second list of tensors to be transmitted; the data collection interface transmits the second list of tensors to be transmitted to the second process determined based on the fourth process identifier.

[0039] In some embodiments, the distributed interface includes a data protocol broadcast interface, the first data to be transmitted includes a sixth tensor data to be transmitted, and the second data to be transmitted includes a seventh tensor data to be transmitted; the distributed interface transmits the first data to be transmitted or a second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes, including:

[0040] The data reduction broadcast interface receives the sixth tensor data to be transmitted and the reduction algorithm from the first process;

[0041] The data reduction broadcast interface performs reduction calculations on the sixth tensor data to be transmitted based on the reduction algorithm to obtain the seventh tensor data to be transmitted, and transmits the seventh tensor data to the second process.

[0042] In some embodiments, the data reduction broadcast interface receives the sixth tensor data to be transmitted and the reduction algorithm of the first process, including: the data reduction broadcast interface receives the sixth tensor data to be transmitted, the reduction algorithm and the fifth process group identifier of the first process;

[0043] The data reduction broadcast interface performs reduction calculations on the sixth tensor data to be transmitted based on the reduction algorithm to obtain the seventh tensor data to be transmitted, and transmits the seventh tensor data to the second process, including: the data reduction broadcast interface performs reduction calculations on the fifth tensor data to be transmitted based on the reduction algorithm to obtain the seventh tensor data to be transmitted, and transmits the seventh tensor data to the second process; the second process is all processes in the process group determined based on the fifth process group identifier.

[0044] In some embodiments, the distributed interface includes a data collection and broadcasting interface, the first data to be transmitted includes an eighth tensor data to be transmitted, and the second data to be transmitted includes a third tensor list to be transmitted; the distributed interface transmits the first data to be transmitted or a second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes, including:

[0045] The data collection broadcast interface receives the eighth tensor data to be transmitted from the first process;

[0046] The data collection broadcast interface determines a third list of tensors to be transmitted based on the eighth tensor data to be transmitted, and transmits the third list of tensors to be transmitted to the second process.

[0047] In some embodiments, the data collection broadcast interface receives the eighth tensor data to be transmitted from the first process, including: the data collection broadcast interface receives the eighth tensor data to be transmitted from the first process and the sixth process group identifier;

[0048] The data collection and broadcasting interface determines a third list of tensors to be transmitted based on the eighth tensor data to be transmitted, and transmits the third list of tensors to be transmitted to the second process, including:

[0049] The data collection interface determines a third list of tensors to be transmitted based on the eighth tensor data to be transmitted, and transmits the third list of tensors to be transmitted to the second process; the second process is all processes in the process group to be collected determined based on the sixth process group identifier.

[0050] In some embodiments, the distributed interface includes a communication synchronization interface, and the method further includes:

[0051] The communication synchronization interface responds when the third part of the at least two distributed processes runs to the synchronization node, and the fourth part of the process does not run to the synchronization node, by setting the third part of the process to a waiting state;

[0052] The communication synchronization interface responds to the at least two distributed processes running to the synchronization node, and controls the at least two distributed processes to be in a running state.

[0053] In some embodiments, the distributed interface includes a process ID acquisition interface, and the method further includes:

[0054] The process ID retrieval interface responds to the process ID retrieval request and outputs the process ID of the current process.

[0055] In some embodiments, the distributed interface includes a process count acquisition interface, and the method further includes:

[0056] The process count acquisition interface responds to the process count acquisition interface request and outputs the process count of the current task process.

[0057] On the other hand, embodiments of this disclosure provide a model training apparatus, the apparatus comprising:

[0058] A distributed interface is used to receive the first data to be transmitted from the first process of at least two distributed processes.

[0059] The distributed interface is also used to transmit the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes;

[0060] The first process and the second process are at least one of the at least two distributed processes.

[0061] In another aspect, embodiments of this disclosure provide a computer device including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.

[0062] In another aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.

[0063] In another aspect, embodiments of this disclosure provide a computer program including computer-readable code, which, when executed in a computer device, causes a processor in the computer device to perform some or all of the steps in the above-described method.

[0064] In another aspect, embodiments of this disclosure provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, it implements some or all of the steps in the above method.

[0065] In this embodiment of the disclosure, by setting a distributed interface, during distributed training, the first and second processes in at least two distributed processes can transfer data based on this distributed interface. Thus, based on this distributed interface, data can be transferred among multiple computing devices / nodes, distributing the computation of individual operators across various computing devices / nodes, thereby supporting the training of large datasets and improving overall training efficiency.

[0066] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0067] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0068] Figure 1 A schematic diagram illustrating the implementation process of a model training method provided in this embodiment of the disclosure;

[0069] Figure 2 A schematic diagram illustrating the implementation process of a model training method provided in this embodiment of the disclosure;

[0070] Figure 3 A schematic diagram illustrating the implementation process of a model training method provided in this embodiment of the disclosure;

[0071] Figure 4 A schematic diagram illustrating the implementation process of a model training method provided in this embodiment of the disclosure;

[0072] Figure 5 A schematic diagram illustrating the implementation process of a model training method provided in this embodiment of the disclosure;

[0073] Figure 6 A schematic diagram illustrating the implementation process of a model training method provided in this embodiment of the disclosure;

[0074] Figure 7 A schematic diagram illustrating the implementation process of a model training method provided in this embodiment of the disclosure;

[0075] Figure 8 This is a schematic diagram of the composition structure of a model training device provided in an embodiment of the present disclosure;

[0076] Figure 9 This is a schematic diagram of the composition structure of a model training device provided in an embodiment of the present disclosure;

[0077] Figure 10 This is a schematic diagram of the hardware entity of a model training device provided in an embodiment of this disclosure. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this disclosure. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0079] In the following description, references to "some embodiments" describe a subset of all possible embodiments; however, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein.

[0080] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for descriptive purposes only and is not intended to limit this disclosure.

[0081] This disclosure provides a model training method that can be executed by a processor of a computer device. The computer device can refer to a server, laptop computer, tablet computer, desktop computer, smart TV, set-top box, mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device), or any other device with data processing capabilities.

[0082] Figure 1 This is a schematic diagram illustrating the implementation process of a model training method provided in an embodiment of this disclosure, as shown below. Figure 1 As shown, the method includes the following steps S101 to S103:

[0083] Step S101: The distributed interface receives the first data to be transmitted from the first process of at least two distributed processes.

[0084] In step S102, the distributed interface transmits the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process among the at least two distributed processes.

[0085] In some embodiments, the first process and the second process are at least one of the at least two distributed processes.

[0086] In some embodiments, the first process is the main process, and the second process is at least one of the other processes among the at least two distributed processes besides the main process.

[0087] In some embodiments, the second process is the master process, and the first process is at least one of the other processes among the at least two distributed processes besides the master process.

[0088] In some embodiments, the first process is a first part of the at least two distributed processes, and the first process is a second part of the at least two distributed processes.

[0089] In this context, the first part of the process and the second part of the process are completely identical partial processes; the first part of the process and the second part of the process can also be completely different partial processes; the first part of the process and the second part of the process can have the same overlapping process.

[0090] In some embodiments, the intersection between the first part of the process and the second part of the process is the at least two distributed processes; the intersection between the first part of the process and the second part of the process may also be a part of the at least two distributed processes.

[0091] In some embodiments, the distributed interface includes at least one of the following: a data broadcast interface, a data reduction interface, a data scattering interface, a data collection interface, a data reduction broadcast interface, a data collection broadcast interface, a communication synchronization interface, a process ID acquisition interface, and a process count acquisition interface. Detailed embodiments of each of the above interfaces will be described in detail in the following embodiments.

[0092] In this embodiment of the disclosure, by setting a distributed interface, during distributed training, the first and second processes in at least two distributed processes can transfer data based on this distributed interface. Thus, based on this distributed interface, data can be transferred among multiple computing devices / nodes, distributing the computation of individual operators across various computing devices / nodes, thereby supporting the training of large datasets and improving overall training efficiency.

[0093] Figure 2 This is an optional flowchart illustrating a model training method provided in an embodiment of this disclosure, which can be executed by a processor of a computer device. The distributed interface includes the data broadcast interface, and the first data to be transmitted includes first tensor data to be transmitted; based on... Figure 1 , Figure 1 S101 can be updated to S201, and S102 can be updated to S202, which will be combined Figure 2 The steps shown are explained.

[0094] Step S201: The data broadcast interface receives the first tensor data to be transmitted from the first process and the first process identifier of the first process.

[0095] In step S202, the data broadcast interface transmits the first tensor data to be transmitted to the second process; the second process is all processes in the process group corresponding to the first process identifier.

[0096] In some embodiments, during the process of calling the data broadcast interface to implement the data broadcast function, the data broadcast interface may further include the following optional parameter: first process group identifier.

[0097] In some embodiments, the input parameters of the data broadcast interface include first tensor data to be transmitted and the first process identifier and the first process group identifier of the first process. The above step S201 can be updated to: the data broadcast interface receives the first tensor data to be transmitted, the first process identifier and the first process group identifier of the first process.

[0098] Accordingly, step S202 can be updated to: the data broadcast interface transmits the first tensor data to be transmitted to the second process; the second process is all processes in the process group determined based on the first process group identifier.

[0099] In this embodiment of the disclosure, by setting a data broadcast interface, during distributed training, at least two distributed processes, the first process and the second process, can broadcast data from the first process to the second process based on this data broadcast interface. This increases the number of inter-process data transmission methods and improves the overall training efficiency.

[0100] Figure 3 This is an optional flowchart illustrating a model training method provided in this disclosure embodiment, which can be executed by a computer device's processor. The distributed interface includes a data reduction interface, the first data to be transmitted includes second tensor data to be transmitted, the second data to be transmitted includes third tensor data to be transmitted, based on... Figure 1 , Figure 1 S101 can be updated to S301, and S102 can be updated to S302, which will be combined Figure 3 The steps shown are explained.

[0101] Step S301: The data reduction interface receives the second tensor data to be transmitted and the second process identifier from the first process.

[0102] In step S302, the data protocol interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier.

[0103] In some embodiments, during the process of calling the data reduction interface to implement the data reduction function, the data reduction interface may further include the following optional parameters: reduction algorithm and second process group identifier.

[0104] In some embodiments, the input parameters of the data reduction interface include the second tensor data to be transmitted, the second process identifier, and the reduction algorithm. The above step S301 can be updated to: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, and the reduction algorithm of the first process.

[0105] Accordingly, step 302 above can be updated as follows: the data reduction interface performs reduction calculation on the second tensor data to be transmitted based on the reduction algorithm to obtain the third transmission tensor data, and transmits the third transmission tensor data to the second process determined based on the second process identifier.

[0106] In some embodiments, the input parameters of the data reduction interface include the second tensor data to be transmitted, the second process identifier, and the second process group identifier. The above step S301 can be updated to: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, and the second process group identifier of the process group to be reduced from the first process.

[0107] Accordingly, step 302 above can be updated to: the data reduction interface determines the third tensor data to be transmitted based on the second tensor data to be transmitted of each of the first processes in the process group to be reduced, and transmits the third tensor data to be transmitted to the second process determined based on the second process identifier.

[0108] In some embodiments, the input parameters of the data reduction interface include the second tensor data to be transmitted, the second process identifier, the second process group identifier, and the reduction algorithm. The above step S301 can be updated to: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, the second process group identifier of the process group to be reduced, and the reduction algorithm of the first process.

[0109] Accordingly, step 302 above can be updated as follows: the data reduction interface performs reduction calculation on the second tensor data to be transmitted of each of the first processes in the process group to be reduced based on the reduction algorithm to obtain the third tensor data to be transmitted, and transmits the third tensor data to be transmitted to the second process determined based on the second process identifier.

[0110] In this embodiment of the disclosure, by setting a data reduction interface, during distributed training, at least two distributed processes, the first and second processes, can transmit tensor data from the first process to the second process after reducing it using a reduction method, based on this data reduction interface. This increases the number of inter-process data transmission methods and improves overall training efficiency.

[0111] Figure 4 This is an optional flowchart illustrating a model training method provided in an embodiment of this disclosure, which can be executed by a processor of a computer device. The distributed interface includes a data scattering interface, the first data to be transmitted includes a first list of tensors to be transmitted, and the second data to be transmitted includes a fourth list of tensor data to be transmitted, based on... Figure 1 , Figure 1S101 can be updated to S401, and S102 can be updated to S402 and S403, which will combine Figure 4 The steps shown are explained.

[0112] Step S401: The data scattering interface receives the first list of tensors to be transmitted and the third process identifier of the first process.

[0113] In step S402, the data scattering interface determines at least two second processes in the process group determined based on the third process identifier, and allocates tensor data in the first tensor list to be transmitted to obtain the fourth tensor data to be transmitted corresponding to each second process.

[0114] In step S403, the data scattering interface transmits the fourth tensor data to be transmitted corresponding to each of the second processes to each of the second processes.

[0115] In some embodiments, during the process of calling the data scattering interface to implement the data scattering function, the data specification interface may further include the following optional parameter: third process group identifier.

[0116] In some embodiments, the input parameters of the data scattering interface include a first list of tensors to be transmitted, a third process identifier of the first process, and a third process group identifier. The above step S401 can be updated to: the data scattering interface receives the first list of tensors to be transmitted and the third process identifier of the first process.

[0117] Accordingly, step S402 can be updated as follows: the data scattering interface determines at least two second processes in the process group determined based on the third process group identifier, and allocates tensor data in the first tensor list to be transmitted to obtain the fourth tensor data to be transmitted corresponding to each second process.

[0118] In this embodiment of the disclosure, by setting a data scattering interface, during distributed training, at least two distributed processes, the first process and the second process, can distribute multiple data tensors included in the data tensor list of the first process to the second process based on the data scattering interface. This increases the distributed processing of tensor data, thereby improving the overall training efficiency.

[0119] Figure 5 This is an optional flowchart illustrating a model training method provided in an embodiment of this disclosure, which can be executed by a processor of a computer device. The distributed interface includes a data collection interface, the first data to be transmitted includes a fifth tensor data to be transmitted, and the second data to be transmitted includes a second list of tensors to be transmitted; based on... Figure 1 , Figure 1S101 can be updated to S501, and S102 can be updated to S502, which will be combined Figure 5 The steps shown are explained.

[0120] Step S501: The data collection interface receives the fifth tensor data to be transmitted from the first process and the fourth process identifier from the second process.

[0121] In step S502, the data collection interface determines the second list of tensors to be transmitted based on the fourth tensor data to be transmitted, and transmits the second list of tensors to be transmitted to the second process determined based on the fourth process identifier.

[0122] In some embodiments, during the process of calling the data collection interface to implement the data collection function, the data specification interface may further include the following optional parameter: the fourth process group identifier.

[0123] In some embodiments, the input parameters of the data collection interface include the fifth tensor data to be transmitted, the fourth process identifier of the second process, and the fourth process group identifier. The above step S501 can be updated to: the data collection interface receives the fifth tensor data to be transmitted from the first process, the fourth process identifier of the second process, and the fourth process group identifier; the fourth process group identifier is used to determine the process group to be collected.

[0124] Accordingly, step S502 above can be updated as follows: the data collection interface collects the fifth tensor data to be transmitted from the first process in the process group to be collected, and generates the second tensor list to be transmitted; the data collection interface transmits the second tensor list to be transmitted to the second process determined based on the fourth process identifier.

[0125] In this embodiment of the disclosure, by setting a data collection interface, during distributed training, at least two distributed processes, the first and second processes, can collect multiple data tensors from the first process based on the data scattering interface, obtain a tensor list including multiple data tensors, and transmit it to the second process. In this way, after distributed processing of the tensor data, the tensor list can be regenerated for subsequent calculations, improving the overall training efficiency.

[0126] Figure 6 This is an optional flowchart illustrating a model training method provided in an embodiment of this disclosure. This method can be executed by a processor of a computer device. The distributed interface includes a data reduction broadcast interface, the first data to be transmitted includes a sixth tensor data to be transmitted, and the second data to be transmitted includes a seventh tensor data to be transmitted, based on... Figure 1 , Figure 1S101 can be updated to S601, and S102 can be updated to S602, which will be combined Figure 6 The steps shown are explained.

[0127] Step S601: The data reduction broadcast interface receives the sixth tensor data to be transmitted and the reduction algorithm from the first process.

[0128] In step S602, the data reduction broadcast interface performs reduction calculation on the sixth tensor data to be transmitted based on the reduction algorithm to obtain the seventh tensor data to be transmitted, and transmits the seventh tensor data to the second process.

[0129] In some embodiments, during the process of calling the data protocol broadcast interface to implement the data protocol broadcast function, the data protocol broadcast interface further includes the following optional parameter: the fifth process group identifier.

[0130] In some embodiments, the input parameters of the data collection interface include the sixth tensor data to be transmitted and the reduction algorithm of the first process, and the fifth process group identifier. The above step S601 can be updated to: the data reduction broadcast interface receives the sixth tensor data to be transmitted, the reduction algorithm and the fifth process group identifier of the first process.

[0131] Accordingly, step S602 above can be updated as follows: the data reduction broadcast interface performs reduction calculation on the fifth tensor data to be transmitted based on the reduction algorithm to obtain the seventh tensor data to be transmitted, and transmits the seventh tensor data to be transmitted to the second process; the second process is all processes in the process group determined based on the fifth process group identifier.

[0132] In this embodiment of the disclosure, by setting a data reduction broadcast interface, during distributed training, at least two distributed processes, namely the first process and the second process, can use the reduction method based on the data reduction broadcast interface to reduce the tensor data of the first process, and broadcast the reduced data tensor to the second process. This increases the number of inter-process data transmission methods and improves the overall training efficiency.

[0133] Figure 7 This is an optional flowchart illustrating a model training method provided in an embodiment of this disclosure. The distributed interface includes a data collection and broadcasting interface. The first data to be transmitted includes an eighth tensor data to be transmitted, and the second data to be transmitted includes a third list of tensors to be transmitted. This method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 S107 in the middle can be updated to S701 to S703, which will combine Figure 7 The steps shown are explained.

[0134] Step S701: The data collection broadcast interface receives the eighth tensor data to be transmitted from the first process.

[0135] In step S702, the data collection broadcast interface determines a third list of tensors to be transmitted based on the eighth tensor data to be transmitted, and transmits the third list of tensors to be transmitted to the second process.

[0136] In some embodiments, during the process of calling the data collection broadcast interface to implement the data collection broadcast function, the data collection broadcast interface may further include the following optional parameter: the sixth process group identifier.

[0137] In some embodiments, the input parameters of the data collection broadcast interface include the eighth tensor data to be transmitted of the first process and the sixth process group identifier. The above step S701 can be updated to: the data collection broadcast interface receives the eighth tensor data to be transmitted of the first process and the sixth process group identifier.

[0138] Accordingly, step S702 above can be updated as follows: the data collection interface determines a third list of tensors to be transmitted based on the eighth tensor data to be transmitted, and transmits the data based on the third list of tensors to be transmitted to the second process; the second process is all processes in the process group to be collected determined based on the sixth process group identifier.

[0139] In this embodiment of the disclosure, by setting a data collection broadcast interface, during distributed training, at least two distributed processes, the first and second processes, can collect multiple data tensors from the first process based on this data collection broadcast interface, obtain a tensor list including multiple data tensors, and broadcast it to the second process. This increases the number of inter-process data transmission methods and improves the overall training efficiency.

[0140] In some embodiments, the distributed interface includes a communication synchronization interface, and the method further includes:

[0141] The communication synchronization interface, in response to the third part of the at least two distributed processes running to the synchronization node and the fourth part of the process not running to the synchronization node, sets the third part of the process to a waiting state; the communication synchronization interface, in response to the at least two distributed processes running to the synchronization node, controls the at least two distributed processes to a running state.

[0142] In some embodiments, the distributed interface includes a process ID acquisition interface, and the method further includes:

[0143] The process ID retrieval interface responds to the process ID retrieval request and outputs the process ID of the current process.

[0144] In some embodiments, the distributed interface includes a process count acquisition interface, and the method further includes:

[0145] The process count acquisition interface responds to the process count acquisition interface request and outputs the process count of the current task process.

[0146] The following describes the application of the model training method provided in this embodiment in a real-world scenario.

[0147] Communication in distributed training is typically implemented using underlying communication libraries and standards, such as OpenMPI and NCCL. Upper-layer interfaces encapsulate these underlying communication capabilities to provide users with standardized distributed interfaces required for deep learning training. Based on the communication needs of various processes during distributed training, including information transmission and retrieval, this disclosure defines eight distributed interfaces: a data broadcast interface, a data reduction interface, a data scattering interface, a data collection interface, a data reduction broadcast interface, a communication synchronization interface, a process ID retrieval interface, and a process count retrieval interface.

[0148] In some embodiments, the data broadcast interface (bcast) is used to broadcast data from one process to various processes. The interface parameters of this data broadcast interface are listed in Table 1:

[0149]

[0150]

[0151] Table 1

[0152] The interface parameters corresponding to this data broadcasting interface include the data tensor to be sent (Tensor, type Tensor), the process to be broadcast to (src, type int), and the target process group (group, type int), where the target process group is an optional parameter.

[0153] In some embodiments, if the data broadcast interface operation is successful, no exception information is generated; if the data broadcast interface operation fails, exception information indicating data broadcast failure is generated.

[0154] In some embodiments, the data reduction interface is used to reduce data from other processes to one process. Common reduction calculation methods include addition, multiplication, summation, and averaging. The interface parameter list for this data broadcasting interface is shown in Table 2.

[0155] Parameter type Keywords Parameter Description Is it optional? enter Tensor The data tensors involved in the reduction are of type Tensor. Required enter dst The target process for reduction computation is of type int. Required enter op The specific reduction calculation method is for type str. Optional enter group The process group participating in the specification is of type int. Optional

[0156] Table 2

[0157] The interface parameters corresponding to this data reduction interface include the data tensor (Tensor, type Tensor) to be reduced, the target process for reduction computation (dst, type int), the reduction computation method (op, type str), and the process group to be reduced (group, type int). Among them, the reduction computation method and the process group to be reduced are optional parameters. If the data reduction interface does not receive a reduction computation method, the default addition is used as the reduction computation method.

[0158] In some embodiments, if the data specification interface operation is successful, no exception information is generated; if the data specification interface operation fails, exception information indicating data specification failure is generated.

[0159] In some embodiments, the data scattering interface is used to scatter a set of data from one process to other processes. The interface parameters of this data scattering interface are listed in Table 3.

[0160]

[0161] Table 3

[0162] The interface parameters for this data scattering interface include a list of data tensors participating in the scattering (tensor_list, type list), the process of scattering data (src, type int), the process group participating in the reduction (group, type int), and the data tensor received by the process (Tensor, type Tensor). Among these, the process group participating in the reduction is an optional parameter, while the list of data tensors participating in the scattering, the process of scattering data, and the data tensor received by the process are required parameters. The process group participating in the reduction, the list of data tensors participating in the scattering, and the process of scattering data are input parameters, and the data tensor received by the process is an output parameter.

[0163] In some embodiments, if the data scattering interface operation is successful, no exception information is generated; if the data scattering interface operation fails, exception information indicating data scattering failure is generated.

[0164] In some embodiments, the data collection interface (gather) is used to collect data from various processes into a single process. The interface parameter list for this data collection interface is shown in Table 4.

[0165]

[0166]

[0167] Table 4

[0168] In some embodiments, the data reduction broadcast (all_reduce) interface is used to reduce the data from each process and then broadcast the result to all processes. Common reduction calculation methods include addition, multiplication, and summation / averaging. The interface parameter list for this data reduction broadcast interface is shown in Table 5.

[0169]

[0170] Table 5

[0171] In some embodiments, the data collection broadcast (all_gather) interface is used to collect data from each process and then broadcast the results to all processes. The interface parameter list for this data collection broadcast interface is shown in Table 6:

[0172]

[0173]

[0174] Table 6

[0175] In some embodiments, the communication synchronization interface is used to trigger process synchronization, waiting for all processes to reach that step before proceeding to the next step. The interface parameter list for this communication synchronization interface is shown in Table 7:

[0176] Parameter type Keywords Parameter Description Is it optional? enter group The process group participating in the synchronization is of type int. Optional

[0177] Table 7

[0178] In some embodiments, the process ID retrieval interface (get_rank) is used to retrieve the process ID of the current process. The list of interface parameters for this process ID retrieval interface is shown in Table 8.

[0179] Parameter type Keywords Parameter Description Is it optional? Output rank The process ID of the current process, which is of type int. Optional

[0180] Table 8

[0181] In some embodiments, the process count acquisition (get_world_size) interface is used to obtain the total number of processes currently running in the task. The list of interface parameters for this process count acquisition interface is shown in Table 9:

[0182] Parameter type Keywords Parameter Description Is it optional? Output world_size The total number of processes for the current task, of type int. Optional

[0183] Table 9

[0184] Based on the distributed interface provided in the above embodiments, algorithms and algorithm frameworks can better call distributed-related interfaces to perform distributed training.

[0185] To better understand the embodiments of this disclosure, the computer vision interface model is briefly described below:

[0186] 1) Algorithm Interface Goals

[0187] By defining the interface adaptation layer, the following goals can be achieved: a) the algorithm implementation can use different deep learning frameworks as backends and can be switched; b) it can run on different hardware, such as servers and distributed clusters; c) the algorithm can read different data through a unified data interface.

[0188] 2) Algorithm Interface Model

[0189] This standard describes the relationship between data, algorithms, and models from three levels: system resources, interface adaptation, and algorithm application. The standard primarily specifies the technical requirements for the interface adaptation layer: a) System resource layer: This includes both hardware and software, providing the necessary storage, computation, and inference functions for computer vision systems. The specific implementation of the system resource layer varies among vendors, and this standard does not further describe or specify it; b) Interface adaptation layer: This provides service interfaces such as data interfaces, optimization interfaces, distributed interfaces, and model interfaces to the upper-layer system processes, ensuring efficient and flexible model training and model migration between different frameworks; c) Algorithm application layer: This completes the model training and inference processes, requiring the use of various interfaces defined in this standard during model training.

[0190] The interface requirements for computer vision systems (mandatory and optional requirements, and appropriate usage) are as follows:

[0191] 1. Data and Model Structure

[0192] 1) Data Structure

[0193] Datasets can be in formats such as images, videos, and binary data. The dataset annotation file is in JSON format and contains annotation information for all samples in the dataset. If the annotation results include other auxiliary files, such as mask layer information, the relative paths of these auxiliary files are stored in the JSON file. Common data types are represented as follows:

[0194] Supports category labels, such as integers, where 0 represents background and positive integers represent foreground; supports bounding boxes, such as using the coordinates of the top-left and bottom-right vertices in the order (top-left x-coordinate, top-left y-coordinate, bottom-right x-coordinate, bottom-right y-coordinate); supports annotation files being parsed into lists or arrays after passing through the data reading interface, where each element is a dictionary or key-value pair container containing all relevant information for a sample, accessible by the dataset using an index; supports loading specific dataset formats and custom datasets; supports annotation files in mainstream open-source dataset formats, such as COCO and PASCAL VOC.

[0195] 2) Model Structure

[0196] a) Basic operators should be supported, including but not limited to "+", "-", "*", " / ", convolution operations, etc.; b) The meaning and values ​​of the parameters of the operator should be defined, and the computational logic for obtaining the output from the input during forward operation should be defined; c) Basic operators should support the construction of computational graphs through operator chaining and function nesting; d) The computational graph should support differentiation through backpropagation using the chain rule; e) The computational graph should support construction through conditional judgments or loops; f) User-defined operators should be compiled and allowed to be added to the computational graph; g) Serialization and deserialization of model parameters should be supported.

[0197] 2. Training Interface

[0198] 1) Optimizer Interface

[0199] The optimizer interface should support network model optimization, updating model parameters according to different optimization algorithms such as SGD, Adam, and Momomtum. Specifically, it should implement functions such as gradient calculation and backpropagation, parameter updates, optimization algorithms, and learning rate updates.

[0200] The optimizer interface should implement: a) support implementation as a class whose constructor parameters are the network model or a list of model parameters, and other required parameters, such as the learning rate; b) support implementation of the `step()` function for performing a single parameter optimization update. After this function is called, the model parameters should be updated based on the accumulated gradients.

[0201] 2) Mixed Precision Training Interface

[0202] The mixed-precision training interface provides unified support for mixed-precision training of algorithms, which can reduce memory consumption and improve training speed when the graphics card supports it. This interface should implement the following functions:

[0203] a) Supports precision conversion, converting model parameters to fp16, except for special layers (such as BN layers), while retaining a copy of the fp32 parameters; b) Supports input forward propagation, converting input data to fp16 for forward propagation and loss calculation; c) Supports loss amplification, amplifying the calculated loss with both fixed and dynamic amplification modes; d) Supports gradient calculation, calculating and backpropagating gradients in fp16 mode, then converting them to fp32 and scaling them back to the actual scale proportionally according to the amplification factor in c); e) Supports parameter update, updating parameters in the fp32 parameter copy based on the gradient calculated in d), and then assigning the updated parameters to the fp16 model.

[0204] 3) Distributed Interface

[0205] Through a distributed interface, the framework can complete the data transfer between multiple processes in a distributed training scenario with multiple machines and multiple GPUs. This interface should have the following core functionalities:

[0206] The distributed interface should cover functions such as `bcast()`, `reduce()`, `scatter()`, `gather()`, `allreduce()`, `allgateher()`, and `sync()`. This set of interfaces should implement the following: support broadcasting data from the main process to each process; support reducing data from each process to the main process; support scattering a set of data from the main process to each process; support collecting data scattered across each process into the main process as a set; support reducing data from each process and then broadcasting it to each process; support collecting scattered data from each process into a set and then broadcasting it to each process; and support ensuring that all previously issued communication commands have been completed.

[0207] 4) Quantization training interface

[0208] The quantization training interface enables algorithms to perceive the information loss caused by model quantization during neural network training. During training, quantized weights are approximated using floating-point weights, allowing the quantized model to be simulated during forward propagation. The floating-point error is then calculated and backpropagated to update the weights. Quantization helps accelerate model inference and reduce storage requirements. This interface should implement the following functionalities:

[0209] a) Supports input quantization, converting input from 32-bit floating-point type to 8-bit or custom-bit fixed-point type; b) Supports quantization and dequantization of convolution and addition operators; c) Supports pseudo-quantization nodes, which should include functions for quantizing and dequantizing floating-point weights; d) Supports error backpropagation, calculating the error through b) during forward propagation, updating the floating-point weights, and then quantizing; e) Supports quantized model output, allowing the quantized trained model to be converted into a fixed-point model for storage.

[0210] 5) Data processing interface

[0211] The data interface should support the conversion of data into formats required by the module, such as tensors. The data interface should have both outer and inner interfaces. The framework should be able to prepare continuous training data for algorithm training in an iterative manner. This interface should have the following core functionalities:

[0212] Supports the implementation of an iterable data loader type, with each iteration returning a batch of data;

[0213] It supports sampling data from the dataset according to training requirements; it supports constructing dataset objects based on the dataset path and related parameters; it supports reading part or all of the dataset data from storage devices or services, such as annotation files and data samples; and it supports preprocessing operations on the data, such as image scaling, flipping, and color perturbation.

[0214] 6) Visual Interface

[0215] The training visualization interface provides visualizations of model structure, parameters, gradients, and features during algorithm training. This interface should support: visualization of model structure diagrams; visualization of feature maps; visualization of weight histograms; visualization of scalar changes; and visualization of convolutional kernels.

[0216] 7) Distillation interface

[0217] The distillation interface supports the use of a teacher network to guide the student network during training, thereby improving the training accuracy of the student model: it should support target distillation method; it should support feature distillation method.

[0218] 8) Graphical computation fusion interface

[0219] Graph-computation fusion optimizes the overall network execution time by analyzing and optimizing the existing network computation graph logic, reducing overhead during operator execution intervals, and improving the utilization of device computing resources. It supports operations such as splitting, reorganizing, and merging existing computation logic; and allows enabling graph-computation fusion by modifying the context parameter in the training script.

[0220] 3. Inference Interface

[0221] 1) Process orchestration interface

[0222] The computer vision system supports workflow orchestration, and its interface meets the following requirements:

[0223] It should be able to combine key processes such as image acquisition, image decoding, image scaling, object detection, image cropping, image classification, and serialization;

[0224] It is advisable to support the plug-in approach for key processes, with configurable properties for each plug-in.

[0225] Users should be able to mount metadata.

[0226] It should support configuration file-based process orchestration management and have management components;

[0227] It should support specifying accelerators for particular processes;

[0228] It should support orchestration of multiple request and multiple output processes;

[0229] The following models should be supported for orchestration: YOLOv3, YOLOv3-tiny, ResNet50, Faster R-CNN, YOLOv4, SSD-VGG16, SSD MobileNet v1 FPN, CRNN, YOLOv5, Faster R-CNN-FPN / Cascade R-CNN-FPN, ResNet-18, DeepLabv3+, CTPN, DeepLabv3, BERT-Base (Uncased), DeepLabv3+, U-Net, Mask R-CNN, FaceNet, SSD MobileNet v1 FPN, OpenPose, Unet++, RetinaNet;

[0230] It should support single-input, single-output, multiple-input, and multiple-output orchestration.

[0231] 2) Data processing interface

[0232] The computer vision system has a data processing interface that meets the following requirements:

[0233] It should support reading data from image files and moving it into a pre-configured cache;

[0234] It should support JPG / JPEG / BMP format image decoding, with a resolution range of (32*32, 8192*8192);

[0235] It should at least support JPG image encoding, with a resolution range of (32*32, 8192*8192);

[0236] Image scaling with specified target width and height should be supported, and image width and height should be scaled and aligned to the step size.

[0237] It should support specifying the expansion ratio in the four directions (top, bottom, left, and right) to expand the area of ​​the target bounding box for cropping.

[0238] It should support H264 / H265 video decoding, with a resolution range of (128*128, 4096*4096).

[0239] Width and height scaling alignment should be supported for step-based scaling.

[0240] It should support resolutions ranging from 128*128 to 1920*1920, and both H264 MP and H265 MP.

[0241] Image normalization, center cropping, affine transformation, and rotation should be supported;

[0242] It should support data transfer between key processes and preferably support the multiple distribution of a single input;

[0243] It should support data transfer between processor memory and main memory;

[0244] It should support frame skipping processing of video data;

[0245] Serialization should be supported.

[0246] 3) Plug-in interface

[0247] The computer vision system should support the development and use of visual processing plugins and meet the following requirements: it should support user-developed plugins, registration, and compilation; it should support interfaces for plugin initialization, deinitialization, execution, attribute registration, and retrieval; it should support interfaces for defining variable and immutable ports for plugin input and output; and it should support interfaces for defining and throwing business logic exceptions.

[0248] It should support a streaming plugin interface to achieve the following functions: a) sending data of a specified type or channel to different ports; b) outputting data from multiple ports in sequence through a single port.

[0249] It should support multiple instantiation interfaces for plugins of the same type; it should support plugin caching mechanisms and interfaces to enable the transfer of business data (such as decoded video and image data) between plugins; it should support description interfaces for plugin metadata (such as classification information and target information), and implement the transfer by relying on the plugin cache; it should support single-input, single-output, multi-input, and multi-output plugin interfaces; it should support inference plugin interfaces, supporting target classification, detection, and tensor-based (input) inference; it should support model post-processing plugin interfaces, enabling it to interface with models for target detection, classification, semantic segmentation, text generation, text box detection, pose detection, etc.

[0250] The video analytics plugin interface should be supported to implement the following functions: a) multi-target (including machine, non-human, and face) path recording; b) face alignment (correcting detected face images); c) video quality diagnosis.

[0251] It should support a debugging plugin interface to enable data export (e.g., JSON format) and data loading and restoration; it should also support a screen display plugin interface to enable drawing basic units on images, such as drawing frames, lines, circles, and writing text.

[0252] 4) Block detection interface

[0253] The computer vision system supports a block detection interface and meets the following requirements:

[0254] It supports filtering duplicate targets in overlapping areas after segmentation; it supports user-defined parameters such as the number / size of segments and overlap, and automatically generates target boxes for image segments; it supports merging images of segmented inference results; and during multi-level inference, it supports filtering post-processing results based on the selection of maximum and minimum area, upper and lower area limits, and confidence thresholds.

[0255] 4. Module

[0256] A Module is a fundamental module in a neural network. Neural network modules are built upon this base class to construct graphs. A Module provides the following functionalities:

[0257] 1) Forward computation of the module: a) Interface name: forward; b) Interface function description: The module performs a forward computation and returns the computation result of the module. If it is in training state, a computation graph is constructed during the forward computation process to calculate the gradient of the module parameters.

[0258] 2) Get trainable parameters of the module: a) Interface name: get_parameters; b) Interface function description: Returns the trainable parameters of the module.

[0259] 3) Retrieve Modules and Submodules: a) Interface Name: get_modules; b) Interface Function Description: Optional. This interface returns an iterator that iterates through the module itself and its submodules. Duplicate modules are returned only once.

[0260] 4) Get module state: a) Interface name: get_state_dict; b) Interface function description: Returns the module state in key-value pairs, including the module parameters and buffer.

[0261] 5) Loading module status: a) Interface name: load_state_dict; b) Interface function description: Loading module status, including module parameters and buffers.

[0262] 6) Module Backward Computation: a) Interface Name: backward(grad_input, grad_output); b) Interface Function Description: The module performs a backward computation and returns the computation result. If in training mode, the gradient of the module parameters is calculated during the backward computation process. This function is automatically generated by the computation graph and can also be registered later using register_backward_function.

[0263] 7) Forward Computation of Module: a) Interface Name: register_backward_function; b) Interface Function Description: The module performs a backward computation and returns the computation result of the module. If it is in training state, the gradient of the module parameters is calculated during the backward computation process.

[0264] Based on the foregoing embodiments, this disclosure provides a model training device, which includes the included units and the modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0265] Based on the foregoing embodiments, this disclosure provides a model training device, which includes the included units and the modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0266] Figure 8 This is a schematic diagram of the composition structure of a model training device provided in an embodiment of the present disclosure, as shown below. Figure 8 As shown, the model training device 800 includes: a distributed interface 801, wherein:

[0267] Distributed interface 801 is used to receive the first data to be transmitted from the first process of at least two distributed processes;

[0268] The distributed interface 801 is also used to transmit the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process of the at least two distributed processes.

[0269] The first process and the second process are at least one of the at least two distributed processes.

[0270] In some embodiments, the first process is the main process, and the second process is at least one of the other processes among the at least two distributed processes besides the main process; or...

[0271] The second process is the main process, and the first process is at least one of the other processes among the at least two distributed processes besides the main process; or,

[0272] The first process is a first part of the at least two distributed processes, and the first process is a second part of the at least two distributed processes. The first part process and the second part process are unrelated.

[0273] In some embodiments, the distributed interface 801 includes at least one of the following: a data broadcast interface, a data protocol interface, a data scattering interface, a data collection interface, a data protocol broadcast interface, a data collection broadcast interface, a communication synchronization interface, a process ID acquisition interface, and a process count acquisition interface.

[0274] Please see Figure 9 , Figure 9 This is a schematic diagram of the composition structure of another model training device provided in an embodiment of the present disclosure, as shown below. Figure 9 As shown, the model training device 900 includes: a data broadcast interface 901, a data reduction interface 902, a data scattering interface 903, a data collection interface 904, a data reduction broadcast interface 905, a data collection broadcast interface 906, a communication synchronization interface 907, a process ID acquisition interface 908, and a process count acquisition interface 909, wherein:

[0275] In some embodiments, the distributed interface includes the data broadcast interface 901, and the first data to be transmitted includes first tensor data to be transmitted; the data broadcast interface 901 is used to receive the first tensor data to be transmitted from the first process and the first process identifier of the first process.

[0276] The data broadcast interface 901 is used to transmit the first tensor data to the second process; the second process is all processes in the process group corresponding to the first process identifier.

[0277] In some embodiments, the data broadcast interface 901 is used to receive the first tensor data to be transmitted by the first process, the first process identifier of the first process, and the first process group identifier of the first process;

[0278] The data broadcast interface 901 is used to transmit the first tensor data to the second process; the second process is all processes in the process group determined based on the first process group identifier.

[0279] In some embodiments, the distributed interface includes a data protocol interface 902, the first data to be transmitted includes a second tensor data to be transmitted, and the second data to be transmitted includes a third tensor data to be transmitted; the data protocol interface 902 is used to receive the second tensor data to be transmitted and the second process identifier of the first process.

[0280] The data protocol interface 902 is used to receive the third tensor data to be transmitted determined based on the second tensor data to be transmitted, and to transmit the third tensor data to the second process determined based on the second process identifier.

[0281] In some embodiments, the data reduction interface 902 is used to receive the second tensor data to be transmitted from the first process, the second process identifier, and the reduction algorithm.

[0282] The data reduction interface 902 is used to perform reduction calculation on the second tensor data to be transmitted based on the reduction algorithm to obtain the third tensor data to be transmitted, and to transmit the third tensor data to the second process determined based on the second process identifier.

[0283] In some embodiments, the data protocol interface 902 is used to receive the second tensor data to be transmitted from the first process, the second process identifier, and the second process group identifier of the process group to be regulated.

[0284] The data reduction interface 902 is used to determine the third tensor data to be transmitted based on the second tensor data to be transmitted of each of the first processes in the process group to be reduced, and to transmit the third tensor data to the second process determined based on the second process identifier.

[0285] In some embodiments, the data reduction interface 902 is used to receive the second tensor data to be transmitted from the first process, the second process identifier, the second process group identifier of the process group to be reduced, and the reduction algorithm.

[0286] The data reduction interface 902 is used to perform reduction calculation on the second tensor data to be transmitted of each of the first processes in the process group to be reduced based on the reduction algorithm, to obtain the third tensor data to be transmitted, and to transmit the third tensor data to the second process determined based on the second process identifier.

[0287] In some embodiments, the distributed interface includes a data scattering interface 903, the first data to be transmitted includes a first list of tensors to be transmitted, and the second data to be transmitted includes a fourth list of tensors to be transmitted; the data scattering interface 903 is used to receive the first list of tensors to be transmitted and a third process identifier of the first process;

[0288] The data scattering interface 903 is used to determine at least two second processes in the process group determined based on the third process identifier, and to allocate tensor data in the first tensor list to be transmitted to obtain the fourth tensor data to be transmitted corresponding to each second process.

[0289] The data scattering interface 903 is used to transmit the fourth tensor data to be transmitted corresponding to each of the second processes to each of the second processes.

[0290] In some embodiments, the data scattering interface 903 is used to receive a first list of tensors to be transmitted, a third process identifier of the first process, and a third process group identifier of the process group to be scattered.

[0291] The data scattering interface 903 is used to determine at least two second processes in the process group determined based on the third process group identifier, and to allocate tensor data in the first tensor list to be transmitted to obtain a fourth tensor data to be transmitted corresponding to each second process.

[0292] In some embodiments, the distributed interface includes a data collection interface 904, the first data to be transmitted includes a fifth tensor data to be transmitted, and the second data to be transmitted includes a second tensor list to be transmitted; the data collection interface 904 is used to receive the fifth tensor data to be transmitted from the first process and the fourth process identifier of the second process.

[0293] The data collection interface 904 is used to determine the second list of tensors to be transmitted based on the fourth tensor data to be transmitted, and to transmit the second list of tensors to be transmitted to the second process determined based on the fourth process identifier.

[0294] In some embodiments, the data collection interface 904 is used to receive the fifth tensor data to be transmitted from the first process, the fourth process identifier and the fourth process group identifier of the second process; the fourth process group identifier is used to determine the process group to be collected.

[0295] The data collection interface 904 is used to collect the fifth tensor data to be transmitted from the first process in the process group to be collected, and to generate the second tensor list to be transmitted; the data collection interface transmits the second tensor list to be transmitted to the second process determined based on the fourth process identifier.

[0296] In some embodiments, the distributed interface includes a data reduction broadcast interface 905, the first data to be transmitted includes a sixth tensor data to be transmitted, and the second data to be transmitted includes a seventh tensor data to be transmitted; the data reduction broadcast interface 905 is used to receive the sixth tensor data to be transmitted and the reduction algorithm of the first process;

[0297] The data reduction broadcast interface 905 is used to perform reduction calculation on the sixth tensor data to be transmitted based on the reduction algorithm to obtain the seventh tensor data to be transmitted, and to transmit the seventh tensor data to the second process.

[0298] In some embodiments, the data protocol broadcast interface 905 is used to receive the sixth tensor data to be transmitted, the reduction algorithm, and the fifth process group identifier of the first process.

[0299] The data reduction broadcast interface 905 is used to perform reduction calculation on the fifth tensor data to be transmitted based on the reduction algorithm to obtain the seventh tensor data to be transmitted, and to transmit the seventh tensor data to be transmitted to the second process; the second process is all processes in the process group determined based on the fifth process group identifier.

[0300] In some embodiments, the distributed interface includes a data collection broadcast interface 906, the first data to be transmitted includes an eighth tensor data to be transmitted, and the second data to be transmitted includes a third tensor list to be transmitted; the data collection broadcast interface 906 is used to receive the eighth tensor data to be transmitted from the first process.

[0301] The data collection broadcast interface 906 is used to determine a third list of tensors to be transmitted based on the eighth tensor data to be transmitted, and to transmit the third list of tensors to be transmitted to the second process.

[0302] In some embodiments, the data collection broadcast interface 906 is used to receive the eighth tensor data to be transmitted from the first process and the sixth process group identifier.

[0303] The data collection broadcast interface 906 is used to determine a third list of tensors to be transmitted based on the eighth tensor data to be transmitted, and to transmit the third list of tensors to be transmitted to a second process; the second process is all processes in the process group to be collected determined based on the sixth process group identifier.

[0304] In some embodiments, the distributed interface includes a communication synchronization interface 907, which is configured to set the third part of the process to a waiting state in response to a third part of the at least two distributed processes running to the synchronization node and the fourth part of the process not running to the synchronization node.

[0305] The communication synchronization interface 907 is used to control the at least two distributed processes to be in a running state in response to the fact that both of the distributed processes have run to the synchronization node.

[0306] In some embodiments, the distributed interface includes a process ID acquisition interface 908, which is used to output the process ID of the current process in response to a process ID acquisition request.

[0307] In some embodiments, the distributed interface includes a process count acquisition interface 909, which is used to output the process count of the current task process in response to a process count acquisition interface request.

[0308] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0309] It should be noted that, in the embodiments of this disclosure, if the above-described model training method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk. Thus, the embodiments of this disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0310] This disclosure provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0311] This disclosure provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium may be transient or non-transient.

[0312] This disclosure provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0313] This disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0314] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referenced interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0315] Figure 10 This is a schematic diagram of the hardware entity of a model training device provided in an embodiment of the present disclosure, such as... Figure 10 As shown, the hardware entity of the model training device 1000 includes a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can run on the processor 1001, and the processor 1001 executes the program to implement the steps in the method of any of the above embodiments.

[0316] The memory 1002 stores computer programs that can run on the processor. The memory 1002 is configured to store instructions and applications that can be executed by the processor 1001. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1001 and the various modules in the model training device 1000. It can be implemented by flash memory or random access memory (RAM).

[0317] When the processor 1001 executes the program, it implements the steps of any of the above-mentioned model training methods. The processor 1001 typically controls the overall operation of the model training device 1000.

[0318] This disclosure provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the model training method as described in any of the above embodiments.

[0319] It should be noted that the descriptions of the storage medium and device embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0320] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this disclosure does not specifically limit the specific implementation.

[0321] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0322] It should be understood that the phrase "an embodiment" or "one embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this disclosure. Therefore, "in one embodiment" or "one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this disclosure, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure. The sequence numbers of the above embodiments of this disclosure are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0323] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0324] In the several embodiments provided in this disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0325] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0326] In addition, each functional unit in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0327] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0328] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0329] The above description is merely an embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A model training method, characterized in that, The method includes: The distributed interface receives the first data to be transmitted from the first process of at least two distributed processes. The distributed interface transmits the first data to be transmitted or the second data to be transmitted, determined based on the data to be transmitted, to the second process among the at least two distributed processes. The first process and the second process are at least one of the at least two distributed processes; The distributed interface is a standardized interface that supports tensor data transmission in multiple communication modes. The distributed interface includes at least one of the following: data broadcast interface, data reduction interface, data scattering interface, data collection interface, data reduction broadcast interface, data collection broadcast interface, communication synchronization interface, process ID acquisition interface, and process count acquisition interface.

2. The method according to claim 1, characterized in that, The first process is the main process, and the second process is at least one of the other processes among the at least two distributed processes besides the main process; or, The second process is the main process, and the first process is at least one of the other processes among the at least two distributed processes besides the main process; or, The first process is a first part of the at least two distributed processes, and the first process is a second part of the at least two distributed processes.

3. The method of claim 1, wherein, The distributed interface includes the data broadcast interface, and the first data to be transmitted includes the first tensor data to be transmitted. The distributed interface receives first data to be transmitted from the first process in at least two distributed processes, including: the data broadcast interface receives first tensor data to be transmitted from the first process and the first process identifier of the first process. The distributed interface transmits the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process among the at least two distributed processes, including: the data broadcast interface transmits the first tensor data to be transmitted to the second process; the second process is all processes in the process group corresponding to the first process identifier.

4. The method of claim 3, wherein, The data broadcast interface receives the first tensor data to be transmitted from the first process; including: the data broadcast interface receives the first tensor data to be transmitted from the first process, the first process identifier of the first process, and the first process group identifier of the first process. The data broadcast interface transmits the first tensor data to be transmitted to the second process; the second process is all processes in the process group corresponding to the first process identifier, including: the data broadcast interface transmits the first tensor data to be transmitted to the second process; the second process is all processes in the process group determined based on the first process group identifier.

5. The method according to any one of claims 1 to 4, characterized in that, The distributed interface includes a data reduction interface, the first data to be transmitted includes a second tensor data to be transmitted, and the second data to be transmitted includes a third tensor data to be transmitted. The distributed interface receives first data to be transmitted from the first process in at least two distributed processes, including: the data reduction interface receives second tensor data to be transmitted from the first process and the second process identifier. The distributed interface transmits the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process among the at least two distributed processes, including: the data reduction interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to be transmitted to the second process determined based on the second process identifier.

6. The method of claim 5, wherein, The data reduction interface receives the second tensor data to be transmitted and the second process identifier of the first process, including: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, and the reduction algorithm of the first process; The data reduction interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to be transmitted to the second process determined based on the second process identifier, including: the data reduction interface performs reduction calculation on the second tensor data to be transmitted based on the reduction algorithm to obtain the third tensor data to be transmitted, and transmits the third tensor data to be transmitted to the second process determined based on the second process identifier.

7. The method according to claim 5, characterized in that, The data reduction interface receives the second tensor data to be transmitted and the second process identifier of the first process, including: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, and the second process group identifier of the process group to be reduced from the first process; The data reduction interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier, including: the data reduction interface determines the third tensor data to be transmitted based on the second tensor data to be transmitted of each first process in the process group to be reduced, and transmits the third tensor data to the second process determined based on the second process identifier.

8. The method according to claim 5, characterized in that, The data reduction interface receives the second tensor data to be transmitted and the second process identifier of the first process, including: the data reduction interface receives the second tensor data to be transmitted, the second process identifier, the second process group identifier of the process group to be reduced, and the reduction algorithm of the first process. The data reduction interface receives the third tensor data to be transmitted based on the second tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier, including: the data reduction interface performs reduction calculation on the second tensor data to be transmitted of each first process in the process group to be reduced based on the reduction algorithm to obtain the third tensor data to be transmitted, and transmits the third tensor data to the second process determined based on the second process identifier.

9. The method according to any one of claims 1 to 4, characterized in that, The distributed interface includes a data scattering interface, the first data to be transmitted includes a first list of tensors to be transmitted, and the second data to be transmitted includes a fourth list of tensors to be transmitted. The distributed interface receives first data to be transmitted from the first process in at least two distributed processes, including: the data scattering interface receives the first list of tensors to be transmitted and the third process identifier of the first process; The distributed interface transmits the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process in the at least two distributed processes, including: the data scattering interface determines at least two second processes in the process group determined based on the third process identifier, and allocates the tensor data in the first tensor list to be transmitted to obtain the fourth tensor data to be transmitted corresponding to each second process; The data scattering interface transmits the fourth tensor data to be transmitted for each of the second processes to each of the second processes.

10. The method of claim 9, wherein, The data scattering interface receives the first list of tensors to be transmitted and the third process identifier of the first process, including: the data scattering interface receives the first list of tensors to be transmitted, the third process identifier of the first process, and the third process group identifier of the process group to be scattered; The data scattering interface determines at least two second processes in the process group determined based on the third process identifier, and allocates tensor data in the first tensor list to be transmitted to obtain a fourth tensor data to be transmitted corresponding to each second process. This includes: the data scattering interface determines at least two second processes in the process group determined based on the third process group identifier, and allocates tensor data in the first tensor list to be transmitted to obtain a fourth tensor data to be transmitted corresponding to each second process.

11. The method of claim 10, wherein, The distributed interface includes a data collection interface, the first data to be transmitted includes a fifth tensor data to be transmitted, and the second data to be transmitted includes a second list of tensors to be transmitted. The distributed interface receives first data to be transmitted from the first process in at least two distributed processes, including: the data collection interface receives fifth tensor data to be transmitted from the first process and the fourth process identifier of the second process; The distributed interface transmits the first data to be transmitted or the second data to be transmitted determined based on the data to be transmitted to the second process among the at least two distributed processes, including: the data collection interface determines the second list of tensors to be transmitted based on the fourth tensor data to be transmitted, and transmits the second list of tensors to be transmitted to the second process determined based on the fourth process identifier.

12. The method according to claim 11, characterized in that, The data collection interface receives the fifth tensor data to be transmitted from the first process and the fourth process identifier of the second process, including: the data collection interface receives the fifth tensor data to be transmitted from the first process, the fourth process identifier of the second process, and the fourth process group identifier; the fourth process group identifier is used to determine the process group to be collected; The data collection interface determines the second to-be-transmitted tensor list based on the fourth to-be-transmitted tensor data and transmits the second to-be-transmitted tensor list to a second process determined based on the fourth process identifier, including: the data collection interface collects fifth to-be-transmitted tensor data of the first process in the to-be-collected process group and generates the second to-be-transmitted tensor list; and the data collection interface transmits the second to-be-transmitted tensor list to a second process determined based on the fourth process identifier.

13. The method according to any one of claims 1 to 4, characterized in that, The distributed interface includes a data protocol broadcast interface, the first to-be-transmitted data includes sixth to-be-transmitted tensor data, and the second to-be-transmitted data includes seventh to-be-transmitted tensor data; The distributed interface receives first to-be-transmitted data of a first process in at least two distributed processes, including: the data protocol broadcast interface receives sixth to-be-transmitted tensor data and a protocol algorithm of the first process; The distributed interface transmits the first to-be-transmitted data or second to-be-transmitted data determined based on the to-be-transmitted data to a second process in the at least two distributed processes, including: the data protocol broadcast interface performs protocol calculation on the sixth to-be-transmitted tensor data based on the protocol algorithm to obtain seventh to-be-transmitted tensor data, and transmits the seventh to-be-transmitted tensor data to the second process.

14. The method of claim 11, wherein The data protocol broadcast interface receives sixth to-be-transmitted tensor data and a protocol algorithm of the first process, including: the data protocol broadcast interface receives sixth to-be-transmitted tensor data, a protocol algorithm, and a fifth process group identifier of the first process; The data protocol broadcast interface performs protocol calculation on the sixth to-be-transmitted tensor data based on the protocol algorithm to obtain seventh to-be-transmitted tensor data, and transmits the seventh to-be-transmitted tensor data to the second process, including: the data protocol broadcast interface performs protocol calculation on the fifth to-be-transmitted tensor data based on the protocol algorithm to obtain seventh to-be-transmitted tensor data, and transmits the seventh to-be-transmitted tensor data to the second process; the second process is all processes in a process group determined based on the fifth process group identifier.

15. The method according to any one of claims 1 to 4, characterized in that, The distributed interface includes a data collection broadcast interface, the first to-be-transmitted data includes eighth to-be-transmitted tensor data, and the second to-be-transmitted data includes a third to-be-transmitted tensor list; The distributed interface receives first to-be-transmitted data of a first process in at least two distributed processes, including: the data collection broadcast interface receives eighth to-be-transmitted tensor data of the first process; The distributed interface transmits the first to-be-transmitted data or second to-be-transmitted data determined based on the to-be-transmitted data to a second process in the at least two distributed processes, including: the data collection broadcast interface determines a third to-be-transmitted tensor list based on the eighth to-be-transmitted tensor data and transmits the third to-be-transmitted tensor list to the second process.

16. The method of claim 15, wherein The data collection broadcast interface receives eighth to-be-transmitted tensor data of the first process, including that the data collection broadcast interface receives the eighth to-be-transmitted tensor data of the first process and sixth process group identification; The data collection broadcast interface determines a third to-be-transmitted tensor list based on the eighth to-be-transmitted tensor data, and transmits to the second process based on the third to-be-transmitted tensor list, including: The data collection interface determines a third to-be-transmitted tensor list based on the eighth to-be-transmitted tensor data, and transmits to the second process based on the third to-be-transmitted tensor list; the second process is all processes in a to-be-collected process group determined based on the sixth process group identification.

17. The method according to any one of claims 1 to 4, characterized in that, The distributed interface includes a communication synchronization interface, and the method further includes: The communication synchronization interface sets the third part of processes to a waiting state in response to that the third part of processes in the at least two distributed processes run to a synchronization node and the fourth part of processes do not run to the synchronization node; The communication synchronization interface controls the at least two distributed processes to be in a running state in response to that the at least two distributed processes all run to the synchronization node.

18. The method according to any one of claims 1 to 4, characterized in that, The distributed interface includes a process number acquisition interface, and the method further includes: The process number acquisition interface outputs a process number of a current process in response to a process number acquisition request.

19. The method according to any one of claims 1 to 4, characterized in that, The distributed interface includes a process number acquisition interface, and the method further includes: The process number acquisition interface outputs a process number of a current task process in response to a process number acquisition interface request.

20. A model training apparatus, comprising: including: A distributed interface is configured to receive first to-be-transmitted data of a first process in at least two distributed processes; The distributed interface is further configured to transmit the first to-be-transmitted data or second to-be-transmitted data determined based on the to-be-transmitted data to a second process in the at least two distributed processes; The first process and the second process are at least one of the at least two distributed processes; The distributed interface is a standardized interface, and supports tensor data transmission in multiple communication modes; The distributed interface includes at least one of a data broadcast interface, a data protocol interface, a data scattering interface, a data collection interface, a data protocol broadcast interface, a data collection broadcast interface, a communication synchronization interface, a process number acquisition interface, and a process number acquisition interface. 21.A computer device, comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, and the computer device is characterized in that, The processor executes the program to implement the steps in the method of any one of claims 1 to 19.

22. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps in the method of any one of claims 1 to 19.

23. A computer program product, the computer program product comprising a non-transitory computer readable storage medium storing a computer program, the computer program being read and executed by a computer to implement the steps in the method of any one of claims 1 to 19.

Citation Information

Patent Citations

  • Execution configuration information generation method, model training method and device

    CN114202027A

  • Parameter updating method and device, electronic equipment and storage medium

    CN114356540A