Methods for updating network models, image classification, and language modeling

By determining the global synchronization period in distributed training and combining global and local model parameters synchronization, the Gossip SGD algorithm takes into account both the convergence speed and communication speed in large-scale distributed training, achieving faster training convergence and higher model accuracy.

CN114492152BActive Publication Date: 2025-05-13ALIBABA GROUP HOLDING LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202011157840.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-26
Publication Date
2025-05-13
Estimated Expiration
2040-10-26

AI Technical Summary

Technical Problem

In large-scale distributed training, the convergence speed of the Gossip SGD algorithm is affected by the topology and the number of nodes, and cannot take into account both the communication speed and the convergence. The existing technology cannot take into account both the convergence of large-scale distributed training tasks while ensuring the communication speed of distributed systems.

Method used

By determining the global synchronization cycle, controlling the update timing of the global model parameters of the network model to be trained in distributed training, combining global and local model parameters synchronization, and using a decentralized distributed training algorithm that accelerates convergence in global synchronization, adaptively synchronizes global or local model parameters.

Benefits of technology

On the basis of ensuring the communication speed of distributed systems, the convergence and training efficiency of large-scale distributed training tasks are improved, training time is reduced, and convergence and communication speed are better than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492152B_ABST
    Figure CN114492152B_ABST
Patent Text Reader

Abstract

The present application discloses a method for updating a network model, a method for image classification, and a method for language modeling. The method for updating a network model includes: determining a global synchronization period, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, the local model parameters of the network model to be trained are synchronized, and when the global synchronization period is reached, the global model parameters of the network model to be trained are synchronized. The present application solves the technical problem in the related art that when training a neural network model, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for updating a network model, a method for image classification, and a method for language modeling. Background Art

[0002] In related technologies, due to the progress of deep learning and the explosive growth of data sets, the performance of big data models and big data sets in tasks in scenarios such as computer vision and natural language processing has been significantly improved. Traditional single-machine training is far from meeting the current data volume requirements. Large-scale distributed training (multi-machine parallel) using the parallel SGD algorithm is widely used in industry and academia. Although All-Reduce-based parameter synchronization solves the central communication congestion of the parameter server, a single All-Reduce performs fine pipeline communication between all nodes, which easily leads to a large amount of synchronization overhead and delay. For these reasons, there is still a lot of room for improvement in the scalability of the distributed training system of parallel SGD in large-scale computing clusters.

[0003] In order to improve the scalability of the distributed training system of the parallel SGD algorithm, the processing method of the GossipSGD algorithm proposed in the relevant technology has received widespread attention. Compared with the parallel SGD algorithm based on All-Reduce, the Gossip SGD algorithm only requires synchronization information between nodes and neighbors. Under a specific topological structure, it can also fully utilize the bandwidth like all-reduce, thereby reducing the amount and number of single information communications.

[0004] However, since the convergence speed of the Gossip SGD algorithm is affected by the topology and the number of nodes, it often suffers from precision loss compared to the All-Reduce-based parallel SGD algorithm in large-scale training, and cannot take into account the convergence of large-scale distributed training tasks; that is, there is a technical problem in the relevant technology that it cannot take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0005] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0006] The embodiments of the present application provide a method for updating a network model, a method for image classification, and a method for language modeling, so as to at least solve the technical problem in the related art that, when training a neural network model, it is impossible to balance the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0007] According to one aspect of an embodiment of the present application, a method for updating a network model is provided, comprising: determining a global synchronization period, wherein the global synchronization period is used to control the update timing of global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, performing local model parameter synchronization on the network model to be trained, and when the global synchronization period is reached, performing global model parameter synchronization on the network model to be trained.

[0008] According to another aspect of an embodiment of the present application, a method for image classification is also provided, including: obtaining an image to be classified; inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating a network model to be trained in distributed training, and the updating timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0009] According to another aspect of an embodiment of the present application, another image classification method is provided, including: receiving an image to be classified from a client; inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating a network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; and the image classification result is fed back to the client.

[0010] According to another aspect of an embodiment of the present application, another image classification method is provided, including: sending an image to be classified to a server; receiving an image classification result fed back by the server, wherein the image classification result is obtained by the server inputting the image to be classified into a neural network model for image classification processing, the neural network model is a model obtained by updating a network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0011] According to another aspect of an embodiment of the present application, a language modeling method is also provided, including: obtaining a word sequence to be processed; inputting the above-mentioned word sequence to be processed into a neural network model for language modeling processing, and outputting a language modeling result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through a global synchronization cycle.

[0012] According to another aspect of an embodiment of the present application, another language modeling method is provided, including: receiving a word sequence to be processed from a client; inputting the above-mentioned word sequence to be processed into a neural network model for language modeling processing, and outputting a language modeling result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through a global synchronization cycle; and the above-mentioned language modeling result is fed back to the above-mentioned client.

[0013] According to another aspect of an embodiment of the present application, there is provided another method of language modeling, comprising: sending a word sequence to be processed to a server; receiving a language modeling result fed back by the server, wherein the language modeling result is obtained by the server inputting the word sequence to be processed into a neural network model for language modeling processing, the neural network model is a model obtained by updating a network model to be trained in distributed training, and the updating timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0014] According to another aspect of an embodiment of the present application, a non-volatile storage medium is also provided, wherein the non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute any one of the above-mentioned methods for updating the network model, the above-mentioned image classification method, and the above-mentioned language modeling method.

[0015] According to another aspect of an embodiment of the present application, a device for updating a network model is also provided, comprising: a processor; and a memory, connected to the above-mentioned processor, for providing the above-mentioned processor with instructions for processing the following processing steps: determining a global synchronization period, wherein the above-mentioned global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the above-mentioned global synchronization period is not reached, synchronizing the local model parameters of the above-mentioned network model to be trained, and when the above-mentioned global synchronization period is reached, synchronizing the global model parameters of the above-mentioned network model to be trained.

[0016] In an embodiment of the present application, a global synchronization period is determined, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, the local model parameters of the network model to be trained are synchronized, and when the global synchronization period is reached, the global model parameters of the network model to be trained are synchronized.

[0017] The embodiment of the present application proposes a decentralized distributed training algorithm that combines global synchronization to accelerate convergence. By determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, a neural network model is trained, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm. Moreover, under the same number of iterations, the method for updating the network model proposed in the embodiment of the present application has better convergence for large-scale distributed training tasks than the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce training time and improve communication speed.

[0018] Therefore, the embodiment of the present application achieves the purpose of improving the convergence and training time of distributed training tasks by determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, thereby achieving the technical effect of taking into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system, and further solves the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for updating a network model is shown;

[0021] Figure 2 is a flow chart of a method for updating a network model according to an embodiment of the present application;

[0022] Figure 3 is a flow chart of an image classification method according to an embodiment of the present application;

[0023] Figure 4 is a flowchart of another image classification method according to an embodiment of the present application;

[0024] Figure 5 is a flowchart of another image classification method according to an embodiment of the present application;

[0025] Figure 6 is a flow chart of a method for language modeling according to an embodiment of the present application;

[0026] Figure 7 is a flowchart of another language modeling method according to an embodiment of the present application;

[0027] Figure 8 is a flowchart of another language modeling method according to an embodiment of the present application;

[0028] Fig. 9 is a structural diagram of a device for updating a network model according to an embodiment of the present application;

[0029] Fig.10 is a structural schematic diagram of an image classification device according to an embodiment of the present application;

[0030] Fig.11 is a structural schematic diagram of a language modeling device according to an embodiment of the present application;

[0031] Fig.12 is a structural diagram of a device for updating a network model according to an embodiment of the present application;

[0032] Fig.13 It is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0035] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:

[0036] All-reduce: A commonly used communication operator (primitives) in distributed training. Different nodes use certain topological rules (Ring / Tree) to achieve full synchronization of information, such as averaging the gradients of different nodes.

[0037] Gossip: In large-scale training, the gossip algorithm refers to the situation in which each node exchanges information only with its immediate neighbors in one round of iterations. Compared with all-reduce / Parameter Server and other methods, the information between nodes is not completely synchronized. From a global perspective, the gossip algorithm achieves global synchronization through multiple local synchronizations.

[0038] Imagenet-1K: refers to a commonly used public dataset in the field of image classification, with 1,000 classifications and a training data volume of approximately 1.28 million images.

[0039] Wikipedia, Bookcorps: English language libraries commonly used in natural language processing, often used for tasks such as language modeling pre-training.

[0040] SGD: refers to a common optimization algorithm in machine learning, namely the asynchronous stochastic gradient descent algorithm, which is widely used in deep learning.

[0041] Local SGD: Also known as model average. In distributed training, each node performs a global parameter synchronization after performing several local parameter updates. Its advantage is that it reduces the frequency of communication, but its disadvantage is that it has poor convergence and the final trained model has poor accuracy.

[0042] Gossip SGD: refers to an optimization algorithm based on the rumor algorithm. Gossip SGD only requires nodes to synchronize information with their neighbors. Under a specific topology, it can fully utilize bandwidth similar to All-reduce, thereby reducing the amount and number of single information communications. However, its convergence speed is affected by the topology and the number of nodes, and in large-scale training, it often has a loss of accuracy compared to the parallel SGD algorithm based on All-reduce.

[0043] Parallel SGD: In distributed training, each node synchronizes gradients / parameters in each iteration, which has good convergence, but slow speed and poor scalability of distributed systems.

[0044] Example 1

[0045] According to an embodiment of the present application, a method embodiment for updating a network model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0046] The method embodiment provided in Embodiment 1 of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for updating a network model is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0047] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0048] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for updating the network model in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the above-mentioned method for updating the network model is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0049] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0050] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0051] It should be noted that compared with the method of updating the network model by performing data parallel training through the All-Reduce algorithm in related technologies, in Gossip training, the training nodes are connected through a certain logical topological structure. Assuming there are n nodes, an n×n topological matrix W can be used to describe the connection relationship between the nodes, that is, w ij represents the weight of information flow from node j to node i. ij If it is 0, it means that no information flows from node j to node i, which specifically refers to the transfer of model parameters in the training scenario of the embodiment of the present application.

[0052] For the kth Gossip iteration, for node i, firstly, the gradient is obtained by forward and reverse training of the model as in the common data parallel training In Gossip training, first according to the gradient Update the local model: After updating the model, the model of the current node is sent to some nodes according to the current topological matrix W, and the model parameters sent by some nodes are accepted, and the weighted sum is performed to obtain the final model of this round of training.

[0053] Among them, the above in, Represents the set of neighboring nodes of node i, including itself.

[0054] In each iteration, each node performs local communication between nodes according to the topology matrix W, which leads to differences in model parameters between different nodes. The sparser the communication topology between nodes, the more communication nodes there are, and the larger the model variance between nodes. The existence of variance between models will slow down the convergence rate. Therefore, under the same hyperparameter settings (training cycle, learning rate adjustment, etc.), the network model trained by the Gossip SGD algorithm has a problem of decreased accuracy on the verification set.

[0055] In the above operating environment, in order to solve the technical problem that the above Gossip SGD algorithm cannot achieve the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system, this application provides the following Figure 2 A method for updating a network model is shown. Figure 2 is a flow chart of a method for updating a network model according to an embodiment of the present application. Figure 2 As shown, the method for updating the network model includes:

[0056] Step S202, determining a global synchronization period, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training;

[0057] Step S204, when the global synchronization period is not reached, local model parameter synchronization is performed on the network model to be trained; and when the global synchronization period is reached, global model parameter synchronization is performed on the network model to be trained.

[0058] In an embodiment of the present application, a global synchronization period is determined, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, the local model parameters of the network model to be trained are synchronized, and when the global synchronization period is reached, the global model parameters of the network model to be trained are synchronized.

[0059] The embodiment of the present application proposes a decentralized distributed training algorithm that combines global synchronization to accelerate convergence. By determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, a neural network model is trained, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm. Moreover, under the same number of iterations, the method for updating the network model proposed in the embodiment of the present application has better convergence for large-scale distributed training tasks than the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce training time and improve communication speed.

[0060] Therefore, the embodiment of the present application achieves the purpose of improving the convergence and training time of distributed training tasks by determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, thereby achieving the technical effect of taking into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system, and further solves the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0061] In an optional embodiment, the above method is suitable for updating network models of different nodes in a distributed training scenario, and can be widely used in various distributed training scenarios, such as: object classification task scenarios in the field of image classification; language modeling task scenarios in natural language processing.

[0062] Through the embodiments of the present application, the update timing of the global model parameters of the network model to be trained is controlled in distributed training through the global synchronization cycle. The convergence speed can be improved by increasing the global synchronization cycle. The superiority of the algorithm proposed in the embodiments of the present application has been verified in tasks such as object classification in the field of image classification (Imagenet) and language modeling in natural language processing (Wikipedia).

[0063] In the implementation of this application, by introducing the idea of ​​periodic global synchronization, although the time taken for a round of periodic global synchronization is equivalent to that of the All-reduce algorithm, since the global synchronization period can be relatively large and after introducing adaptation for local model parameter synchronization or global model parameter synchronization, the rounds of full synchronization can be reduced, in the training time dimension, it is not much slower than the pure Gossip SGD algorithm. For example, in the task scenario of object classification, the introduction of global synchronization only takes 7% more time to complete the same number of iterations, but it makes up for the accuracy loss caused by the training of the pure Gossip SGD algorithm.

[0064] It should be noted that, although appropriately extending the training cycle can also make up for the loss of accuracy caused by the slow convergence speed, based on the experimental results of the embodiments of the present application, the total training time of the training method that introduces adaptive periodic global synchronization is shorter. The embodiments of the present application introduce an adaptive periodic global synchronization method in decentralized distributed training to accelerate convergence, which can prove the advantages of the method for updating the network model provided by the embodiments of the present application in theory and experiments.

[0065] In an optional embodiment, synchronizing local model parameters of the network model to be trained includes:

[0066] Step S302, using a first stochastic gradient descent algorithm to synchronize local model parameters of nodes that perform local communication according to a topology matrix in the network model to be trained, wherein the first stochastic gradient descent algorithm is an asynchronous stochastic gradient descent algorithm based on a rumor algorithm.

[0067] As an optional embodiment, the parallel stochastic gradient descent algorithm (Algorithm 1) Gossip-PGA (Periodic Global Averaging) based on the communication algorithm in distributed training is as follows:

[0068]

[0069] Based on the above Gossip SGD algorithm training, the embodiment of the present application introduces a hyperparameter H, i.e., the global synchronization period, which is used to control the update timing of the global model parameters of the network model to be trained in distributed training. For example, after every H-1 Gossip SGD algorithm communication, the global model parameter synchronization is performed through the All-Reduce algorithm to train the neural network model: This neural network model can eliminate the model variance between nodes, thereby accelerating convergence (as shown in lines 4-7 of the code of Algorithm 1 above).

[0070] In an optional embodiment, synchronizing global model parameters of the network model to be trained includes:

[0071] Step S402: Using a second stochastic gradient descent algorithm to synchronize global model parameters of the network model to be trained, wherein the second stochastic gradient descent algorithm is a parallel stochastic gradient descent algorithm based on a communication algorithm in distributed training.

[0072] The inventors of the embodiments of the present application observed in actual training that since the variance between models is large in the early stages of training, as the learning rate decreases and the models gradually converge, the variance between models gradually decreases. As mentioned above, the model variance will reduce the convergence speed, so the frequency of global model parameter synchronization should be more frequent in the early stages of training. As the training progresses, the frequency of global model parameter synchronization can be gradually reduced.

[0073] Therefore, the embodiment of the present application further proposes an asynchronous stochastic gradient descent algorithm (Algorithm 2) Gossip-AGA (Adaptive Global Averaging) based on the rumor algorithm to avoid artificially selecting the optimal full communication cycle H, and to minimize the frequency of full communication without affecting the convergence of the model. The Algorithm 2 is as follows:

[0074]

[0075]

[0076] In an optional embodiment, the above method further includes:

[0077] Step S502, performing an initial loss function estimation in a preset number of iterations to obtain a first estimation result;

[0078] Step S504, estimating the current loss function using the current model parameters and the current sampled data to obtain a second estimation result;

[0079] Step S506: using the first estimation result, the second estimation result and the initial value of the global synchronization period, to obtain a target value of the global synchronization period.

[0080] In the early stage of training, the present embodiment will select a preset number of iterations K w The initial loss function F is estimated (as shown in lines 10-11 of the code of Algorithm 2 above) to obtain a first estimation result; the current loss function is estimated by the current model parameter x_i and the current sampled data ξ_i to obtain a second estimation result; using the above first estimation result, the above second estimation result and the initial value of the above global synchronization period, in subsequent training, the iterative global synchronization period H is adaptively adjusted according to the following calculation formula:

[0081] in represents rounding up, H_init is the initial set iteration period and F_init is the model loss function value obtained in the hot start phase, and the remaining iteration nodes communicate based on Gossip according to the topology structure (as shown in line 15 of the code of Algorithm 2 above).

[0082] In an optional embodiment, the above method further includes:

[0083] Step S602, setting the count value of the counter;

[0084] Step S604, comparing the count value with the target value;

[0085] Step S606, when the above-mentioned count value is not equal to the above-mentioned target value, the first stochastic gradient descent algorithm is used to synchronize the local model parameters of the nodes that perform local communication according to the topology matrix in the above-mentioned network model to be trained; when the above-mentioned count value is equal to the above-mentioned target value, the second stochastic gradient descent algorithm is used to synchronize the global model parameters of the above-mentioned network model to be trained, and the above-mentioned count value is cleared.

[0086] Optionally, the above-mentioned counter is used to count the number of Gossip communications since the last global synchronization, that is, the current number of Gossip communications. Only when the count value is equal to the target value, the second stochastic gradient descent algorithm is used to synchronize the global model parameters of the above-mentioned network model to be trained, and the above-mentioned count value is cleared.

[0087] In the embodiment of the present application, a counter C is used to count the number of Gossip communications since the last global synchronization (as shown in lines 2 and 7 of the code of Algorithm 2 above), and the global synchronization period H is initialized to a relatively small value H init (As shown in lines 2-4 of the code of Algorithm 2 above), when the count value of the counter is equal to the current cycle number (ie, the target value) of the global synchronization cycle H, the embodiment of the present application will perform a global synchronization (as shown in lines 6-13 of the code of Algorithm 2 above).

[0088] For example, when the above-mentioned count value is not equal to the above-mentioned target value, the first stochastic gradient descent algorithm is used to synchronize local model parameters of the nodes that perform local communication according to the topology matrix in the above-mentioned network model to be trained; when the above-mentioned count value is equal to the above-mentioned target value, the second stochastic gradient descent algorithm is used to synchronize global model parameters of the above-mentioned network model to be trained, and the above-mentioned count value is cleared to zero.

[0089] Under the above operating environment, this application provides Figure 3 A method for image classification is shown in FIG. Figure 3 is a flow chart of an image classification method according to an embodiment of the present application. Figure 3 As shown, the image classification method includes:

[0090] Step S702, obtaining an image to be classified;

[0091] Step S704, input the above-mentioned image to be classified into the neural network model for image classification processing, and output the image classification result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in the distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through the global synchronization cycle.

[0092] In an optional embodiment, the above method can be applied to but is not limited to the following scenarios: object classification task scenarios in the field of image classification, controlling the update timing of the global model parameters of the network model to be trained in distributed training through the global synchronization cycle, and increasing the global synchronization cycle to improve the convergence speed. In the task scenario of object classification (Imagnet) in the field of image classification, the superiority of the image classification algorithm provided by the embodiment of the present application is verified.

[0093] The embodiment of the present application proposes a decentralized distributed training algorithm that combines global synchronization to accelerate convergence. By determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, a neural network model is trained to solve the problem of slow convergence caused by the decentralized distributed training algorithm. Moreover, under the same number of iterations, the convergence of the embodiment of the present application for large-scale distributed training tasks is better than that of the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce training time and improve communication speed. The above-mentioned image to be classified is input into the neural network model for image classification processing, which can improve the accuracy of the output image classification result and reduce the time of image classification processing.

[0094] Therefore, the embodiments of the present application achieve the purpose of improving the convergence and training time of distributed training tasks when training neural network models, thereby achieving the technical effect of improving the accuracy of output image classification results and reducing the time of image classification processing, and further solving the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0095] Under the above operating environment, this application provides Figure 4 Another image classification method shown in Figure 4 is a flowchart of another image classification method according to an embodiment of the present application. Figure 4 As shown, the image classification method includes:

[0096] Step S802, receiving an image to be classified from a client;

[0097] Step S804, inputting the above-mentioned image to be classified into the neural network model for image classification processing, and outputting the image classification result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in the distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through the global synchronization cycle;

[0098] Step S806: Feedback the image classification result to the client.

[0099] It should be noted that the executor of the image classification method provided in the above steps S802 to S806 is the server. The image classification method can be but is not limited to being applicable to the following scenarios: object classification task scenarios within the field of image classification. The update timing of the global model parameters of the network model to be trained is controlled in distributed training through the global synchronization cycle. The convergence speed can be improved by increasing the global synchronization cycle. In the task scenario of object classification (Imagnet) in the field of image classification, the superiority of the image classification algorithm provided by the embodiment of the present application is verified.

[0100] The embodiment of the present application proposes a decentralized distributed training algorithm combined with global synchronous accelerated convergence. The server receives the image to be classified from the client, inputs the image to be classified into the neural network model for image classification processing, and feeds back the output image classification result to the client.

[0101] Since the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the updating timing of the global model parameters of the above-mentioned network model to be trained is controlled by the global synchronization cycle in the above-mentioned distributed training, the neural network model is trained by determining the global synchronization cycle and adaptively performing global model parameter synchronization or local model parameter synchronization, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm.

[0102] Moreover, under the same number of iterations, the convergence of the embodiment of the present application for large-scale distributed training tasks is better than that of the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce the training time and improve the communication speed. Then, the above-mentioned images to be classified are input into the neural network model for image classification processing, which can improve the accuracy of the output image classification results and reduce the time of image classification processing.

[0103] Therefore, the embodiments of the present application achieve the purpose of improving the convergence and training time of distributed training tasks when training neural network models, thereby achieving the technical effect of improving the accuracy of output image classification results and reducing the time of image classification processing, and further solving the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0104] Under the above operating environment, this application provides Figure 5 Another image classification method shown in Figure 5 is a flowchart of another image classification method according to an embodiment of the present application. Figure 5 As shown, the image classification method includes:

[0105] Step S902, sending the image to be classified to the server;

[0106] Step S904, receiving the image classification result fed back by the server, wherein the image classification result is obtained by the server inputting the image to be classified into the neural network model for image classification processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through the global synchronization cycle.

[0107] It should be noted that the executor of the image classification method provided in the above steps S902 to S906 is the client. The image classification method can be but is not limited to being applicable to the following scenarios: object classification task scenarios within the field of image classification. The update timing of the global model parameters of the network model to be trained is controlled in distributed training through the global synchronization cycle. The convergence speed can be improved by increasing the global synchronization cycle. In the task scenario of object classification (Imagnet) in the field of image classification, the superiority of the image classification algorithm provided by the embodiment of the present application is verified.

[0108] The embodiment of the present application proposes a decentralized distributed training algorithm combined with global synchronous accelerated convergence, where the client sends an image to be classified to the server; the server inputs the image to be classified into a neural network model for image classification processing, and feeds back the output image classification result to the client.

[0109] Since the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the updating timing of the global model parameters of the above-mentioned network model to be trained is controlled by the global synchronization cycle in the above-mentioned distributed training, the neural network model is trained by determining the global synchronization cycle and adaptively performing global model parameter synchronization or local model parameter synchronization, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm.

[0110] Moreover, under the same number of iterations, the convergence of the embodiment of the present application for large-scale distributed training tasks is better than that of the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce the training time and improve the communication speed. Then, the above-mentioned images to be classified are input into the neural network model for image classification processing, which can improve the accuracy of the output image classification results and reduce the time of image classification processing.

[0111] Therefore, the embodiments of the present application achieve the purpose of improving the convergence and training time of distributed training tasks when training neural network models, thereby achieving the technical effect of improving the accuracy of output image classification results and reducing the time of image classification processing, and further solving the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0112] It should also be noted that the preferred implementation of the above-mentioned image classification method embodiment can refer to the relevant description in the method embodiment of updating the network model, which will not be repeated here.

[0113] Under the above operating environment, this application provides Figure 6 A language modeling method is shown in Figure 6 is a flow chart of a method for language modeling according to an embodiment of the present application, such as Figure 6 As shown, the language modeling method includes:

[0114] Step S1002, obtaining a word sequence to be processed;

[0115] Step S1004, input the above-mentioned word sequence to be processed into the neural network model for language modeling processing, and output the language modeling result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through the global synchronization cycle.

[0116] In an optional embodiment, the above method is applicable to the following scenarios: language modeling task scenarios in natural language processing, in which the update timing of the global model parameters of the network model to be trained is controlled in distributed training through the global synchronization cycle, and the convergence speed can be improved by increasing the global synchronization cycle. In the task scenario of language modeling (Wikipedia) in natural language processing, the superiority of the language modeling algorithm provided by the embodiment of the present application is verified.

[0117] The embodiment of the present application proposes a decentralized distributed training algorithm that combines global synchronization to accelerate convergence. By determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, a neural network model is trained to solve the problem of slow convergence caused by the decentralized distributed training algorithm. Moreover, under the same number of iterations, the convergence of the embodiment of the present application for large-scale distributed training tasks is better than that of the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce training time and improve communication speed. The above-mentioned word sequence to be processed is input into the neural network model for language modeling processing, which can improve the accuracy of the output language modeling results and reduce the time of language modeling processing.

[0118] Therefore, the embodiments of the present application achieve the purpose of improving the convergence and training time of distributed training tasks when training neural network models, thereby achieving the technical effect of improving the accuracy of language modeling results and reducing the time of language modeling processing, and further solving the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0119] Under the above operating environment, this application provides Figure 7 Another approach to language modeling is shown in Figure 6 is a flowchart of another language modeling method according to an embodiment of the present application. Figure 7 As shown, the language modeling method includes:

[0120] Step S1102, receiving a word sequence to be processed from a client;

[0121] Step S1104, inputting the above-mentioned word sequence to be processed into the neural network model for language modeling processing, and outputting the language modeling result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in the distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through the global synchronization cycle;

[0122] Step S1106, feeding back the above language modeling result to the above client.

[0123] It should be noted that the execution subject of the language modeling method provided by the above steps S1102 to S1106 is the server. The language modeling method can be applied to but is not limited to the following scenarios: language modeling task scenarios in natural language processing. The update timing of the global model parameters of the network model to be trained is controlled in distributed training through the global synchronization cycle. The convergence speed can be improved by increasing the global synchronization cycle. In the task scenario of language modeling (Wikipedia) in natural language processing, the superiority of the language modeling algorithm provided by the embodiment of the present application is verified.

[0124] The embodiment of the present application proposes a decentralized distributed training algorithm combined with global synchronization accelerated convergence. The server receives a word sequence to be processed from a client; the word sequence to be processed is input into a neural network model for language modeling, and the output language modeling result is fed back to the client.

[0125] Since the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the updating timing of the global model parameters of the above-mentioned network model to be trained is controlled by the global synchronization cycle in the above-mentioned distributed training, the neural network model is trained by determining the global synchronization cycle and adaptively performing global model parameter synchronization or local model parameter synchronization, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm.

[0126] Moreover, under the same number of iterations, the convergence of the embodiment of the present application for large-scale distributed training tasks is better than that of the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce the training time and improve the communication speed. Then, the above-mentioned word sequence to be processed is input into the neural network model for language modeling processing, which can improve the accuracy of the output language modeling results and reduce the time of language modeling processing.

[0127] Therefore, the embodiments of the present application achieve the purpose of improving the convergence and training time of distributed training tasks when training neural network models, thereby achieving the technical effect of improving the accuracy of language modeling results and reducing the time of language modeling processing, and further solving the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0128] Under the above operating environment, this application provides Figure 8 Another language modeling method shown is Figure 8 is a flowchart of another language modeling method according to an embodiment of the present application. Figure 8As shown, the language modeling method includes:

[0129] Step S1202, sending the word sequence to be processed to the server;

[0130] Step S1204, receiving the language modeling result fed back by the server, wherein the language modeling result is obtained by the server inputting the word sequence to be processed into the neural network model for language modeling processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through the global synchronization cycle.

[0131] It should be noted that the executor of the language modeling method provided in the above steps S1202 to S1206 is the client. The language modeling method can be applied to but is not limited to the following scenarios: language modeling task scenarios in natural language processing. The update timing of the global model parameters of the network model to be trained is controlled in distributed training through the global synchronization cycle. The convergence speed can be improved by increasing the global synchronization cycle. In the task scenario of language modeling (Wikipedia) in natural language processing, the superiority of the language modeling algorithm provided in the embodiment of the present application is verified.

[0132] The embodiment of the present application proposes a decentralized distributed training algorithm combined with global synchronization accelerated convergence, where the client sends a word sequence to be processed to the server; the server inputs the above-mentioned word sequence to be processed into a neural network model for language modeling processing, and feeds back the output language modeling results to the client.

[0133] Since the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the updating timing of the global model parameters of the above-mentioned network model to be trained is controlled by the global synchronization cycle in the above-mentioned distributed training, the neural network model is trained by determining the global synchronization cycle and adaptively performing global model parameter synchronization or local model parameter synchronization, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm.

[0134] Moreover, under the same number of iterations, the convergence of the embodiment of the present application for large-scale distributed training tasks is better than that of the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce the training time and improve the communication speed. Then, the above-mentioned word sequence to be processed is input into the neural network model for language modeling processing, which can improve the accuracy of the output language modeling results and reduce the time of language modeling processing.

[0135] Therefore, the embodiments of the present application achieve the purpose of improving the convergence and training time of distributed training tasks when training neural network models, thereby achieving the technical effect of improving the accuracy of language modeling results and reducing the time of language modeling processing, and further solving the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0136] It should also be noted that the preferred implementation of the above-mentioned language modeling method embodiment can refer to the relevant description in the method embodiment of updating the network model, which will not be repeated here.

[0137] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0138] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a non-volatile storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the above methods of each embodiment of the present application.

[0139] Example 2

[0140] According to an embodiment of the present application, a device embodiment for implementing the above-mentioned method for updating the network model is also provided. Fig. 9 is a schematic diagram of a structure of a device for updating a network model according to an embodiment of the present application, such as Fig. 9 As shown, the device for updating the network model includes: a determination module 500, an update module 502, wherein:

[0141] The determination module 500 is used to determine the global synchronization period, wherein the above-mentioned global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; the update module 502 is used to synchronize the local model parameters of the network model to be trained when the above-mentioned global synchronization period is not reached, and synchronize the global model parameters of the network model to be trained when the above-mentioned global synchronization period is reached.

[0142] It should be noted that the above-mentioned determination module 500 and update module 502 correspond to steps S202 to S204 in Example 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computer terminal 10 provided in Example 1.

[0143] According to an embodiment of the present application, a device embodiment for implementing the above-mentioned image classification method is also provided. Fig.10 is a schematic diagram of the structure of an image classification device according to an embodiment of the present application, such as Fig.10 As shown, the image classification device includes: an image acquisition module 600, a classification processing module 602, wherein:

[0144] The image acquisition module 600 is used to acquire the image to be classified; the classification processing module 602 is used to input the above-mentioned image to be classified into the neural network model for image classification processing, and output the image classification result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in the distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through the global synchronization cycle.

[0145] It should be noted that the above-mentioned image acquisition module 600 and classification processing module 602 correspond to steps S702 to S704 in Example 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computer terminal 10 provided in Example 1.

[0146] According to an embodiment of the present application, a device embodiment for implementing the above-mentioned language modeling method is also provided. Fig.11 is a structural diagram of a language modeling device according to an embodiment of the present application, such as Fig.11 As shown, the language modeling device includes: a sequence acquisition module 700, a modeling processing module 702, wherein:

[0147] The sequence acquisition module 700 is used to acquire the word sequence to be processed; the modeling processing module 702 is used to input the above-mentioned word sequence to be processed into the neural network model for language modeling processing, and output the language modeling result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in the distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through the global synchronization cycle.

[0148] It should be noted that the above sequence acquisition module 700 and modeling processing module 702 correspond to steps S1002 to S1004 in Example 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0149] It should also be noted that the preferred implementation of this embodiment can refer to the relevant description in Example 1, which will not be repeated here.

[0150] Example 3

[0151] According to an embodiment of the present application, an embodiment of a device for updating a network model is also provided. The device for updating a network model may be any computing device in a computing device group. Fig.12 is a structural diagram of a device for updating a network model according to an embodiment of the present application, such as Fig.12 As shown, the device for updating the network model includes: a processor 800 and a memory 802, wherein:

[0152] Processor 800; and memory 802, connected to the processor 800, for providing instructions for the processor to process the following processing steps: determining a global synchronization period, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, synchronizing the local model parameters of the network model to be trained, and when the global synchronization period is reached, synchronizing the global model parameters of the network model to be trained.

[0153] In an embodiment of the present application, a global synchronization period is determined, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, the local model parameters of the network model to be trained are synchronized, and when the global synchronization period is reached, the global model parameters of the network model to be trained are synchronized.

[0154] The embodiment of the present application proposes a decentralized distributed training algorithm that combines global synchronization to accelerate convergence. By determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, a neural network model is trained, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm. Moreover, under the same number of iterations, the method for updating the network model proposed in the embodiment of the present application has better convergence for large-scale distributed training tasks than the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce training time and improve communication speed.

[0155] Therefore, the embodiment of the present application achieves the purpose of improving the convergence and training time of distributed training tasks by determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, thereby achieving the technical effect of taking into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system, and further solves the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0156] It should also be noted that the preferred implementation of this embodiment can refer to the relevant description in Example 1, which will not be repeated here.

[0157] Example 4

[0158] According to an embodiment of the present application, an embodiment of a computer terminal is also provided, and the computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above-mentioned computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0159] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.

[0160] In this embodiment, the computer terminal can execute the following program codes of the method for updating the network model: determining a global synchronization period, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, performing local model parameter synchronization on the network model to be trained, and when the global synchronization period is reached, performing global model parameter synchronization on the network model to be trained.

[0161] Optionally, Fig.13 is a structural block diagram of a computer terminal according to an embodiment of the present application, such as Fig.13As shown, the computer terminal may include: one or more (only one is shown in the figure) processors 902 , a memory 904 , and a peripheral interface 906 .

[0162] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for updating the network model in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned method for updating the network model. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0163] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: determine the global synchronization period, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, synchronize the local model parameters of the network model to be trained, and when the global synchronization period is reached, synchronize the global model parameters of the network model to be trained.

[0164] Optionally, the processor may also execute the program code of the following steps: using a first stochastic gradient descent algorithm to synchronize local model parameters of nodes that perform local communication according to a topology matrix in the network model to be trained, wherein the first stochastic gradient descent algorithm is an asynchronous stochastic gradient descent algorithm based on a rumor algorithm.

[0165] Optionally, the processor may also execute the program code of the following steps: using a second stochastic gradient descent algorithm to perform global model parameter synchronization on the network model to be trained, wherein the second stochastic gradient descent algorithm is a parallel stochastic gradient descent algorithm based on the communication algorithm in distributed training.

[0166] Optionally, the processor may also execute the program code of the following steps: performing initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization period to obtain the target value of the global synchronization period.

[0167] Optionally, the processor may also execute the program code of the following steps: setting the count value of the counter; comparing the count value with the target value; when the count value is not equal to the target value, using a first stochastic gradient descent algorithm to perform local model parameter synchronization on nodes that perform local communication according to a topology matrix in the network model to be trained; when the count value is equal to the target value, using a second stochastic gradient descent algorithm to perform global model parameter synchronization on the network model to be trained, and clearing the count value.

[0168] Optionally, the processor may also execute the program code of the following steps: obtaining an image to be classified; inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the updating timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0169] Optionally, the processor may also execute the program code of the following steps: obtaining a word sequence to be processed; inputting the word sequence to be processed into a neural network model for language modeling processing, and outputting a language modeling result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the timing of updating the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0170] By adopting the embodiment of the present application, a scheme for updating a network model is provided. By determining a global synchronization period, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, the local model parameters of the network model to be trained are synchronized, and when the global synchronization period is reached, the global model parameters of the network model to be trained are synchronized.

[0171] The embodiment of the present application proposes a decentralized distributed training algorithm that combines global synchronization to accelerate convergence. By determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, a neural network model is trained, which can solve the problem of slow convergence caused by the decentralized distributed training algorithm. Moreover, under the same number of iterations, the method for updating the network model proposed in the embodiment of the present application has better convergence for large-scale distributed training tasks than the decentralized distributed training algorithm. Compared with the synchronous distributed training algorithm, it can not only achieve similar convergence, but also greatly reduce training time and improve communication speed.

[0172] Therefore, the embodiment of the present application achieves the purpose of improving the convergence and training time of distributed training tasks by determining the global synchronization period and adaptively performing global model parameter synchronization or local model parameter synchronization, thereby achieving the technical effect of taking into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system, and further solves the technical problem in the related technology that when training neural network models, it is impossible to take into account the convergence of large-scale distributed training tasks while ensuring the communication speed of the distributed system.

[0173] Optionally, the processor may also call the information and application programs stored in the memory through the transmission device to perform the following steps: receiving an image to be classified from a client; inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; and the image classification result is fed back to the client.

[0174] Optionally, the processor may also call the information and application programs stored in the memory through the transmission device to perform the following steps: sending the image to be classified to the server; receiving the image classification result fed back by the server, wherein the image classification result is obtained by the server inputting the image to be classified into the neural network model for image classification processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through the global synchronization cycle.

[0175] Optionally, the processor may also call the information and application programs stored in the memory through the transmission device to perform the following steps: receiving a word sequence to be processed from the client; inputting the word sequence to be processed into the neural network model for language modeling processing, and outputting a language modeling result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; and the language modeling result is fed back to the client.

[0176] Optionally, the processor may also call the information and application programs stored in the memory through the transmission device to perform the following steps: sending the word sequence to be processed to the server; receiving the language modeling result fed back by the server, wherein the language modeling result is obtained by the server inputting the word sequence to be processed into the neural network model for language modeling processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through the global synchronization cycle.

[0177] It can be understood by those skilled in the art that Fig.13 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig.13 It does not limit the structure of the above electronic device. For example, the computer terminal may also include Fig.13 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.13 Different configurations shown.

[0178] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable non-volatile storage medium. The non-volatile storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0179] Example 5

[0180] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium can be used to store the program code executed by the method for updating the network model, the method for image classification, and the method for language modeling provided in the above embodiment 1.

[0181] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0182] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: determining a global synchronization period, wherein the global synchronization period is used to control the update timing of the global model parameters of the network model to be trained in distributed training; when the global synchronization period is not reached, performing local model parameter synchronization on the network model to be trained, and when the global synchronization period is reached, performing global model parameter synchronization on the network model to be trained.

[0183] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: using a first stochastic gradient descent algorithm to synchronize local model parameters of nodes that perform local communication according to a topology matrix in the above-mentioned network model to be trained, wherein the above-mentioned first stochastic gradient descent algorithm is an asynchronous stochastic gradient descent algorithm based on a rumor algorithm.

[0184] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: using a second stochastic gradient descent algorithm to synchronize global model parameters of the above-mentioned network model to be trained, wherein the above-mentioned second stochastic gradient descent algorithm is a parallel stochastic gradient descent algorithm based on the communication algorithm in distributed training.

[0185] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation using current model parameters and current sampling data to obtain a second estimation result; and using the above-mentioned first estimation result, the above-mentioned second estimation result and the initial value of the above-mentioned global synchronization period to obtain the target value of the above-mentioned global synchronization period.

[0186] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: setting a count value of a counter; comparing the count value with the target value; when the count value is not equal to the target value, using a first stochastic gradient descent algorithm to perform local model parameter synchronization on nodes in the network model to be trained that perform local communication according to a topology matrix; when the count value is equal to the target value, using a second stochastic gradient descent algorithm to perform global model parameter synchronization on the network model to be trained, and clearing the count value.

[0187] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: obtaining an image to be classified; inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0188] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: obtaining a word sequence to be processed; inputting the above-mentioned word sequence to be processed into a neural network model for language modeling processing, and outputting a language modeling result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the above-mentioned network model to be trained is controlled in the above-mentioned distributed training through a global synchronization cycle.

[0189] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: receiving an image to be classified from a client; inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; and the image classification result is fed back to the client.

[0190] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: sending an image to be classified to a server; receiving an image classification result fed back by the server, wherein the image classification result is obtained by the server inputting the image to be classified into a neural network model for image classification processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0191] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: receiving a word sequence to be processed from a client; inputting the above-mentioned word sequence to be processed into a neural network model for language modeling processing, and outputting a language modeling result, wherein the above-mentioned neural network model is a model obtained by updating the network model to be trained in distributed training, and controlling the update timing of the global model parameters of the above-mentioned network model to be trained in the above-mentioned distributed training through a global synchronization cycle; and feeding back the above-mentioned language modeling result to the above-mentioned client.

[0192] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: sending a word sequence to be processed to a server; receiving a language modeling result fed back by the server, wherein the language modeling result is obtained by the server inputting the word sequence to be processed into a neural network model for language modeling processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle.

[0193] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0194] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0195] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the above-mentioned units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0196] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0197] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0198] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the above-mentioned methods of each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0199] The above are only preferred implementations of the present application. It should be pointed out that ordinary technicians in this technical field can make several improvements and modifications without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for updating a network model, characterized in that: include: Determine a global synchronization cycle, wherein the global synchronization cycle is used to control the update timing of the global model parameters of the network model to be trained in distributed training, wherein the neural network model obtained after the network model to be trained is updated is used to process the image to be classified in the object classification task scenario in the image classification field, or the word sequence to be processed in the language modeling task scenario in natural language processing; When the global synchronization period is not reached, local model parameter synchronization is performed on the network model to be trained, and when the global synchronization period is reached, global model parameter synchronization is performed on the network model to be trained; The method further comprises: obtaining a target value of the global synchronization cycle; setting a count value of a timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.

2. The method according to claim 1, characterized in that Synchronizing local model parameters of the network model to be trained includes: A first stochastic gradient descent algorithm is used to synchronize local model parameters of nodes that perform local communication according to a topology matrix in the network model to be trained, wherein the first stochastic gradient descent algorithm is an asynchronous stochastic gradient descent algorithm based on a rumor algorithm.

3. The method according to claim 1, characterized in that Synchronizing global model parameters of the network model to be trained includes: A second stochastic gradient descent algorithm is used to synchronize global model parameters of the network model to be trained, wherein the second stochastic gradient descent algorithm is a parallel stochastic gradient descent algorithm based on a communication algorithm in distributed training.

4. The method according to claim 1, characterized in that The method further comprises: When it is determined according to the comparison result that the count value is not equal to the target value, a first stochastic gradient descent algorithm is used to synchronize local model parameters of nodes that perform local communication according to the topology matrix in the network model to be trained; when it is determined according to the comparison result that the count value is equal to the target value, a second stochastic gradient descent algorithm is used to synchronize global model parameters of the network model to be trained, and the count value is cleared.

5. A method for image classification, characterized in that: include: Get the image to be classified; Inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; Among them, controlling the update timing of the global model parameters of the network model to be trained in the distributed training through the global synchronization cycle includes: obtaining the target value of the global synchronization cycle; setting the count value of the timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.

6. A method for image classification, characterized in that: include: Receive the image to be classified from the client; Inputting the image to be classified into a neural network model for image classification processing, and outputting an image classification result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; Among them, controlling the update timing of the global model parameters of the network model to be trained in the distributed training through the global synchronization cycle includes: obtaining the target value of the global synchronization cycle; setting the count value of the timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Feeding back the image classification result to the client; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.

7. A method for image classification, characterized in that: include: Send the image to be classified to the server; Receive the image classification result fed back by the server, wherein the image classification result is obtained by the server inputting the image to be classified into a neural network model for image classification processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; Among them, controlling the update timing of the global model parameters of the network model to be trained in the distributed training through the global synchronization cycle includes: obtaining the target value of the global synchronization cycle; setting the count value of the timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.

8. A method for language modeling, characterized in that: include: Get the word sequence to be processed; Inputting the word sequence to be processed into a neural network model for language modeling processing, and outputting a language modeling result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; Among them, controlling the update timing of the global model parameters of the network model to be trained in the distributed training through the global synchronization cycle includes: obtaining the target value of the global synchronization cycle; setting the count value of the timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.

9. A method for language modeling, characterized in that: include: Receive a word sequence to be processed from a client; Inputting the word sequence to be processed into a neural network model for language modeling processing, and outputting a language modeling result, wherein the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; Among them, controlling the update timing of the global model parameters of the network model to be trained in the distributed training through the global synchronization cycle includes: obtaining the target value of the global synchronization cycle; setting the count value of the timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Feeding back the language modeling result to the client; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.

10. A method for language modeling, characterized in that: include: Send the word sequence to be processed to the server; Receive the language modeling result fed back by the server, wherein the language modeling result is obtained by the server inputting the word sequence to be processed into a neural network model for language modeling processing, the neural network model is a model obtained by updating the network model to be trained in distributed training, and the update timing of the global model parameters of the network model to be trained is controlled in the distributed training through a global synchronization cycle; Among them, controlling the update timing of the global model parameters of the network model to be trained in the distributed training through the global synchronization cycle includes: obtaining the target value of the global synchronization cycle; setting the count value of the timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein, when the program is running, the device where the non-volatile storage medium is located is controlled to execute the method for updating the network model described in any one of claims 1 to 4, the method for image classification described in any one of claims 5 to 7, and the method for language modeling described in any one of claims 8 to 10.

12. A device for updating a network model, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Determine a global synchronization cycle, wherein the global synchronization cycle is used to control the update timing of the global model parameters of the network model to be trained in distributed training, wherein the neural network model obtained after the network model to be trained is updated is used to process the image to be classified in the object classification task scenario in the image classification field, or the word sequence to be processed in the language modeling task scenario in natural language processing; When the global synchronization period is not reached, local model parameter synchronization is performed on the network model to be trained, and when the global synchronization period is reached, global model parameter synchronization is performed on the network model to be trained; The instruction further includes: obtaining a target value of the global synchronization cycle; setting a count value of a timer; comparing the count value with the target value to obtain a comparison result; when it is determined according to the comparison result that the count value is equal to the target value, updating the global model parameters of the network model to be trained; Among them, obtaining the target value of the global synchronization cycle includes: performing an initial loss function estimation within a preset number of iterations to obtain a first estimation result; performing a current loss function estimation through current model parameters and current sampling data to obtain a second estimation result; and using the first estimation result, the second estimation result and the initial value of the global synchronization cycle to obtain the target value of the global synchronization cycle.