Serialization model parallel training method based on fixed point iteration
By adopting fixed point iteration and parallel prefixes and algorithms in the sequence model, the bottleneck problems existing in the parallel training of the sequence model are solved, efficient parallel training is achieved, and the training speed and adaptability of the model are improved.
Patent Information
- Application Number
- CN202510194842.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
AI Technical Summary
Sequential models such as RNN and Neural ODE have bottlenecks in parallel training, resulting in limited training speed and efficiency. In particular, the serial characteristics of the nonlinear recursive layer make parallelization difficult to achieve.
Using a method based on fixed point iteration and parallel prefix and algorithm, the steady-state solution of the sequence model is iteratively approximates the sequence calculation task and decomposes the sequence calculation task into independent subtasks, efficient parallel training of the sequence model is achieved.
It effectively reduces the training time of the sequence model, improves the convergence speed of the model, and adapts to more models, not only optimizes the training performance of RNN and Neural ODE, but also provides new ideas for the efficient training and application of future sequence models.
Smart Images

Figure CN120087439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and model parallelization, and in particular to a serial model parallel training method based on fixed-point iteration. Background Art
[0002] In the past decade, significant progress has been made in the field of deep learning, and parallelization technology is undoubtedly a key force driving this progress. The core of parallelization lies in the efficient allocation and execution of computing tasks through dedicated hardware accelerators, such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs). These hardware accelerators can quickly process the matrix multiplication operations that are abundant in deep learning, thus greatly shortening the time for model training and evaluation. This efficiency enables researchers to conduct trial-and-error experiments more quickly, accelerating the iteration and development of deep learning technology. However, the application of parallelization technology in the field of deep learning is not without limitations.
[0003] For sequence models, such as Recurrent Neural Networks (RNNs) and Neural Ordinary Differential Equations (Neural ODEs), the effect of parallelization is not ideal. This is because the inherent characteristics of these models determine that they need to serially evaluate each time step of the sequence. In sequence models, the calculation of the current time step depends on the result of the previous time step, and this dependency makes it difficult to achieve parallelization. Therefore, sequential evaluation has become a bottleneck in training sequence deep learning models, limiting the training speed and efficiency of the models. Due to the challenges of sequence models in parallelization, the research focus has gradually shifted in recent years. Researchers have started to pay more attention to models that can better utilize the advantages of parallelization. For example, in the field of natural language processing, the attention mechanism and its related variants have gradually replaced traditional RNN models and become the mainstream choice for language modeling. An important advantage of the attention mechanism is that it can process all elements in the input sequence in parallel, thus significantly improving the training efficiency of the model. This parallelization feature makes the attention mechanism perform well in dealing with large-scale datasets and complex tasks, further promoting its wide application in the field of natural language processing. At the same time, as a Continuous Normalizing Flow (CNF) model, Neural ODE is also undergoing a similar transformation. Traditional Neural ODE models need to simulate Ordinary Differential Equations (ODEs) to achieve training, which is computationally complex and difficult to parallelize. To overcome this bottleneck, researchers have started to explore new training methods, attempting to develop training strategies that do not rely on ODE simulation. These new methods aim to improve the training efficiency of Neural ODE models and enable them to better adapt to the parallelization requirements of modern deep learning.
[0004] Despite the many challenges that sequence models face in parallelization, there are still researchers working on solving this problem. Some recent studies have attempted to improve the parallelization of RNNs, mainly focusing on linear recurrent layers that can be evaluated in parallel with prefix sums. By parallelizing the linear recurrent layers, the training speed of the RNN model can be increased to some extent. However, this improvement does not apply to non-linear recurrent layers. Due to the computational dependence of non-linear recurrent layers on the length of the sequence, their inherent serial nature makes parallelization difficult to achieve. Therefore, despite some progress in parallelizing RNNs, non-linear recurrent layers remain a difficult problem to solve.
[0005] In summary, parallelization technology has provided strong impetus for the development of deep learning in the past decade, but there are still many challenges in the application of sequence models. These challenges have prompted researchers to continuously explore new methods and technologies to overcome the parallelization bottleneck of sequence models, and have also promoted the development of the deep learning field towards other models and directions with greater parallelization potential. Summary of the Invention
[0006] The purpose of the invention is to provide a parallel training method for serialized models based on fixed-point iteration in view of the deficiencies of the prior art. By using the fixed-point iteration method and the parallel prefix sum algorithm, it realizes the efficient parallel training of sequence models, optimizes the training performance of serialized models, effectively reduces the training time of sequence models, and greatly improves the convergence speed of the models. The fixed-point iteration approximates the steady-state solution of the sequence model through iteration, and the prefix sum algorithm decomposes the sequence calculation task into multiple independent subtasks, avoiding direct dependencies between time steps, thus enabling parallel computing. This method makes full use of the parallel capabilities of hardware such as GPUs / TPUs and can adaptively adjust the algorithm parallelism according to the hardware state. It is difficult for ordinary sequence model training methods to make full use of the parallel capabilities of hardware. However, due to the difficulty of parallelizing non-linearity, the linear recurrence formula obtained by using the fixed-point recurrence formula can use the parallel Prefix sums algorithm to accelerate the forward propagation of the model. The serialized model of the present invention does not require the model to be linear, which enables the system to adapt to more models. It not only optimizes the training performance of RNNs and Neural ODEs, but also provides new ideas and methods for the efficient training and application of future sequence models, promoting the further development of sequence models in the field of deep learning and effectively solving the problem of difficult parallelization of sequence models such as recurrent neural networks (RNNs) and neural ordinary differential equations (Neural ODEs).
[0007] The object of the present invention is achieved as follows: A parallel training method for a serialized model based on fixed-point iteration, characterized by using the fixed-point iteration method and the parallel prefix sum algorithm to achieve parallel training of linear or non-linear serialized models. The fixed-point iteration uses a fixed-point recurrence formula for model iteration based on the fixed point, propagates the sequence model forward, and obtains a linear recurrence model according to the general form of the Neural ODE or RNN sequence model; the parallel prefix sum algorithm uses an adaptive parallel Prefix Sums algorithm to automatically select the optimal hyperparameter configuration, calculates the gradients of each part according to the result of one-step forward propagation of the model, and updates the model parameters through backpropagation.
[0008] The parallel training method for the serialized model based on fixed-point iteration specifically includes the following steps:
[0009] Step 1: Analyze and solve the fixed-point recurrence formula
[0010] 1-1: For Neural ODE, its general form is: , and the initial value is known . The recurrence formula for the general form of the model is expressed by the following formula (a):
[0011] (a).
[0012] Among them, the operators L and G are both linear operators, f represents the non-linear function determined by the model, x represents the input variable, represents the model parameters included in f.
[0013] According to the comparison of the above formula (a), it can be obtained that: , where is time. Comparing the calculation logic in the recurrence formula for the general form of the model, add to both ends of the equation, which is determined by the last line of formula (a). And denote the right side of the equation as , then the equation of Neural ODE shown in the following formula (b) can be obtained:
[0014] (b).
[0015] Then, solve the equation by the method of variation of constants and discretize it to obtain the recurrence formula for Neural ODE shown in the following formula (c):
[0016] (c).
[0017] For RNN, its general form is: , such a general form not only includes the basic RNN, but also includes LSTM and GRU. Then, according to the comparison with the general form expressed by equation (a), it can be obtained that: . Its general form can be transformed into the RNN equation expressed by equation (d), where represents the output of the i-th step of the RNN. Solving this equation is equivalent to solving the recurrence formula for Neural ODE.
[0018] (d).
[0019] Step 2: Forward propagation of the parallel algorithm
[0020] Taking Neural ODE as an example, denote , with the initial value being , and define the operation symbols as follows:
[0021] .
[0022] Then, only need to multiply from to (denote the total length of the sequence as n) to obtain from the second part of the result, and this multiplication process can use the parallel Prefix Sums algorithm. There are many kinds of parallel algorithms, and the common ones are the balanced binary tree algorithm, the Odd / Even algorithm based on the divide-and-conquer strategy, and the Upper / Lower algorithm based on the divide-and-conquer strategy.
[0023] Step 3: Calculate the gradient backpropagation
[0024] According to all the evaluation values obtained in Step 2 and the true value calculate the loss value , and then find the gradient of L with respect to the model parameters . After that, update the model parameters according to the gradient descent method:
[0025] .
[0026] The fixed-point iteration is used to initialize the hidden state of the sequence model, and approximates the steady-state solution of the sequence model through iteration, providing a theoretical basis for parallel processing; the parallel prefix sum algorithm decomposes the sequence calculation task into multiple independent subtasks, avoiding the direct dependence between time steps, thereby realizing parallel computing, and the prefix sum algorithm decomposes the calculation of each time step into independent incremental updates, making full use of the parallel computing power of hardware accelerators (such as GPUs or TPUs).
[0027] The present invention has the following beneficial technical effects and significant technological progress compared with the prior art:
[0028] 1) Parallelize the training process of the sequence model, make full use of the hardware parallel capabilities, improve the training speed of the sequence model, and do not require the model to be linear, enabling the system to adapt to more models;
[0029] 2) Effectively solve the problem that the training efficiency of sequence models such as RNN and Neural ODE is limited by the sequential dependence between time steps. For example, the output of each time step of an RNN depends on the hidden state of the previous time step, and Neural ODE needs to solve ordinary differential equations through numerical methods, which makes parallelization difficult to achieve;
[0030] 3) The model continues to advance subsequent parallelization according to the fixed-point recurrence function, and the linear recurrence formula obtained by the fixed-point recurrence formula uses the parallel Prefix sums algorithm to accelerate the forward propagation of the model; 4) Can select appropriate hyperparameters according to the hardware state and adaptively adjust the algorithm parallelism;
[0031] 5) The fixed-point iteration approximates the steady-state solution of the sequence model through iteration, providing a theoretical basis for parallel processing. The prefix sum algorithm decomposes the sequence calculation task into multiple independent subtasks, avoiding direct dependence between time steps, thereby achieving parallel computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Schematic diagram of the parallel sequence model training system constructed for the present invention;
[0033] Figure 2 Schematic diagram of the intermediate state of the parallel Prefix Sums algorithm after adaptive optimization. DETAILED DESCRIPTION OF THE INVENTION
[0034] The present invention adopts the fixed-point iteration method to give a parallel processing based on the Prefix Sums algorithm for the problem that sequence models such as RNN and Neural ODE are difficult to parallelize, and optimize their training performance. The technical solutions of the present invention will be described in detail below in conjunction with the embodiments and the drawings.
[0035] Embodiment 1
[0036] Refer to Figure 1 , for the parallel sequence model training system constructed according to the present invention, which does not require the model to be linear, its steps can be divided into three steps: 1) Analyze and solve the fixed-point recurrence formula; 2) Use the parallel algorithm for forward propagation; 3) Calculate the gradient backpropagation.
[0037] Step 1: Solve the recurrence formula
[0038] 1-1: For Neural ODE, its general form is: , and the initial value is known . The recurrence formula for the general form of the model is expressed by the following formula (a):
[0039] (a).
[0040] Among them, the operators L and G are both linear operators.
[0041] According to the comparison of the above formula (a), it can be obtained that: . Comparing the calculation logic in the recurrence formula of the general form of the model, add to both ends of the equation, and denote the right side of the equation as , then the equation of Neural ODE shown in the following formula (b) can be obtained:
[0042] (b).
[0043] Then, by using the method of variation of constants to solve the equation and discretize it, the recurrence formula for Neural ODE shown in the following formula (c) can be obtained:
[0044] (c).
[0045] For RNN, its general form is: , such a general form not only includes the basic RNN, but also includes LSTM and GRU. Then, according to the comparison of the general form expressed by formula (a), it can be obtained that: . Its general form can be transformed into the RNN equation expressed by formula (d), and solving this equation is equivalent to solving the recurrence formula for Neural ODE.
[0046] (d).
[0047] Step 2: Select the parallel algorithm for forward propagation
[0048] The recurrence formula obtained in Step 1 is still non-linear, but it can be solved by the following method:
[0049] Taking Neural ODE as an example, denote , the initial value is , and define the operation symbols as follows:
[0050] .
[0051] Then, only need to multiply from to (denote the total length of the sequence as n), then the second part of the result can be obtained This cumulative multiplication process can be implemented using the parallel Prefix Sums algorithm. There are many types of parallel algorithms, and common ones include the balanced binary tree algorithm, the Odd / Even algorithm based on the divide-and-conquer strategy, and the Upper / Lower algorithm based on the divide-and-conquer strategy.
[0052] The present invention optimizes and expands the Upper / Lower algorithm based on the divide-and-conquer strategy. When facing a model with an extremely long sequence (the sequence length far exceeds the parallelism of the hardware), the algorithm divides the sequence into k parts (as a hyperparameter for algorithm adaptation) instead of two parts, which reduces the depth of the parallel algorithm from to , where k is a hyperparameter, which is determined by the maximum parallelism of the hardware and can be determined by the following formula:
[0053] .
[0054] Among them, T is the maximum parallelism of the hardware, which is determined by the number of cores of the hardware.
[0055] Refer to Figure 2 , for example, the logical core number of the CPU or the CUDA core number of the GPU. Generally, the smallest k value that satisfies the inequality can be taken. Taking k = 3 as an example, the intermediate transfer part is as shown in Figure 2 .
[0056] In addition to optimizing and expanding the Upper / Lower algorithm, the present invention also implements the matrix calculation part using a loop instead of a recursive algorithm. The loop implementation method can greatly reduce the frequent matrix object construction process during function parameter passing in the recursive algorithm. This optimized algorithm first uses a recursive form to calculate the subscripts required for each operation step, but does not actually perform matrix operations. After calculating all the subscripts, it calculates from the deepest layer of the source recursion to the first layer of the source recursion in a loop form according to the subscript information to obtain the result.
[0057] Step 3: Calculate the gradient backpropagation
[0058] This step is consistent with the backpropagation of the normal model training method. First, according to all the evaluation values obtained in Step 2 and the true value , calculate the loss value , then find the gradient of L with respect to the model parameters , and then update the model parameters according to the gradient descent method:
[0059] .
[0060] Other optimizers such as the Adam optimizer can also be used.
[0061] In view of the problem that sequence models such as recurrent neural networks (RNNs) and neural ordinary differential equations (Neural ODEs) are difficult to parallelize, the present invention proposes an optimized parallel processing solution based on the fixed-point iteration method and the prefix sums algorithm, significantly optimizing the training performance of these models. Fixed-point iteration provides a theoretical basis for parallel processing by iteratively approaching the steady-state solution of the sequence model. The prefix sums algorithm, on the other hand, decomposes the sequence calculation task into multiple independent subtasks, avoiding direct dependencies between time steps and thus enabling parallel computing. Specifically, fixed-point iteration is used to initialize the hidden state of the sequence model, while the prefix sums algorithm decomposes the calculation at each time step into independent incremental updates, making full use of the parallel computing power of hardware accelerators such as GPUs or TPUs. The present invention effectively reduces the training time of the sequence model and improves the convergence speed of the model. This method not only optimizes the training performance of RNNs and Neural ODEs, but also provides new ideas and methods for the efficient training and application of future sequence models, promoting the further development of sequence models in the field of deep learning.
[0062] The above is only a further illustration of the present invention and is not intended to limit the present invention. Equivalent implementations without departing from the spirit and scope of the inventive concept shall be included within the scope of the claims of the present invention.
Claims
1. A method for parallel training of serialized models based on fixed point iteration, characterized in that: Fixed point iteration and Prefix Sums algorithm are used to realize parallel training of linear or nonlinear sequence models. The fixed point iteration adopts the fixed point recursive formula of model iteration based on the fixed point, forward propagates the sequence model, and obtains the linear recursive model according to the general form of the sequence model; the parallel prefix sum algorithm adopts the adaptive parallel Prefix Sums algorithm to automatically select the optimal hyperparameter configuration, obtain the result of one-step forward propagation of the model, then calculate the gradient of each part, and update the model parameters through back propagation.
2. The method for parallel training of serialized models based on fixed point iteration according to claim 1, characterized in that: The serialization model is a linear or nonlinear model.
3. The method for parallel training of serialized models based on fixed point iteration according to claim 1, characterized in that: The fixed point recursion formula continues to advance the subsequent parallelization of the sequence model according to the fixed point recursion function, and uses the parallel Prefix sums algorithm to accelerate the forward propagation of the model.
4. The method for parallel training of serialized models based on fixed point iteration according to claim 1 or claim 3, characterized in that: The parallel Prefix Sums algorithm selects appropriate hyperparameters according to hardware adaptiveness, and uses the parallel capabilities of GPU / TPU hardware to adaptively adjust the algorithm parallelism.