A distributed processing system, a task scheduling method, and a parameter determination method

Through alternate parallel processing of neural network accelerator and processing units in distributed processing systems, the problem of Transformer model in GPU memory limitation is solved, and the computing efficiency is improved.

CN118394495BActive Publication Date: 2025-07-22BEIJING QINGCHENG JIZHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410254588.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-07-22
Estimated Expiration
2044-03-06

AI Technical Summary

Technical Problem

In the prior art, the Transformer model is limited in memory capacity when running on GPU, resulting in limited computing efficiency, making it difficult to efficiently handle large-scale sequence generation tasks.

Method used

A distributed processing system is adopted to optimize the calculation process to improve efficiency through alternating parallel processing of neural network accelerator and processing unit, combining task scheduling and parameter determination methods.

Benefits of technology

By storing feature vectors to the processing unit, the memory usage of the neural network accelerator is reduced, the number of sequence generation tasks that can be processed in a single batch is improved, and the computing efficiency of the distributed processing system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118394495B_ABST
    Figure CN118394495B_ABST
Patent Text Reader

Abstract

The present invention provides a distributed processing system, a task scheduling method, and a parameter determination method. The distributed processing system includes a first module and a second module, and the first module is communicatively connected to each second module; the first module includes a neural network accelerator, and the second module includes at least one processing unit; the neural network accelerator is configured to calculate and obtain a first feature vector based on an input feature vector of a sequence generation task and send the first feature vector to the at least one processing unit; the neural network accelerator is further configured to calculate and obtain a corresponding output result based on a second feature vector; the at least one processing unit is configured to store the first feature vector, calculate and obtain a second feature vector based on the first feature vector, and send the second feature vector to the neural network accelerator. The distributed processing system, the task scheduling method, and the parameter determination method provided by the embodiments of the present invention improve the computing efficiency of the distributed processing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a distributed processing system, a task scheduling method, a parameter determination method and an apparatus. Background Art

[0002] With the development of artificial intelligence technology, machine learning models have been more and more widely used, and machine learning algorithms can be applied to natural language processing, image recognition, equipment fault prediction, etc.

[0003] Transformer is a neural network model structure that can be applied to application scenarios such as answering questions and text generation. In the calculation process of the Transformer model, the main calculation mode is matrix-vector multiplication, and the above calculation mode has very low efficiency when running on a Graphics Processing Unit (GPU for short). In order to improve the calculation efficiency, multiple requests are usually packed into a matrix and matrix-matrix multiplication operation is performed with the parameter matrix. As the packing size increases, the running efficiency of the GPU gradually improves, but due to the limited memory capacity of the GPU and the limited data that can be stored, the calculation efficiency of the GPU is limited during model calculation. Summary of the Invention

[0004] In view of the problems in the prior art, embodiments of the present invention provide a distributed processing system, a task scheduling method, a parameter determination method and an apparatus, which can at least partially solve the problems existing in the prior art.

[0005] In a first aspect, the present invention proposes a distributed processing system, including a first module and a second module, wherein:

[0006] The first module is communicatively connected to each second module; the first module includes a neural network accelerator, and the second module includes at least one processing unit;

[0007] The neural network accelerator is configured to calculate a first feature vector according to the input feature vector of the sequence generation task and send the first feature vector to the at least one processing unit; the neural network accelerator is further configured to calculate a corresponding output result according to the second feature vector;

[0008] The at least one processing unit is configured to store the first feature vector, calculate a second feature vector according to the first feature vector and send the second feature vector to the neural network accelerator.

[0009] Further, the number of the neural network accelerators is multiple.

[0010] Further, the neural network accelerator adopts a graphics processing unit.

[0011] Further, the processing unit employs a central processing unit, a field programmable gate array, or a graphics processing unit.

[0012] In a second aspect, the present invention provides a task scheduling method for a distributed processing system according to any one of the above embodiments, including:

[0013] Receiving at least one sequence generation task;

[0014] Dividing the sequence generation task into a first part task and a second part task according to the number of sequences of the sequence generation task;

[0015] Performing task scheduling on the neural network accelerator and the at least one processing unit, such that the neural network accelerator and the at least one processing unit alternately and parallelly process the first part task and the second part task.

[0016] In a third aspect, the present invention provides a task scheduling method for a distributed processing system according to any one of the above embodiments, including:

[0017] Receiving a plurality of sequence generation tasks; evenly dividing each sequence generation task into a first sequence set and a second sequence set according to the number of sequences of the sequence generation task;

[0018] Obtaining a splitting number according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and a scheduling parameter; wherein, the scheduling parameter is preset; the generation length of the sequence generation task is preset; scheduling the current processing sequences of the neural network accelerator and the at least one processing unit according to the scheduling parameter and the splitting number, so as to shorten the waiting time of the neural network accelerator and the at least one processing unit during alternately and parallelly processing the first sequence set and the second sequence set of the plurality of sequence generation tasks.

[0019] Further, the obtaining a splitting number according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and a scheduling parameter includes:

[0020] Calculating to obtain a splitting number C according to the formula C = [BF / (2S)], where [] represents rounding, B represents the number of sequences of the sequence generation task, F represents the scheduling parameter, and S represents the generation length.

[0021] In a fourth aspect, the present invention provides a parameter determination method for a distributed processing system according to any one of the above embodiments, including:

[0022] Obtain the number of sequences of the sequence generation task according to the constraints of the single - calculation duration function, generation length, total delay time, and number of sequences corresponding to the neural network accelerator; wherein, the single - calculation duration function corresponding to the neural network accelerator is obtained by testing the calculation duration of the first module for processing a single sequence generation task alone; the single - calculation duration function corresponding to the neural network accelerator is a relationship function between the calculation duration of the sequence generation task and the number of sequences.

[0023] Obtain the number of processing units according to the number of sequences, the generation length, the unit calculation duration, and the single - calculation duration function corresponding to the neural network accelerator; wherein, the unit calculation duration is obtained by testing the calculation duration of a single processing unit for processing a single element with a sequence length of 1.

[0024] Further, the constraint condition for the number of sequences is: T(B) < L / (2NS), where B represents the number of sequences, T(B) represents the single - calculation duration function corresponding to the neural network accelerator, N represents the number of layers of the model, and S represents the generation length.

[0025] Further, the obtaining the number of processing units according to the number of sequences, the generation length, the unit calculation duration, and the single - calculation duration function corresponding to the neural network accelerator includes:

[0026] Calculate the number of processing units m according to the formula m=(BSR) / (2T(B)), where B represents the number of sequences, S represents the generation length, R represents the unit calculation duration, and T(B) represents the single - calculation duration function corresponding to the neural network accelerator.

[0027] In a fifth aspect, the present invention proposes a task scheduling device, including:

[0028] A first receiving unit, configured to receive at least one sequence generation task;

[0029] A partitioning unit, configured to partition the sequence generation task into a first - part task and a second - part task according to the number of sequences of the sequence generation task;

[0030] A first scheduling unit, configured to perform task scheduling on the neural network accelerator and the at least one processing unit, so that the neural network accelerator and the at least one processing unit alternately and parallelly process the first - part task and the second - part task.

[0031] In a sixth aspect, the present invention proposes a task scheduling device, including:

[0032] A second receiving unit, configured to receive a batch of sequence generation tasks;

[0033] A first obtaining unit, configured to obtain a splitting number according to a generation length of the batch sequence generation task, a sequence number of each sequence generation task, and a scheduling parameter, where the scheduling parameter is preset, and the generation length of the batch sequence generation task is preset;

[0034] A second scheduling unit, configured to perform task scheduling on the neural network accelerator and the at least one processing unit according to the scheduling parameter and the splitting number, so as to shorten the waiting time of the neural network accelerator and the at least one processing unit.

[0035] In a seventh aspect, the present invention provides a parameter determination device, including:

[0036] A second obtaining unit, configured to obtain a sequence number of a sequence generation task according to a single calculation duration function corresponding to the neural network accelerator, a generation length, a total delay time, and a constraint condition of the sequence number, where the single calculation duration function corresponding to the neural network accelerator is obtained by testing the calculation duration of the first module for separately processing a single sequence generation task, and the single calculation duration function corresponding to the neural network accelerator is a relationship function between the calculation duration of the sequence generation task and the sequence number;

[0037] A third obtaining unit, configured to obtain the number of the processing units according to the sequence number, the generation length, a unit calculation duration, and the single calculation duration function corresponding to the neural network accelerator, where the unit calculation duration is obtained by testing the calculation duration of a single processing unit for processing a single element with a sequence length of 1.

[0038] In an eighth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the program to implement the method according to any one of the foregoing embodiments.

[0039] In a ninth aspect, the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the method according to any one of the foregoing embodiments is implemented.

[0040] In a tenth aspect, the present invention provides a computer program product, including a computer program / instruction, and when the computer program / instruction is executed by a processor, the method according to any one of the foregoing embodiments is implemented.

[0041] The distributed processing system, task scheduling method, parameter determination method and device provided by the embodiments of the present invention include a first module and a second module. The first module is communicatively connected to each second module. The first module includes a neural network accelerator, and the second module includes at least one processing unit. The neural network accelerator is configured to calculate a first feature vector based on the input feature vector of the sequence generation task and send the first feature vector to the at least one processing unit. The neural network accelerator is further configured to calculate a corresponding output result based on the second feature vector. The at least one processing unit is configured to store the first feature vector, calculate a second feature vector based on the first feature vector and send the second feature vector to the neural network accelerator. Since the first feature vector is stored in the processing unit, the memory occupancy of the neural network accelerator is saved, the number of sequences of the sequence generation tasks that can be processed in a single batch is increased, and the computing efficiency of the distributed processing system is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0043] Figure 1 FIG. 1 is a schematic diagram of the principle of the distributed processing system provided by the first embodiment of the present invention.

[0044] Figure 2 FIG. 2 is a schematic structural diagram of the distributed processing system provided by the second embodiment of the present invention.

[0045] Figure 3 FIG. 3 is a schematic diagram of the task scheduling timing provided by the third embodiment of the present invention.

[0046] Figure 4 FIG. 4 is a schematic diagram of the task scheduling timing provided by the fourth embodiment of the present invention.

[0047] Figure 5 FIG. 5 is a schematic flowchart of the task scheduling method provided by the fifth embodiment of the present invention.

[0048] Figure 6 FIG. 6 is a schematic flowchart of the task scheduling method provided by the sixth embodiment of the present invention.

[0049] Figure 7 FIG. 7 is a schematic flowchart of the parameter determination method provided by the seventh embodiment of the present invention.

[0050] Figure 8It is a schematic structural diagram of a task scheduling device provided by the eighth embodiment of the present invention.

[0051] Figure 9 It is a schematic structural diagram of a task scheduling device provided by the ninth embodiment of the present invention.

[0052] Figure 10 It is a schematic structural diagram of a parameter determination device provided by the tenth embodiment of the present invention.

[0053] Figure 11 It is a schematic physical structure diagram of an electronic device provided by the eleventh embodiment of the present invention. Detailed implementation manners

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other arbitrarily.

[0055] To facilitate the understanding of the technical solutions provided in this application, the relevant content of the technical solutions in this application will be described below first.

[0056] Token: A single element in a sequence generation task. For example, a word in English text, or a character or word in Chinese text.

[0057] Feature Vector: An array composed of several floating-point numbers, which is used to represent the features of an element in a neural network for processing by the neural network.

[0058] Sequence generation task (Batch): A computing task that processes several different requests after packing them together. Neural network accelerator: A computing acceleration hardware, often used for the calculation of neural network models.

[0059] The Transformer model may include multiple sequentially connected Transformer modules. Each Transformer module includes an attention module (Attention) and a multi-layer perceptron module (MLP). The attention module is connected to the multi-layer perceptron module. Both the attention module and the multi-layer perceptron module are specific neural network modules.

[0060] For the attention module, given the input feature vector of the i-th token as Xi, the calculation process of its output feature vector is as follows:

[0061]

[0062] A i= Normalize{Q i ·K j (j = 1, ..., i - 1)} (2)

[0063]

[0064] Y i = W o O i (4)

[0065] Among them, V i , K i , Q i represent the feature vectors corresponding to the i-th token in a certain layer in the sequence generation task; A ij represents the j-th element in vector A i ; Normalize represents the normalization process; W q , W k , W v and W o are model parameter matrices, i is a positive integer, and j is a positive integer.

[0066] In the multi-layer perceptron module, the feature vectors corresponding to each token are processed separately and there is no dependency relationship between them.

[0067] In the sequence generation task, since the tokens are generated one by one, the V i , K i , Q i corresponding to the previous tokens remain unchanged. Only the feature vectors corresponding to the latest token need to be processed. Therefore, for the V i , K i , Q i in the calculations corresponding to formulas (2) and (3), they need to be stored. For the sequence generation task, using the stored V i , K i , Q i can avoid repeated calculations, greatly reducing the total computational amount, thereby improving the computational efficiency.

[0068] Figure 1 is a schematic diagram of the principle of the distributed processing system provided by the first embodiment of the present invention. As Figure 1 shown, the first node includes a GPU. For the input data, the GPU calculates the first feature vectors V i , K i , Q i through formula (1), and then the first feature vectors V i , K i , Q iThe processing unit sent to each second node.

[0069] The processing unit of the second node stores the first feature vector V i , K i , Q i , and calculates the second feature vector according to formulas (2) and (3), then returns the second feature vector to the GPU. The GPU calculates the calculation result corresponding to the input data through formula 4, and then outputs the calculation result. The GPU can also perform calculations on the multi-layer perceptron module.

[0070] By storing the first feature vector V i , K i , Q i in the processing unit, the memory occupancy of the GPU is saved, so that the GPU can process a larger number of sequences at one time. If the first feature vector V i , K i , Q i is sent to the processing unit for storage, and the second feature vector is not calculated by the processing unit according to formulas (2) and (3), but is calculated by the GPU according to formulas (2) and (3). The processing unit needs to send the first feature vector V i , K i , Q i to the GPU, and it can be seen from formula (2) that as i and j increase, the sent first feature vectors K i , Q i will also increase. Therefore, calculating the second feature vector by the processing unit according to formulas (2) and (3) can reduce the communication volume compared with calculating the second feature vector by the GPU according to formulas (2) and (3), thereby improving the calculation efficiency of the model.

[0071] Figure 2 is a schematic structural diagram of the distributed processing system provided by the second embodiment of the present invention. As Figure 2 shown, the distributed processing system provided by the embodiment of the present invention includes a first module 1 and a second module 2, wherein:

[0072] The first module 1 is communicatively connected to each second module 2; the first module 1 includes a neural network accelerator 11, and the second module 2 includes at least one processing unit 21;

[0073] The neural network accelerator 11 is used to calculate the first feature vector according to the input feature vector of the sequence generation task and send the first feature vector to the at least one processing unit 21; the neural network accelerator 11 is also used to calculate the corresponding output result according to the second feature vector;

[0074] The at least one processing unit 21 is configured to store a first feature vector, calculate a second feature vector based on the first feature vector, and send the second feature vector to the neural network accelerator 11.

[0075] Specifically, after the first module 1 receives a sequence generation task, the neural network accelerator 11 can calculate a first feature vector based on the input feature vector of the sequence generation task, and then send the first feature vector to the at least one processing unit 21. The at least one processing unit 21 stores the received first feature vector, and calculates a second feature vector based on the first feature vector. The at least one processing unit 21 sends the second feature vector to the neural network accelerator 11, and the neural network accelerator 11 calculates an output result corresponding to the input feature vector based on the second feature vector.

[0076] Among them, the first module 1 can be a computer, a server, or a server cluster, and the neural network accelerator 11 can be a GPU. The second module can be a computer, a server, or a server cluster. The processing unit 21 can be a Central Processing Unit (CPU), a Field Programmable Gate Array (FPGA), or a GPU. The number of processing units is selected according to actual needs, and is not limited in the embodiments of the present invention.

[0077] The distributed processing system provided by the embodiments of the present invention includes a first module and a second module. The first module is communicatively connected to each second module. The first module includes a neural network accelerator, and the second module includes at least one processing unit. The neural network accelerator is configured to calculate a first feature vector based on the input feature vector of the sequence generation task and send the first feature vector to the at least one processing unit. The neural network accelerator is further configured to calculate a corresponding output result based on the second feature vector. The at least one processing unit is configured to store the first feature vector, calculate a second feature vector based on the first feature vector, and send the second feature vector to the neural network accelerator. Since the first feature vector is stored in the processing unit, the memory occupancy of the neural network accelerator is saved, the number of sequences of the sequence generation tasks that can be processed in a single batch is increased, and the computing efficiency of the distributed processing system is improved.

[0078] Based on the above embodiments, further, the number of neural network accelerators 11 is multiple. The number of neural network accelerators 11 is set according to actual needs, and is not limited in the embodiments of the present invention.

[0079] Based on the above embodiments, further, the neural network accelerator uses a GPU.

[0080] Based on the above embodiments, further, the processing unit 21 uses a CPU, an FPGA, or a GPU.

[0081] It can be understood that the distributed processing system provided by the embodiments of the present invention can be used for model calculation of a trained model. The distributed processing system provided by the embodiments of the present invention is particularly suitable for Transformer model calculation.

[0082] Figure 3 It is a schematic diagram of task scheduling timing provided by the third embodiment of the present invention. As Figure 3 shown, the neural network accelerator and the processing unit will take turns to perform calculations. Therefore, when the neural network accelerator is working to process tasks, the processing unit is in an idle state, waiting for the calculation results of the neural network accelerator. When the processing unit is working to process tasks, the neural network accelerator has to wait for the calculation results of the processing unit.

[0083] To improve the calculation efficiency, as Figure 4 shown, the sequence generation task is divided into two parts. The neural network accelerator first calculates the first part of the task. After calculating the first feature vector of the first part of the task, it sends the first feature vector of the first part of the task to the processing unit. When the processing unit calculates the second feature vector of the first part of the task based on the first feature vector of the first part of the task, the neural network accelerator can calculate the first feature vector of the second part of the task. After the processing unit calculates and obtains the second feature vector of the first part of the task, it sends it to the neural network accelerator; at the same time, after the neural network accelerator calculates the first feature vector of the second part of the task, it sends the first feature vector of the second part of the task to the processing unit. When the neural network accelerator calculates the corresponding output result based on the second feature vector of the first part of the task, the processing unit calculates the second feature vector of the second part of the task based on the first feature vector of the second part of the task. Through the above process, it can be seen that when the neural network accelerator and the processing unit can alternately process the first part of the task and the second part of the task, they can perform partial calculations in parallel, reducing the idle waiting time, thereby improving the calculation efficiency of the sequence generation task.

[0084] Figure 5 It is a schematic flowchart of the task scheduling method provided by the fifth embodiment of the present invention. As Figure 5 shown, the task scheduling method provided by the embodiments of the present invention is applied to the distributed processing system described in any of the above embodiments, and includes:

[0085] S501. Receive at least one sequence generation task;

[0086] Specifically, the first module can receive a sequence generation task and obtain the number of sequences of the sequence generation task. The number of sequences of the sequence generation task is preset and can be set according to actual needs. The first module can include a CPU, and the task scheduling method is executed through the CPU of the first module.

[0087] S502. Divide the sequence generation task into a first part task and a second part task according to the number of sequences of the sequence generation task;

[0088] Specifically, the first module divides the sequence generation task into two parts according to the number of sequences of the sequence generation task: a first part task and a second part task. The first part task can include a first number of elements, and the second part task can include a second number of elements. The sum of the first number and the second number is equal to the number of sequences.

[0089] For example, the sequence generation task is evenly divided into a first part task and a second part task according to the number of sequences, and the first part task and the second part task respectively include elements with half of the number of sequences.

[0090] S503. Perform task scheduling on the neural network accelerator and the at least one processing unit so that the neural network accelerator and the at least one processing unit alternately and parallelly process the first part task and the second part task.

[0091] Specifically, the first module performs task scheduling on the neural network accelerator and the at least one processing unit so that the neural network accelerator and the at least one processing unit can alternately and parallelly process the first part task and the second part task, that is, the neural network accelerator first calculates the first part task. After calculating the first feature vector of the first part task, the neural network accelerator sends the first feature vector of the first part task to the processing unit. When the processing unit calculates the second feature vector of the first part task according to the first feature vector of the first part task, the neural network accelerator can calculate the first feature vector of the second part task. After the processing unit calculates and obtains the second feature vector of the first part task, it sends it to the neural network accelerator; at the same time, after the neural network accelerator calculates the first feature vector of the second part task, it sends the first feature vector of the second part task to the processing unit. When the neural network accelerator calculates the corresponding output result according to the second feature vector of the first part task, the processing unit calculates the second feature vector of the second part task according to the first feature vector of the second part task. The processing unit sends the second feature vector of the second part task to the neural network accelerator, and the neural network accelerator calculates and obtains the corresponding output result according to the second feature vector of the second part task, thereby completing the processing of the sequence generation task.

[0092] Since the neural network accelerator and the processing unit can process some of the computational processes of the first part of the tasks and the second part of the tasks in parallel, the processing efficiency of the sequence generation task is improved.

[0093] Suppose that a number of sequence generation tasks are currently being processed, and each sequence generation task among the number of sequence generation tasks is divided into a first part of the task and a second part of the task. It is set that generating a token for all the sequences currently being processed is one step of work. In this step of work, the working time of the neural network accelerator is proportional to the number of sequences currently being processed, and the working time of the processing unit is proportional to the total length of the generated tokens of the sequences currently being processed. The working process of the distributed processing system is to repeat the above "one step of work" until each processed sequence among the number of sequence tasks has generated S tokens, that is, the number of sequence tasks is completed. Among them, the sequence generation task includes a number of sequences, and the generated length of each sequence is S.

[0094] For a sequence generation task, it is evenly divided into a first part of the task and a second part of the task according to the number of sequences. After S steps, each sequence has obtained S tokens and ends the processing. Then the processing of the next sequence generation task is carried out. Similarly, after S steps, the first part of the task ends. And so on. In this case, the working time of the processing unit changes greatly, resulting in a longer overall working time of the distributed processing system.

[0095] The following takes the processing process of a sequence generation task with a sequence number of 12 and a generated length of 6 as an example to illustrate the scheduling scheme of the neural network accelerator and the processing unit. Suppose the working time of the neural network accelerator is 3 times the number of sequences currently being processed. The working time of the processing unit is the total number of generated tokens of the sequences currently being processed. The total working time of one step of work takes the larger value between the working time of the neural network accelerator and the working time of the processing unit.

[0096] The sequence generation task includes 12 sequences, namely A1, A2, A3, A4, A5, A6, A7, A8, A9, A10, A11, and A12. The generation length of each sequence is 6. The sequence task is divided into the first part task and the second part task. The first part task includes six sequences A1, A2, A3, A4, A5, and A6, and the second part task includes six sequences A7, A8, A9, A10, A11, and A12. For the first part task, as shown in Table 1, the currently processed sequences are A1, A2, A3, A4, A5, and A6. At step 1, the first token (represented by *) of each sequence is generated. The working time of the neural network accelerator is 18, the working time of the processor is 6, and the total working time of a single step is 18. At the second step, the second token of each sequence is generated. The working time of the neural network accelerator is 18, the working time of the processor is 12, and the total working time of a single step is 18. At the third step, the third token of each sequence is generated. The working time of the neural network accelerator is 18, the working time of the processor is 18, and the total working time of a single step is 18. At the fourth step, the fourth token of each sequence is generated. The working time of the neural network accelerator is 18, the working time of the processor is 24, and the total working time of a single step is 24. At the fifth step, the fifth token of each sequence is generated. The working time of the neural network accelerator is 18, the working time of the processor is 30, and the total working time of a single step is 30. At the sixth step, the sixth token of each sequence is generated. The working time of the neural network accelerator is 18, the working time of the processor is 36, and the total working time of a single step is 36. Since the number of sequences in the currently processed sequences is 6, the working time of the neural network accelerator is fixed at 18, and the working time of the processor increases as the number of generated tokens in the currently processed sequences increases. For the first part of the scheduling task, the currently processed sequences are A7, A8, A9, A10, A11, and A12. For the second part task, as shown in Table 2, the currently processed sequences are A7, A8, A9, A10, A11, and A12. The working time of the neural network accelerator, the working time of the processor, and the total working time of a single step for each step sequence are shown in Table 2 in detail and will not be elaborated here. Since the first part task and the second part are processed alternately, the completion times of the first part task and the second part task are the same.

[0097] Table 1 Scheduling Scheme for the First Part Task

[0098]

[0099] Table 2 Scheduling Scheme for the Second Part Task

[0100]

[0101] Since the working duration of the neural network accelerator is proportional to the number of tasks in the sequence generation task, and the working duration of the processing unit is proportional to the total length of the currently generated sequences in the sequence generation task, where the currently generated sequences refer to the sequences that are currently being processed and have not been completely processed. During the processing of the sequence generation task, as the sequences continuously grow, the calculation delay of the processing unit increases. Due to the different time-consuming of the neural network accelerator and the processing unit when performing parallel calculations, one party waits for the other party, reducing the processing efficiency of the sequence generation task. Therefore, the embodiments of the present invention introduce scheduling parameters to schedule the current number of processed sequences of the neural network accelerator and the processing unit, so as to reduce the mutual waiting time between the neural network accelerator and the processing unit.

[0102] Figure 6 is a schematic flowchart of the task scheduling method provided by the sixth embodiment of the present invention, as Figure 6 shown, the task scheduling method of the distributed processing system provided by the embodiments of the present invention is applied to the distributed processing system described in any of the above embodiments, and includes:

[0103] S601. Receive multiple sequence generation tasks;

[0104] Specifically, the first module receives multiple sequence generation tasks. The first module may include a CPU, and the task scheduling method is executed through the CPU of the first module.

[0105] S602. Evenly divide each sequence generation task into a first sequence set and a second sequence set according to the number of sequences of the sequence generation task;

[0106] Specifically, the first module evenly divides each sequence generation task into a first sequence set and a second sequence set according to the number of sequences of the sequence generation task. The number of sequences included in the first sequence set is equal to that included in the second sequence set, and is equal to half of the number of sequences.

[0107] S603. Obtain the splitting number according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and the scheduling parameter; wherein, the scheduling parameter is preset; the generation length of the sequence generation task is preset;

[0108] Specifically, the first module may calculate and obtain the splitting number according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and the scheduling parameter. Among them, the scheduling parameter is preset and is set according to actual experience, which is not limited in the embodiments of the present invention. The generation length of the sequence generation task is preset and is set according to actual needs, which is not limited in the embodiments of the present invention.

[0109] For example, if the generation length of the sequence generation task is S, the number of sequences of the sequence generation task is B, and the scheduling parameter is F, then the splitting number C can be calculated according to the splitting number calculation formula C = [BF / (2S)], where [] represents rounding down.

[0110] S604. Schedule the current processing sequences of the neural network accelerator and at least one processing unit according to the scheduling parameter and the splitting number, so as to shorten the waiting time of the neural network accelerator and the at least one processing unit when alternately and parallelly processing the first sequence set and the second sequence set of the multiple sequence generation tasks.

[0111] Specifically, the first module determines the number of current processing sequences according to the scheduling parameter, and then every the scheduling parameter steps, the number of current processing sequences increases by the splitting number, so as to schedule the current processing sequences of the neural network accelerator and the at least one processing unit, which can shorten the mutual waiting time of the neural network accelerator and the at least one processing unit when alternately and parallelly processing the first sequence set and the second sequence set of the multiple sequence generation tasks, thereby improving the processing efficiency of the multiple sequence generation tasks.

[0112] Taking the processing process of multiple sequence generation tasks with the number of sequences being 12 and the generation length being 6 as an example, the scheduling scheme of the neural network accelerator and the processing unit is described below. Assume that the working time of the neural network accelerator is 3 times the number of sequences of the current processing sequence. The working time of the processing unit is the total number of generated tokens of the current processing sequence.

[0113] For the convenience of description, the sequences in the first sequence set of the multiple sequence generation tasks are represented by A, and the sequences in the second sequence set are represented by B. The first sequence set of the first sequence generation task is represented as A1, the first sequence set of the second sequence generation task is represented as A2, the first sequence set of the third sequence generation task is represented as A3, and so on. Similarly, the second sequence set of the first sequence generation task is represented as B1, the second sequence set of the second sequence generation task is represented as B2, the second sequence set of the third sequence generation task is represented as B3, and so on. The sequences in the first sequence set of the first sequence generation task are respectively represented as A1-1, A1-2, A1-3, A1-4, A1-5, A1-6. The sequences in the second sequence set of the first sequence generation task are respectively represented as B1-1, B1-2, B1-3, B1-4, B1-5, B1-6. The representations of the second sequence generation task and subsequent sequence generation tasks are similar to those of the first sequence generation task, and are not elaborated here.

[0114] As shown in Table 3, when the scheduling parameter is set to 2, then C = [BF / 2S] = [12x2 / (2x6)] = 2. Starting from the first sequence set of the first sequence generation task, the current processing sequences are A1-1 and A1-2. At step 1, the first token (denoted by *) of each sequence is generated. The working time of the neural network accelerator is 6, the working time of the processor is 2, and the total working time per step is 6. At step 2, the second tokens of sequences A1-1 and A1-2 are generated. The working time of the neural network accelerator is 6, the working time of the processor is 2, and the total working time per step is 6. At step 3, 2 more sequences need to be added to the current processing sequences. The third tokens of sequences A1-1 and A1-2, and the first tokens of A1-3 and A1-4 are generated. The working time of the neural network accelerator is 12, the working time of the processor is 8, and the total working time per step is 12. At step 4, the fourth tokens of sequences A1-1 and A1-2, and the second tokens of A1-3 and A1-4 are generated. The working time of the neural network accelerator is 12, the working time of the processor is 12, and the total working time per step is 12. At step 5, 2 more sequences need to be added to the current processing sequences. The fifth tokens of sequences A1-1 and A1-2, the third tokens of A1-3 and A1-4, and the first tokens of A1-5 and A1-6 are generated. The working time of the neural network accelerator is 18, the working time of the processor is 18, and the total working time per step is 18. At step 6, the sixth tokens of sequences A1-1 and A1-2, the fourth tokens of A1-3 and A1-4, and the second tokens of A1-5 and A1-6 are generated. The working time of the neural network accelerator is 18, the working time of the processor is 24, and the total working time per step is 24. At step 7, sequences A1-1 and A1-2 of the first sequence generation task are completed. 2 more sequences need to be added to the current processing sequences. The fifth tokens of sequences A1-3 and A1-4, the third tokens of A1-5 and A1-6, and the first tokens of A2-1 and A2-2 of the second sequence generation task are generated. The working time of the neural network accelerator is 18, the working time of the processor is 18, and the total working time per step is 18. At step 8, the sixth tokens of sequences A1-3 and A1-4, the fourth tokens of A1-5 and A1-6, and the second tokens of A2-1 and A2-2 of the second sequence generation task are generated. The working time of the neural network accelerator is 18, the working time of the processor is 24, and the total working time per step is 24.At the ninth step, sequences A1-3 and A1-4 of the first sequence generation task are completed. The current processing sequences need to add 2 sequences, generate the fifth token of sequences A1-5 and A1-6, generate the third token of sequences A2-1 and A2-2 of the second sequence generation task, and generate the first token of sequences A2-3 and A2-4. The working time of the neural network accelerator is 18, the working time of the processor is 18, and the total working time of a single step is 18. At the tenth step, generate the sixth token of sequences A1-5 and A1-6, generate the fourth token of sequences A2-1 and A2-2 of the second sequence generation task, and generate the second token of sequences A2-3 and A2-4. The working time of the neural network accelerator is 18, the working time of the processor is 24, and the total working time of a single step is 24. At the eleventh step, sequences A1-5 and A1-6 of the first sequence generation task are completed. The current processing sequences need to add 2 sequences, generate the fifth token of sequences A2-1 and A2-2 of the second sequence generation task, generate the third token of sequences A2-3 and A2-4, and generate the first token of sequences A2-5 and A2-6. The working time of the neural network accelerator is 18, the working time of the processor is 18, and the total working time of a single step is 18. At the twelfth step, generate the sixth token of sequences A2-1 and A2-2 of the second sequence generation task, generate the fourth token of sequences A2-3 and A2-4, and generate the second token of sequences A2-5 and A2-6. The working time of the neural network accelerator is 18, the working time of the processor is 24, and the total working time of a single step is 24. At the thirteenth step, sequences A2-1 and A2-2 of the second sequence generation task are completed. The current processing sequences need to add 2 sequences, generate the fifth token of sequences A2-3 and A2-4 of the second sequence generation task, generate the third token of sequences A2-5 and A2-6, and generate the first token of sequences A3-1 and A3-2 of the third sequence generation task. The working time of the neural network accelerator is 18, the working time of the processor is 18, and the total working time of a single step is 18. At the fourteenth step, generate the sixth token of sequences A2-3 and A2-4 of the second sequence generation task, generate the fourth token of sequences A2-5 and A2-6, and generate the second token of sequences A3-1 and A3-2 of the third sequence generation task. The working time of the neural network accelerator is 18, the working time of the processor is 24, and the total working time of a single step is 24. And so on, the subsequent processing processes will not be elaborated.

[0115] Table 3 Scheduling Schemes for Multiple First Sequence Sets

[0116]

[0117] For the sequence scheduling process of the second sequence set for each sequence generation task, as shown in Table 4, the working time of each step neural network accelerator, the working time of the processor, and the total working time of a single step are detailed in Table 4 and will not be elaborated here.

[0118] As can be seen from Table 3, the working time of the neural network accelerator finally stabilizes at 18, and the working time of the processing unit finally varies between 18 and 24. Compared with the working time of the processing unit in Table 1, which varies between 6 - 36, the waiting time between the neural network accelerator and the processing unit can be reduced, thereby improving the processing efficiency of multiple sequence generation tasks, especially when there are many sequence generation tasks.

[0119] Table 4 Scheduling scheme for multiple second sequence sets

[0120]

[0121] Based on the above embodiments, further, obtaining the splitting quantity according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and the scheduling parameter includes:

[0122] Calculating the splitting quantity C according to the formula C = [BF / (2S)], where [] represents taking the integer, B represents the number of sequences of the sequence generation task, F represents the scheduling parameter, and S represents the generation length.

[0123] In a distributed processing system, for the neural network accelerator configured in the first module, if the number of processing units configured in the second module is small, then the computing speed of the processing units in the second module cannot match that of the neural network accelerator, resulting in the neural network accelerator being idle due to waiting for the calculation results of the processing units in the second module, leading to waste of resources of the neural network accelerator. On the contrary, if the number of processing units configured in the second module is excessive, then the neural network accelerator in the first module cannot match the computing speed of the processing units in the second module, resulting in the processing units configured in the second module being idle due to waiting for the calculation results of the neural network accelerator, leading to waste of resources of the processing units. In addition, if there are requirements for the latency of the distributed system processing model, then the number of sequences of the sequence generation task needs to be controlled.

[0124] It can be understood that the execution subject of the parameter determination method provided in the embodiments of the present invention includes but is not limited to a computer.

[0125] Figure 7 It is a schematic flowchart of the parameter determination method provided in the seventh embodiment of the present invention. As Figure 7 shown, the parameter determination method provided in the embodiments of the present invention is applied to the distributed processing system described in any of the above embodiments and includes:

[0126] S701. Obtain the number of sequences of the sequence generation task according to the constraints of the single - calculation duration function, generation length, total delay time, and number of sequences corresponding to the neural network accelerator; wherein, the single - calculation duration function corresponding to the neural network accelerator is obtained by testing the calculation duration of the first module processing a single sequence generation task alone; the single - calculation duration function corresponding to the neural network accelerator is a relational function between the calculation duration of the sequence generation task and the number of sequences.

[0127] Specifically, given a model and the number n of neural network accelerators, taking the number of sequences of the sequence generation task as the dependent variable, test the first module of the distributed processing system alone, test the calculation duration of the n neural network accelerators of the first module processing a single sequence generation task, and obtain the single - calculation duration function T(B) corresponding to the neural network accelerator. When testing, the number n of neural network accelerators can be enumerated. The number of sequences of the sequence generation task can be obtained according to the constraints of the single - calculation duration function, generation length, total delay time, and number of sequences corresponding to the neural network accelerator. Among them, testing the first module alone means completing the calculation only through the first module without communicating with the second module. The constraint condition of the number of sequences is obtained in advance; the generation length and total delay time are set according to actual needs, which are not limited in the embodiments of the present invention. n is a positive integer.

[0128] For example, when the constraint condition of the number of sequences is satisfied, select the largest number of sequences as the number of sequences of the sequence generation task.

[0129] S702. Obtain the number of processing units according to the number of sequences, the generation length, the unit calculation duration, and the single - calculation duration function corresponding to the neural network accelerator; wherein, the unit calculation duration is obtained by testing the calculation duration of a single processing unit processing a single element with a sequence length of 1.

[0130] Specifically, given a model and the number of sequences of the obtained sequence generation task, test the calculation duration of a single processing unit processing a single element with a sequence length of 1, and obtain the unit calculation duration R. The number of processing units m can be obtained according to the number of sequences, the generation length, the unit calculation duration, and the single - calculation duration function corresponding to the neural network accelerator. The number of processing units m matches the number n of neural network accelerators. Corresponding to the distributed processing system of a given model, the first module includes n neural network accelerators, and the second module includes m processing units, which can reduce resource waste. m is a positive integer.

[0131] The parameter determination method provided by the embodiments of the present invention can reasonably determine the number of sequences of the sequence generation task and the number of processing units, reducing resource waste.

[0132] Based on the above embodiments, further, the constraint condition for the number of sequences is: T(B) < L / (2NS), where B represents the number of sequences, T(B) represents the single calculation duration function corresponding to the neural network accelerator, N represents the number of model layers, and S represents the generation length.

[0133] Specifically, the constraint condition for the number of sequences can limit the value range of the number of sequences B. Since the larger B is, the more efficiently the neural network accelerator can be utilized, so the largest B that satisfies the constraint condition for the number of sequences is taken as the number of sequences of the sequence generation task. Among them, for a given model, the number of model layers is known.

[0134] Based on the above embodiments, further, the obtaining of the number of processing units according to the number of sequences, the generation length, the unit calculation duration, and the single calculation duration function corresponding to the neural network accelerator includes:

[0135] According to the formula m = (BSR) / (2T(B)), the number of processing units m is calculated, where B represents the number of sequences, S represents the generation length, R represents the unit calculation duration, and T(B) represents the single calculation duration function corresponding to the neural network accelerator.

[0136] Figure 8 It is a schematic structural diagram of the task scheduling device provided by the eighth embodiment of the present invention. As Figure 8 shown, the task scheduling device provided by the embodiments of the present invention includes a first receiving unit 801, a first partitioning unit 802, and a first scheduling unit 803, where:

[0137] The first receiving unit 801 is configured to receive at least one sequence generation task; the first partitioning unit 802 is configured to partition the sequence generation task into a first part of tasks and a second part of tasks according to the number of sequences of the sequence generation task; the first scheduling unit 803 is configured to perform task scheduling on the neural network accelerator and the at least one processing unit, so that the neural network accelerator and the at least one processing unit alternately and parallelly process the first part of tasks and the second part of tasks.

[0138] Specifically, the first receiving unit 801 can receive the sequence generation task and obtain the number of sequences of the sequence generation task. Among them, the number of sequences of the sequence generation task is preset, and the number of sequences of the sequence generation task is set according to actual needs.

[0139] The first partitioning unit 802 divides the sequence generation task into two parts: a first part task and a second part task according to the number of sequences of the sequence generation task. The first part task may include a first number of elements, and the second part task may include a second number of elements. The sum of the first number and the second number is equal to the number of sequences.

[0140] The first scheduling unit 803 schedules tasks for the neural network accelerator and the at least one processing unit, so that the neural network accelerator and the at least one processing unit can alternately and parallelly process the first part task and the second part task, that is, the neural network accelerator first calculates the first part task, and after calculating the first feature vector of the first part task, sends the first feature vector of the first part task to the processing unit. When the processing unit calculates the second feature vector of the first part task according to the first feature vector of the first part task, the neural network accelerator can calculate the first feature vector of the second part task. After the processing unit calculates the second feature vector of the first part task, it sends it to the neural network accelerator; at the same time, after the neural network accelerator calculates the first feature vector of the second part task, it sends the first feature vector of the second part task to the processing unit. When the neural network accelerator calculates the corresponding output result according to the second feature vector of the first part task, the processing unit calculates the second feature vector of the second part task according to the first feature vector of the second part task. The processing unit sends the second feature vector of the second part task to the neural network accelerator, and the neural network accelerator calculates the corresponding output result according to the second feature vector of the second part task, thereby completing the processing of the sequence generation task.

[0141] Figure 9 It is a schematic structural diagram of the task scheduling device provided in the ninth embodiment of the present invention, as Figure 9 shown, the task scheduling device provided in the embodiment of the present invention includes a second receiving unit 901, a second partitioning unit 902, a first obtaining unit 903, and a second scheduling unit 904, where:

[0142] The second receiving unit 901 is configured to receive a plurality of sequence generation tasks; the second partitioning unit 902 is configured to evenly partition each sequence generation task into a first sequence set and a second sequence set according to the number of sequences of the sequence generation task; the first obtaining unit 903 is configured to obtain a splitting number according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and a scheduling parameter; wherein, the scheduling parameter is preset; the generation length of the sequence generation task is preset; the second scheduling unit 904 is configured to schedule the current processing sequences of the neural network accelerator and the at least one processing unit according to the scheduling parameter and the splitting number, so as to shorten the waiting time of the neural network accelerator and the at least one processing unit when alternately and parallelly processing the first sequence set and the second sequence set of the plurality of sequence generation tasks.

[0143] Specifically, the second receiving unit 901 receives a plurality of sequence generation tasks.

[0144] The second partitioning unit 902 evenly partitions each sequence generation task into a first sequence set and a second sequence set according to the number of sequences of the sequence generation task. The number of sequences included in the first sequence set is equal to that included in the second sequence set, and is equal to half of the number of sequences.

[0145] The first obtaining unit 903 can calculate and obtain the splitting number according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and the scheduling parameter. Wherein, the scheduling parameter is preset and is set according to actual experience, and is not limited in the embodiments of the present invention. The generation length of the sequence generation task is preset and is set according to actual needs, and is not limited in the embodiments of the present invention.

[0146] The second scheduling unit 904 determines the number of current processing sequences according to the scheduling parameter, and then increases the number of current processing sequences by the splitting number every the scheduling parameter steps, so as to schedule the current processing sequences of the neural network accelerator and the at least one processing unit, and can shorten the mutual waiting time between the neural network accelerator and the at least one processing unit when the neural network accelerator and the at least one processing unit alternately and parallelly process the first sequence set and the second sequence set of the plurality of sequence generation tasks, thereby improving the processing efficiency of the plurality of sequence generation tasks.

[0147] Figure 10 It is a schematic structural diagram of a parameter determination device provided by the tenth embodiment of the present invention. As Figure 10 shown, the parameter determination device provided by the embodiments of the present invention includes a second obtaining unit 1001 and a third obtaining unit 1002, wherein:

[0148] The second obtaining unit 1001 is configured to obtain the number of sequences of the sequence generation task according to the constraints of the single-computation duration function, generation length, total delay time, and number of sequences corresponding to the neural network accelerator; wherein, the single-computation duration function corresponding to the neural network accelerator is obtained by testing the computation duration of the first module for separately processing a single sequence generation task; the single-computation duration function corresponding to the neural network accelerator is a relationship function between the computation duration of the sequence generation task and the number of sequences. The third obtaining unit 1002 is configured to obtain the number of processing units according to the number of sequences, the generation length, the unit computation duration, and the single-computation duration function corresponding to the neural network accelerator; wherein, the unit computation duration is obtained by testing the computation duration of a single processing unit for processing a single element with a sequence length of 1.

[0149] Specifically, given a model and the number n of neural network accelerators, the first module of the distributed processing system is separately tested with the number of sequences of the sequence generation task as the dependent variable, and the computation duration of the n neural network accelerators of the first module for processing a single sequence generation task is tested to obtain the single-computation duration function T(B) corresponding to the neural network accelerator. When testing, the number n of neural network accelerators can be enumerated. The second obtaining unit 1001 can obtain the number of sequences of the sequence generation task according to the constraints of the single-computation duration function, generation length, total delay time, and number of sequences corresponding to the neural network accelerator. Among them, separately testing the first module means that the computation is only completed through the first module without communicating with the second module. The constraint condition of the number of sequences is obtained in advance; the generation length and the total delay time are set according to actual needs, and the embodiments of the present invention do not make limitations. n is a positive integer.

[0150] Given a model and the number of sequences of the sequence generation task, the computation duration of a single processing unit for processing a single element with a sequence length of 1 is tested to obtain the unit computation duration R. The third obtaining unit 1002 can obtain the number m of processing units according to the number of sequences, the generation length, the unit computation duration, and the single-computation duration function corresponding to the neural network accelerator. The number m of processing units matches the number n of neural network accelerators. Corresponding to the distributed processing system of the given model, the first module includes n neural network accelerators, and the second module includes m processing units, which can reduce resource waste. m is a positive integer.

[0151] Based on the above embodiments, further, the constraint condition of the number of sequences is: T(B) < L / (2NS), where B represents the number of sequences, T(B) represents the single-computation duration function corresponding to the neural network accelerator, N represents the number of layers of the model, and S represents the generation length.

[0152] Based on the above embodiments, further, obtaining the number of processing units according to the number of sequences, the generated length, the unit calculation duration, and the single calculation duration function corresponding to the neural network accelerator includes:

[0153] Calculating to obtain the number of processing units m according to the formula m = (BSR) / (2T(B)), where B represents the number of sequences, S represents the generated length, R represents the unit calculation duration, and T(B) represents the single calculation duration function corresponding to the neural network accelerator.

[0154] The embodiments of the device provided by the embodiments of the present invention can specifically be used to execute the processing procedures of the corresponding method embodiments above, and their functions will not be elaborated here. Reference can be made to the detailed descriptions of the above method embodiments.

[0155] Figure 11 is a schematic physical structure diagram of an electronic device provided by the eleventh embodiment of the present invention. As Figure 11 shown, the electronic device may include: a processor 1101, a communications interface 1102, a memory 1103, and a communication bus 1104. Among them, the processor 1101, the communications interface 1102, and the memory 1103 communicate with each other through the communication bus 1104. The processor 1101 can call the logical instructions in the memory 1103 and can execute the methods provided by the above method embodiments.

[0156] In addition, when the logical instructions in the above-mentioned memory 1103 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.

[0157] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, enable the computer to execute the methods provided in the above-described method embodiments.

[0158] This embodiment provides a computer-readable storage medium that stores a computer program, which causes the computer to execute the methods provided in the above-described method embodiments.

[0159] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0160] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1Steps of the functions specified in one or more boxes.

[0163] In the description of this specification, the description with reference to terms such as "one embodiment", "one specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0164] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A distributed processing system, characterized in that, It includes a first module and a second module, where: The first module is communicatively connected to each second module; the first module includes a neural network accelerator, and the second module includes at least one processing unit; The neural network accelerator is used to calculate and obtain a first feature vector based on the input feature vector of the sequence generation task and send the first feature vector to the at least one processing unit; the neural network accelerator is also used to calculate and obtain a corresponding output result based on a second feature vector; The at least one processing unit is used to store the first feature vector, calculate and obtain a second feature vector based on the first feature vector, and send the second feature vector to the neural network accelerator; The first module is used to receive a plurality of sequence generation tasks; evenly divide each sequence generation task into a first sequence set and a second sequence set according to the number of sequences of the sequence generation task; obtain a splitting number according to the generation length of the sequence generation task, the number of sequences of the sequence generation task, and a scheduling parameter, where the scheduling parameter is preset; the generation length of the sequence generation task is preset; schedule the current processing sequences of the neural network accelerator and the at least one processing unit according to the scheduling parameter and the splitting number, so as to shorten the waiting time of the neural network accelerator and the at least one processing unit when alternately and parallelly processing the first sequence set and the second sequence set of the plurality of sequence generation tasks.

2. The distributed processing system according to claim 1, wherein The number of the neural network accelerators is multiple.

3. The distributed processing system according to claim 1, wherein The neural network accelerator adopts a graphics processing unit.

4. The distributed processing system according to claim 1, characterized in that, The first module is specifically used to calculate and obtain a splitting number C according to the formula C = [BF / (2S)], where [] represents rounding, B represents the number of sequences of the sequence generation task, F represents the scheduling parameter, and S represents the generation length.

5. The distributed processing system according to any one of claims 1 to 4, characterized in that The processing unit adopts a central processing unit, a field programmable gate array or a graphics processing unit.

6. A task scheduling method for a distributed processing system according to any one of claims 1 to 5, characterized in that, It includes: Receiving at least one sequence generation task; Dividing the sequence generation task into a first partial task and a second partial task according to the number of sequences of the sequence generation task; Performing task scheduling on the neural network accelerator and the at least one processing unit, so that the neural network accelerator and the at least one processing unit alternately and parallelly process the first partial task and the second partial task.

7. A method for determining parameters of a distributed processing system according to any one of claims 1 to 5, characterized in that, It includes: Obtaining the number of sequences of the sequence generation task according to the constraint conditions of the single calculation duration function corresponding to the neural network accelerator, the generation length, the total delay time, and the number of sequences; where the single calculation duration function corresponding to the neural network accelerator is obtained by testing the calculation duration of the first module processing a single sequence generation task alone; the single calculation duration function corresponding to the neural network accelerator is a relationship function between the calculation duration of the sequence generation task and the number of sequences; Obtaining the number of processing units according to the number of sequences, the generation length, the unit calculation duration, and the single calculation duration function corresponding to the neural network accelerator; where the unit calculation duration is obtained by testing the calculation duration of a single processing unit processing a single element with a sequence length of 1.

8. The method according to claim 7, wherein The constraint condition for the number of sequences is: T(B) < L / (2NS), where B represents the number of sequences, T(B) represents the single-computation duration function corresponding to the neural network accelerator, N represents the number of layers of the model, and S represents the generation length.

9. The method according to claim 7, wherein The obtaining of the number of processing units according to the number of sequences, the generation length, the unit computation duration, and the single-computation duration function corresponding to the neural network accelerator includes: Calculating the number of processing units m according to the formula m = (BSR) / (2T(B)), where B represents the number of sequences, S represents the generation length, R represents the unit computation duration, and T(B) represents the single-computation duration function corresponding to the neural network accelerator.

10. A task scheduling device, characterized in that, Applied to the distributed processing system according to any one of claims 1 to 5, it includes: A receiving unit, configured to receive at least one sequence generation task; A partitioning unit, configured to partition the sequence generation task into a first part task and a second part task according to the number of sequences of the sequence generation task; A scheduling unit, configured to perform task scheduling on the neural network accelerator and the at least one processing unit, so that the neural network accelerator and the at least one processing unit alternately and parallelly process the first part task and the second part task.

11. A parameter determination device, characterized in that, Applied to the distributed processing system according to any one of claims 1 to 5, it includes: A sequence number obtaining unit, configured to obtain the number of sequences of the sequence generation task according to the single-computation duration function corresponding to the neural network accelerator, the generation length, the total delay time, and the constraint condition of the number of sequences; wherein, the single-computation duration function corresponding to the neural network accelerator is obtained by testing the computation duration of the first module for independently processing a single sequence generation task; the single-computation duration function corresponding to the neural network accelerator is a relational function between the computation duration of the sequence generation task and the number of sequences; A processing unit number obtaining unit, configured to obtain the number of processing units according to the number of sequences, the generation length, the unit computation duration, and the single-computation duration function corresponding to the neural network accelerator; wherein, the unit computation duration is obtained by testing the computation duration of a single processing unit for processing a single element with a sequence length of 1.

12. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 6 to 9.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 6 to 9 are implemented.

14. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 6 to 9 are implemented.