Model operator processing method, apparatus, electronic device and storage medium

By evaluating and selectively recalculating model operators based on memory and computation, the method addresses the inefficiencies in existing technologies, enhancing deep learning performance and efficiency through efficient video memory usage.

JP7789155B2Active Publication Date: 2025-12-19BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024161279
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-10-17
Filing Date
2024-09-18
Publication Date
2025-12-19
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

Existing deep learning technologies suffer from inefficient recalculation processes due to bulk recalculating operators, which reduces execution efficiency.

Method used

A method and apparatus for processing model operators by determining an operator set, evaluating memory and computation time for each operator, and identifying a first operator for recalculation based on a recalculation evaluation parameter, allowing for the removal of operators with low cost performance during recalculation.

Benefits of technology

Improves computational performance and efficient video memory usage by selectively recalculating operators with high cost performance, enhancing overall model calculation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007789155000003
    Figure 0007789155000003
  • Figure 0007789155000004
    Figure 0007789155000004
  • Figure 0007789155000005
    Figure 0007789155000005
Patent Text Reader

Abstract

To provide a model operator processing method capable of achieving higher precision operator control by calculating recalculation evaluation parameters of each operator, determining a first operator required to get involved in recalculation on the basis of the recalculation evaluation parameters, removing operators with low recalculation cost performance during recalculation, improving performance by an efficient video memory, and further improving model calculation performance, a model operator processing device, an electronic device, and a storage medium.SOLUTION: A model operator processing method includes the steps of: determining an operator set of model networking including a plurality of operators; determining the storage space occupied by an output tensor of the operator and the calculation time consumed by the operator in forward calculation for each operator in the operator set; and a first operator participating in recalculation in the model from the operator set on the basis of the storage amount and the calculation time of the operator.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of deep learning technology, and in particular to a method, apparatus, electronic device, and storage medium for processing a model operator. [Background technology]

[0002] In the field of deep learning, recalculation is one of the technological paths to realize large-scale model training. In the existing recalculation process, operators are recalculated in bulk, which reduces the execution efficiency of recalculation. Summary of the Invention [Problem to be solved by the invention]

[0003] The present disclosure provides methods, apparatus, electronic devices, and storage media for processing model operators. [Means for solving the problem]

[0004] According to one aspect of the present disclosure, there is provided a method for processing model operators, comprising: determining an operator set for model networking, the operator set including a plurality of operators; for each operator in the operator set, determining an amount of memory occupied by an output tensor of the operator and a computation time required for a forward computation of the operator; and determining a first operator from the operator set that is involved in a recomputation in the model based on the amount of memory and the computation time of the operator.

[0005] According to another aspect of the present disclosure, there is provided a processing apparatus for model operators, the apparatus including: a first determination module for determining an operator set for model networking, the operator set including a plurality of operators; a second determination module for determining, for each operator in the operator set, an amount of memory occupied by an output tensor of the operator and a computation time required for a forward computation of the operator; and a third determination module for determining, from the operator set, a first operator to be involved in recomputation in the model based on the amount of memory and the computation time of the operator.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device including at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform a model operator processing method described in an embodiment of the above aspect.

[0007] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer program / instructions being stored thereon, the computer instructions causing the computer to perform the method for processing a model operator as described in the embodiment of the above aspect.

[0008] According to another aspect of the present disclosure, there is provided a computer program, which, when executed by a processor, implements the method for processing a model operator according to the embodiment of the above aspect.

[0009] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following description. [Brief explanation of the drawings]

[0010] The drawings are used for better understanding of the present technical solution and are not intended to limit the present disclosure. [Figure 1] 1 is a schematic flowchart of a processing method for a model operator provided by an embodiment of the present disclosure; [Figure 2] 10 is a schematic flowchart of another model operator processing method provided by an embodiment of the present disclosure; [Figure 3] 10 is a schematic flowchart of another model operator processing method provided by an embodiment of the present disclosure; [Figure 4] 10 is a schematic flowchart of another model operator processing method provided by an embodiment of the present disclosure; [Figure 5] FIG. 1 is a schematic diagram of merging operators provided by an embodiment of the present disclosure; [Figure 6] 10 is a schematic flowchart of another model operator processing method provided by an embodiment of the present disclosure; [Figure 7] FIG. 1 is a schematic diagram of forward computing operators provided by an embodiment of the present disclosure. [Figure 8] FIG. 10 is a schematic diagram of recalculating operators provided by an embodiment of the present disclosure. [Figure 9] FIG. 1 is a schematic diagram of mixed-parallel model networking provided by an embodiment of the present disclosure. [Figure 10] 10 is a schematic flowchart of another model operator processing method provided by an embodiment of the present disclosure; [Figure 11] FIG. 1 is a schematic block diagram of a processing device for a model operator provided by an embodiment of the present disclosure. [Figure 12] FIG. 1 is a block diagram of an electronic device for implementing a model operator processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] The following describes exemplary embodiments of the present disclosure in conjunction with the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included therein and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, the following description will omit descriptions of well-known functions and structures.

[0012] Hereinafter, a method, apparatus, and electronic device for processing a model operator according to an embodiment of the present disclosure will be described with reference to the drawings.

[0013] Artificial Intelligence (AI) is a field that studies how computers can simulate certain human thought processes and intelligent behaviors (learning, reasoning, thinking, planning, etc.), and includes both hardware and software technologies. AI hardware technology generally includes several aspects such as computer vision technology, speech recognition technology, natural language processing technology and its learning / deep learning, big data processing technology, and knowledge graph technology.

[0014] Natural language processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. NLP is a science that integrates linguistics, computer science, and mathematics. NLP is mainly applied in machine translation, public opinion monitoring, automatic summarization, viewpoint extraction, text classification, question answering, text semantic comparison, and speech recognition.

[0015] Deep learning (DL) is a new research direction in the field of machine learning (ML), introduced to bring machine learning closer to its original goal: artificial intelligence. Deep learning learns the internal rules and representation levels of sample data, and the information acquired during this learning process is highly useful for interpreting data such as text, images, and voice. The ultimate goal is for machines to have human-like analytical learning capabilities and be able to recognize data such as text, images, and voice. Deep learning is a complex machine learning algorithm, and its results in terms of voice and image recognition far exceed those of previous related technologies.

[0016] Machine translation, also known as automatic translation, is the process of using a computer to convert one natural language (the source language) into another natural language (the target language). It is a branch of computational linguistics and one of the goals of artificial intelligence.

[0017] FIG. 1 is a schematic flowchart of a model operator processing method provided by an embodiment of the present disclosure.

[0018] As shown in FIG. 1, the processing method of this model operator can include the following steps S101 to S103.

[0019] In S101, an operator set of the model networking is determined, and the operator set includes a plurality of operators.

[0020] Note that the execution entity of the model operator processing method according to the embodiment of the present disclosure may be a hardware device having data processing capabilities and / or software required to drive the operation of the hardware device. Optionally, the execution entity may include a server, a user terminal, and other smart devices. Optionally, the user terminal may include, but is not limited to, a mobile phone, a personal computer, an intelligent voice interaction device, etc. Optionally, the server may include, but is not limited to, a network server, an application server, a distributed system server, or a server combined with a blockchain, etc. The embodiment of the present disclosure does not specifically limit this.

[0021] In some implementations, the operator set of the model networking can be determined based on information such as the application scenario and structure of the model networking. Different application scenarios and structures correspond to different operator sets. Optionally, the model networking is applied to scenarios such as image processing, target detection, speech recognition, and text generation. Optionally, the structure of the model networking can be a mixed parallel + Transformer structure.

[0022] Alternatively, mixed parallelism may include, but is not limited to, data parallelism and Transformer, model parallelism and Transformer, pipeline parallelism and Transformer, and model and pipeline parallelism.

[0023] In some implementations, the set of operators in the model networking may be determined in combination with the performance of the operators. For example, if a convolution operator is used for image processing, the set of operators in the model networking for an image processing scenario may include multiple convolution operators.

[0024] In the exemplary explanation, Transformer networking is taken as an example, and its operator set includes operators such as FA, Matmul, ReduceScatter, and Allgather.

[0025] In S102, for each operator in the operator set, the amount of memory occupied by the output tensor of the operator and the calculation time required for the forward calculation of the operator are determined.

[0026] In deep learning frameworks, the input and output of operators are usually represented in the form of tensors, which indicate the shape and content of the data. Here, tensors are multidimensional arrays, which can be considered as an extension of scalars, vectors, and matrices, and can be understood as high-dimensional matrices.

[0027] In some implementations, the amount of memory occupied by the output tensor of an operator can be determined based on parameters of the operator input. Optionally, the amount of memory occupied by the output tensor of an operator can be calculated based on parameters such as the batch size (batch_size, B), the sequence length (sequence_len, S), the dimension of the hidden state (hidden_size, H), the model parallelism mp_degree (M), the number of heads (num_head, A), and the dimension of the hidden layer of the feedforward neural network (ffn_hidden_size, H').

[0028] In some implementations, the computation time taken to compute an operator in forward direction can be determined by analyzing the progress of the forward computation of the operator. The progress of the forward computation of the operator can also be monitored to obtain the time from the start to the end of the forward computation of the operator, which can then be used to determine the computation time taken to compute the operator in forward direction.

[0029] In S103, a first operator involved in recalculation in the model is determined from the operator set based on the memory amount and calculation time of the operator.

[0030] The memory size and calculation time of an operator can reflect the cost performance of the operator's involvement in recalculation. The larger the memory size and the shorter the calculation time of the operator, the higher the cost performance of the operator's involvement in recalculation. Conversely, the smaller the memory size and the longer the calculation time of the operator, the lower the cost performance of the operator's involvement in recalculation.

[0031] In some implementations, operators in the operator set may be recalculated and filtered to improve the efficiency of the operator's video memory usage and improve computational performance. Optionally, the ratio between the storage size of the output tensor of an operator and the computation time of the forward computation of this operator may be calculated, and the first operator from the operator set to participate in the recalculation in the model may be determined based on this ratio.

[0032] Optionally, an operator that is involved in recalculation and has high cost performance can be determined from the set of operators based on the ratio between memory amount and calculation time, and set as the first operator.

[0033] According to the model operator processing method provided by the embodiment of the present disclosure, a set of operators for model networking is determined, and for each operator in the set of operators, the amount of memory occupied by the output tensor of the operator and the calculation time required for the forward calculation of the operator are obtained, and the first operator that needs to be involved in recalculation is determined. This makes it possible to remove operators with low recalculation cost performance during recalculation, thereby improving performance through efficient video memory and increasing calculation performance per unit video memory exchange. Efficient use of video memory further improves the calculation performance of the model. Operator refinement control can be achieved for each type of operator in the model.

[0034] FIG. 2 is a schematic flowchart of a model operator processing method provided by an embodiment of the present disclosure.

[0035] As shown in FIG. 2, the processing method of this model operator can include the following steps S201 to S203.

[0036] In S201, an operator set of the model networking is determined, and the operator set includes a plurality of operators.

[0037] The details of step S201 can be referred to in the above embodiment, and the description will be omitted here.

[0038] In S202, for each operator in the operator set, a recalculation evaluation parameter for the operator is determined.

[0039] The recalculation value parameter of an operator can represent the video memory size saved by the recalculation of the operator per unit time, and is used to describe the efficiency of performance improvement by the video memory of the operator. The larger the recalculation value parameter of an operator, the more frequently the operator needs to be recalculated.

[0040] In some implementations, operators in an operator set can be recalculated and filtered based on the recalculation evaluation parameters of the operators to improve the efficiency of video memory usage of the operators and improve computational performance. Optionally, the recalculation evaluation parameters of the operators can be determined by analyzing the computation and video memory efficiency of each operator.

[0041] In some implementations, the ratio can be determined as a recalculation evaluation parameter by calculating the ratio between the storage size of the output tensor of an operator and the computation time of the forward computation of this operator.

[0042] In S203, a first operator involved in the recalculation in the model is determined from the operator set based on the recalculation evaluation parameter.

[0043] In some implementations, for an operator set, the operators in the operator set can be sorted based on the recalculation evaluation parameter, and the first operator involved in the recalculation in the model can be determined based on the sorting result. Also, the first operator can be determined based on a set threshold by comparing the magnitude of the recalculation evaluation parameter with the magnitude of the set threshold.

[0044] Optionally, when the recalculation evaluation parameter is sorted in descending order, the operator at the front of the sorting can be determined as the first operator. For example, when the operators in the operator set are sorted in descending order of the recalculation evaluation parameter, the three operators at the front of the sorting result can be determined as the first operator.

[0045] Optionally, when the recalculation evaluation parameters are sorted in ascending order, the operator at the end of the sorting can be determined as the first operator. For example, when the operators in the operator set are sorted in ascending order of the recalculation evaluation parameters, the three operators at the end of the sorting result can be determined as the first operator.

[0046] Optionally, a first operator involved in the recalculation in the model is determined from the set of operators by comparing the magnitude of the recalculation evaluation parameter with the magnitude of a set threshold, the first operator being an operator whose recalculation evaluation parameter is equal to or greater than the set threshold.

[0047] According to the model operator processing method provided by the embodiment of the present disclosure, a set of operators in the model networking is determined, and for each operator in the set of operators, its recalculation evaluation parameter is calculated, and then the first operator that needs to be involved in the recalculation is determined based on the recalculation evaluation parameter. This makes it possible to eliminate operators with low recalculation cost performance during recalculation, thereby improving performance through efficient video memory and increasing the calculation performance per unit video memory exchange. Efficient use of video memory further improves the calculation performance of the model. Operator refinement control can be achieved for each type of operator in the model.

[0048] FIG. 3 is a schematic flowchart of a model operator processing method provided by an embodiment of the present disclosure.

[0049] As shown in FIG. 3, the processing method of this model operator can include the following steps S301 to S305.

[0050] In S301, an operator set of the model networking is determined, and the operator set includes a plurality of operators.

[0051] The content of step S301 can be referred to in the above embodiment, and the description thereof will be omitted here.

[0052] In S302, the amount of memory occupied by the output tensor of the operator is determined.

[0053] In deep learning frameworks, the input and output of operators are usually represented in the form of tensors, which indicate the shape and content of the data. Here, tensors are multidimensional arrays, which can be considered as an extension of scalars, vectors, and matrices, and can be understood as high-dimensional matrices.

[0054] In some implementations, the amount of memory occupied by the output tensor of the operator can be determined based on parameters of the operator input. Optionally, the amount of memory occupied by the output tensor of the operator can be calculated based on parameters such as the batch size (batch_size, B), the sequence length (sequence_len, S), the dimension of the hidden state (hidden_size, H), the model parallelism mp_degree (M), the number of heads (num_head, A), and the dimension of the hidden layer of the feedforward neural network (ffn_hidden_size, H').

[0055] In S303, the calculation time required for the forward calculation of the operator is determined.

[0056] In some implementations, the computation time taken to compute an operator in forward direction can be determined by analyzing the progress of the forward computation of the operator. The progress of the forward computation of the operator can also be monitored to obtain the time from the start to the end of the forward computation of the operator, which can then be used to determine the computation time taken to compute the operator in forward direction.

[0057] In S304, the recalculation evaluation parameters of the operators are determined based on the memory amount and calculation time.

[0058] In some implementations, we take the ratio between the amount of memory occupied by the output tensor of the operator and the computation time taken during forward computation, and determine this ratio as the recalculation evaluation parameter of the operator. Optionally, the formula for calculating the recalculation evaluation parameter is as follows:

number

[0059] η represents the recalculation evaluation parameter of the operator, Mem represents the amount of memory occupied by the output tensor of the operator, and Time represents the calculation time required for the forward calculation of the operator.

[0060] In some implementations, the recalculation value parameter represents the video memory size saved by recalculation per unit time, i.e., the larger the recalculation value parameter of an operator, the more likely the operator is to be involved in recalculation.

[0061] In an exemplary explanation, for an operator such as Layernorm, the calculation is fast and the video memory is large, i.e., the calculation time is short and the memory amount is large, so the corresponding recalculation evaluation parameter is large, so Layernorm recalculation is necessary. For the MatMul operator, the calculation amount is large, the calculation time is long and the recalculation efficiency is low, so MatMul is unlikely to be involved in recalculation.

[0062] In some implementations, offline analysis and calculations can be performed on the amount of memory occupied by the operator's output tensor and the computation time required for forward calculations. These offline analysis and calculations are performed to obtain the recalculation evaluation parameters for the operator.

[0063] In S305, a first operator involved in the recalculation in the model is determined from the operator set based on the recalculation evaluation parameter.

[0064] In some implementations, the operators in the operator set can be sorted based on the recalculation evaluation parameter, and a first operator can be selected from the operator set based on the sorting result. During recalculation, it can be realized that an operator with high recalculation cost performance can be used, thereby realizing performance improvement through efficient video memory.

[0065] Optionally, the operators in the operator set are sorted in descending order of recalculation evaluation parameters, and the first N operators are selected and set as the first operator. Optionally, the operators in the operator set are sorted in descending order of recalculation evaluation parameters, and the last N operators are selected and set as the first operator, where N is a natural number greater than or equal to 1.

[0066] For example, if N is equal to 5, the first 5 operators sorted from largest to smallest can be selected to be the first operator, or the last 5 operators sorted from smallest to largest can be selected to be the first operator.

[0067] In some implementations, a threshold value for the recalculation evaluation parameter may be set to select a first operator from the set of operators based on the threshold value. During recalculation, an operator with high recalculation cost performance may be used, thereby achieving improved performance with efficient video memory.

[0068] Optionally, the recalculation evaluation parameter of the operators in the operator set is compared with a set threshold, and an operator whose recalculation evaluation parameter is equal to or greater than the set threshold is selected as the first operator.

[0069] According to the model operator processing method provided by the embodiment of the present disclosure, a set of operators for model networking is determined, and for each operator in the operator set, the amount of memory occupied by the output tensor of the operator and the calculation time required for the forward calculation of the operator are determined, and a recalculation evaluation parameter is then calculated. Based on the recalculation evaluation parameter, the first operator that needs to be involved in the recalculation is determined. This allows operators with low recalculation cost performance to be removed during recalculation, thereby improving performance through efficient video memory and increasing calculation performance per unit video memory exchange. Efficient use of video memory further improves the calculation performance of the model. Operator refinement control can be achieved for each type of operator in the model.

[0070] FIG. 4 is a schematic flowchart of a model operator processing method provided by an embodiment of the present disclosure.

[0071] As shown in FIG. 4, the processing method of this model operator can include the following steps S401 to S404.

[0072] In S401, candidate operators required for model networking are determined.

[0073] In S402, operators of the target class are determined from the candidate operators, and a merging process is performed on the operators of the target class to obtain an operator set.

[0074] In some implementations, the required candidate operators can be determined based on information such as the application scenario, structure, etc. of the model networking, combined with the performance of the operators.

[0075] Note that if a class of operators does not require forward inputs and outputs when computed backward, then this class of operators requires special processing: it can be merged with the previous operator, and the merged operator can be considered as one operator, further improving the performance of the model computation.

[0076] In some implementations, candidate operators whose forward inputs and outputs are not required during backward computation can be determined as operators of the target class, operators of the target class can be determined from the candidate operators, and a merging operation can be performed on the operators of the target class to obtain an operator set.

[0077] Optionally, for any candidate operator belonging to the target class of operators, a previous candidate operator adjacent to the candidate operator can be determined, and the candidate operator can be combined with the previous candidate operator to obtain a merged operator. Further, an operator set is obtained based on the merged operator and the remaining candidate operators. For example, the candidate operators include Operator A, Operator B, Operator C, Operator D, and Operator E, and Operator C is merged with Operator B to obtain Merged Operator B', and Operator A, Operator D, and Operator E are the remaining candidate operators, which together with Merged Operator B' constitute an operator set.

[0078] In an exemplary explanation, when the model networking is a mixed parallel + Transformer structure, there is a class of operator ReduceScatter communication, and its backward calculation does not need forward input and output, and it is preceded by a RowLN matrix, so RowLN + ReduceScatter is merged to obtain a merged operator RowLN + ReduceScatter, that is, the two operators RowLN_0 and ReduceScatter as a whole are one merged operator.

[0079] Figure 5 shows a schematic diagram of operator merging. RowLN_0 and ReduceScatter are merged to obtain the merged operator, and the operator in the dashed box in Figure 5 is the merged operator. Here, BSHReduceScatter represents the parameter of the operator, and BSH / M represents the output size of the ReduceScatter operator. Here, B represents batch_size, S represents sequence_len, H represents hidden_size, and M represents mp_degree.

[0080] In S403, for each operator in the operator set, a recalculation evaluation parameter for the operator is determined.

[0081] In some implementations, since the merged operator is composed of multiple operators, the recalculation evaluation parameters of the merged operator can be determined based on the memory amount and forward calculation time of the output tensor of each candidate operator included in the merged operator.

[0082] Optionally, by obtaining the memory amount of the output tensor of each candidate operator and the calculation time required for forward calculation, the recalculation evaluation parameters of the merged operator are calculated based on the above formula (1).

[0083] In S404, a first operator involved in the recalculation in the model is determined from the operator set based on the recalculation evaluation parameter.

[0084] The content of step S404 can be referred to in the above embodiment, and the description thereof will be omitted here.

[0085] According to the model operator processing method provided by the embodiment of the present disclosure, candidate operators required for model networking are obtained, operators of a target class are determined from the candidate operators, and a merging process is performed on the operators of the target class to obtain an operator set. For each operator in the operator set, its recalculation evaluation parameter is calculated, and the first operator that needs to be involved in the recalculation is determined based on the recalculation evaluation parameter. This makes it possible to remove operators with low recalculation cost performance during recalculation, thereby improving performance through efficient video memory and increasing the calculation performance per unit video memory exchange. Efficient use of video memory further improves the calculation performance of the model. Operator refinement control can be achieved for each type of operator in the model.

[0086] FIG. 6 is a schematic flowchart of a processing method of a model operator provided by an embodiment of the present disclosure.

[0087] As shown in FIG. 6, the processing method of this model operator can include the following steps S601 to S605.

[0088] In S601, an operator set of the model networking is determined, and the operator set includes a plurality of operators.

[0089] In S602, for each operator in the operator set, a recalculation evaluation parameter for the operator is determined.

[0090] In S603, a first operator involved in the recalculation in the model is determined from the operator set based on the recalculation evaluation parameter.

[0091] The above embodiment can be referred to for the details of steps S601 to S605, and the description thereof will be omitted here.

[0092] In S604, the forward calculation of the model is performed according to the forward logical order of the operators in the operator set, and the intermediate result of the first operator is released.

[0093] In some implementations, after determining the first operator, the model networking operators can be recalculated. The recalculation can include forward calculation, forward recalculation, and backward calculation. During the forward calculation process, if the recalculation evaluation parameter of the first operator is high, it indicates that the calculation time of the first operator is short. Since the forward recalculation time is short, the intermediate results of the first operator can be released to reduce the occupation of video memory.

[0094] Figure 7 shows a schematic diagram of forward calculation of operators. The operator set of one layer of the model includes five operators, OP0, OP1, OP2, OP3, and OP4. The operator recalculation evaluation parameters can determine that OP0, OP1, OP3, and OP4 are the first operators, and OP2 is the second operator. Forward calculation is performed on the operators of this layer of the model according to the logical order, i.e., the intermediate results of the operators are calculated in the order of OP0, OP1, OP2, OP3, and OP4, and the intermediate results of OP0, OP1, OP3, and OP4 are released.

[0095] In S605, an intermediate result of the forward calculation of a second operator other than the first operator in the operator set is stored in a video memory, and subsequent recalculation for the second operator is skipped based on the intermediate result of the second operator.

[0096] In some implementations, the first operator refers to an operator in the operator set that needs to be recalculated, and the second operator refers to an operator in the operator set that does not need to be recalculated. When performing the first forward calculation, it is determined based on the recalculation evaluation parameters of the operators that the second operator does not need to be recalculated in subsequent steps. Therefore, during the first forward calculation, the intermediate results of the forward calculation of the second operator are stored in the video memory, and subsequent recalculation of the second operator is skipped based on the intermediate results of the second operator. For example, in the operators shown in FIG. 7, OP2 is the second operator, and the intermediate result of OP2 obtained by the forward calculation is C1, which is stored in the video memory.

[0097] In some implementations, after storing the intermediate results of the second operator in video memory, the forward computation can be performed again according to a first logical order of the forward computation of the operators of the operator set, where the first logical order is the forward logical order.

[0098] Optionally, based on the forward input of the first operator, a forward recalculation is performed on the first operator to obtain an intermediate result of the forward calculation of the first operator. When the forward calculation reaches the second operator, the intermediate result of the second operator is read from the video memory, thereby reducing the forward calculation of the second operator during the forward recalculation, and increasing the video memory can improve calculation performance and further reduce the overall recalculation time.

[0099] Furthermore, a backward calculation of the model is performed based on the intermediate result of the first operator and the intermediate result of the second operator. Optionally, the backward calculation can be performed according to a second logical order of the backward calculation of the operators of the operator set.

[0100] Optionally, for a first operator, the first operator is calculated backward based on the intermediate result of the first operator and the backward output of the previous operator during the backward calculation of the first operator, and the backward output of the first operator is obtained and input to the next operator. When the backward calculation reaches the second operator, the intermediate result of the second operator is read from the video memory, and the backward output of the second operator is determined based on the intermediate result of the second operator and the backward output of the previous operator during the backward calculation, and the backward output of the second operator is input to the next operator of the second operator.

[0101] That is, when the first operator is calculated backward, the intermediate result of the first operator and the backward output of the previous backward calculation of the first operator are input to the first operator to perform the backward calculation. Figure 8 shows a schematic diagram of recalculating operators, taking OP1 as an example. When calculating backward, the input of OP1 includes the intermediate result of the forward calculation of OP1 and the backward output of OP2. As shown by the dashed line in the figure, during forward recalculation, the recalculation of OP2 can be skipped by reading C1 stored in the video memory.

[0102] In some implementations, the forward input of the first operator during the first forward calculation of the model can be stored in the video memory. This allows the input of the first operator to be read from the video memory and input to the first operator for recalculation when forward calculation is performed on the model again. Taking OP0 shown in Figure 7 as an example, OP0 is the first operator, C0 is the forward input of OP0, and OP0 can perform forward calculation based on C0 to obtain the intermediate results of OP0.

[0103] According to the model operator processing method provided by the embodiment of the present disclosure, a set of operators in the model network is determined, and for each operator in the operator set, its recalculation evaluation parameter is calculated, and then the first operator that needs to be involved in the recalculation is determined based on the recalculation evaluation parameter. This makes it possible to eliminate operators with low recalculation cost performance during recalculation, thereby improving performance through efficient video memory and increasing calculation performance per unit video memory exchange. By efficiently using video memory, the calculation performance of the model is further improved. Operator refinement control can be achieved for each type of operator in the model. During recalculation, intermediate results of the second operator are saved in video memory, reducing the forward calculation of the second operator and increasing the video memory, thereby improving calculation performance.

[0104] In the exemplary explanation, the recalculation process will be described using Transformer networking as an example. For Transformer, the relevant operators (OP) mainly include FA, Matmul, ReduceScatter, and Allgather. After the operators are merged, the recalculation evaluation parameters are calculated using the above formula (1) to obtain the recalculation evaluation parameter sort table as shown in Figure 1. [Table 1]

[0105] As can be seen from Table 1, the operator set contains nine operators, and operators numbered 1, 3, and 6 are merged operators. The recalculation evaluation parameters are calculated, and the operators are sorted in ascending order of recalculation evaluation parameters. The last four operators are selected as the first operators. The first five operators are the second operators. Here, the time percentage represents the proportion of the calculation time required to calculate one of the operators in the entire operator set. For example, the calculation time of the FA operator accounts for 13.8% of the total calculation time for the entire operator set. Here, the calculation formula represents the size of the operator's output, and its value is the size of the memory capacity. In the formula, B represents batch_size, S represents sequence_len, H represents hidden_size, M represents mp_degree, A represents num_head, and H' represents ffn_hidden_size.

[0106] Figure 9 shows a schematic diagram of the networking of the mixed parallel model, which includes a MultiHeadAttention network and a Multilayer Perception (MLP) network.

[0107] Here, the MultiHeadAttention network includes a LayerNom operator, an ALLgather operator, a ColumnLN_0 operator, an FA operator, a RowLN_0 operator, and a ReduceScatter operator, and the MLP network includes a LayerNom operator, an ALLgather operator, a ColumnLN_1 operator, a Silu + Ele operator, a RowLN_1 operator, and a ReduceScatter operator.

[0108] Here, BSH / M, BSH, BSH' / M, etc., in the inputs of each operator in Figure 9, represent the operator input parameters. B represents the batch size (batch_size), S represents the sequence length (sequence_len), H represents the dimension of the hidden state (hidden_size), M represents the degree of model parallelism (mp_degree), A represents the number of heads (num_head), and H' represents the dimension of the hidden layer of the feedforward neural network (ffn_hidden_size). The FA operator has two outputs, and LSE represents the name of the output of the FA operator, with its size being BAS / M.

[0109] The operator in the dashed box is the second operator, as shown in Figure 9. Figure 9 further illustrates the forward logical order of the operators.

[0110] During the recalculation, the forward calculation is performed according to the forward logical order of the operators in Figure 9, the intermediate results of the first operator are released, i.e., the intermediate results of the last four operators in Table 1 are released, and the intermediate results of the forward calculation of the second operator are stored in the video memory. Further, according to the forward logical order of the operators in Figure 9, a forward recalculation is performed on the first operator based on the forward input of the first operator to obtain the intermediate results of the forward calculation of the first operator. When the forward recalculation reaches the second operator, the intermediate results of the second operator are read from the video memory, and the backward calculation of the model is performed based on the intermediate results of the first operator and the intermediate results of the second operator.

[0111] The backward calculation process of the first operator takes the Allgather operator as an example. Based on the intermediate result of Allgather and the backward output of ColumnLN_0, Allgather is backward calculated, and the backward output of Allgather is obtained and input into LayerNorm.

[0112] The backward calculation process of the second operator takes ColumnLN_1 as an example, reads the intermediate result of ColumnLN_1 from the video memory, and determines the backward output of ColumnLN_1 based on the intermediate result of ColumnLN_1 and the backward output of Silu+Ele, and inputs the backward output of ColumnLN_1 into Allgather.

[0113] FIG. 10 is a schematic flowchart of a model operator processing method provided by an embodiment of the present disclosure.

[0114] As shown in FIG. 10, the processing method of this model operator may include the following steps S1001 to S1008.

[0115] In S1001, candidate operators required for model networking are determined.

[0116] In S1002, operators of the target class are determined from the candidate operators, and a merging process is performed on the operators of the target class to obtain an operator set.

[0117] In S1003, for each operator in the operator set, the amount of memory occupied by the output tensor of the operator is determined.

[0118] In S1004, the calculation time required for the forward calculation of the operator is determined.

[0119] In S1005, the recalculation evaluation parameters of the operators are determined based on the memory amount and calculation time.

[0120] In S1006, a first operator involved in the recalculation in the model is determined from the operator set based on the recalculation evaluation parameter.

[0121] In S1007, the forward calculation of the model is performed according to the forward logical order of the operators in the operator set, and the intermediate result of the first operator is released.

[0122] In S1008, an intermediate result of the forward calculation of a second operator other than the first operator in the operator set is stored in the video memory, and subsequent recalculation of the second operator is skipped based on the intermediate result of the second operator.

[0123] According to the model operator processing method provided by the embodiment of the present disclosure, a set of operators in the model network is determined, and for each operator in the operator set, its recalculation evaluation parameter is calculated, and then the first operator that needs to be involved in the recalculation is determined based on the recalculation evaluation parameter. This makes it possible to eliminate operators with low recalculation cost performance during recalculation, thereby improving performance through efficient video memory and increasing calculation performance per unit video memory exchange. By efficiently using video memory, the calculation performance of the model is further improved. Operator refinement control can be achieved for each type of operator in the model. During recalculation, intermediate results of the second operator are saved in video memory, reducing the forward calculation of the second operator and increasing the video memory, thereby improving calculation performance.

[0124] Corresponding to the model operator processing method provided by some of the above embodiments, one embodiment of the present disclosure further provides a model operator processing device, and the model operator processing device provided by the embodiments of the present disclosure corresponds to the model operator processing method provided by some of the above embodiments, so the above embodiments of the model operator processing method also apply to the model operator processing device provided by the embodiments of the present disclosure, and descriptions thereof will be omitted in the following embodiments.

[0125] FIG. 11 is a schematic block diagram of a processing device for a model operator provided by an embodiment of the present disclosure.

[0126] As shown in FIG. 11, the processing device 1100 of the model operator in the embodiment of the present disclosure includes a first determining module 1101, a second determining module 1102 and a third determining module 1103.

[0127] A first determination module 1101 determines an operator set of a model networking, where the operator set includes a plurality of operators.

[0128] A second determination module 1102 determines, for each operator in the operator set, the amount of memory occupied by the output tensor of the operator and the computation time taken when computing the operator in forward direction.

[0129] A third determination module 1103 determines a first operator involved in recalculation in the model from the set of operators based on the memory amount and the calculation time of the operator.

[0130] In one embodiment of the present disclosure, the third determination module 1103 further determines recalculation evaluation parameters of the operator based on the memory amount and the calculation time, and determines a first operator from the operator set that will be involved in recalculation in the model based on the recalculation evaluation parameters of the operator.

[0131] In an embodiment of the present disclosure, the third determination module 1103 further obtains a ratio between the memory amount and the calculation time, and determines the ratio as the recalculation evaluation parameter.

[0132] In one embodiment of the present disclosure, the larger the recalculation evaluation parameter of the operator, the more likely the operator is to participate in recalculation.

[0133] In one embodiment of the present disclosure, the third determination module 1103 further sorts the operators in the operator set based on the recalculated evaluation parameters, and selects the first operator from the operator set based on the sorting result.

[0134] In one embodiment of the present disclosure, the third determination module 1103 further compares the recalculation evaluation parameters of the operators in the operator set with a set threshold, and selects the operators whose recalculation evaluation parameters are greater than or equal to the set threshold as the first operators.

[0135] In one embodiment of the present disclosure, the first determination module 1101 further determines candidate operators required during the model networking, determines operators of a target class from the candidate operators, and performs a merging process on the operators of the target class to obtain the operator set.

[0136] In one embodiment of the present disclosure, the first determination module 1101 further determines, for any candidate operator belonging to the target class of operators, a previous candidate operator adjacent to the any candidate operator, combines the any candidate operator with the previous candidate operator to obtain a merged operator, and obtains the operator set based on the merged operator and the remaining candidate operators.

[0137] In one embodiment of the present disclosure, the first determination module 1101 further determines candidate operators whose forward inputs and outputs are not required during backward calculation, to be operators of the target class.

[0138] In one embodiment of the present disclosure, the third determination module 1103 further determines recalculation evaluation parameters of the merged operator based on the memory amount of the output tensor of each candidate operator included in the merged operator and the calculation time.

[0139] In one embodiment of the present disclosure, the third determination module 1103 further performs forward calculation of the model according to the forward logical order of the operators in the operator set, releases the intermediate result of the first operator, stores the intermediate result of the forward calculation of a second operator other than the first operator in the operator set in a video memory, and skips subsequent recalculation for the second operator based on the intermediate result of the second operator.

[0140] In one embodiment of the present disclosure, the third determination module 1103 further performs forward calculation again according to a first logical order of forward calculation of the operators of the operator set, and based on the forward input of the first operator, performs forward recalculation on the first operator to obtain intermediate results of the forward calculation of the first operator; when the forward calculation reaches the second operator, reads the intermediate results of the second operator from the video memory, and performs backward calculation of the model based on the intermediate results of the first operator and the intermediate results of the second operator.

[0141] In one embodiment of the present disclosure, the third determination module 1103 further performs backward calculation according to a second logical order of backward calculation of the operators in the operator set, and for the first operator, performs backward calculation of the first operator based on the intermediate result of the first operator and the backward output of the previous operator during the backward calculation of the first operator, obtains the backward output of the first operator and inputs it into the next operator; when the backward calculation reaches the second operator, reads the intermediate result of the second operator from the video memory, determines the backward output of the second operator based on the intermediate result of the second operator and the backward output of the previous operator during the backward calculation, and inputs the backward output of the second operator into the next operator of the second operator.

[0142] In one embodiment of the present disclosure, the device further stores in the video memory a forward input of a first operator during a first forward calculation of the model, and when performing forward calculation again on the model, reads the input of the first operator from the video memory and inputs it into the first operator to perform recalculation.

[0143] In one embodiment of the present disclosure, the third determination module 1103 further sorts the operators in the operator set in descending order of the recalculation evaluation parameter and selects the first N operators to be the first operator, or sorts the operators in the operator set in descending order of the recalculation evaluation parameter and selects the last N operators to be the first operator, where N is a natural number greater than or equal to 1.

[0144] The model operator processing device provided by the embodiment of the present disclosure determines an operator set for the model networking, calculates a recalculation evaluation parameter for each operator in the operator set, and then determines a first operator that needs to be involved in the recalculation based on the recalculation evaluation parameter. This makes it possible to remove operators with low recalculation cost performance during recalculation, thereby improving performance through efficient video memory and increasing calculation performance per unit video memory exchange. By efficiently using video memory, the calculation performance of the model can be further improved. Operator refinement control can be achieved for each type of operator in the model. During recalculation, intermediate results of the second operator are saved in video memory, reducing the forward calculation of the second operator and increasing the video memory, thereby improving calculation performance.

[0145] Furthermore, in the technical solution disclosed herein, the acquisition, storage, application, etc. of relevant user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and morals.

[0146] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium. According to an embodiment of the present disclosure, the present disclosure provides a computer program, which, when executed by a processor, implements the method for processing a model operator proposed by the present disclosure.

[0147] 12 is a schematic block diagram of an exemplary electronic device 1200 for implementing embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the description herein and / or the practice of the present disclosure as sought.

[0148] 12, electronic device 1200 includes a computing unit 1201 that performs various appropriate operations and processes in accordance with computer programs / instructions stored in a read-only memory (ROM) 1202 or loaded from a storage unit 1206 into a random access memory (RAM) 1203. The RAM 1203 may also store various programs and data necessary for the operation of electronic device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are connected to one another via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0149] The components of the electronic device 1200 are connected to an I / O interface 1205, which includes an input unit 1206 such as a keyboard, a mouse, etc., an output unit 1207 such as various types of displays, speakers, etc., a storage unit 1208 such as a magnetic disk, an optical disk, etc., and a communication unit 1209 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 enables the electronic device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunications networks.

[0150] The computing unit 1201 may be various general-purpose and / or special-purpose processing components having processing and computation capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various machine driving learning model algorithm computing units, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs each of the methods and processes described above, e.g., the model operator processing method. For example, in some embodiments, the model operator processing method may be implemented as a computer software program tangibly embodied in a machine-readable medium such as the storage unit 1206, and some or all of the computer program / instructions may be loaded and / or installed into the electronic device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program / instructions are loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the model operator processing method described above may be performed. Alternatively, in other embodiments, the computation unit 1201 may be configured in any other suitable manner (eg, via firmware) to perform the processing methods of the model operators.

[0151] Various embodiments of the systems and techniques described herein above may be realized in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being embodied in one or more computer programs / instructions that may be executed and / or interpreted by a programmable system including at least one programmable processor, which may be an application specific or general purpose programmable processor, and that may receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.

[0152] Program code for carrying out the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus such that, when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine, partially on a remote machine, or entirely on a remote machine or server.

[0153] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include one or more line-based electrical connections, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0154] To provide interaction with a user, the systems and techniques described herein can be executed on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices can also provide interaction with a user; for example, the feedback provided to the user can be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback) and can receive input from the user in any form (including acoustic, speech, or tactile input).

[0155] The systems and techniques described herein can be implemented on a computing system including a back-end component (e.g., a data server), or a computing system including a middleware component (e.g., an application server), or a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such back-end, middleware, and front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0156] The computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between client and server is created by computer programs / instructions running on corresponding computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server incorporating a blockchain.

[0157] It should be understood that steps can be rearranged, added, or deleted using the various types of flows shown above. For example, the steps described in the present disclosure may be performed in parallel, sequentially, or in a different order, but the present specification does not limit the scope of the present disclosure as long as the technical solutions disclosed in the present disclosure can achieve the desired results.

[0158] The above specific embodiments do not limit the scope of protection of the present disclosure. It should be understood that those skilled in the art can make various modifications, combinations, subcombinations, and substitutions according to design requirements and other factors. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. A method for processing a model operator, comprising: determining an operator set for the model networking, the operator set including a plurality of operators; For each operator in the set of operators, determining the amount of memory occupied by the output tensor of the operator and the computation time taken when computing the operator in a forward direction; determining a first operator from the set of operators that will be involved in recalculation in the model based on the memory amount and the calculation time of the operator; determining a first operator involved in recalculation in the model from the set of operators based on the memory amount and the calculation time of the operator, determining a recalculation evaluation parameter for the operator based on the memory amount and the calculation time; determining a first operator from the set of operators that will participate in a recalculation in the model based on a recalculation evaluation parameter of the operator; How to handle model operators.

2. The step of determining a recalculation evaluation parameter of the operator based on the memory amount and the calculation time includes: obtaining a ratio between the memory amount and the calculation time, and determining the ratio as the recalculation evaluation parameter; The method of claim 1 .

3. the larger the recalculation evaluation parameter of said operator, the more likely said operator is to participate in recalculation; The method of claim 1 .

4. determining a first operator from the set of operators that will participate in recalculation in the model based on a recalculation evaluation parameter of the operator, sorting the operators in the operator set based on the recalculation evaluation parameter, and selecting the first operator from the operator set based on a sorting result; A method for processing a model operator according to any one of claims 1 to 3.

5. determining a first operator from the set of operators that will participate in recalculation in the model based on a recalculation evaluation parameter of the operator, comparing the recalculated evaluation parameters of the operators in the operator set with a set threshold; selecting an operator whose recalculated evaluation parameter is equal to or greater than the set threshold value as the first operator, A method for processing a model operator according to any one of claims 1 to 3.

6. The step of determining a set of operators of the model networking includes: determining candidate operators required during said model networking; determining operators of a target class from the candidate operators, and performing a merging process on the operators of the target class to obtain the operator set; A method for processing a model operator according to any one of claims 1 to 3.

7. The step of performing a merging process on the operators of the target class to obtain the operator set includes: for any candidate operator belonging to the target class of operators, determining a previous candidate operator adjacent to the any candidate operator, and combining the any candidate operator with the previous candidate operator to obtain a merged operator; obtaining the set of operators based on the merged operators and remaining candidate operators; 7. A method for processing a model operator according to claim 6.

8. Determining an operator of a target class from among the candidate operators includes: determining candidate operators whose forward inputs and outputs are not required during the backward computation to be operators of said target class; 7. A method for processing a model operator according to claim 6.

9. The method comprises: determining a recalculation evaluation parameter of the merged operator based on the memory amount of the output tensor of each candidate operator included in the merged operator and the calculation time; 8. A method for processing a model operator according to claim 7.

10. After the step of determining a first operator from the set of operators that is involved in a recalculation in the model, performing a forward computation of the model according to a forward logical order of the operators of the operator set and releasing intermediate results of the first operator; storing an intermediate result of a forward calculation of a second operator other than the first operator in the operator set in a video memory, and skipping a subsequent recalculation for the second operator based on the intermediate result of the second operator; A method for processing a model operator according to any one of claims 1 to 3.

11. skipping a subsequent recalculation for the second operator based on an intermediate result of the second operator includes: - performing a forward computation again according to a first logical order of forward computation of the operators of said operator set; performing a forward recalculation on the first operator based on a forward input of the first operator to obtain an intermediate result of the forward calculation of the first operator; when the forward calculation reaches the second operator, reading an intermediate result of the second operator from the video memory; performing a backward calculation of the model based on intermediate results of the first operator and intermediate results of the second operator; 11. The method of claim 10.

12. performing a backward calculation of the model based on an intermediate result of the first operator and an intermediate result of the second operator, performing backward computations according to a second logical order of backward computations of the operators of the operator set; For the first operator, performing a backward calculation of the first operator based on an intermediate result of the first operator and a backward output of a previous operator during the backward calculation of the first operator, to obtain a backward output of the first operator and input it to a next operator; When the backward calculation reaches the second operator, reading an intermediate result of the second operator from the video memory, determining a backward output of the second operator based on the intermediate result of the second operator and a backward output of a previous operator during the backward calculation, and inputting the backward output of the second operator to a next operator of the second operator.

12. A method for processing a model operator according to claim 11.

13. The method comprises: storing in said video memory a forward input of a first operator in a first forward calculation of said model; When performing forward calculation again on the model, the input of the first operator is read from the video memory, and the input is input to the first operator for recalculation.

12. A method for processing a model operator according to claim 11.

14. the step of sorting the operators in the operator set based on the recalculation evaluation parameter, and selecting the first operator from the operator set based on the sorting result, Sorting the operators in the operator set in descending order of the recalculation evaluation parameter, and selecting the first N operators as the first operator; or a step of sorting the operators in the operator set in ascending order of the recalculation evaluation parameter, selecting the last N operators, and setting them as the first operator; The N is a natural number equal to or greater than 1.

5. A method for processing a model operator according to claim 4.

15. 1. A processing device for a model operator, comprising: a first determination module for determining an operator set of a model networking, the operator set including a plurality of operators; a second determination module for determining, for each operator in the set of operators, the amount of memory occupied by the output tensor of the operator and the computation time taken when computing the operator in a forward direction; a third determination module for determining a first operator involved in recalculation in the model from the set of operators based on the memory amount and the calculation time of the operator; The third determination module further comprises: determining a recalculation evaluation parameter for the operator based on the memory amount and the calculation time; determining a first operator from the set of operators that will participate in a recalculation in the model based on a recalculation evaluation parameter of the operator; Model operator processing unit.

16. The third determination module further comprises: obtaining a ratio between the memory amount and the calculation time, and determining the ratio as the recalculation evaluation parameter; 16. A model operator processing device according to claim 15.

17. the larger the recalculation evaluation parameter of said operator, the more likely said operator is to participate in recalculation; 16. A model operator processing device according to claim 15.

18. The third determination module further comprises: sorting the operators in the operator set based on the recalculation evaluation parameter, and selecting the first operator from the operator set based on the sorting result; A device for processing a model operator according to any one of claims 15 to 17.

19. The third determination module further comprises: comparing the recalculated evaluation parameters of the operators in the operator set with a set threshold; an operator whose recalculation evaluation parameter is equal to or greater than the set threshold is selected as the first operator; A device for processing a model operator according to any one of claims 15 to 17.

20. The first determination module further comprises: determining candidate operators required during said model networking; determining operators of a target class from the candidate operators, and performing a merging process on the operators of the target class to obtain the operator set; A device for processing a model operator according to any one of claims 15 to 17.

21. The first determination module further comprises: For any candidate operator belonging to the target class of operators, determine a previous candidate operator adjacent to the any candidate operator, and combine the any candidate operator with the previous candidate operator to obtain a merged operator; obtaining the set of operators based on the merged operators and the remaining candidate operators; 21. A model operator processing device according to claim 20.

22. The first determination module further comprises: determining candidate operators whose forward inputs and outputs are not required during the backward computation, and making them operators of the target class; 21. A model operator processing device according to claim 20.

23. The third determination module further comprises: determining a recalculation evaluation parameter for the merged operator based on the memory amount of the output tensor of each candidate operator included in the merged operator and the calculation time; 22. A model operator processing device according to claim 21.

24. The third determination module further comprises: performing a forward computation of the model according to a forward logical order of the operators of the operator set and releasing intermediate results of the first operator; storing an intermediate result of a forward calculation of a second operator other than the first operator in the operator set in a video memory, and skipping a subsequent recalculation of the second operator based on the intermediate result of the second operator; A device for processing a model operator according to any one of claims 15 to 17.

25. The third determination module further comprises: performing a forward computation again according to a first logical order of forward computation of the operators of said operator set; performing a forward recalculation on the first operator based on a forward input of the first operator to obtain an intermediate result of the forward calculation of the first operator; When the forward calculation reaches the second operator, reading an intermediate result of the second operator from the video memory; performing a backward calculation of the model based on the intermediate results of the first operator and the intermediate results of the second operator; 25. A model operator processing device according to claim 24.

26. The third determination module further comprises: performing backward computations according to a second logical order of backward computations of the operators of the operator set; For the first operator, perform a backward calculation of the first operator based on an intermediate result of the first operator and a backward output of a previous operator during the backward calculation of the first operator, to obtain a backward output of the first operator and input it into a next operator; When the backward calculation reaches the second operator, read an intermediate result of the second operator from the video memory, determine a backward output of the second operator based on the intermediate result of the second operator and a backward output of a previous operator during the backward calculation, and input the backward output of the second operator to a next operator of the second operator.

26. A model operator processing device according to claim 25.

27. The apparatus further comprises: storing a forward input of a first operator in a first forward calculation of the model in the video memory; When performing forward calculation again on the model, the input of the first operator is read from the video memory, and the input is input to the first operator for recalculation.

26. A model operator processing device according to claim 25.

28. The third determination module further comprises: The operators in the operator set are sorted in descending order of the recalculation evaluation parameter, and the first N operators are selected and used as the first operator; or or The operators in the operator set are sorted in ascending order of the recalculation evaluation parameter, and the last N operators are selected and set as the first operator; The N is a natural number equal to or greater than 1.

20. A model operator processing device according to claim 18.

29. 1. An electronic device comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform a method according to any one of claims 1 to 3. Electronic devices.

30. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to perform the method according to any one of claims 1 to 3. A non-transitory computer-readable storage medium.

31. A computer program comprising: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is realized. Computer program.

Citation Information

Patent Citations

  • Method for optimizing hierarchical neural network and device therefor

    JP1993197821A

  • Optimization device, optimization method, and program

    JP2020135748A

  • Memory-efficient backpropagation through time

    US20190188572A1