Model operator processing method, device, electronic device and storage medium

By analyzing the set of operators in the deep learning model, operators with higher storage amount and calculation time are determined as priority objects for recomputation, the problem of low recomputation efficiency in the prior art is solved, and the memory usage efficiency and computing performance are improved.

CN117455005BActive Publication Date: 2025-06-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311345985.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-17
Publication Date
2025-06-03
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

In the recomputation process in the field of deep learning, unified recomputation of operators leads to low execution efficiency.

Method used

By determining the set of operators for the model networking, the storage amount occupied by its output tensor and the calculation time during forward calculation are calculated for each operator, and the priority operator participating in the recalculation is then determined.

Benefits of technology

It realizes the removal of low-cost-performance operators during recalculation, and improves the efficiency of video memory usage and computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117455005B_ABST
    Figure CN117455005B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for processing model operators, relating to the technical field of deep learning. The specific implementation scheme is as follows: determine an operator set of a model network architecture, where the operator set includes multiple operators; for each operator in the operator set, determine the storage amount occupied by the output tensor of the operator and the computation time consumed during the forward computation of the operator; according to the storage amount and computation time of the operator, determine a first operator participating in recomputation in the model from the operator set. Thus, this solution realizes fine-grained control of operators by determining the operator set of the model network architecture and calculating the recomputation evaluation parameters of each operator. And according to the recomputation evaluation parameters, determine the first operator that needs to participate in recomputation. During the recomputation process, operators with low recomputation cost performance are eliminated to achieve high-efficiency video memory for performance, thereby improving the computation performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning technologies, and in particular, to a method, apparatus, electronic device, and storage medium for processing model operators. Background Art

[0002] In the field of deep learning, recomputation is one of the technical paths for training large models. In the existing recomputation process, operators are uniformly recomputed, resulting in low execution efficiency of recomputation. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for processing model operators.

[0004] According to one aspect of the present disclosure, there is provided a method for processing model operators, including: determining an operator set of a model network architecture, where the operator set includes multiple operators; for each operator in the operator set, determining the storage amount occupied by the output tensor of the operator and the computing time consumed during the forward computation of the operator; and determining, from the operator set, a first operator participating in recomputation in the model according to the storage amount and the computing time of the operator.

[0005] According to another aspect of the present disclosure, there is provided a device for processing model operators, including: a first determination module for determining an operator set of a model network architecture, where the operator set includes multiple operators; a second determination module for determining, for each operator in the operator set, the storage amount occupied by the output tensor of the operator and the computing time consumed during the forward computation of the operator; and a third determination module for determining, from the operator set, a first operator participating in recomputation in the model according to the storage amount and the computing time of the operator.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method for processing model operators according to the embodiment of the above aspect.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where computer programs / instructions are stored thereon, and the computer instructions are used to cause a computer to execute the method for processing model operators according to the embodiment of the above aspect.

[0008] According to another aspect of the present disclosure, there is provided a computer program product including computer programs / instructions which, when executed by a processor, implement the method for processing model operators described in the above-mentioned embodiment of one aspect.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic flowchart of a method for processing model operators provided by an embodiment of the present disclosure;

[0012] Figure 2 is a schematic flowchart of another method for processing model operators provided by an embodiment of the present disclosure;

[0013] Figure 3 is a schematic flowchart of another method for processing model operators provided by an embodiment of the present disclosure;

[0014] Figure 4 is a schematic flowchart of another method for processing model operators provided by an embodiment of the present disclosure;

[0015] Figure 5 is a schematic diagram of merging operators provided by an embodiment of the present disclosure;

[0016] Figure 6 is a schematic flowchart of another method for processing model operators provided by an embodiment of the present disclosure;

[0017] Figure 7 is a schematic diagram of forward calculation of operators provided by an embodiment of the present disclosure;

[0018] Figure 8 is a schematic diagram of recomputation of operators provided by an embodiment of the present disclosure;

[0019] Figure 9 is a schematic diagram of networking of a hybrid parallel model provided by an embodiment of the present disclosure;

[0020] Figure 10 is a schematic flowchart of another method for processing model operators provided by an embodiment of the present disclosure;

[0021] Figure 11 is a schematic structural diagram of a device for processing model operators provided by an embodiment of the present disclosure;

[0022] Figure 12 A block diagram of an electronic device for implementing the processing method of the model operator according to an embodiment of the present disclosure. Detailed implementation manners

[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0024] The following describes a method, an apparatus, and an electronic device for processing a model operator according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0025] Artificial Intelligence (AI) is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level technologies and software-level technologies. Artificial Intelligence hardware technologies generally include several aspects such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0026] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Natural language processing is mainly applied in aspects such as machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, and speech recognition.

[0027] Deep Learning (DL) is a new research direction in the field of Machine Learning (ML). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies.

[0028] Machine translation, also known as automatic translation, is the process of using a computer to convert one natural language (source language) into another natural language (target language). It is a branch of computational linguistics and one of the goals of artificial intelligence.

[0029] Figure 1 It is a schematic flowchart of a method for processing model operators provided by an embodiment of the present disclosure.

[0030] As Figure 1 shown, the method for processing the model operator may include:

[0031] S101, determining an operator set for model networking, where the operator set includes multiple operators.

[0032] It should be noted that the execution subject of the method for processing model operators in the embodiments of the present disclosure may be a hardware device with data processing capabilities and / or the necessary software for driving the hardware device to work. Optionally, the execution subject may include a server, a user terminal, and other intelligent devices. Optionally, the user terminal includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, etc. Optionally, the server includes, but is not limited to, a web server, an application server, and may also be a server of a distributed system, or a server combined with a blockchain, etc. The embodiments of the present disclosure do not make specific limitations.

[0033] In some implementations, the operator set for model networking may be determined based on information such as the application scenario and structure of the model networking. Different application scenarios and structures correspond to different operator sets. Optionally, model networking is applicable to scenarios such as image processing, object detection, speech recognition, and text generation. Optionally, the structure of the model networking may be a hybrid parallel + Transformer structure.

[0034] Optionally, the hybrid parallel may include, but is not limited to, data parallel and Transformer, model parallel and Transformer, pipeline parallel and Transformer, and model and pipeline parallel.

[0035] In some implementations, the operator set for model networking may also be determined in combination with the performance of the operators. For example, if a convolution operator is used for image processing, then the operator set for model networking in the image processing scenario includes multiple convolution operators.

[0036] Exemplarily, taking Transformer networking as an example, its operator set includes operators such as FA, Matmul, ReduceScatter, and Allgather.

[0037] S102. For each operator in the operator set, determine the storage occupied by the output tensor of the operator and the computing time consumed during the forward computation of the operator.

[0038] It can be understood that in a deep learning framework, the inputs and outputs of an operator are usually represented in the form of tensors. The input and output tensors of an operator describe the shape and content of the data. Among them, a tensor is a multi-dimensional array, which can be regarded as an extension of scalars, vectors, and matrices, and can be understood as a high-dimensional matrix.

[0039] In some implementations, based on the parameters of the operator input, the storage occupied by the output tensor of the operator can be determined. Optionally, based on parameters such as batch size (batch_size, B), sequence length (sequence_len, S), dimension of the hidden state (hidden_size, H), model parallelism mp_degree (M), number of heads (num_head, A), dimension of the hidden layer of the feed-forward neural network (ffn_hidden_size, H'), etc., the storage occupied by the output tensor of the operator can be calculated.

[0040] In some implementations, the computing time consumed during the forward computation of the operator can be determined by analyzing the process of the operator's forward computation. The process of the operator's forward computation can also be monitored to obtain the time from the start to the end of the operator's forward computation as the computing time consumed during the operator's forward computation.

[0041] S103. According to the storage and computing time of the operator, determine the first operator in the operator set that participates in recomputation in the model.

[0042] It can be understood that the storage and computing time of an operator can reflect the cost performance of the operator participating in recomputation. If the storage of the operator is larger and the computing time is smaller, it means that the cost performance of the operator participating in recomputation is higher; if the storage of the operator is smaller and the computing time is larger, it means that the cost performance of the operator participating in recomputation is lower.

[0043] In some implementations, the operators in the operator set can be screened for recomputation to improve the video memory usage efficiency of the operators and enhance the computing performance. Optionally, by calculating the ratio of the storage size of the operator's output tensor to the computing time of the operator's forward computation, and based on this ratio, determine the first operator in the operator set that participates in recomputation in the model.

[0044] Optionally, based on the ratio of the storage and computing time, determine the operators with high cost performance for participating in recomputation from the operator set as the first operator.

[0045] According to the method for processing model operators provided by the embodiments of the present disclosure, by determining the operator set of the model network architecture, and for each operator in the operator set, obtaining the storage amount occupied by the output tensor of the operator and the computing time consumed during the forward calculation of the operator, and then determining the first operator that needs to participate in recomputation. It can be realized that during the recomputation process, operators with low recomputation cost performance are excluded, so as to achieve high-efficiency video memory for performance, enabling more computing performance to be exchanged per unit of video memory. By efficiently using the video memory, the computing performance of the model is further improved. For each type of operator in the model, refined control of the operator can be realized.

[0046] Figure 2 FIG. is a schematic flowchart of a method for processing model operators provided by the embodiments of the present disclosure.

[0047] As Figure 2 shown, the method for processing model operators may include:

[0048] S201, determine the operator set of the model network architecture, where the operator set includes multiple operators.

[0049] For the relevant content of step S201, reference may be made to the above embodiments, and details are not described herein again.

[0050] S202, for each operator in the operator set, determine the recomputation evaluation parameter of the operator.

[0051] It can be understood that the recomputation evaluation parameter of the operator can represent the size of the video memory saved by the operator per unit time during recomputation, and is used to describe the efficiency of the operator's video memory for performance. The larger the recomputation evaluation parameter of the operator, the more the operator needs to be recomputed.

[0052] In some implementations, based on the recomputation evaluation parameter of the operator, filtering the operators in the operator set for recomputation can improve the video memory usage efficiency of the operator and enhance the computing performance. Optionally, the recomputation evaluation parameter of the operator can be determined by analyzing the computing and video memory efficiency of each operator.

[0053] In some implementations, the ratio of the storage size of the operator output tensor to the computing time of the forward calculation of the operator can be calculated, and the ratio can be determined as the recomputation evaluation parameter.

[0054] S203, according to the recomputation evaluation parameter, determine the first operator in the model that participates in recomputation from the operator set.

[0055] In some implementations, for the operator set, each operator in the operator set can be sorted according to the recomputation evaluation parameter, and the first operator in the model that participates in recomputation can be determined according to the sorting result. The first operator can also be determined by comparing the recomputation evaluation parameter with a set threshold based on the set threshold.

[0056] Optionally, the recomputation evaluation parameters can be sorted in descending order, and then the operator ranked higher in the sorting is determined as the first operator. For example, each operator in the operator set can be sorted in descending order according to the recomputation evaluation parameters, and the top 3 operators in the sorting result can be determined as the first operator.

[0057] Optionally, the recomputation evaluation parameters can be sorted in ascending order, and then the operator ranked lower in the sorting is determined as the first operator. For example, each operator in the operator set can be sorted in ascending order according to the recomputation evaluation parameters, and the last 3 operators in the sorting result can be determined as the first operator.

[0058] Optionally, by comparing the recomputation evaluation parameters with a set threshold, the first operator participating in recomputation in the model is determined from the operator set, where the first operator is the operator whose recomputation evaluation parameter is greater than or equal to the set threshold.

[0059] According to the method for processing model operators provided by the embodiments of the present disclosure, by determining the operator set of the model network architecture and calculating the recomputation evaluation parameter for each operator in the operator set, and then determining the first operator that needs to participate in recomputation according to the recomputation evaluation parameter. It is possible to eliminate operators with low recomputation cost performance during the recomputation process, thereby achieving high-efficiency video memory for performance, enabling more computing performance to be exchanged per unit of video memory. By efficiently using the video memory, the computing performance of the model is further improved. For each type of operator in the model, fine-grained control of the operator can be achieved.

[0060] Figure 3 It is a schematic flowchart of a method for processing model operators provided by the embodiments of the present disclosure.

[0061] As Figure 3 shown, the method for processing model operators may include:

[0062] S301, determine the operator set of the model network architecture, and the operator set includes multiple operators.

[0063] For the relevant content of step S301, reference can be made to the above embodiments, which will not be elaborated here.

[0064] S302, determine the storage amount occupied by the output tensor of the operator.

[0065] It can be understood that in a deep learning framework, the input and output of an operator are usually represented in the form of tensors (Tensor), and the input and output tensors of the operator describe the shape and content of the data. Among them, a tensor is a multi-dimensional array, which can be regarded as an extension of a scalar, a vector, and a matrix, and can be understood as a high-dimensional matrix.

[0066] In some implementations, the storage occupied by the output tensor of an operator can be determined based on the parameters input to the operator. Optionally, the storage occupied by the output tensor of the operator can be calculated based on parameters such as batch size (batch_size, B), sequence length (sequence_len, S), dimension of the hidden state (hidden_size, H), model parallelism degree mp_degree (M), number of heads (num_head, A), dimension of the hidden layer of the feed-forward neural network (ffn_hidden_size, H'), etc.

[0067] S303. Determine the computation time consumed during the forward computation of the operator.

[0068] In some implementations, the computation time consumed during the forward computation of the operator can be determined by analyzing the process of the operator's forward computation. The process of the operator's forward computation can also be monitored to obtain the time from the start to the end of the operator's forward computation as the computation time consumed during the operator's forward computation.

[0069] S304. Determine the recomputation evaluation parameter of the operator based on the storage and the computation time.

[0070] In some implementations, the ratio of the storage occupied by the output tensor of the operator to the computation time consumed during the forward computation is obtained, and the ratio is determined as the recomputation evaluation parameter of the operator. Optionally, the formula for calculating the recomputation evaluation parameter is as follows:

[0071]

[0072] where η represents the recomputation evaluation parameter of the operator, Mem represents the storage occupied by the output tensor of the operator, and Time represents the computation time consumed during the forward computation of the operator.

[0073] In some implementations, the recomputation evaluation parameter represents the size of the video memory saved by recomputation per unit time. That is to say, the larger the recomputation evaluation parameter of the operator, the greater the possibility that the operator participates in recomputation.

[0074] Exemplarily, for operators such as Layernorm, the computation is fast and the video memory is large. That is to say, the computation time is small and the storage is large, so the corresponding recomputation evaluation parameter is large, and then Layernorm needs to be recomputed. For the MatMul operator, due to the heavy computation and long computation time, the recomputation efficiency is low, so the possibility of MatMul participating in recomputation is small.

[0075] In some implementations, the storage occupied by the output tensor of the operator and the computation time consumed during the forward computation can be analyzed and calculated offline. And the storage and the computation time are analyzed and calculated offline to obtain the recomputation evaluation parameter of the operator.

[0076] S305. Determine a first operator involved in recomputation in the model from the set of operators according to the recomputation evaluation parameter.

[0077] In some implementations, the operators in the set of operators can be sorted according to the recomputation evaluation parameter, and the first operator can be selected from the set of operators based on the sorting result. It can be realized that during the recomputation process, operators with high recomputation cost performance are used, so as to achieve high-efficiency video memory for performance.

[0078] Optionally, the operators in the set of operators are sorted in descending order according to the recomputation evaluation parameter, and the first N operators are selected as the first operator. Optionally, the operators in the set of operators are sorted in ascending order according to the recomputation evaluation parameter, and the last N operators are selected as the first operator. Wherein, N is a natural number greater than or equal to 1.

[0079] For example, when N is equal to 5, the first 5 operators sorted from large to small can be selected as the first operator. The last 5 operators sorted from small to large can also be selected as the first operator.

[0080] In some implementations, the first operator can also be selected from the set of operators by setting a set threshold of the recomputation evaluation parameter and based on the set threshold. It can be realized that during the recomputation process, operators with high recomputation cost performance are used, so as to achieve high-efficiency video memory for performance.

[0081] Optionally, the recomputation evaluation parameter of the operators in the set of operators is compared with the set threshold. The operators whose recomputation evaluation parameter is greater than or equal to the set threshold are selected as the first operator.

[0082] According to the method for processing model operators provided by the embodiments of the present disclosure, by determining the set of operators for model networking, and for each operator in the set of operators, determining the storage amount occupied by the output tensor of the operator and the computing time consumed during the forward computation of the operator, and then calculating the recomputation evaluation parameter. And according to the recomputation evaluation parameter, determine the first operator that needs to participate in the recomputation. It can be realized that during the recomputation process, operators with low recomputation cost performance are eliminated, so as to achieve high-efficiency video memory for performance, so that more computing performance can be exchanged per unit of video memory. By efficiently using the video memory, the computing performance of the model is further improved. For each type of operator in the model, refined control of the operator can be realized.

[0083] Figure 4 It is a schematic flowchart of a method for processing model operators provided by the embodiments of the present disclosure.

[0084] As Figure 4 shown, the method for processing model operators may include:

[0085] S401. Determine the candidate operators required for model networking.

[0086] S402. Determine the target class operators from the candidate operators, and perform merging processing on the target class operators to obtain an operator set.

[0087] In some implementations, the candidate operators required during determination can be determined based on information such as the application scenario and structure of model networking, combined with the performance of the operators.

[0088] It can be understood that if a type of operator does not require the forward input and output during reverse calculation, this type of operator needs special processing. This operator can be merged into the previous operator, and the merged operator is regarded as one operator, thereby improving the performance of model calculation.

[0089] In some implementations, the candidate operators that do not require forward input and forward output during the reverse calculation process can be determined as the target class operators, and the target class operators can be determined from the candidate operators, and then the merging process can be performed on the target class operators to obtain an operator set.

[0090] Optionally, for any candidate operator belonging to the target class operators, the previous candidate operator adjacent to any candidate operator can be determined, and any candidate operator is combined with the previous candidate operator to obtain a merged operator. Further, based on the merged operator and the remaining candidate operators, an operator set is obtained. For example, the candidate operators include operator A, operator B, operator C, operator D, and operator E. Among them, operator C is merged into operator B to obtain a merged operator B'. Then, operator A, operator D, and operator E are the remaining candidate operators, which form an operator set with the merged operator B'.

[0091] Exemplarily, if the model networking is a hybrid parallel + Transformer structure, there is a type of communication operator ReduceScatter communication, whose reverse calculation does not require forward input and output, and it is preceded by RowLN matrix multiplication. Therefore, RowLN + ReduceScatter is merged to obtain a merged operator RowLN + ReduceScatter, that is, the two operators RowLN_0 and ReduceScatter are regarded as a whole merged operator.

[0092] Such as Figure 5 The schematic diagram of operator merging shown. Merge RowLN_0 and ReduceScatter to obtain a merged operator. Figure 5The operator within the dashed box is a merging operator. Among them, BSH represents the parameter of the ReduceScatter operator, and BSH / M represents the output size of the ReduceScatter operator. Among them, B represents batch_size, S represents sequence_len, H represents hidden_size, and M represents mp_degree.

[0093] S403. For each operator in the operator set, determine the recomputation evaluation parameter of the operator.

[0094] In some implementations, since the merging operator consists of multiple operators, the recomputation evaluation parameter of the merging operator can be determined based on the storage amount and forward computation time of the output tensor of each candidate operator included in the merging operator.

[0095] Optionally, by obtaining the storage amount of the output tensor of each candidate operator and the computation time consumed during forward computation, and based on the above formula (1), calculate the recomputation evaluation parameter of the merging operator.

[0096] S404. According to the recomputation evaluation parameter, determine the first operator participating in recomputation in the model from the operator set.

[0097] For the relevant content of step S404, reference can be made to the above embodiments and will not be elaborated here.

[0098] According to the method for processing model operators provided by the embodiments of the present disclosure, by obtaining candidate operators required for model networking, determining target type operators from the candidate operators, and performing merging processing on the target type operators to obtain an operator set. For each operator in the operator set, calculate its recomputation evaluation parameter, and then determine the first operator that needs to participate in recomputation according to the recomputation evaluation parameter. It can be realized that during the recomputation process, operators with low recomputation cost performance are eliminated, thereby realizing high-efficiency video memory for performance, enabling more computing performance to be exchanged per unit of video memory. By efficiently using video memory, the computing performance of the model is further improved. For each type of operator in the model, fine-grained control of the operator can be realized.

[0099] Figure 6 It is a schematic flowchart of a method for processing model operators provided by the embodiments of the present disclosure.

[0100] As Figure 6 shown, the method for processing model operators may include:

[0101] S601. Determine the operator set for model networking, where the operator set includes multiple operators.

[0102] S602. For each operator in the operator set, determine the recomputation evaluation parameter of the operator.

[0103] S603. Determine a first operator involved in recomputation in the model from the set of operators according to the recomputation evaluation parameter.

[0104] For the relevant content of steps S601 - S603, reference may be made to the above embodiments and will not be elaborated here.

[0105] S604. Perform forward computation of the model in the forward logical order of the operators in the set of operators, and release the intermediate results of the first operator.

[0106] In some implementations, after determining the first operator, recomputation can be performed on the operators in the model network. Recomputation includes forward computation, forward recomputation, and backward computation. During the forward computation process, since the recomputation evaluation parameter of the first operator is high, indicating that the computation time of the first operator is small, the time consumed during forward recomputation is small, and the intermediate results of the first operator can be released, thereby reducing the occupancy of video memory.

[0107] As Figure 7 shown in the schematic diagram of forward computation of operators. For the set of operators in one layer (Layer) of the model, there are 5 operators: OP0, OP1, OP2, OP3, and OP4. Through the recomputation evaluation parameter of the operators, OP0, OP1, OP3, and OP4 can be determined as the first operators, and OP2 as the second operator. Perform forward computation on the operators in this layer of the model in the logical order, that is, in the order of OP0, OP1, OP2, OP3, and OP4, calculate the intermediate results of the operators, and release the intermediate results of OP0, OP1, OP3, and OP4.

[0108] S605. Store the intermediate results of the forward computation of the second operator other than the first operator in the set of operators in video memory, and skip the subsequent recomputation of the second operator based on the intermediate results of the second operator.

[0109] In some implementations, the first operator refers to the operator in the set of operators that needs to be recomputed, and the second operator refers to the operator in the set of operators that does not need to be recomputed. During the first forward computation, since it is determined based on the recomputation evaluation parameter of the operator that the second operator does not need to be recomputed in the subsequent process, it is necessary to store the intermediate results of the forward computation of the second operator in video memory during the first forward computation, and then skip the subsequent recomputation of the second operator based on the intermediate results of the second operator. For example, Figure 7 for the operator shown in , OP2 is the second operator, and the intermediate result of OP2 obtained through forward computation is C1, and C1 is saved in video memory.

[0110] In some implementations, after storing the intermediate result of the second operator in the video memory, the forward calculation can be executed again according to the first logical order of the forward calculations of the operators in the operator set. Here, the first logical order is also the forward logical order.

[0111] Optionally, based on the forward input of the first operator, perform a forward recomputation on the first operator to obtain the intermediate result of the forward calculation of the first operator. When the forward calculation reaches the second operator, read the intermediate result of the second operator from the video memory, so that in the forward recomputation, the forward calculation of the second operator is reduced, and the computing performance is traded for increased video memory, thereby reducing the overall time consumption of the recomputation.

[0112] Furthermore, based on the intermediate result of the first operator and the intermediate result of the second operator, perform the backward calculation of the model. Optionally, the backward calculation can be performed according to the second logical order of the backward calculations of the operators in the operator set.

[0113] Optionally, for the first operator, based on the intermediate result of the first operator and the backward output of the previous operator during the backward calculation of the first operator, perform a backward calculation on the first operator to obtain the backward output of the first operator and input it into the next operator. When the backward calculation reaches the second operator, read the intermediate result of the second operator from the video memory, determine the backward output of the second operator based on the intermediate result of the second operator and the backward output of the previous operator during the backward calculation, and input the backward output of the second operator into the next operator of the second operator.

[0114] That is to say, when performing the backward calculation on the first operator, input the intermediate result of the first operator and the backward output of the previous backward calculation of the first operator into the first operator for backward calculation. As Figure 8 shown in the schematic diagram of recomputing the operator, taking OP1 as an example, when performing the backward calculation, the input of OP1 includes the intermediate result of the forward calculation of OP1 and the backward output of OP2. During the forward recomputation process, by reading C1 stored in the video memory, the recomputation of OP2 can be skipped, as shown by the dotted part in the figure.

[0115] In some implementations, it is also possible to store the forward input of the first operator during the first forward calculation of the model in the video memory. To achieve reading the input of the first operator from the video memory when performing the forward calculation on the model again and inputting it into the first operator for recomputation. Taking Figure 7 OP0 in as an example, OP0 is the first operator, C0 is the forward input of OP0, and OP0 can perform a forward calculation based on C0 to obtain the intermediate result of OP0.

[0116] According to the method for processing model operators provided by the embodiments of the present disclosure, by determining the operator set of the model network architecture, and for each operator in the operator set, calculating its recomputation evaluation parameter, and then according to the recomputation evaluation parameter, determining the first operator that needs to participate in recomputation. It can be realized that during the recomputation process, operators with low recomputation cost performance are eliminated, so as to achieve high-efficiency video memory for performance, enabling more computing performance to be exchanged per unit of video memory. By efficiently using the video memory, the computing performance of the model is further improved. For each type of operator in the model, fine-grained control of the operator can be realized. During the recomputation process, by storing the intermediate result of the second operator in the video memory, the forward calculation of the second operator can be reduced, and computing performance is exchanged by increasing the video memory.

[0117] For example, taking the Transformer network architecture as an example, the process of recomputation is described. For Transformer, the related operators mainly include operators such as FA, Matmul, ReduceScatter, and Allgather. After the merge operator, according to the above formula (1), the recomputation evaluation parameter is calculated, and the recomputation evaluation parameter sorting table is obtained as shown in Table 1.

[0118] Table 1

[0119]

[0120] As can be seen from Table 1, there are 9 operators in the operator set, among which the operators numbered 1, 3, and 6 are merge operators. By calculating the recomputation evaluation parameter, sorting them in ascending order according to the recomputation evaluation parameter, and selecting the last 4 operators as the first operators. Then the first 5 operators are the second operators. Among them, the time ratio represents the computing time consumed by calculating any operator, and its proportion in the total computing time of the entire operator set. For example, the computing time of the FA operator accounts for 13.8% of the total computing time of the entire operator set. Among them, the calculation formula represents the size of the operator output, and its value is the size of the storage volume. In the formula, B represents batch_size, S represents sequence_len, H represents hidden_size, M represents mp_degree, A represents num_head, and H' represents ffn_hidden_size.

[0121] As Figure 9 shown is the network architecture diagram of the hybrid parallel model, where the hybrid parallel model includes a multi-head attention (MultiHeadAttention) network and a multi-layer perception (Multilayer Perception, MLP) network.

[0122] Among them, the MultiHeadAttention network includes: LayerNom operator, ALLgather operator, ColumnLN_0 operator, FA operator, RowLN_0 operator, and ReduceScatter operator. The MLP network includes: LayerNom operator, ALLgather operator, ColumnLN_1 operator,

[0123] Silu+Ele operator, RowLN_1 operator, and ReduceScatter operator.

[0124] Among them, Figure 9 BSH / M, BSH, BSH′ / M, etc. of the input of each operator in represent the parameters of the operator input. B represents the batch size batch_size, S represents the sequence length sequence_len, H represents the dimension of the hidden state hidden_size, M represents the model parallelism mp_degree, A represents the number of heads num_head, and H' represents the dimension of the hidden layer of the feed-forward neural network ffn_hidden_size. The FA operator has two outputs, and LSE represents the name of the output of the FA operator, with a size of BAS / M.

[0125] As Figure 9 shown, the operators within the dashed box are the second operators. Figure 9 The forward logical order of the operators is also shown in.

[0126] During the recomputation process, forward computation is performed according to the forward logical order of the operators in, and the intermediate results of the first operator are released, that is, the intermediate results of the last 4 operators in Table 1 are released. At the same time, the intermediate results of the forward computation of the second operator are stored in the video memory. Then, according to the forward logical order of the operators in, based on the forward input of the first operator, forward recomputation is performed on the first operator to obtain the intermediate results of the forward computation of the first operator. When the forward recomputation reaches the second operator, the intermediate results of the second operator are read from the video memory, and the backward computation of the model is performed based on the intermediate results of the first operator and the intermediate results of the second operator. Figure 9 Figure 9 Taking the Allgather operator as an example, the backward computation process of the first operator performs backward computation on Allgather based on the intermediate results of Allgather and the backward output of ColumnLN_0, obtains the backward output of Allgather, and inputs it into LayerNorm.

[0127]

[0128] Taking ColumnLN_1 as an example, the reverse calculation process of the second operator reads the intermediate result of ColumnLN_1 from the video memory, determines the reverse output of ColumnLN_1 based on the intermediate result of ColumnLN_1 and the reverse output of Silu+Ele, and inputs the reverse output of ColumnLN_1 into Allgather.

[0129] Figure 10 It is a schematic flowchart of a method for processing a model operator provided by an embodiment of the present disclosure.

[0130] As Figure 10 shown, the method for processing the model operator may include:

[0131] S1001, determine candidate operators required when the model is networked.

[0132] S1002, determine target type operators from the candidate operators, and perform merging processing on the target type operators to obtain an operator set.

[0133] S1003, for each operator in the operator set, determine the storage amount occupied by the output tensor of the operator.

[0134] S1004, determine the computing time consumed during the forward calculation of the operator.

[0135] S1005, determine the recomputation evaluation parameter of the operator according to the storage amount and the computing time.

[0136] S1006, according to the recomputation evaluation parameter, determine the first operator participating in recomputation in the model from the operator set.

[0137] S1007, perform the forward calculation of the model in the forward logical order of the operators in the operator set, and release the intermediate result of the first operator.

[0138] S1008, store the intermediate result of the forward calculation of the second operator other than the first operator in the operator set in the video memory, and perform subsequent recomputation of the second operator based on the intermediate result of the second operator.

[0139] According to the method for processing model operators provided by an embodiment of the present disclosure, by determining the operator set of model networking, and for each operator in the operator set, calculating its recomputation evaluation parameter, and then according to the recomputation evaluation parameter, determining the first operator that needs to participate in recomputation. It can be realized that during the recomputation process, operators with low recomputation cost performance are excluded, thereby realizing high-efficiency video memory for performance, enabling more computing performance to be exchanged per unit of video memory. By efficiently using video memory, the computing performance of the model is further improved. For each type of operator in the model, fine-grained control of the operator can be realized. During the recomputation process, by storing the intermediate result of the second operator in the video memory, the forward calculation of the second operator can be reduced, and computing performance is exchanged by increasing the video memory.

[0140] Corresponding to the method for processing model operators provided by the above several embodiments, an embodiment of the present disclosure further provides a device for processing model operators. Since the device for processing model operators provided by the embodiment of the present disclosure corresponds to the method for processing model operators provided by the above several embodiments, the implementation manners of the above method for processing model operators are also applicable to the device for processing model operators provided by the embodiment of the present disclosure, and will not be described in detail in the following embodiments.

[0141] Figure 11 It is a schematic structural diagram of a device for processing model operators provided by an embodiment of the present disclosure.

[0142] As Figure 11 shown, the device 1100 for processing model operators according to an embodiment of the present disclosure includes a first determination module 1101, a second determination module 1102, and a third determination module 1103.

[0143] The first determination module 1101 is configured to determine the operator set of model networking, and the operator set includes multiple operators.

[0144] The second determination module 1102 is configured to, for each operator in the operator set, determine the storage amount occupied by the output tensor of the operator and the computing time consumed during the forward calculation of the operator.

[0145] The third determination module 1103 is configured to, according to the storage amount and the computing time of the operator, determine the first operator in the model that participates in recomputation from the operator set.

[0146] In an embodiment of the present disclosure, the third determination module 1103 is further configured to: determine the recomputation evaluation parameter of the operator according to the storage amount and the computing time; determine the first operator in the model that participates in recomputation from the operator set according to the recomputation evaluation parameter of the operator.

[0147] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: obtain a ratio of the storage amount to the calculation time, and determine the ratio as the recomputation evaluation parameter.

[0148] In one embodiment of the present disclosure, the greater the recomputation evaluation parameter of the operator, the greater the possibility that the operator participates in recomputation.

[0149] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: sort the operators in the operator set according to the recomputation evaluation parameter, and select the first operator from the operator set based on the sorting result.

[0150] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: compare the recomputation evaluation parameter of the operator in the operator set with a set threshold; select the operator whose recomputation evaluation parameter is greater than or equal to the set threshold as the first operator.

[0151] In one embodiment of the present disclosure, the first determination module 1101 is further configured to: determine candidate operators required for the model networking; determine target class operators from the candidate operators, and perform a merging process on the target class operators to obtain the operator set.

[0152] In one embodiment of the present disclosure, the first determination module 1101 is further configured to: for any candidate operator belonging to the target class operator, determine a previous candidate operator adjacent to the any candidate operator, combine the any candidate operator with the previous candidate operator to obtain a merged operator; obtain the operator set based on the merged operator and the remaining candidate operators.

[0153] In one embodiment of the present disclosure, the first determination module 1101 is further configured to: determine candidate operators that do not require forward input and forward output during the reverse calculation process as the target class operators.

[0154] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: determine the recomputation evaluation parameter of the merged operator based on the storage amount and the calculation time of the output tensor of each candidate operator included in the merged operator.

[0155] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: perform forward calculation of the model in the forward logical order of the operators in the operator set, release the intermediate result of the first operator; store the intermediate result of the forward calculation of the second operator other than the first operator in the operator set in the video memory, and skip subsequent recomputation of the second operator based on the intermediate result of the second operator.

[0156] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: perform forward calculation again in accordance with the first logical order of forward calculation of the operators in the operator set; perform forward re-calculation on the first operator based on the forward input of the first operator to obtain an intermediate result of the forward calculation of the first operator; when the forward calculation reaches the second operator, read the intermediate result of the second operator from the video memory; and perform backward calculation of the model based on the intermediate result of the first operator and the intermediate result of the second operator.

[0157] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: perform backward calculation in accordance with the second logical order of backward calculation of the operators in the operator set; for the first operator, perform backward calculation on the first operator based on the intermediate result of the first operator and the backward output of the previous operator during the backward calculation of the first operator to obtain the backward output of the first operator and input it into the next operator; when the backward calculation reaches the second operator, read the intermediate result of the second operator from the video memory, determine the backward output of the second operator based on the intermediate result of the second operator and the backward output of the previous operator during the backward calculation, and input the backward output of the second operator into the next operator of the second operator.

[0158] In one embodiment of the present disclosure, the apparatus further includes: storing the forward input of the first operator during the first forward calculation of the model in the video memory; and when performing forward calculation on the model again, reading the input of the first operator from the video memory and inputting it into the first operator for re-calculation.

[0159] In one embodiment of the present disclosure, the third determination module 1103 is further configured to: sort the operators in the operator set in descending order according to the re-calculation evaluation parameter and select the first N operators as the first operator; or sort the operators in the operator set in ascending order according to the re-calculation evaluation parameter and select the last N operators as the first operator; where N is a natural number greater than or equal to 1.

[0160] According to the processing device of the model operator provided by the embodiments of the present disclosure, by determining the operator set of the model network architecture, and for each operator in the operator set, calculating its recomputation evaluation parameter, and then according to the recomputation evaluation parameter, determining the first operator that needs to participate in recomputation. It can achieve that in the recomputation process, operators with low recomputation cost performance are excluded, so as to realize high-efficiency video memory for performance, enabling more computing performance to be exchanged per unit of video memory. By efficiently using the video memory, the computing performance of the model is further improved. For each type of operator in the model, fine-grained control of the operator can be achieved. During the recomputation process, by storing the intermediate result of the second operator in the video memory, the forward calculation of the second operator can be reduced, and computing performance is exchanged by increasing the video memory.

[0161] In the technical solution of the present disclosure, the acquisition, storage, application, etc. of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0162] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0163] Figure 12 FIG. shows a schematic block diagram of an exemplary electronic device 1200 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0164] As Figure 12 shown, the device 1200 includes a computing unit 1201, which can execute various appropriate actions and processes according to the computer programs / instructions stored in the read-only memory (ROM) 1202 or the computer programs / instructions loaded from the storage unit 1206 into the random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.

[0165] Multiple components in device 1200 are connected to I / O interface 1205, including: input unit 1206 such as a keyboard, mouse, etc.; output unit 1207, such as various types of displays, speakers, etc.; storage unit 1208, such as a disk, optical disc, etc.; and communication unit 1209, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0166] Computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 1201 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1201 executes the various methods and processes described above, such as the processing method of model operators. For example, in some embodiments, the processing method of model operators can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1206. In some of these embodiments, part or all of the computer program / instructions can be loaded and / or installed onto device 1200 via ROM 1202 and / or communication unit 1209. When the computer program / instructions are loaded into RAM 1203 and executed by computing unit 1201, one or more steps of the processing method of model operators described above can be executed. Alternatively, in other embodiments, computing unit 1201 can be configured to execute the processing method of model operators in any other suitable manner (e.g., by means of firmware).

[0167] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs / instructions, the one or more computer programs / instructions can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0168] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0169] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0170] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0171] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.

[0172] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs / instructions running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.

[0173] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. There is no limitation herein.

[0174] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for processing model operators, wherein, the method includes: Determine the operator set of the model network architecture, where the operator set includes multiple operators; For each operator in the operator set, determine the storage amount occupied by the output tensor of the operator and the computing time consumed during the forward calculation of the operator; Obtain the ratio of the storage amount to the computing time, and determine the ratio as the recomputation evaluation parameter of the operator. The greater the recomputation evaluation parameter of the operator, the greater the possibility that the operator participates in recomputation; According to the recomputation evaluation parameter of the operator, determine the first operator in the model that participates in recomputation from the operator set, where the first operator is the operator in the operator set that needs to be recomputed.

2. The method according to claim 1, wherein, the step of determining the first operator in the model that participates in recomputation from the operator set according to the recomputation evaluation parameter of the operator includes: Sort the operators in the operator set according to the recomputation evaluation parameter, and select the first operator from the operator set based on the sorting result.

3. The method according to claim 1, wherein, the step of determining the first operator in the model that participates in recomputation from the operator set according to the recomputation evaluation parameter of the operator includes: Compare the recomputation evaluation parameter of the operator in the operator set with a set threshold; Select the operator whose recomputation evaluation parameter is greater than or equal to the set threshold as the first operator.

4. The method according to claim 1, wherein, the step of determining the operator set of the model network architecture includes: Determine the candidate operators required for the model network architecture; Determine the candidate operators that do not require forward input and forward output during the reverse calculation process as target class operators, and perform merging processing on the target class operators to obtain the operator set.

5. The method according to claim 4, wherein, the step of performing merging processing on the target class operators to obtain the operator set includes: For any candidate operator belonging to the target class operators, determine the previous candidate operator adjacent to the any candidate operator, and combine the any candidate operator with the previous candidate operator to obtain a merged operator; Based on the merged operator and the remaining candidate operators, obtain the operator set.

6. The method according to claim 5, wherein, the method further includes: Based on the storage amount of the output tensor and the computing time of each candidate operator included in the merged operator, determine the recomputation evaluation parameter of the merged operator.

7. The method according to claim 1, wherein, after determining the first operator in the model that participates in recomputation from the operator set, it further includes: Execute the forward calculation of the model in the forward logical order of the operators in the operator set, and release the intermediate result of the first operator; Store the intermediate results of the forward calculation of the second operators in the operator set except the first operator in the video memory, and skip the subsequent recomputation of the second operators based on the intermediate results of the second operators.

8. The method according to claim 7, wherein, skipping subsequent recomputation of the second operator based on the intermediate result of the second operator includes: performing forward computation again in accordance with the first logical order of forward computation of the operators in the operator set; performing forward recomputation on the first operator based on the forward input of the first operator to obtain the intermediate result of the forward computation of the first operator; when the forward computation reaches the second operator, reading the intermediate result of the second operator from the video memory; performing backward computation of the model based on the intermediate result of the first operator and the intermediate result of the second operator.

9. The method according to claim 8, wherein, performing backward computation of the model based on the intermediate result of the first operator and the intermediate result of the second operator includes: performing backward computation in accordance with the second logical order of backward computation of the operators in the operator set; for the first operator, performing backward computation on the first operator based on the intermediate result of the first operator and the backward output of the previous operator during the backward computation of the first operator, and inputting the backward output of the first operator into the next operator; when the backward computation reaches the second operator, reading the intermediate result of the second operator from the video memory, determining the backward output of the second operator based on the intermediate result of the second operator and the backward output of the previous operator during the backward computation, and inputting the backward output of the second operator into the next operator of the second operator.

10. The method according to claim 8, wherein, the method further includes: storing the forward input of the first operator during the first forward computation of the model in the video memory; when performing forward computation on the model again, reading the input of the first operator from the video memory and inputting it into the first operator for recomputation.

11. The method according to claim 2, wherein, sorting the operators in the operator set according to the recomputation evaluation parameter, and selecting the first operator from the operator set based on the sorting result includes: sorting the operators in the operator set in descending order according to the recomputation evaluation parameter, and selecting the first N operators as the first operator; or, sorting the operators in the operator set in ascending order according to the recomputation evaluation parameter, and selecting the last N operators as the first operator; wherein, N is a natural number greater than or equal to 1.

12. A processing device for model operators, wherein, the device includes: a first determination module for determining an operator set of a model network architecture, the operator set including a plurality of operators; a second determination module for determining, for each operator in the operator set, the storage amount occupied by the output tensor of the operator and the computation time consumed during the forward computation of the operator; a third determination module for determining, from the operator set, the first operator participating in recomputation in the model according to the storage amount and the computation time of the operator; the third determination module is further configured to: Obtain the ratio of the storage amount to the calculation time, and determine the ratio as the recalculation evaluation parameter of the operator. The larger the recalculation evaluation parameter of the operator, the greater the possibility that the operator participates in recalculation; According to the recalculation evaluation parameter of the operator, determine the first operator participating in recalculation in the model from the operator set, where the first operator is the operator that needs to be recalculated in the operator set.

13. The device according to claim 12, wherein, The third determination module is further configured to: Sort the operators in the operator set according to the recalculation evaluation parameter, and select the first operator from the operator set based on the sorting result.

14. The device according to claim 12, wherein, The third determination module is further configured to: Compare the recalculation evaluation parameter of the operator in the operator set with a set threshold; Select the operator whose recalculation evaluation parameter is greater than or equal to the set threshold as the first operator.

15. The device according to claim 12, wherein, The first determination module is further configured to: Determine the candidate operators required for the model networking; Determine the candidate operators that do not require forward input and forward output during the reverse calculation process as the target type operators, and perform merging processing on the target type operators to obtain the operator set.

16. The device according to claim 15, wherein, The first determination module is further configured to: For any candidate operator belonging to the target type operator, determine the previous candidate operator adjacent to the any candidate operator, and combine the any candidate operator with the previous candidate operator to obtain a merged operator; Based on the merged operator and the remaining candidate operators, obtain the operator set.

17. The device according to claim 16, wherein, The third determination module is further configured to: Based on the storage amount and the calculation time of the output tensor of each candidate operator included in the merged operator, determine the recalculation evaluation parameter of the merged operator.

18. The device according to claim 12, wherein, The third determination module is further configured to: Perform the forward calculation of the model in the forward logical order of the operators in the operator set, and release the intermediate result of the first operator; Store the intermediate result of the forward calculation of the second operator other than the first operator in the operator set in the video memory, and skip the subsequent recalculation of the second operator based on the intermediate result of the second operator.

19. The device according to claim 18, wherein, The third determination module is further configured to: Perform the forward calculation again in the first logical order of the forward calculation of the operators in the operator set; Perform forward recalculation on the first operator based on the forward input of the first operator to obtain the intermediate result of the forward calculation of the first operator; When the forward calculation reaches the second operator, read the intermediate result of the second operator from the video memory; Perform the reverse calculation of the model based on the intermediate result of the first operator and the intermediate result of the second operator.

20. The apparatus according to claim 19, wherein, the third determination module is further configured to: perform reverse calculation according to a second logical order of reverse calculation of the operators in the operator set; for the first operator, perform reverse calculation on the first operator based on the intermediate result of the first operator and the reverse output of the previous operator during the reverse calculation of the first operator, and input the reverse output of the first operator into the next operator; when the reverse calculation reaches the second operator, read the intermediate result of the second operator from the video memory, determine the reverse output of the second operator based on the intermediate result of the second operator and the reverse output of the previous operator during the reverse calculation, and input the reverse output of the second operator into the next operator of the second operator.

21. The apparatus according to claim 19, wherein, the apparatus further comprises: storing the forward input of the first operator during the first forward calculation of the model in the video memory; when performing forward calculation on the model again, reading the input of the first operator from the video memory and inputting it into the first operator for re-calculation.

22. The apparatus according to claim 13, wherein, the third determination module is further configured to: sort the operators in the operator set in descending order according to the re-calculation evaluation parameter, and select the first N operators as the first operator; or, sort the operators in the operator set in ascending order according to the re-calculation evaluation parameter, and select the last N operators as the first operator; wherein, N is a natural number greater than or equal to 1.

23. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

25. A computer program product comprising computer programs / instructions, characterized in that, when the computer programs / instructions are executed by a processor, the method according to any one of claims 1-11 is implemented.

Citation Information

Patent Citations

  • Graph data weight calculation method and device and electronic equipment

    CN110688610A

  • Video uploading server deployment method and device

    CN115361379A