An acceleration method based on an acceleration library and a method for generating images
By analyzing the neural network model file in the acceleration library and using the output of the target upper-layer operator to obtain weight information, the problem of insufficient compatibility of the acceleration library for variable weight operators is solved, and efficient acceleration of the neural network is achieved and cost is reduced.
Patent Information
- Application Number
- CN202210199246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-03-02
AI Technical Summary
The existing acceleration library has limited compatibility when dealing with convolution and deconvolution operators of variable weights, and cannot complete the acceleration of neural networks normally, especially for the acceleration of networks such as style generative adversarial networks (stylegan2).
By analyzing the neural network model file and determining that the target type operator is missing weight information, the connection relationship is used to obtain the output of the target upper-layer operator as weight information, and the format conversion and fusion rule optimization are performed in the acceleration library to achieve acceleration of the target type operator.
Improves the compatibility of the acceleration library with variable weight operators, can normally complete the acceleration of neural networks, reduces the acceleration cost, and does not need to modify the original neural network model.
Smart Images

Figure CN114841330B_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and in particular, to an acceleration method based on an acceleration library and a method for generating images. Background Art
[0002] Neural networks are generally deployed in hardware to complete inference (inference: forward propagation) through the hardware. However, the hardware resources of the hardware (such as computing cores, the number of threads, etc.) are limited. During the deployment of neural networks, it is necessary to reasonably allocate hardware resources to each operator included in the neural network through software acceleration methods under limited hardware resources to complete the acceleration of the neural network.
[0003] In related technologies, acceleration libraries developed by some hardware companies are used to achieve the acceleration of neural networks. By using the acceleration library, in various hardware resource allocation scenarios, a hardware resource allocation scenario with good effects can be selected to complete the acceleration of the neural network.
[0004] However, for some types of operators included in neural networks (these types of operators are referred to as target type operators in this application), the compatibility of these acceleration libraries is limited. That is, only when the weight information required by the target type operator is fixed, these acceleration libraries can implement the operation of the target type operator and complete the acceleration. If the weight information is variable, these acceleration libraries cannot support the operation of the target type operator, that is, the acceleration of the neural network cannot be completed normally. Among them, the target type operator can be a basic operation operator such as a convolution operator, a deconvolution operator, a pooling operator, or a complex operation operator constructed based on the basic operation operator, etc.
[0005] Taking the neural network as the stylegan2 (the second-generation style-based generative adversarial network) network as an example. The target type operators included in the stylegan2 network are convolution operators and deconvolution operators with variable weights. The acceleration library is the TensorRT library developed by NVIDIA Corporation. The TensorRT library only supports the operation of convolution operators and deconvolution operators with fixed weights, and does not support the operation of convolution operators and deconvolution operators with variable weights. Therefore, the TensorRT library does not support the acceleration of the stylegan2 network. Summary of the Invention
[0006] In view of this, the present application discloses at least one acceleration method based on an acceleration library. The method may include: obtaining a model file of a neural network model to be accelerated; the neural network model includes target type operators that perform operations based on weight information; parsing the model file to obtain each operator included in the neural network model, the input information, output information corresponding to each operator, and the connection relationship between the operators; in the case where the input information corresponding to the target type operator does not include weight information, determining a target upper-layer operator connected to the target type operator according to the connection relationship; the target upper-layer operator is used to generate the weight information required by the target type operator; determining the weight information based on the output of the target upper-layer operator, so as to utilize the weight information to implement the target type operator during the process of accelerating the neural network model based on the acceleration library.
[0007] In some embodiments, after parsing to obtain each operator included in the neural network model, it further includes: performing format conversion on each operator according to the preset operator conversion method in the acceleration library to obtain the operators after format conversion.
[0008] In some embodiments, the input information includes the input size information of the input data; the method further includes: for each operator among the operators: determining the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator; and allocating a first storage space for storing the input data and a second storage space for storing the output data for the operator according to the input size information and the output size information.
[0009] In some embodiments, the determining the weight information based on the output of the target upper-layer operator connected to the target type operator includes: obtaining the output data stored in the second storage space corresponding to the target upper-layer operator; and determining the output data as the weight information.
[0010] In some embodiments, the determining the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator includes: in the case where the input information corresponding to the target type operator includes weight information, determining the output size information of the output data of the target type operator according to the input size information and the size information of the weight information included in the input information; in the case where the input information corresponding to the target type operator does not include weight information, determining the output size information of the output data of the target type operator according to the input size information and the size information of the output data output by the target upper-layer operator.
[0011] In some embodiments, the acceleration library includes preset fusion rules; accelerating the neural network model based on the acceleration library includes: fusing at least two operators indicated by the preset fusion rules according to the preset fusion rules to obtain a fused operator; determining the remaining operators except the at least two operators among the operators; and running the remaining operators and the fused operator to complete the acceleration of the neural network model.
[0012] In some embodiments, at least one implementation scheme corresponding to each of the operators; running the remaining operators and the fused operator to complete the acceleration of the neural network model includes: generating at least one implementation scheme corresponding to the fused operator according to at least one implementation scheme corresponding to the at least two operators; and running the remaining operators and the fused operator to determine a final implementation scheme corresponding to each of the remaining operators and the fused operator among at least one implementation scheme corresponding to each of the remaining operators and the fused operator.
[0013] In some embodiments, running the remaining operators and the fused operator to determine a final implementation scheme corresponding to each of the remaining operators and the fused operator among at least one implementation scheme corresponding to each of the remaining operators and the fused operator includes: generating all implementation scheme combinations according to at least one implementation scheme corresponding to each of the remaining operators and the fused operator; the implementation scheme combinations include one implementation scheme corresponding to each of the remaining operators and the fused operator; for each implementation scheme combination among all the implementation scheme combinations, running the remaining operators and the fused operator according to the implementation schemes in the implementation scheme combination to obtain an inference duration corresponding to the implementation scheme combination; and determining a final implementation scheme for the remaining operators and the fused operator based on the target implementation scheme combination corresponding to the shortest inference duration among the inference durations.
[0014] In some embodiments, running the remaining operators and the fused operator to obtain an inference duration corresponding to the implementation scheme combination includes: running the remaining operators and the fused operator to obtain the running durations of the remaining operators and the fused operator; and determining the sum of the running durations of the remaining operators and the fused operator as the inference duration.
[0015] In some embodiments, when the target type operator has fixed weights, the remaining operator and / or the fused operator does not include a Reformat operator corresponding to the weights of the target type operator.
[0016] The present application also provides a method for generating an image. The method may include: obtaining a model file of a generation image model for generating an image; the generation image model includes a target type operator that performs operations based on weight information; the target type operator includes a convolution operator and a deconvolution operator; parsing the model file to obtain each operator included in the generation image model, the input information, output information corresponding to each operator, and the connection relationship between the operators; in a case where the input information corresponding to the target type operator does not include weight information, determining, according to the connection relationship, a target upper-layer operator connected to the target type operator; the target upper-layer operator is used to generate the weight information required by the target type operator; determining the weight information based on the output of the target upper-layer operator; accelerating the generation image model based on the acceleration library; wherein, the target type operator is implemented by using the weight information; deploying the generation image model that has completed acceleration; inputting parameters for generating a target image into the deployed generation image model to generate a target image.
[0017] The present application also provides an acceleration device based on an acceleration library. The device may include: an obtaining module that obtains a model file of a neural network model to be accelerated; the neural network model includes a target type operator that performs operations based on weight information; a parsing module that parses the model file to obtain each operator included in the neural network model, the input information, output information corresponding to each operator, and the connection relationship between the operators; a first determination module that, in a case where the input information corresponding to the target type operator does not include weight information, determines, according to the connection relationship, a target upper-layer operator connected to the target type operator; the target upper-layer operator is used to generate the weight information required by the target type operator; a second determination module that determines the weight information based on the output of the target upper-layer operator, so as to implement the target type operator by using the weight information during the process of accelerating the neural network model based on the acceleration library.
[0018] The present application also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor realizes the acceleration method or the image generation method based on the acceleration library shown in any of the foregoing embodiments by running the executable instructions.
[0019] The present application also provides a computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to make a processor execute the acceleration method or the image generation method based on the acceleration library shown in any of the foregoing embodiments.
[0020] In the foregoing solution, in the case where the weight information is not included in the parsed input information, the target upper-layer operator connected to the target type operator can be determined according to the connection relationship, and the output of the target upper-layer operator can be determined as the weight information, so that the weight information can be normally obtained for the target type operator to complete the acceleration of the neural network. Compared with the related art, it is equivalent to adding a path for obtaining weight information to the acceleration library, enabling the acceleration library to normally obtain variable weight information, normally implement the operation of the target type operator, and thus normally complete the acceleration of the neural network, improving the compatibility of the acceleration library.
[0021] In addition, this solution does not require any modification to the original neural network model, reducing the acceleration cost.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in one or more embodiments of this application or the related art, the following will briefly introduce the drawings required for the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments recorded in one or more embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a schematic flowchart of an acceleration method based on an acceleration library shown in an embodiment of this application;
[0025] Figure 2 It is a schematic flowchart of a model parsing method shown in an embodiment of this application;
[0026] Figure 3 It is a schematic flowchart of a model optimization method shown in an embodiment of this application;
[0027] Figure 4 It is a schematic flowchart of a method for determining the final implementation solution shown in an embodiment of this application;
[0028] Figure 5 It is a schematic flowchart of a model parsing method shown in an embodiment of this application;
[0029] Figure 6 It is a schematic flowchart of a model optimization method shown in an embodiment of this application;
[0030] Figure 7 It is a schematic flowchart of a method for generating an image shown in an embodiment of this application;
[0031] Figure 8 A structural schematic diagram of an acceleration device based on an acceleration library shown in an embodiment of the present application;
[0032] Figure 9 A hardware structural schematic diagram of an electronic device shown in an embodiment of the present application. Detailed implementation manners
[0033] Exemplary embodiments will be described in detail below, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0034] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items. It should also be understood that the word "if" used herein can be interpreted as "when", "while", or "in response to determining" depending on the context.
[0035] Based on this, the present application proposes an acceleration method based on an acceleration library. The method can add a method for obtaining variable weights to the acceleration library, so that weights can also be obtained for target type operators in the case of variable weights to complete the acceleration of target type operators.
[0036] The method may include:
[0037] Obtain a model file of a neural network model to be accelerated; the neural network includes target type operators that perform operations based on weight information; parse the model file to obtain each operator included in the neural network, the input information corresponding to each operator, and the connection relationship between the operators; in the case where the input information corresponding to the target type operator does not include weight information, determine a target upper-level operator connected to the target type operator according to the connection relationship; the target upper-level operator is used to generate the weight information required by the target type operator; based on the output of the target upper-level operator, determine the weight information, so as to utilize the weight information to implement the target type operator during the process of accelerating the neural network model based on the acceleration library.
[0038] In the foregoing method, in the case where the weight information is not included in the parsed input information, the target upper-layer operator connected to the target type operator can be determined according to the connection relationship, and the output of the target upper-layer operator can be determined as the weight information, so that the weight information can be normally obtained for the target type operator to complete the acceleration of the neural network. Compared with the related art, it is equivalent to adding a path for obtaining weight information to the acceleration library, enabling the acceleration library to normally obtain variable weight information, normally implement the operation of the target type operator, and thus normally complete the acceleration of the neural network, improving the compatibility of the acceleration library.
[0039] In addition, this method does not require any modification to the original neural network model, reducing the acceleration cost.
[0040] The following is an example description with reference to the accompanying drawings. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an acceleration method based on an acceleration library shown in an embodiment of the present application.
[0041] Figure 1 The acceleration method based on the acceleration library (hereinafter referred to as the acceleration method) shown can be applied to an electronic device. Among them, the electronic device can execute the acceleration method by loading software logic corresponding to the acceleration method. The type of the electronic device can be a laptop computer, a computer, a server, a mobile phone, a personal digital assistant (PDA), etc. The type of the electronic device is not particularly limited in the present application. The electronic device can also be a client device or a server device, which is not particularly limited here.
[0042] As Figure 1 shown, the method can include S102 - S108. Unless otherwise specified, the execution order of these steps is not particularly limited in the present application.
[0043] Among them, S102, obtain the model file of the neural network model to be accelerated.
[0044] The neural network model can be composed of any type of neural network. The neural network model includes a target type operator that performs operations based on weight information. The target type operator can be a basic operation operator such as a convolution operator, a deconvolution operator, a pooling operator, or a complex operation operator constructed based on the basic operation operator.
[0045] For example, the neural network model can be an image classification model, a generated image model, a synthetic image model, etc., and the target type operator can be a convolution and deconvolution operator with fixed weights in the image classification model. For another example, the neural network model can be a stylegan2 model. The target type operator can be a convolution and deconvolution operator with variable weights in the image classification model.
[0046] The neural network model is generally saved in the form of a model file. By parsing the model file, relevant information of the neural network model can be obtained. The relevant information can include input information, output information of each operator included in the model, the connection relationship between operators, etc. The input information of the operator can include various types of inputs required for the operator to perform operations, such as weight information and input features.
[0047] For example, the formula of a certain 2D convolution operator is: f(x) = w1*x1 + w2*x2 + w3*x3 + w4*x4; where, w1, w2, w3, and w4 are weight information (which can also be referred to as convolution weights, etc.), x1, x2, x3, and x4 are input features, and f(x) is output feature information. Among them, in some ways, the input features and the output features can be saved in the form of a feature map.
[0048] In this step, the model file saved for the neural network model can be obtained.
[0049] S104. Parse the model file to obtain each operator included in the neural network model, the input information corresponding to each operator, and the connection relationship between the operators.
[0050] This step can use the parsing method carried by the acceleration library to parse the model file to obtain the aforementioned relevant information included in the model file.
[0051] In some ways, the acceleration library can also perform a legality check on the model file. For example, the model file should include each operator included in the neural network model and the input information required for each operator. If the corresponding input information is not parsed out or some information is missing in the corresponding input information (such as missing weight information) when parsing a certain operator, it will be considered that the legality check fails. If the legality check fails, the acceleration of the neural network model cannot continue.
[0052] In this application, the parsing process of the acceleration library can be modified so that the legality check failure caused by the lack of weight information can be ignored, so that the subsequent acceleration can be performed normally.
[0053] S106. When the input information corresponding to the target type operator does not contain weight information, determine the target upper-layer operator connected to the target type operator; the target upper-layer operator is used to generate the weight information required by the target type operator.
[0054] S108. Based on the output of the target upper-layer operator, determine the weight information, so as to utilize the weight information to implement the target type operator during the process of accelerating the neural network model based on the acceleration library.
[0055] The acquisition path of the weight information is preset in the acceleration library, that is, the weight information is acquired from the model file. In the case where the weight information is immutable, the weight information will be pre-stored in the model file, and according to the acquisition path, the weight information can be normally acquired to complete the operation of the target type operator. However, in the case where the weight information is mutable, the weight information is calculated by other upper-layer operators during the inference process and will not be pre-stored in the model file. According to the original acquisition path, the weight information cannot be normally acquired.
[0056] For the case where the weight information cannot be normally acquired, the weight information can be determined based on the output of the target upper-layer operator connected to the target type operator according to the connection relationship.
[0057] Among them, the target upper-layer operator is used to generate the weight information required by the target type operator.
[0058] The connection relationship is used to indicate the sequential operation relationship between operators. For example, for operators A and B with a connection relationship, where A is the upper-layer operator. That is, it is necessary to operate A first and then operate B, and the output of A is the input of B.
[0059] In some ways, for each operator, the function corresponding to each input of the operator can be maintained. For example, for a convolution operator, it can be maintained whether each input of the convolution operator is for input features or for input weights.
[0060] In S106, the upper-layer operator connected to the target type operator can be determined according to the connection relationship, and then the target upper-layer operator for inputting weights can be determined from the upper-layer operators.
[0061] After determining the target upper-layer operator, the weight information can be determined based on the output of the target upper-layer operator.
[0062] After determining the weight information of the target type operator, during the process of accelerating the neural network model based on the acceleration library, the weight information can be utilized to implement the target type operator, so that the acceleration of the neural network model can be normally completed by using the acceleration library.
[0063] In the foregoing solution, in the case where the weight information is not included in the parsed input information, the target upper-layer operator connected to the target type operator can be determined according to the connection relationship, and the output of the target upper-layer operator can be determined as the weight information, so that the weight information can be normally obtained for the target type operator to complete the acceleration of the neural network. Compared with the related art, it is equivalent to adding a path for obtaining weight information to the acceleration library, so that the acceleration library can normally obtain variable weight information, normally implement the operation of the target type operator, and thus normally complete the acceleration of the neural network, improving the compatibility of the acceleration library.
[0064] In addition, this method does not require any modification to the original neural network model, reducing the acceleration cost.
[0065] The following is an example description in combination with the scenario of accelerating a neural network model using an acceleration library.
[0066] The acceleration may include two parts: model parsing and model optimization. The following will separately describe these two parts.
[0067] Model parsing:
[0068] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for model parsing shown in an embodiment of the present application. As Figure 2 shown, the method may include S202 - S206. Unless otherwise specified, the execution order of these steps is not limited in the present application.
[0069] S202, obtain the model file of the neural network model.
[0070] S204, parse the model file to obtain each operator included in the neural network model, the input information corresponding to each operator, and the connection relationship between the operators.
[0071] The relevant descriptions of S202 - S204 can refer to the foregoing embodiments and will not be elaborated here.
[0072] S206, for each operator among the operators: determine the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator; allocate a first storage space for storing the input data and a second storage space for storing the output data for the operator according to the input size information and the output size information.
[0073] In the acceleration library, it is necessary to determine fixed storage spaces for the inputs and outputs of each operator of the neural network model. Thus, during the neural network inference process, there is no need to frequently apply for storage spaces for the inputs and outputs, saving hardware resources and improving the inference efficiency. It can be understood that for some operators with connection relationships, the output of the upper-layer operator is the input of the lower-layer operator. Therefore, the second storage space allocated for the output of the upper-layer operator and the first storage space corresponding to the input of the lower-layer operator can be the same storage space, thus avoiding waste of storage space.
[0074] Among the input information parsed in S204, input size information is included, and the output size information of each operator needs to be obtained based on the input size information. In some ways, for different operators in the acceleration library, the determination logic for determining its output size information based on the input size information is integrated. In S206, it is only necessary to call these determination logics for each operator to complete the determination of the output size information.
[0075] This application needs to add the above-mentioned determination logic for the target type operator so that the acceleration library can be compatible with the target type operator with fixed weights and variable weights to determine the output size information.
[0076] Among them, in the case where the weights of the target type operator are fixed, the input information corresponding to the target type operator may include weight information. At this time, the output size information of the output data of the target type operator can be determined according to the input size information and the size information of the weight information included in the input information.
[0077] For example, the target type operator is a two-dimensional operation, and the size information corresponding to the input feature of the target type operator is [n, c, h, w], where n is the preset batch size, c is the number of channels of the input feature, h is the height of the input feature, and w is the width of the input feature. The size information of the weight information indicates that the convolution kernel is K[2], the boundary padding is P[2], the stride is S[2], and the number of output feature maps is Cout.
[0078] When the target type operator is a two-dimensional convolution, according to the formula for determining the output size in two-dimensional convolution, the height of the output is (h - K[0] + 2 * P[0]) / S[0] + 1, and the width is (w - K[1] + 2 * P[1]) / S[1] + 1.
[0079] Thus, the size information of the output of the target type operator can be obtained as:
[0080] [n, Cout, (h - K[0] + 2 * P[0]) / S[0] + 1, (w - K[1] + 2 * P[1]) / S[1] + 1].
[0081] When the target type operator is a two-dimensional transposed convolution, according to the formula for determining the output size in two-dimensional transposed convolution, the height of the output is (h - 1) * S[0] + 2 * P[0] - K[0] + 2, and the width is (w - 1) * S[1] + 2 * P[1] - K[1] + 2.
[0082] From this, the size information of the output of the target type operator can be obtained as:
[0083] [n, Cout, (h - 1) * S[0] + 2 * P[0] - K[0] + 2, (w - 1) * S[1] + 2 * P[1] - K[1] + 2].
[0084] For another example, the target type operator is a three-dimensional operation, and the size information of the input feature corresponding to the target type operator is [n, c, d, h, w], where n is the preset batch size, c is the number of channels of the input feature, d is the depth of the input feature, h is the height of the input feature, and w is the width of the input feature. The size information of the weight information indicates that the convolution kernel is L[3], the boundary padding is M[3], the stride is N[3], and the number of output feature maps is Cout.
[0085] When the target type operator is a three-dimensional convolution, according to the formula for determining the output size in three-dimensional convolution, the depth of the output is (d - L[0] + 2 * M[0]) / N[0] + 1, the height is (h - L[1] + 2 * M[1]) / N[1] + 1, and the width is (w - L[2] + 2 * M[2]) / N[2] + 1.
[0086] From this, the size information of the output of the target type operator can be obtained as:
[0087] [n, Cout, (d - L[0] + 2 * M[0]) / N[0] + 1, (h - L[1] + 2 * M[1]) / N[1] + 1, (w - L[2] + 2 * M[2]) / N[2] + 1].
[0088] When the target type operator is a three-dimensional transposed convolution, according to the formula for determining the output size in three-dimensional transposed convolution, the depth of the output is (d - 1) * N[0] + 2 * M[0] - L[0] + 2, the height is (h - 1) * N[1] + 2 * M[1] - L[1] + 2, and the width is (w - 1) * N[2] + 2 * M[2] - L[2] + 2.
[0089] From this, the size information of the output of the target type operator can be obtained as:
[0090] [n, Cout, (d - 1)*N[0]+2*M[0]-L[0]+2, (h - 1)*N[1]+2*M[1]-L[1]+2, (w - 1)*N[2]+2*M[2]-L[2]+2.
[0091] In the case where the weights of the target type operator are variable, the input information corresponding to the target type operator may not contain weight information. At this time, the output size information of the output data of the target type operator can be determined according to the input size information and the size information of the output data output by the target upper-layer operator.
[0092] That is, in the case where the weights are variable, the size information of the weights can be determined based on the size information of the output data output by the target upper-layer operator, so that the size information of the output of the target type operator can be determined based on the formula for determining the output size of the foregoing two-dimensional or three-dimensional convolution and transposed convolution.
[0093] For example, the target type operator is a two-dimensional operation, and the size information corresponding to the input feature of the target type operator is [n, c, h, w], where n is the preset batch size, c is the number of channels of the input feature, h is the height of the input feature, and w is the width of the input feature. The size information of the output data output by the target upper-layer operator (i.e., the size information of the weight information) is [n2, c2, h2, w2]. Among them, c2 can be regarded as the number of feature maps, h2 can be regarded as the height of the convolution kernel, that is, K[0], and w2 can be regarded as the width of the convolution kernel, that is, K[1]. The boundary padding in other size parameters of the target type operation is P[2], and the stride is S[2].
[0094] Specifically, when the target type operator is a two-dimensional convolution, its output is:
[0095] [n, c2, (h - h2 + 2*P[0]) / S[0]+1, (w - w2 + 2*P[1]) / S[1]+1].
[0096] When the target type operator is a two-dimensional transposed convolution, its output is:
[0097] [n, c2, (h - 1)*S[0]+2*P[0]-h2+2, (w - 1)*S[1]+2*P[1]-w2+2].
[0098] The target type operator is a three-dimensional operation. The size information corresponding to the input feature of the target type operator is [n, c, d, h, w], where n is the preset batch size, c is the number of channels of the input feature, h is the height of the input feature, and w is the width of the input feature. The size information of the output data output by the target upper-layer operator (i.e., the size information of the weight information) is [n2, c2, d2, h2, w2]. Among them, c2 can be regarded as the number of feature maps, d2 can be regarded as the depth of the convolutional kernel, that is, L[0], h2 can be regarded as the height of the convolutional kernel, that is, L[1], and w2 can be regarded as the width of the convolutional kernel, that is, L[2]. Among the other size parameters of the target type operation, the boundary padding is M[3], and the stride is N[3].
[0099] Specifically, when the target type operator is a three-dimensional convolution, its output is:
[0100] [n, n2, (d - d2 + 2 * M[0]) / N[0] + 1, (h - h2 + 2 * M[1]) / N[1] + 1, (w - w2 + 2 * M[2]) / N[2] + 1].
[0101] When the target type operator is a two-dimensional transposed convolution, its output is:
[0102] [n, c2, (d - 1) * N[0] + 2 * M[0] - d2 + 2, (h - 1) * N[1] + 2 * M[1] - h2 + 2, (w - 1) * N[2] + 2 * P[2] - w2 + 2].
[0103] It can be understood that when the target type operator is a pooling operator, the output size can be determined by referring to the aforementioned convolutional operator. Only the formula for determining the output size needs to be replaced with the formula for determining the output size for the pooling operator. When the target type operator is a complex operation operator constructed based on basic operation operators such as pooling operators and / or convolutional operators, the size information of its output can be determined based on the method of determining the output size of the basic operation operator.
[0104] After determining the size information of the output of each operator included in the neural network model, the first storage space for storing the input data and the second storage space for storing the output data can be allocated for each operator according to the input size information and output size information of each operator. Thus, during the neural network inference process, there is no need to frequently apply for storage space for input and output, saving hardware resources and improving inference efficiency.
[0105] In some embodiments, neural network models generated under different frameworks (such as frameworks like caffe, tensorflow, pytorch, onnx, etc.) have different operator formats. In the acceleration library, it is necessary to unify the neural network operator formats under various frameworks. The format may include operator names, operator symbols, etc.
[0106] In this example, after parsing each operator included in the neural network model, it further includes:
[0107] According to the preset operator conversion method in the acceleration library, perform format conversion on each of the operators to obtain the operators after format conversion.
[0108] In this step, after completing the parsing of the model file, the generation framework of the neural network model can be determined. Then, according to the operator conversion method corresponding to the generation framework, format conversion can be completed for each operator. This facilitates the acceleration library to perform relevant acceleration processing.
[0109] Based on the foregoing S202 - S206, the step of model parsing for the neural network model is completed.
[0110] Model optimization:
[0111] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a model optimization method shown in an embodiment of the present application. As Figure 3 shown, the method may include S302 - S306. Unless otherwise specified, the present application does not limit the execution order of these steps.
[0112] S302, according to the preset fusion rules, fuse at least two operators indicated by the fusion rules to obtain a fused operator.
[0113] In a neural network, many operators can be fused together, thereby improving the overall computational efficiency. The fusion may include horizontal fusion and vertical fusion. Among them, horizontal fusion refers to fusing convolution operators or deconvolution operators with the same convolution kernel size but different weights, and improving the operation efficiency by using one convolution kernel. Vertical fusion refers to reducing the consumption of Kernel Launch and avoiding the video memory read and write operations between layers by fusing operations in the same order. For example, Conv, BatchNorm, and Relu operators can be fused.
[0114] The acceleration library will include some preset fusion rules. These fusion rules indicate the operators that can be fused. In S302, the operators that can be fused indicated by these fusion rules can be fused to obtain a fused operator.
[0115] For example, for a convolution operation with variable weights, the Reshape operation for the input of the convolution operator and the convolution operator can be fused into one convolution operator, thereby reducing the model inference time for model acceleration.
[0116] For another example, for a transposed convolution operation with variable weights, the Reshape operation for the first input of the transposed convolution operator, the Reshape, Reshape, Transpose, Reshape operations for the second input, and the transposed convolution operator can be fused into one transposed convolution operator, thereby reducing the model inference time for model acceleration.
[0117] In some embodiments, an output cache can also be pre-allocated for the concat operator included in the neural network model, thereby improving the operation efficiency of the model for model acceleration.
[0118] S304. Determine the remaining operators in the various operators except the at least two operators.
[0119] Some operators in the neural network model will not participate in the fusion. In this step, these remaining operators that do not participate in the fusion can be determined.
[0120] S306. Run the remaining operators and the fusion operator to complete the acceleration of the neural network model.
[0121] Each type of operator corresponding to the neural network model corresponds to at least one implementation scheme.
[0122] The implementation scheme refers to the scheme for allocating hardware resources to the operator.
[0123] The computing hardware needs to provide hardware resources for the operator to complete the operation of the operator. Since the neural network includes multiple operators, there are many schemes for the computing hardware to allocate hardware resources to the operator. That is, the operator corresponds to at least one implementation scheme.
[0124] How to reasonably allocate hardware resources for these operators, that is, how to determine the final implementation scheme for these operators, is also the content that the acceleration library needs to optimize. The final implementation scheme is the most reasonable and efficient running scheme.
[0125] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for determining the final implementation scheme shown in the embodiments of the present application. Figure 4 The steps shown are supplementary explanations for S306. As Figure 4 shown, the method may include S402 - S404. Unless otherwise specified, the present application does not limit the execution order of these steps.
[0126] S402. Generate at least one implementation solution corresponding to the fused operator according to at least one implementation solution corresponding to each of the at least two operators.
[0127] Since the acceleration library will perform the operator fusion step, that is, it is necessary to determine the implementation solution corresponding to the fused operator obtained after fusion.
[0128] In some ways, the implementation solution of the fused operator can continue the implementation solution of any operator before fusion. For example, operator A corresponds to implementation solutions 1, 2, 3; operator B corresponds to implementation solutions 4, 5, 6, then the implementation solutions after the fusion of operator A and operator B can be 1, 2, 3 or 4, 5, 6.
[0129] In some ways, the implementation solution of the fused operator can be the summary result of the implementation solutions of each operator before fusion. For example, operator A corresponds to implementation solutions 1, 2, 3; operator B corresponds to implementation solutions 4, 5, 6, then the implementation solutions after the fusion of operator A and operator B can be 1, 2, 3, 4, 5, 6.
[0130] S404. Run the remaining operators and the fused operator to determine the final implementation solutions corresponding to the remaining operators and the fused operator respectively among at least one implementation solution corresponding to each of the remaining operators and the fused operator.
[0131] In some ways, all implementation solution combinations can be generated first according to at least one implementation solution corresponding to each of the remaining operators and the fused operator.
[0132] Each implementation solution combination includes one implementation solution corresponding to each of the remaining operators and the fused operator. In this step, one implementation solution can be taken from at least one implementation solution corresponding to each of the remaining operators and the fused operator multiple times for combination to obtain all implementation solution combinations.
[0133] For example, the remaining operator C corresponds to implementation solutions 1, 2; the fused operator D corresponds to implementation solutions 3, 4. That is, all the obtained implementation solution combinations include four implementation solution combinations: 1 and 3; 1 and 4; 2 and 3; 2 and 4.
[0134] After obtaining all the implementation solution combinations, for each implementation solution combination in all the implementation solution combinations, run the remaining operators and the fused operator according to each implementation solution in the implementation solution combination to obtain the inference duration corresponding to the implementation solution combination.
[0135] The inference duration is the inference duration of the neural network model.
[0136] In some ways, for each of the implementation plan combinations, the remaining operator and the fusion operator can be run to obtain the running durations of the remaining operator and the fusion operator, and then the sum of the running durations of the remaining operator and the fusion operator is determined as the inference duration.
[0137] It should be noted that some operators can be run in advance before model inference, that is, the running durations of these operators can be excluded from the model's inference duration, so that a more accurate inference duration can be determined, and then the model can be optimized accurately.
[0138] In the case where the target type operator has fixed weights, the Reformat operator for the weights can be run in advance before model inference, so the running duration of this Reformat operator can be excluded from the inference duration. That is, in the case where the target type operator has fixed weights, the remaining operator and / or the fusion operator do not include the Reformat operator corresponding to the weights of the target type operator. Thus, the running duration of this Reformat operator can be excluded from the inference duration, so as to obtain a more accurate inference duration for accurate model optimization.
[0139] After obtaining the inference duration corresponding to each implementation plan combination respectively, based on the target implementation plan combination corresponding to the shortest inference duration among the inference durations, the final implementation plans can be determined for the remaining operator and the fusion operator.
[0140] In some ways, the inference durations corresponding to each implementation plan combination can be compared to find the shortest inference duration among them, and then the implementation plans within the target implementation plan combination corresponding to the shortest inference duration can be obtained, and the obtained implementation plans can be determined as the final implementation plans corresponding to the remaining operator and the fusion operator.
[0141] After determining the final implementation plans, the final implementation plans corresponding to each fusion operator and remaining operator of the neural network model can be saved to save the optimization and acceleration results of the neural network model. When the neural network model is used later, the optimized neural network can be deployed, so that when the neural network model is inferred, efficient inference can be performed according to the saved final implementation plans.
[0142] The following takes the acceleration library as the TensorRT library developed by NVIDIA Corporation and the neural network model as stylegan2 as an example for an embodiment description. Among them, stylegan2 includes convolutional operators and transposed convolutional operators with variable weights.
[0143] Model parsing:
[0144] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a method for model parsing shown in an embodiment of this application. As Figure 5 shown, the method may include S501-S506. Unless otherwise specified, this application does not limit the execution order of these steps.
[0145] S501, obtain the initial model file of stylegan2.
[0146] S502, parse the initial model file to obtain a parsing result.
[0147] The parsing result includes the generation framework tensorflow corresponding to stylegan2, each operator of stylegan2, the input information of each operator, and the connection relationship between each operator.
[0148] S503, perform a legality verification according to the parsed initial model file.
[0149] In this step, it can be determined whether the initial model file is a file saved under the tensorflow framework. If so, it is considered to pass the legality verification, otherwise it is considered not to pass the legality verification.
[0150] In this step, the legality of the input information of each operator can also be verified. If there is a situation of missing information, it can be considered not to pass the legality verification. However, in this application, a special processing for the convolutional operator with variable weights is performed, and it is considered to pass the legality verification even if there is a situation of missing weight information.
[0151] In this step, for the convolutional operator and the transposed convolutional operator with variable weights (hereinafter collectively referred to as the target operator), the input weight required for the target operator can also be determined according to the connection relationship and the output of the target upper-layer operator connected to the target operator, so as to normally complete the acceleration of stylegan2.
[0152] If the legality verification is passed, the subsequent process can be continued. If the legality verification is not passed, the acceleration process can be terminated and an alarm can be issued.
[0153] S504, perform a format conversion on each parsed operator to obtain the prefabricated operators in the acceleration library.
[0154] In this step, the conversion method of the operator corresponding to the tensorflow framework stored in the acceleration library can be used to unify the operator names, which is convenient for subsequent optimization processing of the acceleration library.
[0155] S505, determine the input size information and output size information of each prefabricated operator.
[0156] In this step, the output size information of each operator can be determined based on the input size information included in the input information of each operator. Among them, the output size information can be determined according to the output size determination formula stored in the acceleration library. The methods for determining the output sizes of the convolution operator and the deconvolution operator with variable weights can refer to the foregoing embodiments and will not be elaborated here.
[0157] After obtaining the input size information and output size information of each operator, the input size information and output size information can be determined as the input size information and output size information of the corresponding prefabricated operator.
[0158] S506. According to the input size information and output size information corresponding to each prefabricated operator, allocate a first storage space for storing input data and a second storage space for storing output data for each prefabricated operator.
[0159] Among them, for some operators with a connection relationship, the output of the upper-layer operator is the input of the lower-layer operator. The second storage space allocated for the output of the upper-layer operator and the first storage space corresponding to the input of the lower-layer operator can be the same storage space, thereby avoiding waste of storage space.
[0160] In the acceleration library, it is necessary to determine fixed storage spaces for the inputs and outputs of each prefabricated operator, so that during the neural network inference process, there is no need to frequently apply for storage spaces for the inputs and outputs, saving hardware resources and improving the inference efficiency.
[0161] Through S501-S506, the model parsing for stylegan2 is completed.
[0162] Model optimization:
[0163] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of a model optimization method shown in an embodiment of the present application. As Figure 6 shown, the method may include S601-S306. Unless otherwise specified, the present application does not limit the execution order of these steps.
[0164] S601. According to a preset fusion rule, fuse at least two operators indicated by the fusion rule to obtain a fused operator.
[0165] In this step, for the convolution operation with variable weights, the Reshape operation for the input of the convolution operator and the convolution operator can be fused into one convolution operator, thereby reducing the model inference time for model acceleration.
[0166] For the deconvolution operation with variable weights, the Reshape operation for the first input of the deconvolution operator, the Reshape, Reshape, Transpose, Reshape operations for the second input, and the deconvolution operator can be fused into one deconvolution operator, thereby reducing the model inference time for model acceleration.
[0167] Output caches can also be pre-allocated for the concat operators included in the neural network model, thereby improving the operation efficiency of the model for model acceleration.
[0168] Each operator in StyleGAN2 corresponds to multiple implementation schemes. In this step, multiple implementation schemes corresponding to the fused operator can also be determined according to the implementation schemes respectively corresponding to the operators before fusion.
[0169] S602. Determine the remaining operators among the various operators excluding the at least two operators.
[0170] S603. Generate all combinations of implementation schemes according to at least one implementation scheme respectively corresponding to the remaining operators and the fused operator.
[0171] In this step, one implementation scheme can be taken from at least one implementation scheme respectively corresponding to the remaining operators and the fused operator multiple times for combination until all combinations of implementation schemes are obtained.
[0172] S604. For each combination of implementation schemes among all the combinations of implementation schemes, run the remaining operators and the fused operator according to the implementation schemes in the combination of implementation schemes, and obtain the inference duration corresponding to the combination of implementation schemes.
[0173] In this step, for each combination of implementation schemes, a model inference can be performed to obtain the running durations respectively corresponding to the remaining operators and the fused operator. It should be particularly noted that, in order to perform model inference normally, for the convolution operator and deconvolution operator with variable weights (hereinafter collectively referred to as the target operator), according to the connection relationship, the output of the target upper-layer operator connected to the target operator can also be determined as the input weight required by the target operator, so that StyleGAN2 inference can be completed normally.
[0174] In some ways, the output data stored in the second storage space corresponding to the target upper-layer operator can be obtained; then the output data can be determined as the input weight.
[0175] S605. Based on the target combination of implementation schemes corresponding to the shortest inference duration among the inference durations, determine the final implementation schemes for the remaining operators and the fused operator.
[0176] In this step, the shortest inference duration can be determined by comparing the inference durations, and then the target implementation solution combination can be determined according to the shortest inference duration. After that, the final implementation solutions can be determined for the remaining operators and the fusion operator according to the target implementation solution combination.
[0177] S606. Generate and save an optimized model file according to the determined final implementation solution.
[0178] According to S601 - S606, the model optimization step is completed. The inference of stylegan2 can be efficiently completed according to the optimized model file.
[0179] This application also proposes a method for generating images. In this method, the image generation model can be accelerated based on an acceleration library.
[0180] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a method for generating images shown in an embodiment of this application. As Figure 7 shown, the method may include S701 - S707. Unless otherwise specified, this application does not particularly limit the execution order of these steps.
[0181] S701. Obtain the model file of the image generation model for generating images.
[0182] The image generation model includes target type operators that perform operations based on weight information; the target type operators include convolution operators and deconvolution operators.
[0183] For example, the image generation model may be a Gan model. The target type operators may be convolution operators and deconvolution operators with fixed or variable weights.
[0184] S702. Parse the model file to obtain each operator included in the image generation model, the input information, output information corresponding to each operator, and the connection relationship between the operators.
[0185] S703. In the case where the input information corresponding to the target type operator does not include weight information, determine the target upper - layer operator connected to the target type operator according to the connection relationship; the target upper - layer operator is used to generate the weight information required by the target type operator.
[0186] S704. Determine the weight information based on the output of the target upper - layer operator.
[0187] S705. Accelerate the image generation model based on the acceleration library; wherein, the target type operator is implemented using the weight information.
[0188] For the description of S701 - S705, reference can be made to the foregoing embodiments and will not be elaborated here.
[0189] S706, deploy the generated image model that has completed acceleration.
[0190] In some embodiments, based on the generated image model that has completed acceleration, an optimized model file corresponding to it can be generated. When performing model deployment, relevant deployment can be carried out based on the optimized model file, so that when using this generated image model to generate images, relevant operations can be carried out based on the optimized model file, improving the image generation speed.
[0191] S707, input the parameters for generating the target image into the deployed generated image model to generate the target image.
[0192] The parameters can affect the detailed parts of the target image. For example, if the target image is a human head image, the parameters can affect details such as the hair and wrinkles of the human head. In some ways, the parameters are latent variables that affect the aforementioned detailed parts. By adjusting the values of the latent variables, target images with different details can be obtained.
[0193] In the foregoing method, before deploying the generated image model for generation, the generated image model is first accelerated through an acceleration library, so that the operation speed of the generated image model can be improved, and then the speed of generating the target image can be improved.
[0194] In addition, in the case where the weight information is not included in the parsed input information, according to the connection relationship, the target upper - layer operator connected to the target - type operator can be determined, and the output of the target upper - layer operator is determined as the weight information, so that the weight information can be normally obtained for the target - type operator to complete the acceleration of the neural network. Compared with the related technology, it is equivalent to adding a path for obtaining weight information to the acceleration library, enabling the acceleration library to normally obtain variable weight information, normally implement the operation of the target - type operator, and thus normally complete the acceleration of the generated image model, improving the compatibility of the acceleration library. And the foregoing solution does not require any modification to the original generated image model, reducing the acceleration cost.
[0195] The following discloses related embodiments for accelerating the generated image model. The related descriptions of these embodiments can refer to the embodiments for describing the acceleration method above and will not be elaborated here.
[0196] In some embodiments, after parsing each operator included in the generated image model, it further includes:
[0197] According to the preset operator conversion method in the acceleration library, perform format conversion on each operator to obtain the operators after format conversion.
[0198] In some embodiments, the input information includes input size information of input data; the method further includes:
[0199] For each operator among the operators:
[0200] Determine output size information of the output data of the operator according to the input size information of the input data corresponding to the operator;
[0201] Allocate a first storage space for storing the input data and a second storage space for storing the output data for the operator according to the input size information and the output size information.
[0202] In some embodiments, the determining the weight information based on the output of the target upper-layer operator includes:
[0203] Obtain the output data stored in the second storage space corresponding to the target upper-layer operator;
[0204] Determine the output data as the weight information.
[0205] In some embodiments, the determining the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator includes:
[0206] In a case where weight information is included in the input information corresponding to the target type operator, determine the output size information of the output data of the target type operator according to the input size information and the size information of the weight information included in the input information;
[0207] In a case where weight information is not included in the input information corresponding to the target type operator, determine the output size information of the output data of the target type operator according to the input size information and the size information of the output data output by the target upper-layer operator.
[0208] In some embodiments, the acceleration library includes preset fusion rules;
[0209] Accelerating the generation image model based on the acceleration library includes:
[0210] According to the preset fusion rules, fuse at least two operators indicated by the fusion rules to obtain a fused operator;
[0211] Determine the remaining operators other than the at least two operators among the operators;
[0212] Run the remaining operators and the fused operator to complete the acceleration of the generation image model.
[0213] In some embodiments, at least one implementation solution corresponding to each operator;
[0214] Running the remaining operators and the fusion operator to complete the acceleration of the generated image model includes:
[0215] Generating at least one implementation solution corresponding to the fusion operator according to at least one implementation solution corresponding to the at least two operators;
[0216] Running the remaining operators and the fusion operator to determine the final implementation solutions corresponding to the remaining operators and the fusion operator respectively among at least one implementation solution corresponding to the remaining operators and the fusion operator respectively.
[0217] In some embodiments, running the remaining operators and the fusion operator to determine the final implementation solutions corresponding to the remaining operators and the fusion operator respectively among at least one implementation solution corresponding to the remaining operators and the fusion operator respectively includes:
[0218] Generating all implementation solution combinations according to at least one implementation solution corresponding to the remaining operators and the fusion operator respectively; the implementation solution combinations include one implementation solution corresponding to the remaining operators and the fusion operator respectively;
[0219] For each implementation solution combination in all the implementation solution combinations, running the remaining operators and the fusion operator according to each implementation solution in the implementation solution combination to obtain the inference duration corresponding to the implementation solution combination;
[0220] Based on the target implementation solution combination corresponding to the shortest inference duration among the inference durations, determining the final implementation solutions for the remaining operators and the fusion operator.
[0221] In some embodiments, running the remaining operators and the fusion operator to obtain the inference duration corresponding to the implementation solution combination includes:
[0222] Running the remaining operators and the fusion operator to obtain the running durations of the remaining operators and the fusion operator;
[0223] Determining the sum of the running durations of the remaining operators and the fusion operator as the inference duration.
[0224] In some embodiments, when the target type operator has fixed weights, the remaining operator and / or the fusion operator does not include a Reformat operator corresponding to the weights of the target type operator.
[0225] Corresponding to any of the above embodiments, the present application also proposes an acceleration device based on an acceleration library.
[0226] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an acceleration device based on an acceleration library shown in an embodiment of the present application. As Figure 8 shown, the acceleration device 800 based on the acceleration library may include:
[0227] An acquisition module 810, which acquires a model file of a neural network model to be accelerated; the neural network model includes target type operators that perform operations based on weight information;
[0228] An analysis module 820, which analyzes the model file to obtain each operator included in the neural network model, the input information, output information corresponding to each operator, and the connection relationship between the operators;
[0229] A first determination module 830, in a case where the input information corresponding to the target type operator does not include weight information, determines a target upper-layer operator connected to the target type operator according to the connection relationship; the target upper-layer operator is used to generate the weight information required by the target type operator;
[0230] A second determination module 840, which determines the weight information based on the output of the target upper-layer operator, so as to implement the target type operator by using the weight information during the process of accelerating the neural network model based on the acceleration library.
[0231] In some embodiments, the device 800 further includes:
[0232] A conversion module, after analyzing each operator included in the neural network model, performs format conversion on each operator according to the preset operator conversion method in the acceleration library to obtain operators after format conversion.
[0233] In some embodiments, the input information includes input size information of input data; the device 800 further includes:
[0234] An allocation module, for each operator among the operators:
[0235] Determines the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator;
[0236] Allocates a first storage space for storing input data and a second storage space for storing output data for the operator according to the input size information and the output size information.
[0237] In some embodiments, the second determination module 840 further:
[0238] Obtain the output data stored in the second storage space corresponding to the target upper-layer operator;
[0239] Determine the output data as the weight information.
[0240] In some embodiments, the allocation module further:
[0241] In the case where the input information corresponding to the target type operator includes weight information, determine the output size information of the output data of the target type operator according to the input size information and the size information of the weight information included in the input information;
[0242] In the case where the input information corresponding to the target type operator does not include weight information, determine the output size information of the output data of the target type operator according to the input size information and the size information of the output data output by the target upper-layer operator.
[0243] In some embodiments, the acceleration library includes preset fusion rules; the apparatus 800 further includes:
[0244] An acceleration module that accelerates the neural network model based on the acceleration library;
[0245] The acceleration module further:
[0246] According to the preset fusion rules, fuse at least two operators indicated by the fusion rules to obtain a fused operator;
[0247] Determine the remaining operators except the at least two operators among the operators;
[0248] Run the remaining operators and the fused operator to complete the acceleration of the neural network model.
[0249] In some embodiments, at least one implementation scheme corresponding to each operator;
[0250] The acceleration module further:
[0251] Generate at least one implementation scheme corresponding to the fused operator according to at least one implementation scheme corresponding to the at least two operators;
[0252] Run the remaining operators and the fused operator to determine the final implementation schemes corresponding to the remaining operators and the fused operator respectively among at least one implementation scheme corresponding to the remaining operators and the fused operator respectively.
[0253] In some embodiments, the acceleration module further:
[0254] Generate all implementation scheme combinations according to at least one implementation scheme corresponding to the residual operator and the fusion operator respectively; the implementation scheme combination includes one implementation scheme corresponding to the residual operator and the fusion operator respectively.
[0255] For each implementation scheme combination in all the implementation scheme combinations, run the residual operator and the fusion operator according to the respective implementation schemes in the implementation scheme combination to obtain an inference duration corresponding to the implementation scheme combination.
[0256] Based on the target implementation scheme combination corresponding to the shortest inference duration among the inference durations, determine the final implementation schemes for the residual operator and the fusion operator.
[0257] In some embodiments, the acceleration module further:
[0258] Run the residual operator and the fusion operator to obtain the running durations of the residual operator and the fusion operator.
[0259] Determine the sum of the running durations of the residual operator and the fusion operator as the inference duration.
[0260] In some embodiments, when the target type operator has fixed weights, the residual operator and / or the fusion operator does not include a Reformat operator corresponding to the weights of the target type operator.
[0261] In the foregoing solution, in the case where the input information obtained by parsing does not include the weight information, the target upper-layer operator connected to the target type operator can be determined according to the connection relationship, and the output of the target upper-layer operator is determined as the weight information, so that the weight information can be normally obtained for the target type operator to complete the acceleration of the neural network. Compared with the related art, it is equivalent to adding a path for obtaining weight information to the acceleration library, enabling the acceleration library to normally obtain variable weight information, normally implement the operations of the target type operator, and thus normally complete the acceleration of the neural network, improving the compatibility of the acceleration library.
[0262] In addition, this solution does not require any modification to the original neural network model, reducing the acceleration cost.
[0263] This application also proposes an apparatus for generating an image.
[0264] The apparatus for generating an image may include:
[0265] An acquisition module that acquires a model file of a generation image model for generating an image; the generation image model includes a target type operator that performs operations based on weight information; the target type operator includes a convolution operator and a deconvolution operator.
[0266] A parsing module that parses the model file to obtain each operator included in the generated image model, the input information, output information corresponding to each operator, and the connection relationship between the operators.
[0267] A third determination module that, when the input information corresponding to the target type operator does not include weight information, determines a target upper-layer operator connected to the target type operator according to the connection relationship; the target upper-layer operator is used to generate the weight information required by the target type operator.
[0268] A fourth determination module that determines the weight information based on the output of the target upper-layer operator.
[0269] An acceleration module that accelerates the generated image model based on the acceleration library; wherein, the weight information is used to implement the target type operator.
[0270] A deployment module that deploys the accelerated generated image model.
[0271] A generation module that inputs parameters for generating a target image into the deployed generated image model to generate a target image.
[0272] In some embodiments, the device further includes:
[0273] A conversion module that, after parsing to obtain each operator included in the generated image model, performs format conversion on the operators according to the preset operator conversion method in the acceleration library to obtain operators after format conversion.
[0274] In some embodiments, the input information includes input size information of input data; the device further includes:
[0275] An allocation module for each operator:
[0276] Determine the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator.
[0277] According to the input size information and the output size information, allocate a first storage space for storing input data and a second storage space for storing output data for the operator.
[0278] In some embodiments, the fourth determination module further:
[0279] Obtain the output data stored in the second storage space corresponding to the target upper-layer operator.
[0280] Determine the output data as the weight information.
[0281] In some embodiments, the allocation module further:
[0282] When the input information corresponding to the target type operator includes weight information, determine the output size information of the output data of the target type operator according to the input size information and the size information of the weight information included in the input information;
[0283] When the input information corresponding to the target type operator does not include weight information, determine the output size information of the output data of the target type operator according to the input size information and the size information of the output data output by the upper target operator.
[0284] In some embodiments, the acceleration library includes preset fusion rules;
[0285] The acceleration module further:
[0286] According to the preset fusion rules, fuse at least two operators indicated by the fusion rules to obtain a fused operator;
[0287] Determine the remaining operators among the operators except the at least two operators;
[0288] Run the remaining operators and the fused operator to complete the acceleration of the generated image model.
[0289] In some embodiments, at least one implementation scheme corresponding to each operator;
[0290] The acceleration module further:
[0291] Generate at least one implementation scheme corresponding to the fused operator according to at least one implementation scheme corresponding to the at least two operators respectively;
[0292] Run the remaining operators and the fused operator to determine the final implementation schemes corresponding to the remaining operators and the fused operator respectively among at least one implementation scheme corresponding to the remaining operators and the fused operator respectively.
[0293] In some embodiments, the acceleration module further:
[0294] Generate all implementation scheme combinations according to at least one implementation scheme corresponding to the remaining operators and the fused operator respectively; the implementation scheme combinations include one implementation scheme corresponding to the remaining operators and the fused operator respectively;
[0295] For each implementation solution combination among all the implementation solution combinations, run the remaining operators and the fusion operator according to the implementation solutions in the implementation solution combination, and obtain the inference duration corresponding to the implementation solution combination;
[0296] Based on the target implementation solution combination corresponding to the shortest inference duration among the inference durations, determine the final implementation solution for the remaining operators and the fusion operator.
[0297] In some embodiments, the acceleration module further:
[0298] Run the remaining operators and the fusion operator to obtain the running durations of the remaining operators and the fusion operator;
[0299] Determine the sum of the running durations of the remaining operators and the fusion operator as the inference duration.
[0300] In some embodiments, when the target type operator has fixed weights, the remaining operator and / or the fusion operator do not include a Reformat operator corresponding to the weights of the target type operator.
[0301] In the foregoing solution, before deploying the generation of the image generation model, the image generation model is accelerated by an acceleration library first, so that the operation speed of the image generation model can be improved, and further the speed of generating the target image can be improved.
[0302] In addition, in the case where the weight information is not included in the parsed input information, the target upper-layer operator connected to the target type operator can be determined according to the connection relationship, and the output of the target upper-layer operator is determined as the weight information, so that the weight information can be normally obtained for the target type operator to complete the acceleration of the neural network. Compared with the related art, it is equivalent to adding a path for obtaining weight information to the acceleration library, so that the acceleration library can normally obtain variable weight information, normally implement the operation of the target type operator, and thus normally complete the acceleration of the image generation model, improving the compatibility of the acceleration library. And the foregoing solution does not require any modification to the original image generation model, reducing the acceleration cost.
[0303] The embodiments of the acceleration device and / or the image generation device based on the acceleration library shown in this application can be applied to an electronic device. Correspondingly, this application discloses an electronic device, which may include: a processor.
[0304] A memory for storing instructions executable by the processor.
[0305] Among them, the processor is configured to call the executable instructions stored in the memory to implement the acceleration method based on the acceleration library and / or the method of generating an image shown in any of the foregoing embodiments.
[0306] Please refer to Figure 9 , Figure 9 which is a schematic diagram of the hardware structure of an electronic device shown in an embodiment of the present application.
[0307] As Figure 9 shown, the electronic device may include a processor for executing instructions, a network interface for network connection, a memory for storing operation data for the processor, and a non-volatile memory for storing corresponding instructions of an acceleration device based on an acceleration library and / or a device for generating an image.
[0308] Among them, the embodiments of the acceleration device based on the acceleration library and / or the device for generating an image may be implemented by software, or may be implemented by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the electronic device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware level, in addition to Figure 9 the processor, memory, network interface, and non-volatile memory shown, the electronic device where the device in the embodiment is located usually may further include other hardware according to the actual functions of the electronic device, which will not be elaborated here.
[0309] It can be understood that, in order to improve the processing speed, the corresponding instructions of the acceleration device based on the acceleration library and / or the device for generating an image may also be directly stored in the memory, which is not limited herein.
[0310] The present application proposes a computer-readable storage medium, and the storage medium stores a computer program, and the computer program can be used to enable the processor to execute the acceleration method based on the acceleration library and / or the method of generating an image shown in any of the foregoing embodiments.
[0311] Those skilled in the art should understand that one or more embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present application may take the form of a computer program product implemented on one or more computer-usable storage media (which may include but are not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0312] "And / or" in this application means having at least one of the two. For example, "A and / or B" can include three scenarios: A, B, and "A and B".
[0313] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the data processing device, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant parts of the method embodiment for the related content.
[0314] The above describes specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0315] The embodiments of the subject matter and functional operations described in this application can be implemented in the following: digital electronic circuits, tangible computer software or firmware, computer hardware that can include the structures disclosed in this application and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this application can be implemented as one or more computer programs, i.e., one or more modules in computer program instructions encoded on a tangible non-transitory program carrier to be executed by a data processing apparatus or to control the operation of the data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode and transmit information to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0316] The processing and logic flows described in this application can be executed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flows can also be executed by dedicated logic circuits - such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits), and the apparatus can also be implemented as dedicated logic circuits.
[0317] A computer suitable for executing a computer program can include, for example, a general and / or special purpose microprocessor, or any other type of processing unit. Generally, the processing unit will receive instructions and data from a read-only memory and / or a random access memory. The basic components of a computer can include a processing unit for implementing or executing instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also can include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operatively coupled to such mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer is not necessarily required to have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0318] Computer-readable media suitable for storing computer program instructions and data can include all forms of non-volatile memory, media, and memory devices, such as can include semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0319] Although this application contains many specific implementation details, these should not be construed as limiting the scope of any disclosure or the scope of what is claimed, but rather as mainly for describing the features of specific embodiments of a particular disclosure. Certain features described in multiple embodiments in this application can also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features can function in certain combinations and even be initially claimed as such, one or more features from a claimed combination can in some cases be removed from that combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.
[0320] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the embodiments described should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
[0321] Accordingly, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0322] The foregoing are only preferred embodiments of one or more embodiments of the present application, and are not intended to limit one or more embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of the present application shall be included within the scope of protection of one or more embodiments of the present application.
Claims
1. An acceleration method based on an acceleration library, characterized in that, Including: Obtain the model file of the neural network model to be accelerated; the neural network model includes target type operators that perform operations based on weight information. Parse the model file to obtain each operator included in the neural network model, the input information, output information corresponding to each operator, and the connection relationship between the operators. When the input information corresponding to the target type operator does not include weight information, determine the target upper-layer operator connected to the target type operator according to the connection relationship. The target upper-layer operator is used to generate the weight information required by the target type operator. Based on the output of the target upper-layer operator, determine the weight information, so as to use the weight information to implement the target type operator during the process of accelerating the neural network model based on the acceleration library.
2. The method according to claim 1, wherein After parsing to obtain each operator included in the neural network model, it further includes: Perform format conversion on each operator according to the preset operator conversion method in the acceleration library to obtain the operators after format conversion.
3. The method according to claim 1 or 2, characterized in that, The input information includes the input size information of the input data; the method further includes: For each operator among the operators: Determine the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator. Allocate a first storage space for storing the input data and a second storage space for storing the output data for the operator according to the input size information and the output size information.
4. The method according to claim 3, characterized in that The determining the weight information based on the output of the target upper-layer operator includes: Obtain the output data stored in the second storage space corresponding to the target upper-layer operator. Determine the output data as the weight information.
5. The method according to claim 3, wherein The determining the output size information of the output data of the operator according to the input size information of the input data corresponding to the operator includes: When the input information corresponding to the target type operator includes weight information, determine the output size information of the output data of the target type operator according to the input size information and the size information of the weight information included in the input information. When the input information corresponding to the target type operator does not include weight information, determine the output size information of the output data of the target type operator according to the input size information and the size information of the output data output by the target upper-layer operator.
6. The method according to claim 1, wherein The acceleration library includes preset fusion rules. Accelerating the neural network model based on the acceleration library includes: According to the preset fusion rules, fuse at least two operators indicated by the fusion rules to obtain a fused operator. Determine the remaining operators among the operators except the at least two operators. Run the remaining operators and the fused operator to complete the acceleration of the neural network model.
7. The method according to claim 6, characterized in that, At least one implementation scheme corresponding to each operator; The running the remaining operators and the fused operator to complete the acceleration of the neural network model includes: Generate at least one implementation scheme corresponding to the fused operator according to at least one implementation scheme corresponding to the at least two operators. Run the remaining operator and the fusion operator to determine the final implementation schemes corresponding to the remaining operator and the fusion operator respectively among at least one implementation scheme corresponding to the remaining operator and the fusion operator.
8. The method according to claim 7, characterized in that The step of running the remaining operator and the fusion operator to determine the final implementation schemes corresponding to the remaining operator and the fusion operator respectively among at least one implementation scheme corresponding to the remaining operator and the fusion operator includes: Generate all implementation scheme combinations according to at least one implementation scheme corresponding to the remaining operator and the fusion operator respectively; the implementation scheme combinations include one implementation scheme corresponding to the remaining operator and the fusion operator respectively; For each implementation scheme combination among all the implementation scheme combinations, run the remaining operator and the fusion operator according to the implementation schemes in the implementation scheme combination to obtain the inference duration corresponding to the implementation scheme combination; Based on the target implementation scheme combination corresponding to the shortest inference duration among the inference durations, determine the final implementation schemes for the remaining operator and the fusion operator.
9. The method according to claim 8, characterized in that, The step of running the remaining operator and the fusion operator to obtain the inference duration corresponding to the implementation scheme combination includes: Run the remaining operator and the fusion operator to obtain the running durations of the remaining operator and the fusion operator; Determine the sum of the running durations of the remaining operator and the fusion operator as the inference duration.
10. The method according to claim 9, wherein In the case where the target type operator has fixed weights, the remaining operator and / or the fusion operator does not include a Reformat operator corresponding to the weights of the target type operator.
11. A method for generating an image, characterized in that, including: Obtain the model file of the generation image model for generating images; the generation image model includes a target type operator that performs operations based on weight information; the target type operator includes a convolution operator and a deconvolution operator; Parse the model file to obtain each operator included in the generation image model, the input information, output information corresponding to each operator, and the connection relationship between the operators; In the case where the input information corresponding to the target type operator does not include weight information, determine the target upper-layer operator connected to the target type operator according to the connection relationship; The target upper-layer operator is used to generate the weight information required by the target type operator; Determine the weight information based on the output of the target upper-layer operator; Accelerate the generation image model based on the acceleration library; wherein, the target type operator is implemented using the weight information; Deploy the accelerated generation image model; Input the parameters for generating the target image into the deployed generation image model to generate the target image.
12. An acceleration device based on an acceleration library, characterized in that, including: An acquisition module that acquires the model file of the neural network model to be accelerated; the neural network model includes a target type operator that performs operations based on weight information; An analysis module that analyzes the model file to obtain each operator included in the neural network model, the input information, output information corresponding to each operator, and the connection relationship between the operators; The first determination module, in the case that the input information corresponding to the target type operator does not contain weight information, determines a target upper-layer operator connected to the target type operator according to the connection relationship; The target upper-layer operator is used to generate the weight information required by the target type operator; The second determination module determines the weight information based on the output of the target upper-layer operator, so as to implement the target type operator by using the weight information during the acceleration of the neural network model based on the acceleration library.
13. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor realizes the acceleration method based on the acceleration library according to any one of claims 1-10 or the method for generating an image according to claim 11 by running the executable instructions.
14. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to make the processor execute the acceleration method based on the acceleration library according to any one of claims 1-10 or the method for generating an image according to claim 11.
Citation Information
Patent Citations
Network model reasoning acceleration method and device, storage medium and intelligent equipment
CN111340215A
Operation method and device, and related product
CN111695686A
Neural network compiler configuration method and device, computer equipment and storage medium
CN113703741A