Operation execution method and device in target neural network model, and storage medium
By merging the weight matrix of the multi-head attention layer in the target neural network model and performing the operation of the second function, the problem of low computational efficiency is solved, the computational efficiency is improved, and the performance of image recognition and machine translation is enhanced.
Patent Information
- Application Number
- CN202111679495.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The computational efficiency is low during matrix calculations in the multi-head attention layer of the target neural network model.
By reading the merged target weight matrix in the target storage space and performing the operation of the second function in the multi-head attention layer, the number of matrix calculations is reduced and the computational efficiency is improved.
Improves the computational efficiency of multi-head attention layers, reduces storage space requirements, and improves the efficiency of image recognition and machine translation.
Smart Images

Figure CN114358252B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication, in particular to a method and apparatus for operation execution in a target neural network model, and a storage medium. BACKGROUND
[0002] The Transformer is a sequence-to-sequence model that uses a large amount of self-attention mechanism to extract information between sequences. This architecture solves the shortcomings of traditional recurrent neural networks that are difficult to parallelize, can process long sequence data, and avoids problems such as gradient disappearance and gradient explosion. Moreover, due to the powerful feature expression ability of the Transformer, the Transformer architecture has rapidly spread from the field of natural speech processing to other artificial intelligence fields (such as computer vision), and has gradually become a general solution.
[0003] However, in the target neural network model using the Transformer architecture, there is a large amount of matrix operation, for example, in the multi-head attention layer of the Transformer architecture, there are multiple weight matrices that need to be calculated with the input feature matrix, so that the multi-head attention layer has a very large number of matrix calculations in the calculation process, and each matrix calculation needs to access the memory. Further, the existing multi-head attention layer has low calculation efficiency in the process of matrix calculation.
[0004] In view of the related art, there is currently no effective solution to the problem of low calculation efficiency in the process of matrix calculation in the multi-head attention layer of the target neural network model.
[0005] Therefore, it is necessary to improve the related art to overcome the defects in the related art. SUMMARY
[0006] The embodiments of the present application provide a method and apparatus for operation execution in a target neural network model, and a storage medium, to at least solve the problem of low calculation efficiency in the process of matrix calculation in the multi-head attention layer of the target neural network model.
[0007] According to an aspect of an embodiment of the present application, there is provided a method for performing an operation in a target neural network model, comprising: obtaining input parameters of a multi-head attention layer in the target neural network model when the multi-head attention layer performs a target operation, wherein the input parameters of the multi-head attention layer comprise a plurality of feature matrices to be processed, and the target operation is configured to perform an operation of a first function on the input parameters of the multi-head attention layer and a set of weight matrices with predetermined values, and there are a plurality of weight matrices that are allowed to be merged in the set of weight matrices in the first function; reading a target weight matrix in a second function from a target storage space, wherein the second function is obtained after the plurality of weight matrices in the first function are merged, and the target weight matrix is a matrix obtained by merging the plurality of weight matrices; and performing an operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix by the multi-head attention layer to obtain a target operation result.
[0008] According to another aspect of an embodiment of the present application, there is also provided an apparatus for performing an operation in a target neural network model, comprising: an obtaining module configured to obtain input parameters of a multi-head attention layer in the target neural network model when the multi-head attention layer performs a target operation, wherein the input parameters of the multi-head attention layer comprise a plurality of feature matrices to be processed, and the target operation is configured to perform an operation of a first function on the input parameters of the multi-head attention layer and a set of weight matrices with predetermined values, and there are a plurality of weight matrices that are allowed to be merged in the set of weight matrices in the first function; a reading module configured to read a target weight matrix in a second function from a target storage space, wherein the second function is obtained after the plurality of weight matrices in the first function are merged, and the target weight matrix is a matrix obtained by merging the plurality of weight matrices; and an operation module configured to perform an operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix by the multi-head attention layer to obtain a target operation result.
[0009] According to yet another aspect of an embodiment of the present application, there is also provided a computer-readable storage medium having a computer program stored therein, wherein the computer program is configured to perform the method for performing an operation in a target neural network model when executed.
[0010] According to yet another aspect of an embodiment of the present application, there is also provided an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor performs the method for performing an operation in a target neural network model through the computer program.
[0011] By the present application, when the multi-head attention layer in the target neural network model performs the target operation operation, the input parameters of the multi-head attention layer are obtained, and the target weight matrix in the second function is read in the target storage space, and then the multi-head attention layer performs the operation operation of the second function on the input parameters and the target weight matrix, to obtain the target operation result. Since the second function is obtained by merging the plurality of weight matrices in the first function, and the target weight matrix is obtained by merging the plurality of weight matrices in the first function, and then the multi-head attention layer uses the target weight matrix to calculate through the second function, the calculation efficiency is higher than using the plurality of weight matrices through the first function. Furthermore, the above technical solution solves the problem of low calculation efficiency in the process of matrix calculation of the multi-head attention layer of the target neural network model. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the principles of the application. In the drawings:
[0013] Figure 1 is a hardware structure block diagram of a computer terminal of an operation execution method in a target neural network model according to an embodiment of the present application;
[0014] Figure 2 is a flowchart of an operation execution method in a target neural network model according to an embodiment of the present application;
[0015] Figure 3 is a processing schematic diagram of a multi-head attention layer according to an embodiment of the present application (one);
[0016] Figure 4 is a processing schematic diagram of a multi-head attention layer according to an embodiment of the present application (two);
[0017] Figure 5 is a processing schematic diagram of a multi-head attention layer according to an embodiment of the present application (three);
[0018] Figure 6 is a structure block diagram of an operation execution device in a target neural network model according to an embodiment of the present application (one);
[0019] Figure 7 is a structure block diagram of an operation execution device in a target neural network model according to an embodiment of the present application (two). DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0021] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0022] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal or similar computing device. Taking running on a computer terminal as an example, Figure 1 1 is a hardware structure diagram of a computer terminal for executing an operation in a target neural network model according to an embodiment of the present invention. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown in the figure) processor 102 (processor 102 may include but is not limited to a microprocessor (Microprocessor Unit, referred to as MPU) or a programmable logic device (Programmable logic device, referred to as PLD)) and a memory 104 for storing data. In an exemplary embodiment, the computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Equivalent functions or comparisons shown Figure 1 Shown are different configurations with more functionality.
[0023] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as a computer program corresponding to the operation execution method in the target neural network model in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the computer terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0024] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include a wireless network provided by a communication provider of the computer terminal. In one example, the transmission device 106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet in a wireless manner.
[0025] To solve the above-mentioned problems, in the embodiments of the present application, a method for executing an operation in a target neural network model is provided, Figure 2 is a flowchart of the method for executing an operation in a target neural network model according to the embodiments of the present application, which includes the following steps:
[0026] In step S202, when a target operation is performed on a multi-head attention layer in a target neural network model, input parameters of the multi-head attention layer are obtained, wherein the input parameters of the multi-head attention layer include a plurality of feature matrices to be processed, and the target operation is used to perform a first function operation on the input parameters of the multi-head attention layer and a set of weight matrices with predetermined values, and there are a plurality of weight matrices allowing merging in the set of weight matrices in the first function;
[0027] It should be noted that the input parameters include a query matrix Q, a key matrix K, and a value matrix V.
[0028] In step S204, a target weight matrix in a second function is read from a target storage space, wherein the second function is obtained after a plurality of weight matrices in the first function are merged, and the target weight matrix is a matrix obtained by merging the plurality of weight matrices;
[0029] It should be noted that in the multi-head attention layer, there is a mapping relationship between the target operation and the target weight matrix. Specifically, when it is detected that a target operation needs to be performed, the target operation is not performed. Instead, a target address of a target storage space is obtained according to the mapping relationship, and then the target weight matrix is obtained from the target storage space according to the target address. It should be noted that the target storage space includes but is not limited to: memory and hard disk. The target storage space in this embodiment takes memory as an example.
[0030] Step S206: Perform the operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix in the multi-head attention layer to obtain a target operation result.
[0031] That is to say, after obtaining the target weight matrix, the multi-head attention layer will perform the operation of the second function on the input parameters and the target weight matrix, instead of performing the operation of the first function based on the input parameters and the multiple weight matrices.
[0032] Through the above steps, when the multi-head attention layer in the target neural network model performs the target operation, the input parameters of the multi-head attention layer are obtained, and the target weight matrix in the second function is read in the target storage space, and then the multi-head attention layer performs the operation of the second function on the input parameters and the target weight matrix to obtain the target operation result. Since the second function is a function obtained by merging multiple weight matrices in the first function, and the target weight matrix is a matrix obtained by merging multiple weight matrices in the first function, after the multi-head attention layer obtains the input parameters, the calculation efficiency of using the target weight matrix through the second function is higher than that of using multiple weight matrices through the first function. The above technical solution is adopted to solve the problem of low computational efficiency in the process of matrix calculation in the multi-head attention layer of the target neural network model.
[0033] It should be noted that the above-mentioned target neural network models include but are not limited to: machine translation neural network models, image recognition neural network models, classification neural network models, etc.
[0034] In an exemplary embodiment, when the embodiment of the present application is applied to an image recognition neural network model, the computational efficiency of the matrix calculation of the multi-head attention layer in the image recognition neural network model can be improved, thereby improving the recognition efficiency of image recognition.
[0035] For a better understanding, in an exemplary embodiment, before obtaining the input parameters of the multi-head attention layer, it is necessary to obtain a set of weight matrices with predetermined values; merge the multiple weight matrices that are allowed to be merged in the set of weight matrices to obtain the target weight matrix; and store the target weight matrix in the target storage space.
[0036] It should be noted that the above set of weight matrices is determined by the multi-head attention layer during the network training process. After the network training is completed, the multi-head attention layer will merge multiple weight matrices in the set of weight matrices that can be merged to obtain the target weight matrix.
[0037] Specifically, in an optional embodiment, the above-mentioned merging of the plurality of weight matrices that are allowed to be merged in a set of weight matrices to obtain the target weight matrix can be achieved in the following manner: the target weight matrix is obtained by performing the following merging operation: W re_proj =W v W proj , wherein the target weight matrix includes W atten and W re_proj The plurality of feature matrices to be processed in the input parameters include a query matrix, a key matrix, and a value matrix, and the set of weight matrices includes a first weight matrix W corresponding to the query matrix. q , a second weight matrix W corresponding to the key matrix k , a third weight matrix W corresponding to the value matrix v , and the fourth weight matrix W proj , W q and W k is the weight matrix that allows merging, W v and W proj It is the weight matrix that allows merging. It should be noted that the above scale is a preset value.
[0038] It should be noted that, in an exemplary embodiment, before performing the merging operation, it is also necessary to perform the following operations on W q and Adjust the dimensions of Adjust to and will Adjust to Follow the steps below to set W v 、W proj and W re_proj Adjust the dimensions of Adjust to Will Adjust to And Adjusting to Wherein, d model represents the feature vector dimension after the model input vector is subjected to feature embedding, d k represents the feature vector dimension after the key vector is subjected to multi-head attention mapping, n heads represents the number of attention heads in the multi-head attention, d v represents the feature vector dimension after the value vector is subjected to multi-head attention mapping.
[0039] In an exemplary embodiment, the multi-head attention layer performs the operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix to obtain a target operation result, which can be achieved by performing the operation of the second function as follows to obtain the target operation result: Y = ((Query·W atten ·Key T )·Value)·W re_proj , wherein Y represents the target operation result, the plurality of to-be-processed feature matrices in the input parameters include a query matrix, a key matrix, and a value matrix, Query represents the query matrix, Key represents the key matrix, Value represents the value matrix, and the target weight matrix includes W atten and W re_proj , wherein W re_proj = W v ·W proj , the set of weight matrices includes a first weight matrix W q corresponding to the query matrix, a second weight matrix W k corresponding to the key matrix, a third weight matrix W v corresponding to the value matrix, and a fourth weight matrix W proj , W q and W k are weight matrices allowed to be merged, and W v and W proj are weight matrices allowed to be merged; wherein the first function is as follows: wherein scale is a preset value.
[0040] That is, before the multi-head attention layer obtains the input parameters, the set of weight matrices W q , W k , W v , W proj that need to be subjected to the first function calculation are first read out from the memory, and the weight matrices allowed to be merged among them are subjected to a merging operation, specifically, W q , Wk Perform the merge operation to obtain the target weight matrix W atten , W v , W proj Perform the merge operation to obtain the target weight matrix W re_proj , and then save the merged target weight matrix to memory. After the multi-head attention layer obtains the input parameters, it directly obtains the target weight matrix from memory to calculate the second function, rather than reading a set of weight matrices from memory to calculate the first function. This reduces the number of data reads from memory and improves the computational efficiency of the multi-head attention layer. At the same time, the above method also reduces the required storage space.
[0041] For a better understanding, the following is a detailed description: Figure 3 : This is a processing diagram of a multi-head attention layer according to an embodiment of the present invention (I). In the prior art, after obtaining the input: parameter query matrix Query, key matrix Key, value matrix Value, the multi-head attention layer will calculate the first function. Specifically, the following steps will be performed:
[0042] Step S1: Get the weight matrix W corresponding to the query matrix Query from the memory q , and the query matrix Query and the weight matrix W q The operation result 1 is saved in the memory;
[0043] Step S2: Get the weight matrix W corresponding to the key matrix Key from the memory k , and the key matrix Key and the weight matrix W k The operation result 2 is saved in the memory;
[0044] Step S3: Get the weight matrix W corresponding to the value matrix Value from the memory v , and the value matrix Value and the weight matrix W v The result of the operation 3 is saved in the memory;
[0045] Step S4: Obtaining operation result 1 and operation result 2 from the memory, performing operation processing, and then saving the obtained operation result 4 into the memory;
[0046] Step S5: Obtain operation result 3 and operation result 4 from the memory, perform operation processing, and then save the obtained operation result 5 into the memory;
[0047] Step S6: Get the operation result 5 and weight matrix W from the memory proj , and perform calculation processing, and then save the final calculation results into memory.
[0048] It can be seen that the above steps S1-S6 perform 7 read data operations, 6 store data operations and 6 operation operations.
[0049] Figure 4 is a processing schematic diagram of a multi-head attention layer according to an embodiment of the present application (two), if the allowed merging of a set of weight matrices is performed before the input parameters are obtained, the merged target weight matrix is stored in the memory, and then after the multi-head attention layer obtains the input parameters, as shown in Figure 4 , the target weight matrix is read from the memory, and then the calculation of the second function is performed through the target weight matrix and the input parameters, which greatly improves the calculation efficiency. Specifically, the multi-head attention layer performs the following operations:
[0050] Step S1: obtaining the target weight matrix W from the memory atten , and saving the operation result 1 of the query matrix Query, the key matrix Key and the target weight matrix W atten to the memory;
[0051] Step S2: obtaining the operation result 1 from the memory, and saving the operation result 2 of the operation result 1 and the value matrix Value to the memory;
[0052] Step S3: obtaining the operation result 2 from the memory and the target weight matrix W re_proj , and performing operation processing, and then saving the obtained final calculation result to the memory.
[0053] It can be seen that the above steps S1-S3 perform 4 read data operations, 3 store data operations and 3 operation operations.
[0054] Since the above process is only one calculation of the multi-head attention layer, and the target weight matrix and the second function are used to perform the calculation when the multi-head attention layer performs multiple times the above process, the calculation efficiency can be greatly improved.
[0055] In an exemplary embodiment, when the target neural network model is an image recognition neural network model, the original object information to be recognized is also required to be obtained before the input parameters of the multi-head attention layer are obtained; one or more dimensional features of the original object information are obtained; and the plurality of feature matrices to be processed are determined according to the one or more dimensional features.
[0056] Further, in an exemplary embodiment, after the multi-head attention layer performs the operation operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix to obtain a target operation result, the target object information obtained by image recognition on the original object information is determined according to the target operation result.
[0057] That is, the above technical solutions improve the recognition efficiency of the image recognition neural network model in the process of image recognition, and reduce the image recognition time.
[0058] In an exemplary embodiment, the above-mentioned obtaining of the original object information to be subjected to image recognition comprises: obtaining an original image of a target object to be determined. It should be noted that the target object includes but is not limited to: people, animals, objects, etc.
[0059] In an exemplary embodiment, the above-mentioned determining of target object information obtained by image recognition on the original object information according to the target operation result comprises: determining target object information of the original image according to the target operation result.
[0060] Further, in an exemplary embodiment, when the target neural network model is a machine translation neural network model, before obtaining the input parameters of the multi-head attention layer, the original object information to be subjected to machine translation also needs to be obtained; the features of one or more dimensions of the original object information are obtained; and the plurality of feature matrices to be processed are determined according to the features of one or more dimensions.
[0061] Further, in an exemplary embodiment, after the multi-head attention layer performs the operation operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix to obtain the target operation result, the target object information obtained by machine translation on the original object information is also determined according to the target operation result.
[0062] That is, the above technical solutions improve the recognition efficiency of the image recognition neural network model in the process of image recognition, and reduce the image recognition time.
[0063] In an exemplary embodiment, the above-mentioned obtaining of the original object information to be subjected to machine translation comprises: obtaining an original sentence of a semantic to be determined.
[0064] In an exemplary embodiment, the above-mentioned determining of target object information obtained by machine translation on the original object information according to the target operation result comprises: determining semantic information of the original sentence according to the target operation result.
[0065] Obviously, the above-described embodiments are only a part of the embodiments of the present application, not all. In order to better understand the operation execution method in the above-mentioned target neural network model, the above-mentioned process is described below in conjunction with the embodiments, but not used to limit the technical solutions of the embodiments of the present application, specifically:
[0066] In an optional embodiment, the multi-head attention layer in the Transformer architecture is implemented mainly by linearly mapping the query matrix (Query), key matrix (Key), and value matrix (Value) and then solving the dot product. The classic Transformer multi-head attention solution process includes a large number of matrix multiplication operations and linear transformations, resulting in high computational overhead. The algorithm principle is as follows:
[0067] Formula 1:
[0068] Where Q is the query matrix of multi-head mapping Q = (Query·W q ). K is the key matrix of the multi-head mapping K=(Key·W k ). V is the value matrix of multi-head mapping V=(Value·W v ).
[0069] To address this characteristic of the Transformer model, this application proposes a method based on operator fusion that can significantly improve the efficiency of Transformer network inference. By leveraging the associative property of matrix multiplication, the complex multi-head mapping-demapping weights are simply and efficiently fused into two weight matrices before the multi-head attention layer obtains input parameters.
[0070] First, expand the above multi-head attention formula as follows:
[0071] Formula 2:
[0072] Among them, W q 、 W v and W proj is the weight matrix trained by the network and remains unchanged during the inference process. Therefore, it can be combined in advance to reduce the amount of computation during network inference. The multi-head attention layer after operator fusion is as follows:
[0073] Formula 3: ((Query·W atten Key T )·Value)·W re_proj ;
[0074] in, W re_proj =W v W proj This method can significantly reduce the number of multiplication operations in the multi-head attention reasoning process and accelerate the multi-head attention operation.
[0075] Since the original process, (Query W q) After linear mapping, the dimension needs to be changed to split the head dimension Same and (Value·W v ) also needs to undergo the same dimensionality change, which results in resource consumption during the inference process. In this application, similar dimensionality changes can be solved directly before inference. atten and W re_proj When solving W atten During the process, First, W q and Change the dimension so that so In solving W re_proj In the process, But we still need to re_proj Perform a dimensional change so that W re_proj It can play the role of integrating multiple heads: At this point, the dimension change operation can be integrated into the matrix multiplication operation before the inference stage, reducing the waste of computing resources caused by dimension changes during the calculation process.
[0076] In addition, ((Query·W atten Key T )·Value) Operation Comparison This will bring about memory optimization space. Compared with the previous method, when reasoning, the query must first be calculated separately (Query W q )and Then seek and (Value·W v ), and finally multiply to obtain the result. This calculation process is discontinuous, requiring data to be constantly in and out of the cache, resulting in inefficient memory access. However, the improved process makes matrix multiplication continuous, and due to its parallel nature, memory access can be optimized. Figure 5 Schematic diagram of the processing of the multi-head attention layer according to an embodiment of the present invention (III), as shown in FIG. Figure 5 During the calculation process, allocate a block of blocksize×max(len k ,d model ) high-speed memory, and save the calculation results in this memory to reduce the memory IO time consumption.
[0077] Specifically, in an exemplary embodiment, the query matrix Query dimension is [1,300,700], the key matrix Key dimension is [1,600,700], the value matrix Value dimension is [1,600,700], the number of attention heads n_head = 12, the linear mapping length of the query matrix and the key matrix is 400, and the linear mapping length of the value matrix is 500. The comparison of the inference speed of the traditional transformer and the inference speed of the operator fusion transformer is shown in the following table:
[0078] Conventional Operator fusion Inference speed 0.1450 0.0720
[0079] It can be seen that this application has a significant effect on reducing the inference speed of the multi-head attention layer of the Transformer network.
[0080] This approach fuses the weight matrices before the multi-head attention layer performs actual inference, reducing the number of operations required during inference and improving efficiency. Compared to other solutions, this approach reduces the dimensionality of the weight matrices, thereby reducing the storage space required for the inference model.
[0081] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0082] In this embodiment, an operation execution device in a target neural network model is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0083] Figure 6 1 is a structural block diagram of an operation execution device in a target neural network model according to an embodiment of the present invention (I), the device comprising:
[0084] The acquisition module 62 is configured to acquire input parameters of a multi-head attention layer in a target neural network model when the multi-head attention layer performs a target operation, wherein the input parameters of the multi-head attention layer include a plurality of feature matrices to be processed, and the target operation is configured to perform an operation of a first function on the input parameters of the multi-head attention layer and a set of weight matrices with predetermined values, and there are a plurality of weight matrices that allow merging in the set of weight matrices in the first function.
[0085] The reading module 64 is configured to read a target weight matrix in a second function in a target storage space, wherein the second function is obtained after the plurality of weight matrices in the first function are merged, and the target weight matrix is obtained by merging the plurality of weight matrices.
[0086] The operation module 66 is configured to perform an operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix in the multi-head attention layer to obtain a target operation result.
[0087] Through the above device, when the multi-head attention layer in the target neural network model performs the target operation, the input parameters of the multi-head attention layer are acquired, the target weight matrix in the second function is read in the target storage space, and then the operation of the second function is performed on the input parameters and the target weight matrix in the multi-head attention layer to obtain the target operation result. Since the second function is obtained after the plurality of weight matrices in the first function are merged, and the target weight matrix is obtained by merging the plurality of weight matrices in the first function, the calculation efficiency of the multi-head attention layer using the target weight matrix through the second function is higher than that of using the plurality of weight matrices through the first function. Through the above technical solution, the problem of low calculation efficiency in the process of matrix calculation of the multi-head attention layer in the target neural network model is solved.
[0088] Figure 7 FIG. 2 is a structural block diagram of an operation execution device in a target neural network model according to an embodiment of the present application, which includes a processing module 68.
[0089] In an exemplary embodiment, the processing module 68 is configured to acquire the set of weight matrices with predetermined values, merge the plurality of weight matrices that allow merging in the set of weight matrices to obtain the target weight matrix, and store the target weight matrix in the target storage space.
[0090] In an exemplary embodiment, the processing module 68 is further configured to obtain the target weight matrix through the following merging operation: W re_proj= W v · W proj , wherein the target weight matrix comprises W atten and W re_proj , the plurality of feature matrices to be processed in the input parameter comprises a query matrix, a key matrix and a value matrix, the set of weight matrices comprises a first weight matrix W q corresponding to the query matrix, a second weight matrix W k corresponding to the key matrix, a third weight matrix W v corresponding to the value matrix, and a fourth weight matrix W proj , W q and W k are weight matrices allowed to be merged, W v and W proj are weight matrices allowed to be merged.
[0091] In an exemplary embodiment, the processing module 68 is further configured to adjust the dimensions of W q and by adjusting to and adjusting to adjust the dimensions of W v , W proj and W re_proj by adjusting to adjusting to and to wherein d model represents the feature vector dimension after feature embedding of the model input vector, d k represents the feature vector dimension after multi-head attention mapping of the key vector, n heads represents the number of attention heads in the multi-head attention, and d v represents the feature vector dimension after multi-head attention mapping of the value vector.
[0092] In an exemplary embodiment, the operation module 66 is further configured to obtain the target operation result by performing the operation of the following second function: wherein Y represents the target operation result, the plurality of feature matrices to be processed in the input parameter comprises a query matrix, a key matrix and a value matrix, Query represents the query matrix, Key represents the key matrix, Value represents the value matrix, the target weight matrix comprises Watten and W re_proj ; wherein, W re_proj = W v · W proj , the set of weight matrices comprises a first weight matrix W q corresponding to the query matrix, a second weight matrix W k corresponding to the key matrix, a third weight matrix W v corresponding to the value matrix, and a fourth weight matrix W proj , W q and W k are weight matrices that are allowed to be merged, W v and W proj are weight matrices that are allowed to be merged; wherein the first function is a function as follows: wherein scale is a preset value.
[0093] In an example embodiment, the processing module 68 is further configured to, before obtaining the input parameter of the multi-head attention layer, the method further comprises: obtaining original object information to be subjected to image recognition; obtaining one or more dimensional features of the original object information; determining a plurality of feature matrices to be processed according to the one or more dimensional features; after the multi-head attention layer performs the operation of the second function on the input parameter of the multi-head attention layer and the target weight matrix to obtain a target operation result, the method further comprises: determining target object information obtained by performing image recognition on the original object information according to the target operation result.
[0094] In an example embodiment, the processing module 68 is further configured to, the obtaining of the original object information to be subjected to image recognition comprises: obtaining an original image of a target object to be determined; and the determining of the target object information obtained by performing image recognition on the original object information according to the target operation result comprises: determining target object information of the original image according to the target operation result.
[0095] Embodiments of the present application also provide a computer readable storage medium having a computer program stored therein, wherein the computer program is configured to execute the steps in any of the method embodiments described above when running.
[0096] Optionally, in the present embodiment, the storage medium described above can be configured to store a computer program for executing the following steps:
[0097] S1, when a multi-head attention layer in a target neural network model performs a target operation operation, an input parameter of the multi-head attention layer is acquired, wherein the input parameter of the multi-head attention layer includes a plurality of feature matrices to be processed, the target operation operation is used to perform an operation operation of a first function on the input parameter of the multi-head attention layer and a set of weight matrices with predetermined values, and there are a plurality of weight matrices that allow merging in the set of weight matrices in the first function;
[0098] S2, a target weight matrix in a second function is read in a target storage space, wherein the second function is a function obtained after a plurality of weight matrices in the first function are merged, and the target weight matrix is a matrix obtained by performing a merging operation on the plurality of weight matrices;
[0099] S3, the multi-head attention layer performs an operation operation of the second function on the input parameter of the multi-head attention layer and the target weight matrix, and obtains a target operation result.
[0100] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0101] The specific examples in the embodiment can refer to the examples described in the above embodiments and example embodiments, and the embodiment will not be described here.
[0102] The embodiment of the application also provides an electronic device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0103] Optionally, in the embodiment, the processor can be configured to perform the following steps by the computer program:
[0104] S1, when a multi-head attention layer in a target neural network model performs a target operation operation, an input parameter of the multi-head attention layer is acquired, wherein the input parameter of the multi-head attention layer includes a plurality of feature matrices to be processed, the target operation operation is used to perform an operation operation of a first function on the input parameter of the multi-head attention layer and a set of weight matrices with predetermined values, and there are a plurality of weight matrices that allow merging in the set of weight matrices in the first function;
[0105] S2, reading a target weight matrix in a target storage space in a second function, wherein the second function is a function obtained after merging a plurality of weight matrices in the first function, and the target weight matrix is a matrix obtained by performing a merging operation on the plurality of weight matrices;
[0106] S3, performing an operation operation of the second function on the input parameter of the multi-head attention layer and the target weight matrix by the multi-head attention layer to obtain a target operation result.
[0107] In an exemplary embodiment, the electronic device described above can further include a transmission device connected to the processor and an input and output device connected to the processor.
[0108] The specific examples in the present embodiment can refer to the examples described in the above embodiments and exemplary embodiments, which will not be described here again.
[0109] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be realized by general computing devices, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, which can be realized by program codes executable by computing devices, so that they can be stored in storage devices and executed by computing devices, and in some cases, the steps shown or described can be executed in different order, or they can be manufactured into individual integrated circuit modules, or multiple modules or steps can be manufactured into a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.
[0110] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for executing operations in a target neural network model, characterized in that: Applications in image recognition include: When a multi-head attention layer in a target neural network model performs a target operation, obtaining input parameters of the multi-head attention layer, wherein the input parameters of the multi-head attention layer include multiple feature matrices to be processed, and the target operation is used to perform a first function operation on the input parameters of the multi-head attention layer and a set of weight matrices with predetermined values, wherein the set of weight matrices in the first function includes multiple weight matrices that are allowed to be merged; Reading a target weight matrix in a second function in a target storage space, wherein the second function is a function obtained by merging multiple weight matrices in the first function, and the target weight matrix is a matrix obtained by merging the multiple weight matrices; Performing an operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix in the multi-head attention layer to obtain a target operation result; Before obtaining the input parameters of the multi-head attention layer, the method further includes: obtaining a set of weight matrices with predetermined values; merging the plurality of weight matrices allowed to be merged in the set of weight matrices to obtain the target weight matrix; and storing the target weight matrix in the target storage space; Before obtaining the input parameters of the multi-head attention layer, the method further includes: obtaining original object information to be image recognized; obtaining features of one or more dimensions of the original object information; determining the multiple feature matrices to be processed based on the features of the one or more dimensions; after performing the operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix in the multi-head attention layer to obtain a target operation result, the method further includes: determining target object information obtained by performing image recognition on the original object information based on the target operation result; Among them, the obtaining of the original object information to be image recognized includes: obtaining the original image of the target object to be determined; and the determining of the target object information obtained by image recognition of the original object information based on the target operation result includes: determining the target object information of the original image based on the target operation result.
2. The method according to claim 1, characterized in that Merging the plurality of weight matrices allowed to be merged in the set of weight matrices to obtain the target weight matrix includes: The target weight matrix is obtained by the following merging operation: , , The target weight matrix includes and The plurality of feature matrices to be processed in the input parameters include a query matrix, a key matrix, and a value matrix, and the set of weight matrices includes a first weight matrix corresponding to the query matrix , a second weight matrix corresponding to the key matrix , a third weight matrix corresponding to the value matrix , and the fourth weight matrix , and is the weight matrix that allows merging, and is the weight matrix that allows merging.
3. The method according to claim 2, characterized in that Before performing the merging operation, the method further includes: Follow the steps below to and Adjust the dimensions: Will Adjust to , and Adjust to ; in, Represents the dimension of the feature vector after the model input vector is embedded. Represents the dimension of the feature vector after the key vector is multi-headed attention mapped, Indicates the number of attention heads in multi-head attention; Follow the steps below to 、 and Adjust the dimensions: Will Adjust to ,Will Adjust to , and Adjust to ; in, Represents the dimension of the feature vector after multi-head attention mapping of the value vector.
4. The method according to claim 1, wherein The performing the operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix in the multi-head attention layer to obtain a target operation result includes: The target operation result is obtained by performing the following operation of the second function: , Wherein, Y represents the target operation result, the multiple feature matrices to be processed in the input parameters include a query matrix, a key matrix and a value matrix, Query represents the query matrix, Key represents the key matrix, Value represents the value matrix, and the target weight matrix includes and ; in, , , the set of weight matrices includes a first weight matrix corresponding to the query matrix , a second weight matrix corresponding to the key matrix , a third weight matrix corresponding to the value matrix , and the fourth weight matrix , and is the weight matrix that allows merging, and is the weight matrix that allows merging; The first function is as follows: Y= , in, is the default value.
5. An operation execution device in a target neural network model, characterized in that: Applications in image recognition include: an acquisition module, configured to acquire input parameters of the multi-head attention layer when the multi-head attention layer in the target neural network model performs a target operation, wherein the input parameters of the multi-head attention layer include multiple feature matrices to be processed, and the target operation is used to perform a first function operation on the input parameters of the multi-head attention layer and a set of weight matrices with predetermined values, wherein the set of weight matrices in the first function includes multiple weight matrices that are allowed to be merged; a reading module, configured to read a target weight matrix in a second function from a target storage space, wherein the second function is a function obtained by merging multiple weight matrices in the first function, and the target weight matrix is a matrix obtained by merging the multiple weight matrices; an operation module, configured to perform an operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix in the multi-head attention layer to obtain a target operation result; The device further includes: a processing module configured to obtain a set of weight matrices having predetermined values; merge the plurality of weight matrices allowed to be merged in the set of weight matrices to obtain the target weight matrix; and store the target weight matrix in the target storage space; The processing module is configured to obtain original object information to be image-recognized before obtaining input parameters of the multi-head attention layer; obtain features of one or more dimensions of the original object information; determine the plurality of feature matrices to be processed based on the features of the one or more dimensions; and perform the operation of the second function on the input parameters of the multi-head attention layer and the target weight matrix in the multi-head attention layer, and after obtaining a target operation result, determine target object information obtained by performing image recognition on the original object information based on the target operation result. Among them, the processing module is used to obtain the original object information to be image recognized in the following manner: obtaining the original image of the target object to be determined; determining the target object information obtained by image recognition of the original object information according to the target operation result in the following manner: determining the target object information of the original image according to the target operation result.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 4 when executed.
7. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 4 through the computer program.
Citation Information
Patent Citations
CondenseNet algorithm fused with attention selection mechanism
CN111160488A
Data processing method and related equipment
CN112288075A