Reasoning method, electronic device and corresponding device based on neural network
By storing at least part of the Q matrix, K matrix and V matrix to the on-chip memory of the target chip, and performing calculations based on on-chip memory during the inference process of the neural network, the problem of time-consuming operations in the prior art is solved, and the effect of reducing the time-consuming and improving efficiency of the inference process is achieved.
Patent Information
- Application Number
- CN202311866193.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-12-29
AI Technical Summary
In the prior art, the operation process of the Q matrix, K matrix and V matrix takes a long time, resulting in the inference process of the neural network taking a long time and the inference efficiency is low.
By accessing the off-chip memory of the target chip, reading the first matrix, and storing it at least partially from the off-chip memory to the on-chip memory, the Q matrix, the K matrix and the V matrix are operated based on the first matrix stored in the on-chip memory and its available storage space.
The amount of data read to off-chip memory is reduced, the number of times of global imitation is reduced, thereby reducing the time-consuming process in the inference process and improving the inference efficiency.
Smart Images

Figure CN118446306B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a neural network-based reasoning method, electronic equipment, and corresponding devices. Background Art
[0002] With the development of artificial intelligence technology, various types of neural networks have emerged, such as the control network ControlNet. ControlNet can be applied to a variety of scenarios. For example, after receiving a blurred image, ControlNet can output a corresponding clear image through corresponding reasoning operations.
[0003] Some neural networks such as the ControlNet network include an attention module (i.e., Attention module). During the reasoning process, the Attention module needs to operate on the query matrix (i.e., query matrix, which can be abbreviated as Q matrix), the key matrix (i.e., key matrix, which can be abbreviated as K matrix), and the weight matrix (i.e., value matrix, which can be abbreviated as V matrix).
[0004] However, the current process of calculating the Q matrix, the K matrix and the V matrix is time-consuming, which further causes the reasoning process of the neural network to be time-consuming and the reasoning efficiency to be low. Summary of the invention
[0005] In order to solve the problem in the prior art that the calculation process of Q matrix, K matrix and V matrix is time-consuming, resulting in a long reasoning process of the neural network and low reasoning efficiency, the embodiments of the present application provide a reasoning method, electronic device and corresponding device based on a neural network.
[0006] In a first aspect, the present application provides a neural network-based reasoning method, which is applied to a target chip, and the method comprises:
[0007] By accessing the off-chip memory of the target chip, a first matrix stored in the off-chip memory is read, where the first matrix is at least one of a query matrix, a key matrix, and a weight matrix;
[0008] Based on the data amount of the first matrix and the storage capacity of the on-chip memory of the target chip, storing at least part of the first matrix from the off-chip memory to the on-chip memory;
[0009] The query matrix, the key matrix and the weight matrix are operated based on at least a portion of the first matrix stored in the on-chip memory and available storage space of the on-chip memory.
[0010] When reasoning is performed through the solution provided by the embodiment of the present application, at least part of the first matrix stored therein can be read from the on-chip memory, which can reduce the amount of data read from the off-chip memory, that is, reduce the number of global simulations, compared with the existing solution in which the Q matrix, the K matrix, and the V matrix are all stored in the off-chip memory. The speed of reading data from the on-chip memory is faster than the speed of reading data from the off-chip memory. Therefore, the solution provided by the embodiment of the present application can reduce the time consumption in the reasoning process and improve the reasoning efficiency.
[0011] In a feasible design, based on the data amount of the first matrix and the storage capacity of the on-chip memory of the target chip, storing at least part of the first matrix from the off-chip memory to the on-chip memory includes:
[0012] If the storage capacity of a first on-chip memory in the on-chip memory is not less than the sum of the data amounts of the query matrix, the key matrix and the weight matrix, storing the query matrix, the key matrix and the weight matrix from the off-chip memory to the first on-chip memory;
[0013] or,
[0014] If the on-chip memory includes a plurality of second on-chip memories, and the capacity of the on-chip memory is smaller than the data volume of the first matrix, at least part of the matrix blocks split into the first matrix are stored in the corresponding second on-chip memories.
[0015] Through this design, the query matrix, key matrix and weight matrix can be stored in the first on-chip memory, or at least part of the first matrix can be split into matrix blocks and stored in the corresponding second on-chip memory, so as to store at least part of the first matrix in the on-chip memory.
[0016] In a feasible design, if the first matrix includes the query matrix, each first matrix block into which at least part of the query matrix is split is at least one row in the query matrix;
[0017] If the first matrix includes the key matrix, each second matrix block into which at least a portion of the key matrix is split is at least one column of the key matrix;
[0018] If the first matrix includes the weight matrix, each third matrix block into which at least a portion of the weight matrix is split is at least one row in the weight matrix.
[0019] In a feasible design, if the first matrix includes the query matrix and the key matrix, and the complete part of the query matrix and the complete part of the key matrix are both stored in the on-chip memory, the operation on the query matrix, the key matrix and the weight matrix based on at least a part of the first matrix stored in the on-chip memory and the available storage space of the on-chip memory includes:
[0020] Reading the query matrix and the key matrix from the on-chip memory, and performing a product operation on the read query matrix and the key matrix to obtain a second matrix;
[0021] Storing the second matrix in available storage space of the on-chip memory;
[0022] Based on the second matrix read from the available storage space, operation results of the query matrix, the key matrix and the weight matrix are determined.
[0023] Among them, the second matrix is an intermediate operation result. Through this design, the second matrix, an intermediate operation result, can be stored in the on-chip memory, and when the second matrix is needed for subsequent operations, the second matrix can be read by accessing the on-chip memory without global storage.
[0024] In a feasible design, the reading of the query matrix and the key matrix from the on-chip memory and performing a product operation on the read query matrix and the key matrix include:
[0025] Reading the query matrix row by row from the on-chip memory, and reading fourth matrix blocks included in the key matrix one by one from the on-chip memory, each of the fourth matrix blocks including at least one column in the key matrix;
[0026] The products of each row of the query matrix and each of the fourth matrix blocks are calculated respectively, and the matrix formed by the products is the second matrix.
[0027] In a feasible design, if the first matrix also includes a weight matrix, and the complete part of the weight matrix is stored in the on-chip memory, the operation results of the query matrix, the key matrix and the weight matrix are determined based on the second matrix read from the available storage space, including:
[0028] For each submatrix in the second matrix, a first maximum value and a second maximum value of each submatrix are determined respectively, wherein the submatrix is a product of each row of the query matrix and each of the fourth matrix cuts in the key matrix, the first maximum value of the b-th submatrix in the a-th row of the second matrix is the maximum value of each element in the first b-1 submatrices in the a-th row, and the second maximum value of the b-th submatrix in the a-th row of the second matrix is the maximum value of the first maximum value and the maximum value of each element in the b-th submatrix;
[0029] Perform an exponential power operation according to each element in each row of the second matrix and the first maximum value and the second maximum value of the submatrix where the element is located, to obtain an exponential power operation result corresponding to each element in each row of the second matrix;
[0030] Determine the cumulative sum of the exponential power operation results of each row respectively, and determine the inner product matrix of the query matrix and the key matrix based on the exponential power operation results corresponding to each element in the second matrix of each row;
[0031] Reading the weight matrix from the on-chip memory, and calculating a product matrix of the inner product matrix and the weight matrix;
[0032] Determine the ratio of the product matrix to the cumulative sum of the exponential power operation results of each row, the ratio being the operation result of the query matrix, the key matrix and the weight matrix.
[0033] In a feasible design, accessing the off-chip memory of the target chip to read the first matrix stored in the off-chip memory includes:
[0034] The first matrix after format conversion is read by accessing the off-chip memory of the target chip, wherein the data amount of the first matrix before format conversion is greater than the data amount of the first matrix after format conversion.
[0035] In a feasible design, the target chip includes a graphics processing unit GPU;
[0036] The GPU includes at least one computing unit and a local storage unit connected to the computing unit, and the on-chip memory includes the local storage unit;
[0037] Alternatively, the computing unit includes a private storage unit, and the on-chip memory includes the private storage unit.
[0038] In a second aspect, the present application provides an electronic device, comprising: a processor and a memory; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes any one of the methods described in the first aspect.
[0039] In a third aspect, the present application provides a computer storage medium, wherein the computer storage medium stores a computer program or instructions, and when the computer program or instructions are executed, the method described in the first aspect is executed.
[0040] In a fourth aspect, the present application provides a chip system, which includes a processor, wherein the processor is coupled to a memory and is used to execute a computer program or instruction stored in the memory. When the computer program or instruction is executed, the method described in the first aspect is executed.
[0041] The embodiment of the present application provides a neural network-based reasoning method, which can be executed by a target chip. In the method, at least part of the first matrix is stored in the on-chip memory of the target chip, and during the reasoning process, the Q matrix, the K matrix, and the V matrix are operated based on at least part of the first matrix stored in the on-chip memory and the available storage space of the on-chip memory.
[0042] The solution provided by the embodiment of the present application can read at least part of the first matrix stored therein from the on-chip memory when performing reasoning. Compared with the existing solution in which the Q matrix, the K matrix and the V matrix are all stored in the off-chip memory, the amount of data read from the off-chip memory can be reduced, that is, the number of global simulations can be reduced. The speed of reading data from the on-chip memory is faster than the speed of reading data from the off-chip memory. Therefore, the solution provided by the embodiment of the present application can reduce the time consumption in the reasoning process and improve the reasoning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1 This is a structural example diagram of a control network;
[0045] Figure 2 A schematic diagram of the operation flow of a Q matrix, a K matrix and a V matrix;
[0046] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0047] Figure 4 A software structure block diagram of an electronic device provided in an embodiment of the present application;
[0048] Figure 5 A schematic diagram of a workflow of a neural network-based reasoning method provided in an embodiment of the present application;
[0049] Figure 6 A schematic diagram of a workflow of another neural network-based reasoning method provided in an embodiment of the present application;
[0050] Figure 7 A schematic diagram of a workflow of another neural network-based reasoning method provided in an embodiment of the present application;
[0051] FIG8( a) is a schematic diagram of operations in a neural network-based reasoning method provided in an embodiment of the present application;
[0052] FIG8( b ) is a schematic diagram of operations in another neural network-based reasoning method provided in an embodiment of the present application;
[0053] FIG8( c ) is a schematic diagram of operations in another neural network-based reasoning method provided in an embodiment of the present application;
[0054] Fig. 9 A schematic diagram of a workflow of another neural network-based reasoning method provided in an embodiment of the present application;
[0055] Fig.10 A schematic diagram of the structure of a target chip provided in an embodiment of the present application;
[0056] Fig.11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0058] The terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to be used as limitations on the present application. As used in the specification and the appended claims of the present application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the following embodiments of the present application, "at least one", "one or more" refer to one, two or more. The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist; for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0059] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear at different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0060] In order to make the description of the following embodiments clear and concise, a brief introduction to the related technology is first given:
[0061] With the development of artificial intelligence technology, various types of neural networks have emerged, such as the control network ControlNet. In addition, some neural networks such as the ControlNet network include an attention module (i.e., an Attention module).
[0062] Take ControlNet as an example, see Figure 1 As shown in the example diagram, the ControlNet may include: an encoding module (i.e., an Encoder module), a solidification module (i.e., a Text embed), a diffusion model (i.e., a Diffusion model), a control model (i.e., a Control model), and a decoding module (i.e., a Decoder module). In the diffusion model and the control model, an attention module is provided.
[0063] Among them, the encoding module receives the input data, decodes it, and transmits the decoded data to the control model; the solidification module, the attention module and other modules in the control model, and the attention module and other modules in the diffusion model cooperate with each other to perform corresponding reasoning operations, and then transmit the obtained reasoning data to the decoding module; the decoding module decodes the received reasoning data, obtains the decoded data, and outputs the corresponding reasoning results.
[0064] ControlNet can be applied to a variety of scenarios. For example, if the user uses a telephoto lens to shoot, the image captured is sometimes blurry, and the user wants a clearer image, then ControlNet can be used to meet the user's needs. In this case, refer to Figure 1 The input data received by the encoding module may be a first image, which is a blurred image, and the inference result output by the decoding module includes a second image, which is a clear image.
[0065] In addition, during the reasoning process, the attention module in the neural network often needs to operate on the query matrix (i.e., query matrix, which can be referred to as Q matrix), the key matrix (i.e., key matrix, which can be referred to as K matrix) and the weight matrix (i.e., value matrix, which can be referred to as V matrix). Among them, the Q matrix, the K matrix and the V matrix are stored in the memory outside the chip where the operation is performed (i.e., off-chip memory), and the off-chip memory can also be called running memory. In the current technology, the off-chip memory storing the Q matrix, the K matrix and the V matrix is usually a double data rate synchronous dynamic random access memory (DDR SDRAM), and the graphics processing unit (GPU) chip operates on the Q matrix, the K matrix and the V matrix. Among them, the simulated storage operation performed by the GPU on the off-chip memory can be called global simulated storage, that is, the operation of the GPU accessing the DDR SDRAM to obtain the data stored in the DDR SDRAM, and the operation of the GPU storing data in the DDR SDRAM can be called global simulated storage.
[0066] Currently, the attention module operates on the Q matrix, K matrix, and V matrix. Figure 2 The workflow diagram shown generally includes the following steps:
[0067] The first step is to perform transposition operations on the Q matrix and the K matrix respectively to obtain the transposed result QT matrix of the Q matrix and the transposed result KT matrix of the K matrix.
[0068] Since the matrix multiplication operation (ie, MatMul operation) currently used in the operation adopts an outer product scheme, the current operation scheme needs to transpose the Q matrix and the K matrix before the matrix multiplication operation is performed.
[0069] Since both the Q matrix and the K matrix are stored in the off-chip memory, when performing the transposition operation in this step, the GPU first accesses the off-chip memory, reads the Q matrix and the K matrix from the off-chip memory, and performs the transposition operation, and then the GPU stores the transposition result in the off-chip memory.
[0070] The GPU usually reads the matrix from the off-chip memory row by row or column by column. Specifically, the transposition operation of the Q matrix and the K matrix includes the following steps:
[0071] When performing the transposition operation on the Q matrix, the GPU accesses the off-chip memory, reads a row of the Q matrix from the off-chip memory, transposes the read row of the Q matrix, and stores the transposed result in the off-chip memory; then reads another row of the Q matrix from the off-chip memory, transposes the read row of the Q matrix, and stores the transposed result in the off-chip memory, and so on, until the transposition of the Q matrix is completed.
[0072] When performing the transposition operation on the K matrix, the GPU accesses the off-chip memory, reads a column of the K matrix from the off-chip memory, transposes the read column of the K matrix, and stores the transposed result in the off-chip memory; then reads another column of the K matrix from the off-chip memory, transposes the read column of the K matrix, and stores the transposed result in the off-chip memory, and so on, until the transposition of the K matrix is completed.
[0073] The second step is to perform matrix multiplication operation (i.e., MatMul operation) on the QT matrix and the KT matrix to obtain the s matrix.
[0074] Specifically, when performing the operation of this step, the GPU first accesses the off-chip memory, reads the QT matrix and the KT matrix from the off-chip memory, then performs matrix multiplication on the QT matrix and the KT matrix, and then stores the result of the multiplication operation, the s matrix, in the off-chip memory.
[0075] In one example, the Q matrix, K matrix and V matrix can be float type data, and the shape can be 8*4096*4096, 8 refers to the dimension of the matrix, 4096*4096 means that the matrix has 4096 rows and 4096 columns. In this case, the s matrix is 4096*4096, that is, the s matrix has 4096 rows and 4096 columns.
[0076] The third step is to multiply the s matrix element by the scaling factor B, that is, to perform a multiplication operation (i.e., Mul operation) on the s matrix to obtain the sx matrix.
[0077] Specifically, when performing the operation of this step, the GPU first accesses the off-chip memory, reads the s matrix from the off-chip memory, and then multiplies the s matrix element by element by the scaling factor B to obtain the sx matrix, and then the GPU stores the sx matrix in the off-chip memory.
[0078] Among them, if the s matrix is 4096*4096, the sx matrix is also 4096*4096.
[0079] The fourth step is to perform normalization operation (i.e. softmax) on the sx matrix to obtain the p matrix.
[0080] Specifically, when performing the operation of this step, the GPU first accesses the off-chip memory, reads the sx matrix from the off-chip memory, and then normalizes the sx matrix to obtain the p matrix, and then the GPU stores the p matrix in the off-chip memory.
[0081] If the sx matrix is 4096*4096, then the p matrix is also 4096*4096.
[0082] The fifth step is to transpose the p matrix to obtain the pT matrix.
[0083] Specifically, in this step, the GPU first accesses the off-chip memory, reads the p matrix row by row from the off-chip memory, performs a transposition operation on each row of the p matrix, and then stores the result of the transposition operation on each row in the off-chip memory to obtain the pT matrix.
[0084] The sixth step is to perform matrix multiplication on the pT matrix and the V matrix to obtain the o matrix.
[0085] Specifically, in this step, the GPU first accesses the off-chip memory, reads the pT matrix and the V matrix from the off-chip memory, then performs a matrix multiplication operation to obtain the o matrix, and then stores the o matrix in the off-chip memory.
[0086] According to the above introduction to the current calculation method, during the calculation process, the GPU needs to perform multiple simulated storage operations on the off-chip memory, that is, multiple global simulated storage operations are required. Each global access takes a long time, so the current process of calculating the Q matrix, K matrix and V matrix takes a long time, which further leads to a long time for the neural network reasoning process and low reasoning efficiency.
[0087] In order to solve the problem that the prior art consumes a long time to calculate the Q matrix, the K matrix and the V matrix, resulting in low efficiency in neural network reasoning, the embodiments of the present application provide a neural network-based reasoning method, electronic device and corresponding device.
[0088] The neural network-based reasoning method provided in the embodiment of the present application can be applied to an electronic device that supports running a neural network. During the reasoning process, the neural network needs to perform operations on the Q matrix, the K matrix, and the V matrix, and can obtain corresponding reasoning results through reasoning.
[0089] The following embodiments only take the electronic device as a laptop computer as an example to illustrate the method provided in the embodiments of the present application.
[0090] Figure 3 A schematic diagram of the hardware structure of an electronic device 100 provided in an embodiment of the present application.
[0091] like Figure 3As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, a wireless communication module 150, a display screen 160, etc.
[0092] It is to be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0093] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (application processor, AP), a modem processor, a graphics processor (graphics processing unit, GPU), an image signal processor (image signal processor, ISP), a controller, a memory, a video codec, a digital signal processor (digital signal processor, DSP), a baseband processor, and / or a neural network processor (neural-network processing unit, NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0094] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0095] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0096] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an I2C interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, etc.
[0097] It is understandable that the interface connection relationship between the modules shown in this embodiment is only a schematic illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0098] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. While the charging management module 140 charges the battery 142, it can also power the electronic device through the power management module 141.
[0099] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display screen 160, and the wireless communication module 150. In some embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0100] The wireless communication module 150 can provide wireless communication solutions including WLAN (such as Wi-Fi), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. For example, in the embodiment of the present application, the electronic device 100 can establish a Bluetooth connection with a terminal device (such as a wireless headset 100) through the wireless communication module 150.
[0101] The wireless communication module 150 may be one or more devices integrating at least one communication processing module. The wireless communication module 150 receives electromagnetic waves via an antenna, modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 150 may also receive a signal to be sent from the processor 110, modulate the frequency of the signal, amplify the signal, and convert it into electromagnetic waves for radiation via the antenna.
[0102] The electronic device 100 implements the display function through a GPU, a display screen 160, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 160 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.
[0103] The display screen 160 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc.
[0104] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement a data storage function.
[0105] The internal memory 121 may be used to store computer executable program codes, wherein the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 .
[0106] The operating system of the electronic device may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture or a cloud architecture, etc. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device.
[0107] Figure 4 A software architecture diagram of an electronic device 100 provided in an embodiment of the present application.
[0108] like Figure 4 As shown in Figure 1, the layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces.
[0109] In some embodiments, the Android system is divided into an application layer (i.e., APP layer), an application framework layer (i.e., FWK layer), a hardware abstraction layer (HAL) and a driver layer (i.e., Driver layer) from top to bottom, and the Driver layer is connected to the hardware layer (i.e., Hardware layer).
[0110] The application layer (i.e., APP layer) may include a series of application packages, including camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message and other applications.
[0111] The camera can take pictures. If the pictures taken include blurred images, the electronic device can obtain a clear image corresponding to the blurred image by running a neural network. The clear image can be regarded as the inference result of the neural network. In the process of running a neural network, it is often necessary to operate the Q matrix, the K matrix and the V matrix.
[0112] The application framework layer (i.e., FWK layer) provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0113] Among them, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0114] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0115] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0116] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying images.
[0117] The phone manager is used to provide communication functions for electronic devices, such as the management of call status (including answering, hanging up, etc.).
[0118] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0119] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of applications running in the background, or a notification that appears on the screen in the form of a dialog window. For example, a text message is displayed in the status bar, a prompt sound is emitted, an electronic device vibrates, an indicator light flashes, etc.
[0120] HAL is an abstract interface for device kernel drivers, which implements the application program interface that provides access to the underlying device to the higher-level Java API framework. The hardware abstraction layer can include multiple library modules, such as display module, audio module, Bluetooth module, Wi-Fi module, etc. Each module can implement an interface for a specific type of hardware component. When the framework API requires access to the device hardware, the Android system will load the library module for the hardware component.
[0121] See also Figure 4 In an embodiment of the present application, the HAL layer may further include an operation module, which may execute the method provided in the embodiment of the present application to realize operations on the Q matrix, the K matrix and the V matrix.
[0122] The driver layer is the layer between hardware and software. The driver layer may include display driver, camera driver, audio driver, sensor driver, etc.
[0123] Below the driver layer may be a hardware layer (ie, Hardware layer), which may include a touch panel TP and a display screen, and the display screen may include a liquid crystal display (LCD) and the like.
[0124] In the above, the software system of the electronic device adopts the layered architecture of the Android system as an example. Of course, the software system of the electronic device can also adopt other architectures, which is not limited in this application.
[0125] In order to clarify the solution provided by the present application, the solution provided by the present application is introduced and explained through various embodiments in conjunction with the accompanying drawings.
[0126] In order to solve the problem that the prior art consumes a long time to calculate the Q matrix, the K matrix and the V matrix, resulting in low efficiency in neural network reasoning, the embodiments of the present application provide a neural network-based reasoning method, electronic device and corresponding device.
[0127] The neural network-based reasoning method provided in the embodiment of the present application is applied to a target chip, which can run a neural network and obtain reasoning results based on the neural network. Figure 5 , the method may include the following steps:
[0128] Step S11, accessing the off-chip memory of the target chip to read a first matrix stored in the off-chip memory, where the first matrix is at least one of a Q matrix, a K matrix, and a V matrix.
[0129] Among them, the Q matrix is the query matrix, which can also be called the query matrix; the K matrix is the key matrix, which can also be called the key matrix; the V matrix is the value matrix, which can also be called the weight matrix.
[0130] In addition, the first matrix is any one of the Q matrix, K matrix and V matrix, or any two matrices, or the first matrix may also include the Q matrix, K matrix and V matrix at the same time, which is not limited in the embodiment of the present application.
[0131] The Q matrix, K matrix and V matrix are usually stored in the off-chip memory of the target chip, so in this step, the first matrix can be read by accessing the off-chip memory of the target chip. The off-chip memory usually refers to the memory outside the chip, and the on-chip memory usually refers to the memory integrated inside the chip.
[0132] Step S12: based on the data amount of the first matrix and the storage capacity of the on-chip memory of the target chip, storing at least a portion of the first matrix from the off-chip memory to the on-chip memory.
[0133] The target chip performs emulation storage to the on-chip memory faster than the off-chip memory. Therefore, in the embodiment of the present application, if the storage capacity of the on-chip memory of the target chip allows, the Q matrix, the K matrix and the V matrix can all be stored in the on-chip memory. Therefore, in a feasible implementation of step S12, if the storage capacity of the first on-chip memory in the on-chip memory is not less than the sum of the data amounts of the Q matrix, the K matrix and the V matrix, the Q matrix, the K matrix and the V matrix are stored from the off-chip memory to the first on-chip memory.
[0134] In this implementation, the Q matrix, K matrix, and V matrix are all completely stored in the on-chip memory of the target chip. In this case, in the process of operating the Q matrix, K matrix, and V matrix, the Q matrix, K matrix, and V matrix in the on-chip memory can be read by accessing the first on-chip memory, thereby reducing the operation of accessing the off-chip memory during the reasoning process. Accordingly, in this case, the first matrix includes the Q matrix, the K matrix, and the V matrix.
[0135] Alternatively, in another feasible implementation of this step, if the on-chip memory includes multiple second on-chip memories, and the capacity of the on-chip memory is smaller than the amount of data in the first matrix, at least part of the matrix blocks split into the first matrix are stored in corresponding second on-chip memories.
[0136] In this implementation, since the storage capacity of the second on-chip memory is smaller than the data volume of the first matrix, the complete first matrix cannot be stored in a second on-chip memory. In this case, the first matrix can be split to obtain multiple matrix blocks, and different matrices can be switched and stored in different second on-chip memories, wherein the data volume of each matrix block is no larger than the storage space of the second on-chip memory storing the matrix block.
[0137] In addition, if the first matrix is stored in the second on-chip memory in the form of matrix blocks, then in the embodiment of the present application, the target chip will adjust the algorithm used when operating the Q matrix, the K matrix, and the V matrix. In order to facilitate the target chip to operate the Q matrix, the K matrix, and the V matrix through the adjusted algorithm, if the matrix blocks into which at least part of the first matrix is split are stored in the corresponding second on-chip memory, the following method can be used:
[0138] In which, if the first matrix includes a Q matrix, each first matrix block into which at least part of the Q matrix is split is at least one row in the Q matrix, that is, the Q matrix can be split into multiple first matrix blocks by row, and each first matrix block is one or more rows in the Q matrix.
[0139] If the first matrix includes a K matrix, each second matrix block into which at least part of the K matrix is split is at least one column in the K matrix, that is, the K matrix can be split into multiple first matrix blocks by column, and each second matrix block is one or more columns in the K matrix.
[0140] If the first matrix includes a V matrix, each third matrix block into which at least part of the V matrix is split is at least one row in the V matrix, that is, the V matrix can be split into multiple third matrix blocks by row, and each third matrix block is one or more rows in the V matrix.
[0141] Step S13: Based on at least a portion of the first matrix stored in the on-chip memory and the available storage space of the on-chip memory, operate on the Q matrix, the K matrix and the V matrix.
[0142] In the process of performing calculations in step S13, if at least part of the first matrix is needed, at least part of the first matrix stored in the on-chip memory can be read and calculated by accessing the on-chip memory.
[0143] In addition, the intermediate calculation results generated during the calculation process can be stored in the available storage space of the on-chip memory. If the subsequent calculation process needs to apply the intermediate calculation results, the intermediate calculation results can be read by accessing the available storage space of the on-chip memory, and the calculation can be performed using the read intermediate calculation results.
[0144] The embodiment of the present application provides a neural network-based reasoning method, which can be executed by a target chip. In the method, at least part of the first matrix is stored in the on-chip memory of the target chip, and during the reasoning process, the Q matrix, the K matrix, and the V matrix are operated based on at least part of the first matrix stored in the on-chip memory and the available storage space of the on-chip memory.
[0145] The solution provided by the embodiment of the present application can read at least part of the first matrix stored therein from the on-chip memory when performing reasoning. Compared with the existing solution in which the Q matrix, the K matrix and the V matrix are all stored in the off-chip memory, the amount of data read from the off-chip memory can be reduced, that is, the number of global simulations can be reduced. The speed of reading data from the on-chip memory is faster than the speed of reading data from the off-chip memory. Therefore, the solution provided by the embodiment of the present application can reduce the time consumption in the reasoning process and improve the reasoning efficiency.
[0146] Further, in the solution provided in the embodiment of the present application, the Q matrix, the K matrix, and the V matrix can be operated based on at least part of the first matrix stored in the on-chip memory and the available storage space of the on-chip memory. In this case, the intermediate operation results during the operation process can be stored in the available storage space, and if the intermediate operation results are required for subsequent operations, the intermediate operation results can be read from the available storage space, and the operation can be performed based on the read intermediate operation results.
[0147] In the prior art, each time the target chip obtains an intermediate operation result through calculation, it stores it in an off-chip memory. If a subsequent operation requires the intermediate operation result, the target chip accesses the off-chip memory to read the intermediate operation result.
[0148] Therefore, compared with the prior art, the solution provided in the embodiment of the present application can also reduce the number of times intermediate calculation results are stored in the off-chip memory and the number of times intermediate calculation results are read from the off-chip memory, that is, the number of global simulations is further reduced. Accordingly, the time consumed in the reasoning process can be further reduced and the reasoning efficiency can be improved.
[0149] In step S13 provided in the embodiment of the present application, an operation of operating the Q matrix, the K matrix and the V matrix based on at least a portion of the first matrix stored in the on-chip memory and the available storage space of the on-chip memory is disclosed. Figure 6 As shown in the workflow diagram, if the first matrix includes a Q matrix and a K matrix, and the complete part of the Q matrix and the complete part of the K matrix are both stored in the on-chip memory of the target chip, the operation can be implemented by the following steps:
[0150] Step S131, read the Q matrix and the K matrix from the on-chip memory, and perform a product operation on the read Q matrix and K matrix to obtain a second matrix, wherein the product matrix is an intermediate operation result.
[0151] Step S132: store the second matrix into available storage space of the on-chip memory.
[0152] Step S133: Determine the operation results of the Q matrix, the K matrix and the V matrix based on the second matrix read from the available storage space.
[0153] In the process of reasoning through the scheme provided by the above embodiment, the complete parts of the Q matrix and the K matrix are stored in the on-chip memory. The Q matrix and the K matrix can be read by accessing the on-chip memory so as to perform the product operation on the Q matrix and the K matrix without accessing the off-chip memory, thereby reducing the number of global simulations in the reasoning process, reducing the time consumption of the reasoning process, and improving the reasoning efficiency.
[0154] In addition, the solution of the above embodiment stores the second matrix in the available storage space of the on-chip memory, and reads the second matrix from the available storage space when the second matrix is needed subsequently, thereby further reducing the number of global simulations, reducing the time consumption in the reasoning process, and improving the reasoning efficiency.
[0155] The above embodiment assumes that the complete parts of the Q matrix and the K matrix are stored in the on-chip memory. In the actual operation process, only part of the Q matrix or only part of the K matrix may be stored in the on-chip memory.
[0156] Among them, if only part of the Q matrix is stored in the on-chip memory, then during the calculation process, the part of the Q matrix stored in the on-chip memory can be read, and the other part of the Q matrix can be read from the off-chip memory, and the K matrix can be read from the off-chip memory, and the product operation of the Q matrix and the K matrix can be performed.
[0157] In addition, if only part of the K matrix is stored in the on-chip memory, during the calculation process, the part of the K matrix stored in the on-chip memory can be read, and the other part of the K matrix can be read from the off-chip memory, and the Q matrix can be read from the off-chip memory, and the product operation of the Q matrix and the K matrix can be performed.
[0158] Alternatively, if part of the Q matrix and part of the K matrix are stored in the on-chip memory, during the calculation process, the part of the Q matrix and the part of the K matrix stored in the on-chip memory can be read, the other part of the Q matrix and the other part of the K matrix can be read from the off-chip memory, and the product operation is performed on the Q matrix and the K matrix.
[0159] In order to clarify the method of performing the product operation in step S131 of the embodiment of the present application, the embodiment of the present application provides Figure 7 See also Figure 7 The operation of reading the Q matrix and the K matrix from the on-chip memory and performing a product operation on the read Q matrix and the K matrix may include the following steps:
[0160] Step S1311, read the Q matrix row by row from the on-chip memory, and read the fourth matrix blocks included in the K matrix one by one from the on-chip memory.
[0161] Each fourth matrix block includes at least one column of the K matrix. Different fourth matrix blocks include different columns of the K matrix, and all fourth matrix blocks constitute the entire K matrix. Usually, different fourth matrix blocks have the same number of columns.
[0162] Among them, if the Q matrix is stored in the on-chip memory in a block-like manner, and each first matrix block of the Q matrix is at least one row in the Q matrix, the Q matrix can be read row by row by reading the first matrix blocks one by one.
[0163] For example, if the Q matrix is split into multiple first matrix blocks, each first matrix block includes at least one row of the Q matrix, and each first matrix block is stored in a corresponding second on-chip memory, then the first matrix blocks can be read one by one by accessing the second on-chip memory one by one.
[0164] In addition, when determining the product matrix, the embodiment of the present application cuts the K matrix into blocks according to columns to obtain n fourth matrix blocks, where n is a positive integer. If the K matrix is cut into multiple second matrix blocks and stored in the on-chip memory of the target chip, the fourth matrix block may be the same as the second matrix block or different, and the embodiment of the present application does not limit this.
[0165] Among them, if the K matrix is split into multiple second matrix blocks, each second matrix block includes at least one column in the K matrix, and each second matrix block is stored in the corresponding second on-chip memory respectively, then the fourth matrix blocks included in the K matrix can be read one by one by accessing the second on-chip memory one by one.
[0166] Step S1312: Calculate the product of each row of the Q matrix and each fourth matrix block included in the K matrix respectively, and the matrix formed by the products is the second matrix.
[0167] That is to say, the second matrix is obtained by multiplying the Q matrix and the K matrix, and each element in the second matrix is the product of each row of the Q matrix and each fourth matrix block.
[0168] See the example diagram of the product operation shown in Figure 8(a), where Q1 represents the first row in the Q matrix, Q2 represents the second row in the Q matrix, and Q3 represents the third row in the Q matrix.
[0169] K1 represents the first fourth matrix block in the K matrix, and K2 represents the second fourth matrix block in the K matrix. In the process of performing product operations on the Q matrix and the K matrix, it is necessary to calculate the product S11 of Q1 and K1, the product S12 of Q1 and K2, the product S21 of Q2 and K1, the product S22 of Q2 and K2, the product S31 of Q3 and K1, and the product S32 of Q3 and K2. The matrix formed by the products (i.e., the matrix formed by S11, S12, S21, S22, S31, and S32) is the second matrix. Among them, each product is a row vector with the same width as the fourth matrix block, that is, each product (S11, S12, S21, S22, S31, and S32 can be regarded as a row matrix, and the number of columns is the same as the number of columns of the corresponding fourth matrix block).
[0170] In addition, according to the product operation process provided above and Figure 8(a), it can be known that when the Q matrix and the K matrix are multiplied, the Q matrix read row by row and the fourth matrix blocks split by columns of the K matrix are multiplied respectively. In this case, if the Q matrix and the K matrix are stored in the on-chip memory in a block-like manner, the Q matrix is split into at least one first matrix block, each of which is one or more rows in the Q matrix, and the K matrix is split into at least one second matrix block, each of which is one or more columns in the K matrix. This makes it easier for the target chip to read the Q matrix in the on-chip memory row by row, and for the target chip to read the fourth matrix block, thereby further improving the computing speed of the target chip.
[0171] For example, if the second matrix block is the same as the fourth matrix block, and different second matrix blocks are stored in different second on-chip memories, the corresponding fourth matrix block can be read by accessing each second on-chip memory.
[0172] See also Fig. 9 In the flowchart shown, if the first matrix also includes a V matrix, and the complete portion of the V matrix is stored in the on-chip memory, based on the second matrix read from the available storage space, determining the operation results of the Q matrix, the K matrix, and the V matrix includes:
[0173] Step S1313: for each sub-matrix in the second matrix, determine the first maximum value and the second maximum value of each sub-matrix respectively.
[0174] The submatrix is the product of each row of the Q matrix and each fourth matrix block in the K matrix. Referring to the example of FIG8(a), the submatrix of the first row of the product matrix may include S11 and S12.
[0175] The first maximum value of the b-th submatrix in the a-th row of the second matrix is the maximum value of each element in the first b-1 submatrices of the a-th row, and the second maximum value of the b-th submatrix in the a-th row of the second matrix is the maximum value between the first maximum value and the maximum value of each element in the b-th submatrix, and a and b are both positive integers.
[0176] That is, in the embodiment of the present application, the first maximum value of a submatrix in a row of the second matrix may be the second maximum value of a submatrix before the row. And, after determining the first maximum value of the submatrix, the first maximum value may be compared with each element in the submatrix, and the maximum value thereof is the second maximum value of the submatrix.
[0177] In a feasible implementation method of the embodiment of the present application, when determining the first maximum value and the second maximum value in a submatrix in a row of the second matrix, the maximum value of each element in the first submatrix of the second matrix of the row can be first determined, and the maximum value is the first maximum value and the second maximum value of the first submatrix; then, the second maximum value of the first submatrix of the row is used as the first maximum value of the second submatrix of the row, and the maximum value of each element in the second submatrix is determined, and the maximum value of the maximum value of each element in the second submatrix and the first maximum value of the second submatrix is the second maximum value of the second submatrix.
[0178] After determining the first maximum value and the second maximum value of each submatrix in the second matrix, they can be stored in the available storage space of the on-chip memory, and when necessary during the calculation process, the first maximum value and the second maximum value of each submatrix can be read by accessing the available storage space.
[0179] Step S1314: perform an exponential power operation according to each element in each row of the second matrix and the first maximum value and the second maximum value of the submatrix where the element is located, to obtain an exponential power operation result corresponding to each element in each row of the second matrix.
[0180] This step can be achieved by the following formula:
[0181]
[0182] In formula (2), f(x) is the result of the exponential power operation corresponding to each element in a row of the second matrix, x (1) represents the first element in the second matrix of this row, x (2) represents the second element in the second matrix of this row, x (m) represents the mth element in the second matrix of this row, where m is the number of elements in the second matrix of this row, and m is a positive integer. old1 Indicates the first maximum value of the submatrix where the first element in the second matrix of this row is located, mnew1 Indicates the second largest value of the submatrix where the first element in the second matrix of this row is located, m old2 Indicates the first maximum value of the submatrix where the second element in the second matrix of this row is located, m new2 Indicates the second maximum value of the submatrix where the second element in the second matrix of this row is located, m oldm Indicates the first maximum value of the submatrix where the mth element in the second matrix of this row is located, m newm Indicates the second maximum value of the submatrix where the mth element in the second matrix of this row is located.
[0183] In addition, after the exponential power operation result corresponding to each row element is determined, it can be stored in the available storage space of the on-chip memory, and when necessary, the exponential power operation result can be obtained by accessing the available storage space.
[0184] Step S1315, respectively determine the cumulative sum of the exponential power operation results of each row.
[0185] The accumulated sum of the results of the exponential power operations on each row may form an r-row column vector, where r is the number of rows of the second matrix.
[0186] This can be achieved by the following formula:
[0187]
[0188] In formula (2), l(x) represents the cumulative sum of the results of exponential power operations on a row, and f(x) i represents the result of the ith exponential power operation of the row. For example, referring to formula (1), f(x) 1 for
[0189] In addition, after the cumulative sum of the exponential power operation results of each row is determined, it can be stored in the available storage space of the on-chip memory, and when needed, the cumulative sum stored in the available storage space can be read by accessing the available storage space.
[0190] Step S1316: Determine the inner product matrix of the Q matrix and the K matrix based on the exponential power operation results corresponding to each element in each row of the second matrix.
[0191] This can be achieved by the following formula:
[0192] P:=[f 1 (x),f 2 (x),…,f s (x)] Formula (3).
[0193] In the above formula, P represents the inner product matrix of Q matrix and K matrix; f 1(x) represents the exponential power operation result corresponding to each element in the first row of the second matrix, f 2 (x) represents the exponential power operation result corresponding to each element in the second row of the second matrix, f s (x) represents the exponential power operation result corresponding to each element in the s-th row of the second matrix, and s is the number of rows of the second matrix.
[0194] Among them, Figure 8(b) shows that through the operation of the embodiment of the present application, the inner product matrix P can be obtained from the second matrix, and the matrix P is composed of P11, P12, P21, P22, P31 and P32.
[0195] In addition, after the inner product matrix is determined, it can be stored in an available storage space of an on-chip memory, and when necessary, the inner product matrix stored in the available storage space can be read by accessing the available storage space.
[0196] In the embodiment of the present application, there is no strict requirement for the execution order of step S1315 and step S1316. The operation of step S1316 can be executed first, and then the operation of step S1315, or the operations of step S1315 and step S1316 can be executed at the same time. The embodiment of the present application does not limit this.
[0197] Step S1317, read the weight matrix from the on-chip memory, and calculate the product matrix of the inner product matrix and the weight matrix.
[0198] When calculating the product matrix of the inner product matrix and the weight matrix, the inner product matrix and the weight matrix may be multiplied row by row. In this case, if the V matrix is stored in the on-chip memory in a block-by-block manner, the V matrix is split into at least one third matrix block, and each third matrix block is one or more rows in the V matrix, so that the target chip can read the V matrix in the on-chip memory row by row, which can further improve the computing speed of the target chip.
[0199] This can be achieved by the following formula:
[0200] O: = P*V formula (4).
[0201] Among them, O represents the product matrix of the inner product matrix and the weight matrix; V represents the weight matrix.
[0202] Referring to FIG8(c), by operating the matrix P and the matrix V, the matrix O can be obtained.
[0203] After the product of the inner product matrix and the weight matrix is determined, it can be stored in the available storage space of the on-chip memory, and when needed, the product stored in the available storage space can be read by accessing the available storage space.
[0204] In addition, if the V matrix is not stored in the on-chip memory, step S1317 can be adjusted to read the V matrix from the off-chip memory and calculate the product matrix of the inner product matrix and the weight matrix.
[0205] Alternatively, if only part of the V matrix is stored in the on-chip memory, step S1317 can be adjusted to read part of the V matrix from the on-chip memory and read the other part of the V matrix from the off-chip memory, and calculate the product matrix of the inner product matrix and the weight matrix.
[0206] Step S1318, determine the ratio of the product matrix to the cumulative sum of the exponential power operation results of each row, and the ratio is the operation result of the Q matrix, the K matrix and the V matrix.
[0207] This can be achieved by the following formula:
[0208] O': =O / l(x) formula (5).
[0209] In formula (6), O' represents the ratio, that is, the operation result of the Q matrix, the K matrix and the V matrix.
[0210] Through the above operations, the operation results of the Q matrix, the K matrix and the V matrix can be obtained. In addition, during the operation, the Q matrix, the K matrix and the V matrix can be obtained by accessing the on-chip memory, thereby reducing the number of global simulations, correspondingly reducing the time consumption in the reasoning process, and improving the reasoning efficiency.
[0211] Furthermore, during the above-mentioned calculation process, the intermediate calculation results are stored in the available storage space of the on-chip memory. When the intermediate calculation results need to be applied, the intermediate calculation results can be read by accessing the available storage space, thereby further reducing the number of global simulations and correspondingly further reducing the time consumed in the reasoning process.
[0212] In step S11 of the embodiment of the present application, an operation of accessing the off-chip memory of the target chip to read the first matrix stored in the off-chip memory is provided. In a feasible implementation, the operation can be implemented by the following implementation:
[0213] The first matrix after format conversion is read by accessing the off-chip memory of the target chip, and the data amount of the first matrix before the format conversion is greater than the data amount of the first matrix after the format conversion.
[0214] In this implementation, before the target chip accesses the off-chip memory, the first matrix in the off-chip memory has been format-converted, and the data volume of the first matrix before the format conversion is greater than the data volume of the first matrix after the format conversion. In this case, the target chip reads the first matrix after the format conversion, and accordingly, during the operation, the target chip also performs operations using the first matrix after the format conversion. Since the data volume of the first matrix before the format conversion is greater than the data volume of the first matrix after the format conversion, the data volume of the first matrix can be reduced after the format conversion, further improving the reasoning efficiency.
[0215] Exemplarily, if the target chip is a GPU, before the target chip reads the first matrix in the off-chip memory, the central processing unit (CPU) may perform format conversion on the first matrix in the off-chip memory.
[0216] In addition, if the first matrix before conversion is a single-precision floating point number, the first matrix after conversion can be a half-precision floating point number; if the first matrix before conversion is a single-precision floating point number, the first matrix after conversion can be a single-precision floating point number or a half-precision floating point number.
[0217] In the solution provided in the embodiment of the present application, the target chip is a chip that can perform reasoning through a neural network, and during the reasoning process, it is necessary to perform operations on the Q matrix, K matrix and V matrix.
[0218] In a feasible implementation, the target chip includes a GPU, see Fig.10 As shown in the structural diagram of the GPU, the GPU includes at least one computing unit (i.e., Compute Unit) and a local storage unit (i.e., Local Memory) connected to the computing unit. In this case, the on-chip memory may include the local storage unit.
[0219] If the on-chip memory includes a local storage unit, the local storage unit can be used as the second on-chip memory, and each matrix block of a part of the first matrix can be stored in the corresponding local storage unit.
[0220] Alternatively, if the storage space of a local storage unit is not less than the sum of the data amounts of the Q matrix, the K matrix and the V matrix, the local storage unit can also be used as the first on-chip memory to store the Q matrix, the K matrix and the V matrix in the local storage unit.
[0221] In another feasible implementation, the computing unit of the GPU may include a private storage unit (ie, PrivateMemory). In this case, the on-chip memory may include a private storage unit. Fig.10, each computing unit may also include a processing element (Process Element, PE) connected to a private storage unit.
[0222] If the on-chip memory includes a private storage unit, the private storage unit can be used as the second on-chip memory, and each matrix block of a part of the first matrix can be stored in the corresponding private storage unit.
[0223] Alternatively, if the storage space of a private storage unit is not less than the sum of the data amounts of the Q matrix, the K matrix and the V matrix, the private storage unit can also be used as the first on-chip memory to store the Q matrix, the K matrix and the V matrix in the private storage unit.
[0224] Alternatively, in a feasible implementation, the GPU includes both a local storage unit and a private storage unit. In this case, the on-chip memory may include a storage unit with a larger storage space.
[0225] exist Fig.10 In the example shown, N computing units are included, each computing unit includes M private storage units and M processing elements, and each computing unit has a local storage unit connected thereto.
[0226] In addition, if the target chip includes a GPU, in a feasible implementation, the off-chip memory may be a double data rate synchronous dynamic random access memory (DDR). Fig.10 The GPU can read the first matrix by accessing the DDR and store at least part of the first matrix in the on-chip memory.
[0227] The various method embodiments described in this document may be independent solutions or may be combined according to internal logic, and these solutions all fall within the protection scope of this application.
[0228] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0229] Accordingly, the present application embodiment discloses an electronic device, see Fig.11 As shown in the structural schematic diagram, the electronic device comprises:
[0230] Processor 1101 and memory,
[0231] The memory is used to store program instructions;
[0232] The processor 1101 is used to call and execute the program instructions stored in the memory. When the program instructions stored in the memory are executed by the processor 1101, the electronic device executes Figures 5 to 7 ,as well as Fig. 9 All or part of the steps in the corresponding embodiments.
[0233] Furthermore, the electronic device may also include: a transceiver 1102 and a bus 1103 , and the memory includes a random access memory 1104 and a read-only memory 1105 .
[0234] The processor is coupled to the transceiver, random access memory and read-only memory through a bus. When the electronic device needs to be operated, the basic input and output system solidified in the read-only memory or the bootloader boot system in the embedded system is used to start the electronic device and guide it into a normal operating state. After the electronic device enters the normal operating state, the application program and the operating system are run in the random access memory, so that the electronic device executes Figures 5 to 7 ,as well as Fig. 9 All or part of the steps in the corresponding embodiments.
[0235] The electronic device according to the embodiment of the present invention may correspond to the above Figures 5 to 7 ,as well as Fig. 9 The electronic device in the corresponding embodiment, and the processor and storage in the electronic device can implement Figures 5 to 7 ,as well as Fig. 9 For the sake of brevity, the functions of the electronic device in the corresponding embodiment and / or the various steps and methods implemented are not described in detail here.
[0236] In a specific implementation, the embodiment of the present application further provides a computer storage medium, wherein the computer storage medium stores a computer program or instruction, and when the computer program or instruction is executed, the computer can implement the following steps: Figures 5 to 7 ,as well as Fig. 9 All or part of the steps in the corresponding embodiments. The computer-readable storage medium is set in any device, and the arbitrary device may be a random access memory (RAM), and the memory may also include a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); the memory may also include a combination of the above-mentioned types of memory, etc.
[0237] The present application also provides a chip system, which includes a processor coupled to a memory and configured to execute a computer program or instruction stored in the memory. When the computer program or instruction is executed, the chip system can implement the following steps: Figures 5 to 7 ,as well as Fig. 9 All or part of the steps in the corresponding embodiments. The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0238] In a feasible implementation, the chip system includes a target chip, the target chip is used to implement Figures 5 to 7 ,as well as Fig. 9 All or part of the steps in the corresponding embodiments.
[0239] Exemplarily, the target chip may be a GPU, or the target chip may be other chips that can perform operations on the Q matrix, the K matrix, and the V matrix, which is not limited in the embodiments of the present application.
[0240] The steps of the method or algorithm described in the embodiments of the present application can be directly embedded in hardware, a software unit executed by a processor, or a combination of the two. The software unit can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a register, a hard disk, a removable disk, a portable compact disc read-only memory (CD-ROM), or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and can write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be arranged in an ASIC, and the ASIC can be arranged in a user terminal (UE). Optionally, the processor and the storage medium can also be arranged in different components in the UE.
[0241] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium, or a semiconductor medium (e.g., a solid state drive (SSD)), etc.
[0242] The same or similar parts between the various embodiments of this specification can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment part.
[0243] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention or some parts of the embodiments.
[0244] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the embodiment of the road constraint determination device disclosed in this application, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.
[0245] The above-described embodiments of the present invention do not limit the protection scope of the present invention.
Claims
1. A neural network-based reasoning method, It is characterized in that Applied to a target chip, the method comprises: By accessing the off-chip memory of the target chip, a first matrix stored in the off-chip memory is read, where the first matrix is at least one of a query matrix, a key matrix, and a weight matrix; Based on the data amount of the first matrix and the storage capacity of the on-chip memory of the target chip, storing at least part of the first matrix from the off-chip memory to the on-chip memory; Based on at least a portion of the first matrix stored in the on-chip memory and available storage space of the on-chip memory, operating the query matrix, the key matrix, and the weight matrix; If the first matrix includes the query matrix, the key matrix, and the weight matrix, and a complete portion of the query matrix, a complete portion of the key matrix, and a complete portion of the weight matrix are all stored in the on-chip memory, the operation on the query matrix, the key matrix, and the weight matrix includes: For each submatrix in the second matrix, a first maximum value and a second maximum value of each submatrix are determined respectively, the second matrix is a matrix formed by the product of each row of the query matrix and each fourth matrix block, each fourth matrix block includes at least one column in the key matrix, the submatrix is the product of each row of the query matrix and each fourth matrix block in the key matrix, the first maximum value of the b-th submatrix in the a-th row of the second matrix is the maximum value of each element in the first b-1 submatrices in the a-th row, and the second maximum value of the b-th submatrix in the a-th row of the second matrix is the maximum value of the first maximum value and the maximum value of each element in the b-th submatrix; Perform an exponential power operation according to each element in each row of the second matrix and the first maximum value and the second maximum value of the submatrix where the element is located, to obtain an exponential power operation result corresponding to each element in each row of the second matrix; Determine the cumulative sum of the exponential power operation results of each row respectively, and determine the inner product matrix of the query matrix and the key matrix based on the exponential power operation results corresponding to each element in the second matrix of each row; Reading the weight matrix from the on-chip memory, and calculating a product matrix of the inner product matrix and the weight matrix; Determine the ratio of the product matrix to the cumulative sum of the exponential power operation results of each row, the ratio being the operation result of the query matrix, the key matrix and the weight matrix.
2. The method according to claim 1, It is characterized in that Based on the data amount of the first matrix and the storage capacity of the on-chip memory of the target chip, storing at least part of the first matrix from the off-chip memory to the on-chip memory includes: If the storage capacity of a first on-chip memory in the on-chip memory is not less than the sum of the data amounts of the query matrix, the key matrix and the weight matrix, storing the query matrix, the key matrix and the weight matrix from the off-chip memory to the first on-chip memory; or, If the on-chip memory includes a plurality of second on-chip memories, and the capacity of the on-chip memory is smaller than the data volume of the first matrix, at least part of the matrix blocks split into the first matrix are stored in the corresponding second on-chip memories.
3. The method according to claim 2, It is characterized in that If the first matrix includes the query matrix, each first matrix block into which at least a portion of the query matrix is split is at least one row in the query matrix; If the first matrix includes the key matrix, each second matrix block into which at least a portion of the key matrix is split is at least one column of the key matrix; If the first matrix includes the weight matrix, each third matrix block into which at least a portion of the weight matrix is split is at least one row in the weight matrix.
4. The method according to any one of claims 1 to 3, It is characterized in that If the first matrix includes the query matrix and the key matrix, and the complete part of the query matrix and the complete part of the key matrix are stored in the on-chip memory, the operation on the query matrix, the key matrix and the weight matrix based on at least a part of the first matrix stored in the on-chip memory and the available storage space of the on-chip memory includes: Reading the query matrix and the key matrix from the on-chip memory, and performing a product operation on the read query matrix and the key matrix to obtain the second matrix; Storing the second matrix in available storage space of the on-chip memory; Based on the second matrix read from the available storage space, operation results of the query matrix, the key matrix and the weight matrix are determined.
5. The method according to claim 4, It is characterized in that The step of reading the query matrix and the key matrix from the on-chip memory and performing a product operation on the read query matrix and the key matrix comprises: Reading the query matrix row by row from the on-chip memory, and reading the fourth matrix slices included in the key matrix one by one from the on-chip memory; The products of each row of the query matrix and each of the fourth matrix blocks are calculated respectively, and the matrix formed by the products is the second matrix.
6. The method according to any one of claims 1 to 3, It is characterized in that The step of accessing the off-chip memory of the target chip to read the first matrix stored in the off-chip memory includes: The first matrix after format conversion is read by accessing the off-chip memory of the target chip, wherein the data amount of the first matrix before format conversion is greater than the data amount of the first matrix after format conversion.
7. The method according to any one of claims 1 to 3, It is characterized in that The target chip includes a graphics processing unit GPU; The GPU includes at least one computing unit and a local storage unit connected to the computing unit, and the on-chip memory includes the local storage unit; Alternatively, the computing unit includes a private storage unit, and the on-chip memory includes the private storage unit.
8. An electronic device, It is characterized in that include: A processor and a memory; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 7.
9. A computer storage medium, It is characterized in that The computer storage medium stores a computer program or instruction. When the computer program or instruction is executed, the method according to any one of claims 1 to 7 is executed.
10. A chip system, It is characterized in that The chip system includes a processor, which is coupled to a memory and is used to execute a computer program or instruction stored in the memory. When the computer program or instruction is executed, the method as described in any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Matrix storage method, matrix access method and device and electronic equipment
CN111176582A