Data processing method, device, electronic device and computer readable storage medium

By using the coordinate information of the matrix block in the neural network to determine whether it belongs to the masked area and skip invalid calculations, the problem of inefficient calculations is solved, and the effect of improving the computing efficiency is achieved.

CN118227944BActive Publication Date: 2025-05-16SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410331523.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-05-16
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

During the calculation process of neural networks, mask operation causes partial matrix data to be invalid, resulting in inefficient execution.

Method used

By obtaining the coordinate information of the matrix block in the target matrix, it is determined whether it belongs to the masked area. If it belongs, the predetermined calculation operation for the matrix block will be skipped.

Benefits of technology

Reduce invalid memory access and calculation, and improve computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118227944B_ABST
    Figure CN118227944B_ABST
Patent Text Reader

Abstract

A data processing method, a data processing device, an electronic device and a computer-readable storage medium. The data processing method comprises: obtaining coordinate information of a matrix block in a target matrix; determining whether the matrix block belongs to a masked area of ​​the target matrix based on the coordinate information; if the matrix block belongs to the masked area, skipping a predetermined calculation operation on the matrix block. The method can reduce invalid memory access and calculation in the field of artificial intelligence and improve execution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of artificial intelligence, and in particular, to a data processing method, device, electronic device, and computer-readable storage medium. Background Art

[0002] Neural networks can be used in many fields such as speech recognition, image recognition, and natural language recognition. Running neural networks on artificial intelligence chips requires the support of a large number of operators. In some operators, part of the data in the matrix will be masked to a specific value to shield part of the information in the matrix. For example, taking the Chinese-English machine translation task as an example, for a given Chinese sentence, every word can be seen, and every word is visible to every other word. In the process of outputting English, each word can only see the word in front of it, but not the word behind it. By shielding part of the information through masking, the weight of this part can be ignored, reflecting that the characteristics of the following words will not affect the characteristics of the current word. Summary of the invention

[0003] At least one embodiment of the present disclosure provides a data processing method, comprising: obtaining coordinate information of a matrix block in a target matrix; determining whether the matrix block belongs to a masked area of ​​the target matrix based on the coordinate information; if the matrix block belongs to the masked area, skipping a predetermined calculation operation on the matrix block.

[0004] For example, in the data processing method provided by at least one example of the above-mentioned embodiments of the present disclosure, it also includes: if some elements in the matrix block belong to the masked area and the remaining elements belong to the unmasked area of ​​the target matrix, then for each sub-block in the matrix block, based on the coordinate information of the matrix block and the coordinate information of each sub-block, determine whether each sub-block belongs to the masked area; in the process of performing the predetermined calculation operation on the matrix block, for each sub-block belonging to the masked area, skip at least one sub-operation in the predetermined calculation operation.

[0005] For example, in the data processing method provided by at least one example of the above embodiments of the present disclosure, the masked area of ​​the target matrix is ​​the upper triangular area of ​​the target matrix; and the elements of the masked area of ​​the target matrix are set to negative infinity.

[0006] For example, in the data processing method provided by at least one example of the above-mentioned embodiment of the present disclosure, the coordinate information of the matrix block includes the first horizontal coordinate and the first vertical coordinate of the upper left corner element of the matrix block in the target matrix, and the target matrix includes a plurality of the matrix blocks, and the number of rows of elements included in each of the matrix blocks is a first numerical value. Based on the coordinate information, determining whether the matrix block belongs to the masked area of ​​the target matrix includes: determining whether the sum of the first vertical coordinate of the matrix block and the first numerical value is less than or equal to the first horizontal coordinate of the matrix block; if the sum of the first vertical coordinate of the matrix block and the first numerical value is less than or equal to the first horizontal coordinate of the matrix block, then the matrix block belongs to the masked area of ​​the target matrix.

[0007] For example, in the data processing method provided by at least one example of the above-mentioned embodiments of the present disclosure, the coordinate information of each sub-block includes the second horizontal coordinate and the second vertical coordinate of the upper left corner element of the sub-block in the matrix block to which it belongs, and the number of rows of elements included in each sub-block is the second numerical value. Determining whether each sub-block belongs to the masked area includes: for each sub-block, determining whether the sum of the first vertical coordinate of the matrix block to which the sub-block belongs, the second vertical coordinate of the sub-block and the second numerical value is less than or equal to the sum of the first horizontal coordinate of the matrix block to which the sub-block belongs and the second horizontal coordinate of the sub-block; if so, determining that the sub-block belongs to the masked area.

[0008] For example, in the data processing method provided by at least one example of the above-mentioned embodiments of the present disclosure, the predetermined calculation operation includes a normalization calculation operation, and the normalization calculation operation includes a first sub-operation and a second sub-operation; the first sub-operation includes, for each row of the matrix block, finding the maximum element value in the row; the second sub-operation includes, for each row of the matrix block, calculating a plurality of first calculation results corresponding to a plurality of elements in the row, respectively, and calculating the sum of the plurality of first calculation results, wherein the first calculation result corresponding to each element in the row is calculated based on the value of the element and the maximum element data in the row.

[0009] For example, in the data processing method provided by at least one example of the above-mentioned embodiments of the present disclosure, for each sub-block belonging to the masked area, at least one sub-operation in the predetermined computing operation is skipped, including: for each sub-block belonging to the masked area, the second sub-operation on each element in the sub-block is skipped.

[0010] At least one embodiment of the present disclosure provides a data processing device, including an acquisition unit, a determination unit and a calculation unit, wherein the acquisition unit is configured to obtain coordinate information of a matrix block in a target matrix; the determination unit is configured to determine whether the matrix block belongs to a masked area of ​​the target matrix based on the coordinate information; and the calculation unit is configured to skip a predetermined calculation operation on the matrix block if the matrix block belongs to the masked area.

[0011] At least one embodiment of the present disclosure provides an electronic device, comprising a processor; a memory storing one or more computer program modules; wherein the one or more computer program modules are configured to be executed by the processor to implement the data processing method provided by any embodiment of the present disclosure.

[0012] At least one embodiment of the present disclosure provides a computer-readable storage medium storing non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a computer, the data processing method provided by any embodiment of the present disclosure can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, but are not intended to limit the present disclosure.

[0014] Figure 1 shows the calculation process of a fusion operator;

[0015] Figure 2 A flowchart of a data processing method provided by at least one embodiment of the present disclosure is shown;

[0016] Figure 3 A schematic diagram of a target matrix provided by at least one embodiment of the present disclosure is shown;

[0017] Figure 4 A flowchart showing another data processing method provided by at least one embodiment of the present disclosure is shown;

[0018] Figure 5 A schematic diagram showing another target matrix provided by at least one embodiment of the present disclosure is shown;

[0019] Figure 6 A schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure is shown;

[0020] Figure 7 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown;

[0021] Figure 8A schematic block diagram showing another electronic device provided by at least one embodiment of the present disclosure; and

[0022] Fig. 9 A schematic diagram of a computer-readable storage medium provided by at least one embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0024] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, similar words such as "one", "one" or "the" do not indicate quantity restrictions, but indicate that there is at least one. Similar words such as "include" or "comprise" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Similar words such as "connect" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0025] Masking a matrix may be performed by setting part of the data in the matrix to a predetermined value, such as 0, -100, or negative infinity. In the following embodiments, -inf is used to represent negative infinity.

[0026] Taking the Attention fusion operator as an example, the calculation process of the Attention fusion operator is as follows: Figure 1As shown, matrix multiplication operation is performed on matrix Q and matrix K (such as MMA in the figure) to obtain a first matrix (hereinafter also referred to as matrix QK), and the first matrix is ​​masked by using a mask matrix (such as MASK in the figure) to obtain a masked second matrix, in which part of the data is masked. For example, the upper triangular part of the first matrix is ​​masked to negative infinity, and the masked matrix is ​​used as the second matrix. Then, the second matrix is ​​operated by the Softmax function (normalized exponential function) to obtain a third matrix, in which each element of the third matrix is ​​normalized to a value between [0,1]. The third matrix is ​​matrix multiplied with matrix V (such as MMA in the figure) to obtain an output matrix. This Attention-type fusion operator is widely used in large model training and reasoning scenarios, such as MHA (Multi-head attention) and GQA (Grouped-query attention) in large models such as LLAMA (Language Model using Adaptive Attention), GPT (Generative Pre-trained Transformer), and Stable Diffusion.

[0027] In the current Softmax implementation, each row (or column) of data is first traversed to find the maximum value row_max of the current row (or current column). Then, each row (or column) of data is traversed for the second time to find the exponent and sum of each element after subtracting row_max. That is, the exponent exp(x-row_max) is calculated for each element x, and the exponent and exp_sum of the current row are calculated. Finally, each row (or column) of data is traversed for the third time to find the quotient of the exponent exp(x-row_max) corresponding to each element and the exponent and exp_sum.

[0028] The calculation process of the Attention-type fusion operator includes upper triangular masking, which masks the upper triangular part of the matrix QK (first matrix) output by the first MMA to negative infinity. The entire matrix QK, including the upper triangular part masked to negative infinity, still participates in subsequent calculations, including obtaining the result of Softmax and performing matrix multiplication with the matrix V. However, in fact, the data masked to negative infinity in the upper triangular part will calculate exp(-inf) = 0 during the Softmax operation, and the masked part will have a result of 0 in the subsequent second MMA operation. In other words, the calculation of the masked part of the matrix QK is invalid. However, in the actual operation process, the masked part of the data will still be accessed and calculated, resulting in low execution efficiency.

[0029] At least one embodiment of the present disclosure provides a data processing method, a data processing device, an electronic device, and a computer-readable storage medium. The data processing method includes: obtaining coordinate information of a matrix block in a target matrix; determining whether the matrix block belongs to a masked area of ​​the target matrix based on the coordinate information; if the matrix block belongs to the masked area, skipping a predetermined calculation operation on the matrix block.

[0030] In this data processing method, the coordinate information of each matrix block is used to determine whether each matrix block belongs to the masked area of ​​the target matrix, and the predetermined calculation operation on the matrix block in the masked area is skipped. Based on this method, invalid memory access and calculation can be reduced and execution efficiency can be improved.

[0031] Figure 2 A flow chart of a data processing method provided by at least one embodiment of the present disclosure is shown.

[0032] like Figure 2 As shown, the method may include steps S210 to S230.

[0033] Step S210: Obtain coordinate information of a matrix block in the target matrix.

[0034] Step S220: Based on the coordinate information, determine whether the matrix block belongs to the masked area of ​​the target matrix.

[0035] Step S230: If the matrix block belongs to the masked area, skip the predetermined calculation operation on the matrix block.

[0036] For example, the target matrix may be the second matrix in the implementation process of the above-mentioned Attention-type fusion operator, but the present disclosure is not limited thereto. The target matrix may be any matrix that has been masked (i.e., has a masked area) and subsequently requires a predetermined computing operation (such as Softmax). The masked area of ​​the target matrix may be masked to negative infinity (-inf), for example.

[0037] For example, the target matrix may include multiple matrix blocks. In the process of performing a multiplication operation on matrix Q and matrix K to obtain matrix QK, matrix Q may be divided into multiple matrix blocks (chunks) with a step size of chunk Q, and matrix K may be divided into multiple matrix blocks with a step size of chunk K, such as Figure 1As shown, matrix Q and matrix K are used as inputs of matrix multiplication operation, and matrix QK (first matrix) is output. Matrix QK is calculated in units of matrix blocks of size (chunk Q, chunk K). The second matrix obtained by masking the first matrix is ​​also calculated in units of matrix blocks of size (chunk Q, chunk K). If the row direction of the matrix is ​​taken as the width direction of the matrix and the column direction is taken as the height direction of the matrix, then chunk Q represents the height of each matrix block, that is, the number of element rows included in each matrix block, and chunk K represents the width of each matrix block, that is, the number of element columns included in each matrix block.

[0038] Figure 3 A schematic diagram of a target matrix provided by at least one embodiment of the present disclosure is shown.

[0039] like Figure 3 As shown, the target matrix is, for example, a square matrix, that is, the number of rows and columns of the elements of the target matrix is ​​equal. The masked area of ​​the target matrix is, for example, the upper triangular area of ​​the target matrix, the white area in the figure is the masked upper triangular area, and the gray area is the unmasked lower triangular area. The elements of the masked area of ​​the target matrix can be set to negative infinity, that is, the value of the elements in the white area in the figure is negative infinity (-inf). The target matrix 300 may include a plurality of matrix blocks 310, each matrix block 310 may be a square or a rectangle, and each matrix block 310 includes multiple rows and columns of elements. For example, the size of the target matrix is ​​2048 (rows) × 2048 (columns), the size of each matrix block is 128 (rows) × 256 (columns), and the target matrix can be divided into 16*8=128 matrix blocks. In order to distinguish different matrix blocks, in some of the following embodiments, the matrix blocks can be represented by Axx, for example, the first matrix block from the left of the first row can be represented as A00, the second matrix block from the left of the first row can be represented as A01, the third matrix block from the left of the first row can be represented as A02, and so on.

[0040] For example, in some embodiments of the present disclosure, the masked area of ​​the target matrix is ​​an upper triangular area. However, the present disclosure is not limited to this. In other embodiments, the masked area of ​​the target matrix may also be other areas, such as a lower triangular area.

[0041] For example, in some embodiments of the present disclosure, the element value of the masked area is taken as negative infinity (-inf) for description, but the present disclosure is not limited to this. In other embodiments, the elements of the masked area can also be set to other values.

[0042] For example, steps S210 to S230 may be performed on each matrix block in the target matrix. In the following embodiments, one of the matrix blocks (eg, referred to as the current matrix block) is used as an example for description.

[0043] For example, in step S210, the coordinate information of the current matrix block is obtained. The coordinate information of the current matrix block may be the coordinates of one or some elements in the current matrix block, for example, the coordinates of the vertex elements (such as the upper left corner element) of the current matrix block. The coordinates of the elements are used to indicate the position of the elements in the target matrix. For example, the coordinate information of the current matrix block may include the coordinates of the upper left corner element of the current matrix block, with the upper left corner element of the target matrix as the origin. When the size of the target matrix is ​​2048×2048 and the size of each matrix block is 128×256, the upper left corner of the matrix is ​​used as the origin. Figure 3 The coordinates of the upper left corner element of the first matrix block A00 from the left in the first row are (0,0), the coordinates of the upper left corner element of the second matrix block A01 from the left in the first row are (256,0), the coordinates of the upper left corner element of the first matrix block from the left in the second row are (0,128), and so on.

[0044] For example, in other embodiments, the coordinate information of the current matrix block may also be the coordinates of the upper right corner element of the current matrix block, or the coordinates of the lower left corner element, etc.

[0045] For example, in step S220, determining whether the current matrix block belongs to the masked area of ​​the target matrix may include determining whether all elements in the current matrix block belong to the masked area, that is, determining whether the current matrix block completely belongs to the masked area. Whether the current matrix block completely belongs to the masked area can be determined based on the coordinates of one or some elements (such as vertex elements) in the current matrix block, and the specific determination method will be described in the following embodiments. Figure 3 The fully white matrix blocks completely belong to the masked area, the fully gray matrix blocks completely belong to the unmasked area, and some elements of the partially white and partially gray matrix blocks belong to the masked area and the remaining elements belong to the unmasked area of ​​the target matrix. For ease of description, in the following embodiments, the matrix blocks that completely belong to the masked area are referred to as first-type matrix blocks, the matrix blocks that completely belong to the unmasked area are referred to as second-type matrix blocks, and the matrix blocks that partially belong to the masked area and the remaining elements belong to the unmasked area are referred to as third-type matrix blocks.

[0046] For example, in step S230, for the matrix blocks belonging to the masked area (first type of matrix blocks), the predetermined calculation operation on such matrix blocks can be skipped, and the predetermined calculation operation can include, for example, a normalized exponential function operation (Softmax function operation). Since there is no need to perform operations on the first type of matrix blocks, the memory access operation on the first type of matrix blocks can also be skipped, that is, no memory access operation and calculation operation are performed on the first type of matrix blocks.

[0047] According to the data processing method of the embodiment of the present disclosure, the coordinate information of each matrix block is used to determine whether each matrix block belongs to the masked area of ​​the target matrix, and the predetermined calculation operations on the matrix blocks in the masked area are skipped. Based on this approach, invalid memory accesses and operations can be reduced and execution efficiency can be improved.

[0048] For example, the coordinate information of the matrix block includes the horizontal coordinate and the vertical coordinate of the upper left corner element of the matrix block in the target matrix. For the convenience of distinction, the horizontal coordinate of the upper left corner element of the matrix block in the target matrix is ​​called the first horizontal coordinate, represented by w; the vertical coordinate of the upper left corner element of the matrix block in the target matrix is ​​called the first vertical coordinate, represented by h. The target matrix includes multiple matrix blocks, and the number of rows of elements included in each matrix block is the first value, that is, the height of each matrix block is the first value, such as Figure 3 As shown, the first value can be represented by chunk Q. Wherein, w, h and chunk Q are all integers greater than or equal to 0.

[0049] For example, in step S220, determining whether the matrix block belongs to the masked area of ​​the target matrix based on the coordinate information may include: determining whether the sum of the first ordinate and the first numerical value of the matrix block is less than or equal to the first abscissa of the matrix block; if the sum of the first ordinate and the first numerical value of the matrix block is less than or equal to the first abscissa of the matrix block, determining that the matrix block belongs to the masked area of ​​the target matrix.

[0050] For example, if the upper left corner of the target matrix is ​​taken as the origin, when a matrix block satisfies h+chunk Q≤w, the matrix block is considered to belong to the masked area (the first type of matrix block), and the processing of the matrix block can be skipped. When the matrix block satisfies h+chunk Q≤w, the abscissa of the lower left corner element of the matrix block is greater than or equal to the ordinate, and the lower left corner element is located on the diagonal or in the upper triangular area. When the lower left corner element of the matrix block is located on the diagonal or in the upper triangular area, the other elements in the matrix block belong to the masked area, so the matrix block belongs to the masked area.

[0051] For example, the size of the target matrix is ​​2048×2048, and the size of each matrix block of the target matrix is ​​128×256, so chunk Q=128. Figure 3Take the first matrix block A00 from the left of the first row as an example. The coordinates of the upper left corner element of the matrix block A00 are (0, 0), that is, w = 0, h = 0, 0 + 128 > 0. Therefore, h + chunk Q ≤ w is not satisfied. Therefore, the matrix block A00 does not belong to the masked area. Figure 3 Take the second matrix block A01 from the left of the first row as an example. The coordinates of the upper left corner element of the matrix block A01 are (256, 0), that is, h = 0, w = 256, 0 + 128 < 256, so h + chunk Q ≤ w is satisfied. Therefore, the matrix block A01 belongs to the masked area. Figure 3 Take the third matrix block A02 from the left of the first row as an example. The coordinates of the upper left corner element of the matrix block A02 are (512, 0), that is, w = 512, h = 0, 0 + 128 < 512, and h + chunk Q ≤ w. Therefore, the matrix block A02 belongs to the masked area. Figure 3 Taking the first matrix block from the left of the second row as an example, the coordinates of the upper left corner element of the matrix block are (0, 128), that is, w = 0, h = 128, 128 + 128 > 0, which does not satisfy h + chunk Q ≤ w, and thus, the matrix block does not belong to the masked area. By analogy, it can be determined whether each matrix block in the target matrix belongs to the masked area. Based on the above method, it can be quickly and accurately determined whether each matrix block belongs to the masked area, and then, the predetermined calculation operation and memory access operation of the matrix block in the masked area can be skipped.

[0052] For example, for a matrix block that completely belongs to an unmasked area (a second type of matrix block) and a matrix block that partially belongs to a masked area (a third type of matrix block), a predetermined calculation operation can be performed. When a predetermined calculation operation is performed on a second type of matrix block, it can be performed according to a normal process. When a predetermined calculation operation is performed on a third type of matrix block, in some embodiments, it can be performed according to a normal process, but since some elements in the third type of matrix block belong to a masked area, the operation on this part of the data is invalid. In order to further improve the execution efficiency, in the following embodiments of the present disclosure, the execution process is improved for the third type of matrix block.

[0053] Figure 4 A flowchart of another data processing method provided by at least one embodiment of the present disclosure is shown.

[0054] like Figure 4 As shown, the data processing method may include steps S240 to S250 in addition to the above steps S210 to S230.

[0055] Step S240: if some elements in the matrix block belong to the masked area and the remaining elements belong to the unmasked area of ​​the matrix block, then for each sub-block in the matrix block, based on the coordinate information of the matrix block and the coordinate information of each sub-block, determine whether each sub-block belongs to the masked area;

[0056] Step S250: in the process of performing the predetermined calculation operation on the matrix block, for each sub-block belonging to the masked area, skipping at least one sub-operation in the predetermined calculation operation.

[0057] Figure 5 A schematic diagram of another target matrix provided by at least one embodiment of the present disclosure is shown.

[0058] like Figure 5 As shown, the target matrix 500 includes a plurality of matrix blocks 510. To distinguish different matrix blocks, each matrix block can be represented by Axx. In the case where the masked area of ​​the target matrix is ​​an upper triangular area (or a lower triangular area), each matrix block located on the diagonal belongs to the third type of matrix block. Each third type of matrix block can include a plurality of sub-blocks 511. For example, the size of the matrix block 510 is 128 (rows) × 256 (columns), the size of the sub-block 511 is 32 (rows) × 64 (columns), and each matrix block 510 can be divided into 4*4=16 sub-blocks 511. To distinguish different sub-blocks, in some of the following embodiments, the sub-blocks can be represented by Bxx. For example, in each matrix block, the first sub-block from the left of the first row can be represented by B00, the second sub-block from the left of the first row can be represented by B01, the first sub-block from the left of the second row can be represented by B10, and so on.

[0059] For example, steps S240 to S250 may be performed for each third-type matrix block. For each third-type matrix block, in step S240, it may be determined whether each sub-block belongs to the masked area based on the coordinate information of the matrix block and the coordinate information of each sub-block, the coordinate information of the matrix block may include the coordinates (w, h) of the upper left corner element of the matrix block in the target matrix, the coordinate information of the sub-block 511 may be the coordinates of one or some elements in the sub-block 511 in the matrix block, for example, the coordinates of the vertex elements (such as the upper left corner element) of the sub-block 511 in the matrix block, and the coordinates of each element in the sub-block 511 are used to indicate the position of the element in the matrix block.

[0060] For example, the coordinate information of the sub-block 511 may include the horizontal coordinate and the vertical coordinate of the upper left corner element of the sub-block 511 in the corresponding matrix block. In order to distinguish from the above-mentioned first horizontal coordinate and first vertical coordinate, the horizontal coordinate of the sub-block may be referred to as the second horizontal coordinate, represented by sw, and the vertical coordinate of the sub-block may be referred to as the second vertical coordinate, represented by sh, where sw is an integer greater than or equal to 0 and less than or equal to chunk K, and sh is an integer greater than or equal to 0 and less than or equal to chunk Q. When the size of the target matrix is ​​2048 (rows) × 2048 (columns), the size of each matrix block is 128 (rows) × 256 (columns), and the size of each sub-block is 32 (rows) × 64 (columns), Figure 5 In the matrix block A00 shown on the left side, with the upper left corner element of the matrix block A00 as the origin, the second horizontal coordinate and the second vertical coordinate of the upper left corner element of the first sub-block B00 from the first row in the matrix block A00 are both 0, that is, sw=0, sh=0; the second horizontal coordinate of the upper left corner element of the second sub-block B01 from the first row in the matrix block A00 is 64, and the second vertical coordinate is 0, that is, sw=64, sh=0; the second horizontal coordinate of the upper left corner element of the third sub-block B02 from the first row in the matrix block A00 is 128, and the second vertical coordinate is 0, that is, sw=128, sh=0, and so on. The number of rows of elements included in each sub-block is the second value, that is, the height of each sub-block is the second value, and the second value is represented by m. If the size of the sub-block is 32×64, then m=32.

[0061] For example, in other embodiments, the coordinate information of the sub-block may also be the coordinates of the upper right corner element of the sub-block, or the coordinates of the lower left corner element, etc.

[0062] For example, in step S240, the sub-block belonging to the masked region includes that all elements in the sub-block belong to the masked region. In the matrix block, except for the sub-block belonging to the masked region, the remaining sub-blocks belong to the unmasked region, that is, if at least some elements in the sub-block belong to the unmasked region, then the sub-block belongs to the unmasked region. Figure 5 In the matrix block shown on the left, the completely white sub-block is considered to belong to the masked area, and the completely gray sub-block is considered to belong to the unmasked area. For ease of description, in the following embodiments, the sub-blocks belonging to the masked area are referred to as the first type of sub-blocks, and the sub-blocks belonging to the unmasked area are referred to as the second type of sub-blocks. Whether the sub-block belongs to the masked area can be determined based on the coordinate information (w, h) of the matrix block to which the sub-block belongs and the coordinate information (sw, sh) of the sub-block, and the specific determination method will be described in the following embodiments.

[0063] For example, the predetermined calculation operation includes a normalization calculation operation, such as the normalized exponential function operation (Softmax function operation) described above. The predetermined calculation operation includes multiple sub-operations. In step S250, for the first type of sub-block, at least one of the sub-operations can be skipped, that is, there is no need to execute the at least one sub-operation, and there is no need to perform a memory access operation on the first type of sub-block. For the second type of sub-block, multiple sub-operations in the predetermined calculation operation can be executed according to the normal process.

[0064] According to the data processing method of the embodiment of the present disclosure, not only the invalid memory access and calculation of part of the matrix block are reduced at the coarse granularity (matrix block granularity), but also the invalid memory access and calculation of part of the sub-block are reduced at the fine granularity (sub-block granularity), thereby further improving the execution efficiency. For example, in step S240, determining whether each sub-block belongs to the masked area may include: for each sub-block, determining whether the sum of the first ordinate (h) of the matrix block to which the sub-block belongs, the second ordinate (sh) of the sub-block, and the second value (m) is less than or equal to the sum of the first abscissa (w) of the matrix block to which the sub-block belongs and the second abscissa (sw) of the sub-block; if so, the sub-block belongs to the masked area.

[0065] For example, if the upper left corner of the target matrix is ​​taken as the origin of the target matrix, and the upper left corner of the matrix block is taken as the origin of the matrix block, when a sub-block satisfies h+sh+m≤w+sw, it is considered that the sub-block belongs to the masked area (the first type of sub-block), and part of the processing of the sub-block can be skipped. Since h represents the ordinate of the upper left corner of the matrix block where the sub-block is located in the target matrix, and sh represents the ordinate of the upper left corner of the sub-block in the matrix block, then h+sh can represent the ordinate of the upper left corner of the sub-block in the target matrix, and h+sh+m can represent the ordinate of the lower left corner of the sub-block in the target matrix. Since w represents the abscissa of the upper left corner of the matrix block where the sub-block is located in the target matrix, and sw represents the abscissa of the upper left corner of the sub-block in the matrix block, then w+sw can represent the abscissa of the upper left corner of the sub-block in the target matrix, or the abscissa of the lower left corner of the sub-block in the target matrix. When the sub-block satisfies h+sh+m≤w+sw, it can be determined that the ordinate of the lower left corner of the sub-block is less than or equal to the abscissa, and then it can be further determined that the lower left corner of the sub-block is located on the diagonal of the target matrix or in the upper triangular area. When the lower left corner of the sub-block is located on the diagonal of the target matrix or in the upper triangular area, the other elements of the sub-block are all located in the upper triangular area, then it can be determined that the sub-block belongs to the upper triangular area (the masked area).

[0066] For example, the size of each matrix block of the target matrix is ​​128×256, and the size of each sub-block is 32×64 (m=32). Figure 5Take the first sub-block B00 from the left of the first row as an example. The coordinates of the upper left corner element of the matrix block A00 to which the sub-block B00 belongs in the target matrix are (0,0), that is, w=0, h=0. The horizontal and vertical coordinates of the upper left corner element of the sub-block B00 in the matrix block A00 are both 0, that is, sw=0, sh=0. Substituting into the above formula, 0+0+32>0+0, therefore, h+sh+m≤w+sw is not satisfied, and the sub-block B00 does not belong to the masked area. Figure 5 Take the second sub-block B01 from the left of the first row as an example. The coordinates of the upper left corner element of the matrix block A00 to which the sub-block B01 belongs are (0, 0), that is, w=0, h=0. The horizontal coordinate sw of the upper left corner element of the sub-block B01 in the matrix block A00 is 64. Substituting into the above formula, 0+0+32<0+64, therefore, h+sh+m≤w+sw is satisfied, and the sub-block B01 belongs to the masked area. Figure 5 Take the first sub-block B10 from the left of the second row as an example, w=0, h=0 of the matrix block A00 to which it belongs, and the horizontal coordinate sw=0 and the vertical coordinate sh=32 of the upper left corner element of the sub-block B10 in the matrix block A00. Substituting into the above formula, 0+32+32>0+0, therefore, h+sh+m≤w+sw is not satisfied, and the sub-block B10 does not belong to the masked area. Based on the above method, it is possible to quickly and accurately determine whether each sub-block belongs to the masked area, and then, some calculation operations and memory access operations on the sub-blocks of the masked area can be skipped.

[0067] For example, the predetermined calculation operation includes a normalization calculation operation, and the normalization calculation operation includes a first sub-operation and a second sub-operation. The first sub-operation includes, for each row of the matrix block, finding the maximum element value in the row; the second sub-operation includes, for each row of the matrix block, calculating a plurality of first calculation results corresponding to a plurality of elements in the row, and calculating the sum of the plurality of first calculation results, wherein the first calculation result corresponding to each element in the row is calculated based on the value of the element and the maximum element data in the row.

[0068] For example, the first sub-operation may be an operation of performing a first traversal on each row during the Softmax function operation to find the row maximum value row_max (maximum element value). The second sub-operation may be an operation of performing a second traversal on each row during the Softmax function operation to find the exponential sum exp_sum of each element minus row_max, wherein the first calculation result is exp(x-row_max), and the sum of multiple first calculation results is exp_sum.

[0069] For example, in step S250, for each sub-block belonging to the masked area, skipping at least one sub-operation in the predetermined computing operation may include: for each sub-block belonging to the masked area, skipping the second sub-operation on each element in the sub-block.

[0070] For example, for each element in the masked area, the element value is negative infinity. In the second sub-operation, for each element in the masked area, the exponent exp(x-row_max) is equal to 0. In the process of calculating the exponent and exp_sum, adding 0 is equivalent to not adding. Therefore, the second sub-operation for these elements can be omitted, and only the exponential calculation and exponential summation of the elements in the unmasked area need to be performed.

[0071] For example, the predetermined calculation operation may also include a third sub-operation, and the third sub-operation includes, for each row of the matrix block, calculating the quotient of the first calculation result of each element in the row and the above-mentioned sum. The third sub-operation may be an operation of performing a third traversal on each row during the above-mentioned Softmax function operation to calculate the quotient of the exponent exp(x-row_max) corresponding to each element and the exponent sum exp_sum. In the third sub-operation, for each element in the masked area, the exponent exp(x-row_max) is equal to 0, and the quotient of the exponent exp(x-row_max) and the exponent sum exp_sum is also equal to 0. Therefore, the result data corresponding to the elements in the masked area can be directly set to 0, and the third sub-operation for these elements can be omitted.

[0072] For example, in the first traversal process, the maximum value of each row is found. In this process, the result data belonging to the upper triangle area can be directly filled with 0. In this way, in the second traversal process, for sub-blocks that satisfy h+sh+m≤w+sw, the second sub-operation (and the third sub-operation) can be skipped, and there is no need to calculate and access these sub-blocks.

[0073] According to the data processing method of the embodiment of the present disclosure, it is possible to optimize the memory access and calculation levels of the Attention operator. In large model application scenarios, in order to save video memory, the Attention operator is implemented using the flash attention algorithm, that is, the matrices Q, K, and V are divided into multiple matrix blocks (chunks). Only one matrix block is processed at a time. At a coarse granularity, with the matrix block as the granularity, invalid memory access and calculation of some matrix blocks can be skipped. At a fine granularity, with the sub-block in the matrix block as the granularity, invalid memory access and calculation of some sub-blocks can be skipped. In addition to being applied to the Attention operator, the data processing method of the embodiment of the present disclosure can also be applied to other operators.

[0074] Using the data processing method of the embodiment of the present disclosure, at a coarse granularity (matrix block granularity), the time required for memory access and calculation is theoretically only 1 / (2*k_chunk)+1 / 2 of the original memory access and calculation, as shown in the following formula.

[0075] q_chunk=seq_len / chunk Q, k_chunk=seq_len / chunk K

[0076] valid_part=(1+k_chunk)*q_chunk / 2 / (k_chunk*q_chunk)=1 / (2*k_chunk)+1 / 2

[0077] Among them, seq_len is the height and width of the target matrix, chunk Q and chunk K are the height and width of the matrix block, q_chunk indicates the number of matrix blocks divided by the target matrix in the height direction, and k_chunk indicates the number of matrix blocks divided by the target matrix in the width direction. If chunk K = 512, seq_len = 4096, theoretically, 7 / 16 of the calculation can be reduced. Taking the llama2 70B GQA operator as an example, it can be optimized from 20.5577ms to 12.4844ms.

[0078] By using the data processing method of the embodiment of the present disclosure, the time required for memory access and calculation is also reduced at a fine granularity (sub-block granularity), such as the llama2 70B GQA operator can be optimized from 6.7856ms to 6.1425ms.

[0079] Figure 6 A schematic block diagram of a data processing device 600 provided by at least one embodiment of the present disclosure is shown.

[0080] For example, Figure 6 As shown, the data processing device 600 includes an acquisition unit 610, a determination unit 620 and a calculation unit 630. These components are interconnected via a bus system and / or other forms of connection mechanisms (not shown). For example, these modules can be implemented by hardware (e.g., circuit) modules, software modules, or any combination of the two. The following embodiments are the same and will not be repeated here. For example, these units can be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, as well as corresponding computer instructions. It should be noted that Figure 6 The components and structure of the data processing device 600 shown are merely exemplary and not restrictive. The data processing device 600 may also have other components and structures as required.

[0081] The acquisition unit 610 is configured to obtain coordinate information of a matrix block in the target matrix. The acquisition unit 610 may, for example, execute Figure 2 Described step S210.

[0082] The determination unit 620 is configured to determine whether the matrix block belongs to the masked area of ​​the target matrix based on the coordinate information. The determination unit 620 may, for example, perform Figure 2 Described step S220.

[0083] The calculation unit 630 is configured to skip the predetermined calculation operation on the matrix block if the matrix block belongs to the masked area. The calculation unit 630 may, for example, perform Figure 2 Described step S230.

[0084] For example, the acquisition unit 610, the determination unit 620, and the calculation unit 630 may be hardware, software, firmware, or any feasible combination thereof. For example, the acquisition unit 610, the determination unit 620, and the calculation unit 630 may be a dedicated or general circuit, chip, or device, or may be a combination of a processor and a memory. The embodiments of the present disclosure do not limit the specific implementation forms of the above-mentioned units.

[0085] For example, the acquisition unit 610, the determination unit 620, and the calculation unit 630 may include codes and programs stored in a memory; the processor may execute the codes and programs to implement some or all of the functions of the acquisition unit 610, the determination unit 620, and the calculation unit 630 described above. For example, the acquisition unit 610, the determination unit 620, and the calculation unit 630 may be dedicated hardware devices, used to implement some or all of the functions of the acquisition unit 610, the determination unit 620, and the calculation unit 630 described above. For example, the acquisition unit 610, the determination unit 620, and the calculation unit 630 may be a circuit board or a combination of multiple circuit boards, used to implement the functions described above. In an embodiment of the present disclosure, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-temporary memories connected to the processor; and (3) firmware stored in the memory that is executable by the processor.

[0086] It should be noted that in the embodiment of the present disclosure, each unit of the data processing device 600 corresponds to each step of the aforementioned data processing method. For the specific functions of the data processing device 600, please refer to the relevant description of the data processing method, which will not be repeated here. Figure 6The components and structures of the data processing device 600 shown are only exemplary and non-restrictive. The data processing device 600 may also include other components and structures as needed. The data processing device 600 may include more or fewer circuits or units, and the connection relationship between the various circuits or units is not limited and can be determined according to actual needs. The specific configuration of each circuit or unit is not limited and can be composed of analog devices according to circuit principles, or can be composed of digital chips, or can be composed in other applicable ways.

[0087] For example, in the data processing device provided by at least one example of the above-mentioned embodiments of the present disclosure, the determination unit is further configured to: if some elements in the matrix block belong to the masked area and the remaining elements belong to the unmasked area of ​​the target matrix, then for each sub-block in the matrix block, based on the coordinate information of the matrix block and the coordinate information of each sub-block, determine whether the each sub-block belongs to the masked area. The calculation unit is further configured to: in the process of performing the predetermined calculation operation on the matrix block, for each sub-block belonging to the masked area, skip at least one sub-operation in the predetermined calculation operation.

[0088] For example, in a data processing device provided by at least one example of the above embodiments of the present disclosure, the masked area of ​​the target matrix is ​​an upper triangular area of ​​the target matrix; and the elements of the masked area of ​​the target matrix are set to negative infinity.

[0089] For example, in the data processing device provided by at least one example of the above-mentioned embodiments of the present disclosure, the coordinate information of the matrix block includes the first horizontal coordinate and the first vertical coordinate of the upper left corner element of the matrix block in the target matrix, and the target matrix includes a plurality of the matrix blocks, and the number of rows of elements included in each of the matrix blocks is a first numerical value. The determination unit is further configured to: determine whether the sum of the first vertical coordinate of the matrix block and the first numerical value is less than or equal to the first horizontal coordinate of the matrix block; if the sum of the first vertical coordinate of the matrix block and the first numerical value is less than or equal to the first horizontal coordinate of the matrix block, then the matrix block belongs to the masked area of ​​the target matrix.

[0090] For example, in the data processing device provided by at least one example of the above embodiments of the present disclosure, the coordinate information of each sub-block includes the second horizontal coordinate of the upper left corner element of the sub-block in the matrix block to which it belongs. The determination unit is further configured to: for each sub-block, determine whether the sum of the first vertical coordinate of the matrix block to which the sub-block belongs and the first value is less than or equal to the sum of the first horizontal coordinate of the matrix block to which the sub-block belongs and the second horizontal coordinate of the sub-block; if so, determine that the sub-block belongs to the masked area.

[0091] For example, in the data processing device provided by at least one example of the above-mentioned embodiments of the present disclosure, the predetermined calculation operation includes a normalization calculation operation, and the normalization calculation operation includes a first sub-operation and a second sub-operation; the first sub-operation includes, for each row of the matrix block, finding the maximum element value in the row; the second sub-operation includes, for each row of the matrix block, calculating a plurality of first calculation results respectively corresponding to a plurality of elements in the row, and calculating the sum of the plurality of first calculation results, wherein the first calculation result corresponding to each element in the row is calculated based on the value of the element and the maximum element data in the row.

[0092] For example, in the data processing device provided by at least one example of the above embodiments of the present disclosure, the computing unit is further configured to: for each sub-block belonging to the masked area, skip the second sub-operation on each element in the sub-block.

[0093] At least one embodiment of the present disclosure also provides an electronic device, the electronic device includes a processor and a memory, and the memory stores one or more computer program modules. The one or more computer program modules are configured to be executed by the processor to implement the above-mentioned data processing method. The electronic device can use the coordinate information of each matrix block to determine whether each matrix block belongs to the masked area of ​​the target matrix, and skip the predetermined calculation operation of the matrix block in the masked area, thereby reducing invalid memory access and operation and improving execution efficiency.

[0094] Figure 7 A schematic block diagram of an electronic device provided in some embodiments of the present disclosure. Figure 7 As shown, the electronic device 700 includes a processor 710 and a memory 720. The memory 720 stores non-temporary computer-readable instructions (e.g., one or more computer program modules). The processor 710 is used to run non-temporary computer-readable instructions, and the non-temporary computer-readable instructions are executed by the processor 710 to execute one or more steps in the data processing method described above. The memory 720 and the processor 710 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown). For the specific implementation of each step of the data processing method and the related explanation content, please refer to the embodiment of the above-mentioned data processing method, and the repetitions will not be repeated here.

[0095] It should be noted that Figure 7 The components of the electronic device 700 shown are merely exemplary and non-limiting. The electronic device 700 may also have other components according to actual application requirements.

[0096] For example, the processor 710 and the memory 720 may communicate with each other directly or indirectly.

[0097] For example, the processor 710 and the memory 720 may communicate via a network. The network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 710 and the memory 720 may also communicate with each other via a system bus, which is not limited in the present disclosure.

[0098] For example, the processor 710 and the memory 720 may be arranged on a server side (or a cloud side).

[0099] For example, the processor 710 can control other components in the electronic device 700 to perform desired functions. For example, the processor 710 can be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) can be an X86 or ARM architecture, etc. The processor 710 can be a general-purpose processor or a special-purpose processor, and can control other components in the electronic device 700 to perform desired functions.

[0100] For example, the memory 720 may include any combination of one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disk read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor 710 may run one or more computer program modules to implement various functions of the electronic device 700. Various applications and various data, as well as various data used and / or generated by the application, etc. may also be stored in the computer-readable storage medium.

[0101] For example, in some embodiments, the electronic device 700 may be a mobile phone, a computer, a server, etc.

[0102] It should be noted that in the embodiment of the present disclosure, the specific functions and technical effects of the electronic device 700 can refer to the above description of the data processing method, which will not be repeated here.

[0103] Figure 8 This is a schematic block diagram of another electronic device provided in some embodiments of the present disclosure. The electronic device 800 is suitable for implementing the data processing method provided in the embodiments of the present disclosure. The electronic device 800 may be a terminal device, etc. It should be noted that: Figure 8The electronic device 800 shown is only an example and does not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0104] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 810, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 820 or a program loaded from a storage device 880 into a random access memory (RAM) 830. In the RAM 830, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 810, the ROM 820, and the RAM 830 are connected to each other via a bus 840. An input / output (I / O) interface 850 is also connected to the bus 840.

[0105] Typically, the following devices may be connected to the I / O interface 850: an input device 860 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 870 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 880 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 890. The communication device 890 may allow the electronic device 800 to communicate with other electronic devices wirelessly or by wire to exchange data. Although Figure 8 The electronic device 800 is shown to have various devices, but it should be understood that it is not required to implement or possess all of the devices shown, and the electronic device 800 may alternatively implement or possess more or fewer devices.

[0106] For example, according to an embodiment of the present disclosure, the above-mentioned data processing method can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the above-mentioned data processing method. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 890, or installed from the storage device 880, or installed from the ROM 820. When the computer program is executed by the processing device 810, the functions defined in the data processing method provided in the embodiment of the present disclosure can be implemented.

[0107] At least one embodiment of the present disclosure further provides a computer-readable storage medium, which stores non-temporary computer-readable instructions, and when the non-temporary computer-readable instructions are executed by a computer, the above-mentioned data processing method can be implemented. Using the computer-readable storage medium, the coordinate information of each matrix block can be used to determine whether each matrix block belongs to the masked area of ​​the target matrix, and the predetermined calculation operation of the matrix block in the masked area can be skipped, thereby reducing invalid memory access and operation and improving execution efficiency.

[0108] Fig. 9 A schematic diagram of a storage medium provided in some embodiments of the present disclosure. Fig. 9 As shown, the storage medium 900 stores non-transitory computer-readable instructions 910. For example, when the non-transitory computer-readable instructions 910 are executed by a computer, one or more steps in the data processing method described above are performed.

[0109] For example, the storage medium 900 may be applied to the electronic device 700. Figure 7 The memory 720 in the electronic device 700 is shown. For example, the description of the storage medium 900 can refer to Figure 7 The corresponding description of the memory 720 in the electronic device 700 is shown and will not be repeated here.

[0110] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0111] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0112] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

[0113] There are a few points to note about this disclosure:

[0114] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.

[0115] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to obtain new embodiments.

[0116] The above description is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.

Claims

1. A data processing method, comprising: Get the coordinate information of a matrix block in the target matrix; Based on the coordinate information, determining whether the matrix block belongs to a masked area of ​​the target matrix; as well as If the matrix block belongs to the masked area of ​​the target matrix, skipping the predetermined calculation operation on the matrix block; Wherein, the data processing method further includes: If some elements in the matrix block belong to the masked area and the remaining elements belong to the unmasked area of ​​the target matrix, for each sub-block in the matrix block, based on the coordinate information of the matrix block and the coordinate information of each sub-block, determining whether the each sub-block belongs to the masked area of ​​the target matrix; and In the process of performing the predetermined calculation operation on the matrix block, at least one sub-operation in the predetermined calculation operation is skipped for each sub-block belonging to the masked area.

2. The data processing method according to claim 1, wherein: The masked area of ​​the target matrix is ​​the upper triangular area of ​​the target matrix; The elements of the target matrix in the masked region are set to negative infinity.

3. The data processing method according to claim 2, wherein: The coordinate information of the matrix block includes a first horizontal coordinate and a first vertical coordinate of an upper left corner element of the matrix block in the target matrix, the target matrix includes a plurality of the matrix blocks, and the number of rows of elements included in each of the matrix blocks is a first value; Determining whether the matrix block belongs to a masked area of ​​the target matrix based on the coordinate information includes: Determine whether a sum of a first ordinate of the matrix block and the first value is less than or equal to a first abscissa of the matrix block; If the sum of the first ordinate of the matrix block and the first value is less than or equal to the first abscissa of the matrix block, it is determined that the matrix block belongs to the masked area of ​​the target matrix.

4. The data processing method according to claim 3, wherein: The coordinate information of each sub-block includes a second horizontal coordinate and a second vertical coordinate of the upper left corner element of the sub-block in the corresponding matrix block, and the number of rows of elements included in each sub-block is the second value; Determining whether each sub-block belongs to the masked area includes: For each sub-block, determine whether a sum of a first ordinate of a matrix block to which the sub-block belongs, a second ordinate of the sub-block, and the second value is less than or equal to a sum of a first abscissa of the matrix block to which the sub-block belongs and a second abscissa of the sub-block; If so, it is determined that the sub-block belongs to the masked area.

5. The data processing method according to any one of claims 1 to 4, wherein: The predetermined calculation operation includes a normalization calculation operation, and the normalization calculation operation includes a first sub-operation and a second sub-operation; The first sub-operation includes, for each row of elements of the matrix block, finding the maximum element value in the row; The second sub-operation includes calculating, for each row of the matrix block, a first calculation result corresponding to each element in the row, and calculating the sum of each first calculation result, wherein the first calculation result corresponding to each element in the row is calculated based on the numerical value of the element and the maximum element data in the row.

6. The data processing method according to claim 5, wherein: For each sub-block belonging to the masked area, skipping at least one sub-operation in the predetermined computing operation comprises: For each sub-block belonging to the masked region, skipping the second sub-operation on the sub-block.

7. A data processing device, comprising: An acquisition unit configured to obtain coordinate information of a matrix block in a target matrix; a determination unit configured to determine, based on the coordinate information, whether the matrix block belongs to a masked area of ​​the target matrix; a calculation unit configured to skip a predetermined calculation operation on the matrix block if the matrix block belongs to the masked area; The determining unit is further configured to: if some elements in the matrix block belong to the masked area and the remaining elements belong to the unmasked area of ​​the target matrix, then for each sub-block in the matrix block, based on the coordinate information of the matrix block and the coordinate information of each sub-block, determine whether the each sub-block belongs to the masked area of ​​the target matrix; The calculation unit is further configured to: in the process of performing the predetermined calculation operation on the matrix block, for each sub-block belonging to the masked area, skip at least one sub-operation in the predetermined calculation operation.

8. An electronic device comprising: processor; A memory storing one or more computer program modules; The one or more computer program modules are configured to be executed by the processor to implement the data processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing non-transitory computer-readable instructions, which can implement the data processing method according to any one of claims 1 to 6 when executed by a computer.

Citation Information

Patent Citations

  • Scene roaming method and device, equipment and medium

    CN115619986A

  • Matrix operation method, device and unit and electronic equipment

    CN115859011A