Data Processing Method and Apparatus, Electronic Apparatus, and Medium

By adjusting the correspondence between input data and masked data for product calculation, the problem of overfitting in machine learning models is solved, and the generalization ability and operation efficiency of the model are improved.

CN115186815BActive Publication Date: 2025-07-11SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210916696.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-01
Publication Date
2025-07-11
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

In machine learning models, overfitting is prone to occur when there are many parameters and few training samples, resulting in the model's loss function on the training data and the prediction accuracy is high, but the loss function on the test data is relatively large and the prediction accuracy is low.

Method used

A data processing method is adopted to load the input data into the first register, and the masked data is loaded into the second register. By adjusting the correspondence between the input data and the masked data, the bits in the selection instruction select register are used to calculate, reducing the number of instructions and improving operational efficiency.

Benefits of technology

It effectively reduces the number of instructions, greatly improves the operating efficiency, reduces the running time, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115186815B_ABST
    Figure CN115186815B_ABST
Patent Text Reader

Abstract

A data processing method, a data processing device, an electronic device, and a computer-readable storage medium. The data processing method includes: loading input data into a first register and loading masking data into a second register, where the size of the input data is N*M, the size of the masking data is P*Q, each element of the masking data is Z bits, Q*Z = 2*M, each bit of each element of the masking data corresponds to an element of the input data, the storage units of the second register are arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially correspond to i consecutive elements in the same row of the input data, and N, M, P, Q, i, j, and Z are all positive integers; performing a multiplication calculation based on the corresponding relationship between each element of the input data and each bit of each element of the masking data to obtain output data. This method can effectively reduce the number of assembly instructions, greatly improve the running efficiency, and reduce the running time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data processing method, a data processing apparatus, an electronic device, and a computer-readable storage medium. Background Art

[0002] In a machine learning model, if the model has too many parameters and too few training samples, the trained model is prone to overfitting. Overfitting is specifically manifested as follows: the loss function of the model on the training data is small and the prediction accuracy is high; however, the loss function on the test data is relatively large and the prediction accuracy is low. The dropout method can effectively alleviate the occurrence of overfitting and achieve a regularization effect to a certain extent. The dropout method means that during forward propagation, the activation value of a certain neuron stops working with a certain probability, which can make the model more generalizable. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a data processing method, including: loading input data into a first register and loading masking data into a second register, the size of the input data is N*M, the size of the masking data is P*Q, each element of the masking data is Z bits, Q*Z = 2*M, each bit of each element of the masking data corresponds to an element of the input data, the storage units of the second register are arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially correspond to i consecutive elements in the same row of the input data, where N, M, P, Q, i, j, Z are all positive integers; performing a product calculation based on the corresponding relationship between each element of the input data and each bit of each element of the masking data to obtain output data.

[0004] For example, in the data processing method provided in an embodiment of the present disclosure, the first j / 2 columns of the second register are the first group of storage units, and the (j / 2 + 1)-th column to the j-th column of the second register are the second group of storage units; based on the corresponding relationship between each element of the input data and each bit of each element of the masking data, performing a multiplication calculation to obtain output data, including: using a selection instruction to select the bit stored in the s-th column storage unit in the first group of storage units and the bit stored in the s-th column storage unit in the second group of storage units, where the bit stored in the s-th column storage unit in the first group of storage units and the bit stored in the s-th column storage unit in the second group of storage units correspond to the elements in the input data that are in consecutive i columns in every two consecutive rows; based on the corresponding relationship, multiplying the bits in the selected storage units by the corresponding elements of the input data to obtain the elements in the output data that are in consecutive i columns in every two consecutive rows; shifting the first group of storage units and the second group of storage units left by 1 bit or right by 1 bit to obtain the shifted first group of storage units and the shifted second group of storage units, using a selection instruction to select the bit stored in the s-th column storage unit in the shifted first group of storage units and the bit stored in the s-th column storage unit in the shifted second group of storage units, and based on the corresponding relationship, multiplying the bits in the selected storage units by the corresponding elements of the input data, and continuing the shift operation and the multiplication calculation until the selection and multiplication calculation of all columns in the first group of storage units and the second group of storage units are completed.

[0005] For example, in the data processing method provided in an embodiment of the present disclosure, N = 512, M = 1024, P = 256, i = j = Q / 2 = Z = 32.

[0006] For example, in the data processing method provided in an embodiment of the present disclosure, s = 16, the bit stored in the 16th column storage unit in the first group of storage units corresponds to the elements in the input data that are in the 1st column to the 32nd column of the X-th row, and the bit stored in the 16th column storage unit in the second group of storage units corresponds to the elements in the input data that are in the 1st column to the 32nd column of the (X + 1)-th row; the elements in the output data that are in consecutive i columns in every two consecutive rows include the elements in the 1st column to the 32nd column of the X-th row and the (X + 1)-th row in the output data; the bit stored in the 16th column storage unit in the shifted first group of storage units corresponds to the elements in the input data that are in the 33rd column to the 64th column of the X-th row, and the bit stored in the 16th column storage unit in the shifted second group of storage units corresponds to the elements in the input data that are in the 33rd column to the 64th column of the (X + 1)-th row, where X is a positive integer and X is an odd number.

[0007] For example, in the data processing method provided in an embodiment of the present disclosure, the first set of storage units includes 512 bits, the second set of storage units includes 512 bits, the second register includes 2 * 512 bits, and the 2 * 512 bits correspond to the elements in the first column to the 512th column of every two consecutive rows of the input data.

[0008] For example, in the data processing method provided in an embodiment of the present disclosure, the elements in the 513th column to the 1024th column of every two consecutive rows of the input data are in one-to-one correspondence with the 2 * 512 bits stored in the third register. The third register stores masking data that is different from but has the same size as the masking data stored in the second register, and the arrangement of the masking data in the third register is the same as the arrangement of the masking data in the second register.

[0009] For example, in the data processing method provided in an embodiment of the present disclosure, performing a multiplication calculation on the bits in the selected storage unit and the corresponding elements of the input data includes: when the value of the bit in the selected storage unit is 1, using the corresponding element of the input data as the corresponding element in the output data; when the value of the bit in the selected storage unit is 0, using 0 as the corresponding element in the output data.

[0010] For example, in the data processing method provided in an embodiment of the present disclosure, based on the corresponding relationship between the elements of the input data and the elements of the masking data bit by bit, performing a multiplication calculation to obtain the output data further includes: dividing the result obtained from the multiplication calculation by (1 - drop_prob) to obtain the output data, where drop_prob represents the probability that each bit of each element of the masking data is 0.

[0011] For example, the data processing method provided in an embodiment of the present disclosure further includes: storing the output data in the memory according to the corresponding positions of the input data.

[0012] For example, in the data processing method provided in an embodiment of the present disclosure, the format type of the select instruction is the same as the format type of the input data.

[0013] For example, in the data processing method provided in an embodiment of the present disclosure, the format type of the select instruction is BF16, and the format type of the input data is BF16.

[0014] For example, in the data processing method provided in an embodiment of the present disclosure, the data processing method is used for the calculation of the dropout layer of a neural network. The input part of the dropout layer includes input data and masking data. When the value of the bit of the element of the masking data is 1, the corresponding element of the input data is used as the output of the dropout layer; when the value of the bit of the element of the masking data is 0, the corresponding element of the input data is discarded.

[0015] For example, the data processing method provided by an embodiment of the present disclosure further includes: before loading the input data into the first register and loading the masking data into the second register, performing an alignment operation on the input data and the masking data, so that every 2*N elements in the input data correspond to every 1*Z elements in the masking data.

[0016] At least one embodiment of the present disclosure further provides a data processing apparatus, including: a data loading unit configured to load input data into a first register and load masking data into a second register, the size of the input data being N*M, the size of the masking data being P*Q, each element of the masking data being Z bits, Q*Z = 2*M, each bit of each element of the masking data corresponding to one element of the input data, the storage units of the second register being arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially corresponding to i consecutive elements in the same row of the input data, where N, M, P, Q, i, j, and Z are all positive integers; a calculation unit configured to perform a multiplication calculation based on the correspondence between each element of the input data and each bit of each element of the masking data to obtain output data.

[0017] At least one embodiment of the present disclosure further provides a data processing apparatus, including: a processor; and a memory storing computer-executable instructions that, when executed by the processor, implement the data processing method provided by at least one embodiment of the present disclosure.

[0018] At least one embodiment of the present disclosure further provides an electronic device including the data processing apparatus provided by at least one embodiment of the present disclosure.

[0019] At least one embodiment of the present disclosure further provides a computer-readable storage medium for non-transiently storing computer-executable instructions that, when executed by a processor, implement the data processing method provided by at least one embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0021] Figure 1 shows a schematic structural diagram of a register;

[0022] Figure 2 shows a schematic flowchart of a data processing method provided by at least one embodiment of the present disclosure;

[0023] Figure 3shows Figure 2 a schematic flowchart of an example of step S202 in

[0024] Figure 4 shows Figure 3 a schematic diagram of an example of step S301 in

[0025] Figure 5 shows Figure 3 a schematic diagram of an example of step S303 in

[0026] Figure 6 shows a schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure;

[0027] Figure 7 shows a schematic diagram of a data processing device according to an embodiment of the present disclosure;

[0028] Figure 8 shows a schematic diagram of an electronic device according to an embodiment of the present disclosure;

[0029] Figure 9 is a schematic diagram of a storage medium provided by some embodiments of the present disclosure. Detailed Description of the Embodiments

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0031] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms such as "include" or "comprise" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0032] The input part of the dropout layer includes input data and corresponding mask data. Each bit of each element of the mask data corresponds to an element of the input data. If the value of the bit of the element of the mask data is 1, the corresponding element of the input data needs to be retained. If the value of the bit of the element of the mask data is 0, the corresponding element of the input data will be changed to 0. Assume that the size of the input data is [512, 1024] and the data type is BF16, and the size of the mask data is [512, 32] and the data type is FP32. The 32 bits of each FP32 data in the mask data correspond to 32 elements in the same row of the input data. Each row in the mask data has 32 elements, that is, 32 * 32 = 1024 bits. These 1024 bits correspond to a complete row in the input data, that is, 1024 elements of the input data. The mask data will be read from memory into the register, and the schematic structure of the register is as Figure 1 shown. Each register includes 32 channels (Lane0 to Lane31), each channel contains a 32-bit (Bit0 to Bit31) number, and each bit corresponds to an element of the input data. Therefore, a register can just store a complete row of elements in the mask data, that is, 32 elements of FP32.

[0033] Generally speaking, the 32-bit numbers of each channel in the register are corresponding to 32 elements horizontally in the input data. Since the format of the input data is BF16, after being read into the register, it contains two rows, that is, 2 * 32 BF16 data. In order to correspond one by one with the bits of the elements of the mask data, the register storing the input data needs to be split into two groups of independent 1 * 32 FP32 data first. Then, only 32 bits of one channel in the register storing the mask data can be read and stored in the scalar register each time. Then, the masked data is used to perform a selective multiplication operation on the split input data. If the bit of the element of the mask data is 1, the corresponding element of the input data is retained. If the bit of the element of the mask data is 0, the corresponding element of the input data is changed to 0. Then, the calculation results of the two groups of 1 * 32 FP32 types are combined into a group of 2 * 32 BF16 type output data. This calculation process has complex instructions and extremely low efficiency.

[0034] At least one embodiment of the present disclosure provides a data processing method, including: loading input data into a first register and loading masking data into a second register, where the size of the input data is N*M, the size of the masking data is P*Q, each element of the masking data is Z bits, Q*Z = 2*M, each bit of each element of the masking data corresponds to an element of the input data, the storage units of the second register are arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially correspond to i consecutive elements in the same row of the input data, and N, M, P, Q, i, j, and Z are all positive integers; performing a multiplication calculation based on the corresponding relationship between each element of the input data and each bit of each element of the masking data to obtain output data.

[0035] The data processing method provided by the above embodiment of the present disclosure changes the corresponding manner between the masking data and the input data, effectively reduces the number of instructions, greatly improves the running efficiency, and reduces the running time.

[0036] At least one embodiment of the present disclosure further provides a data processing device, an electronic device, and a computer-readable storage medium corresponding to the above data processing method.

[0037] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0038] Figure 2 A schematic flowchart of a data processing method provided by at least one embodiment of the present disclosure is shown.

[0039] As Figure 2 shown, the data processing method includes the following steps S201 to S202.

[0040] Step S201: Loading input data into a first register and loading masking data into a second register.

[0041] For example, the size of the input data is N*M, the size of the masking data is P*Q, each element of the masking data is Z bits, Q*Z = 2*M, each bit of each element of the masking data corresponds to an element of the input data, the storage units of the second register are arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially correspond to i consecutive elements in the same row of the input data, and N, M, P, Q, i, j, and Z are all positive integers.

[0042] For example, N = 512, M = 1024, P = 256, i = j = Q / 2 = Z = 32. That is, the size of the input data is 512 * 1024, the size of the masking data is 256 * 64, each element of the masking data is 32 bits, the number of columns of the masking data (64) multiplied by the number of bits of each element of the masking data (32) is equal to twice the number of columns of the input data (2048), the number of rows of the storage units of the second register (32) is equal to the number of columns of the storage units of the second register (32), and is equal to half the number of columns of the masking data (32), and is equal to the number of bits of each element of the masking data (32). The above specific values are only examples, and the present disclosure does not limit the specific values of N, M, P, Q, i, j, and Z, as long as the above relationships are satisfied.

[0043] Step S202: Based on the correspondence between each element of the input data and each bit of each element of the masking data, perform a multiplication calculation to obtain the output data.

[0044] For example, before step S201, the data processing method provided by the embodiments of the present disclosure may further include: performing an alignment operation on the input data and the masking data, so that every 2 * N elements in the input data correspond to every 1 * Z elements in the masking data.

[0045] For example, N = 512, Z = 32, and the number of bits of each element of the masking data is 32. Then, every 2 * 512 elements in the input data (2 rows and 512 columns, that is, 1024 elements) correspond to every 1 * 32 elements in the masking data (1 row and 32 columns, that is, 32 elements, and the number of bits of 32 elements is 32 * 32 = 1024).

[0046] The present disclosure adjusts the correspondence between each element of the input data and each bit of each element of the masking data, that is, instead of corresponding j bits of each channel in the second register to j consecutive elements horizontally in the input data, it corresponds i bits stored in each column of the storage units in the second register to i consecutive elements horizontally in the input data, so as to facilitate subsequent calculations of each bit of the masking data and each element of the input data. For example, it is convenient to batch-read multiple bits of the masking data in combination with the select instruction, avoiding format adjustment and conversion of the read bits, thereby effectively reducing the number of instructions and improving the running efficiency.

[0047] For example, in some embodiments of the present disclosure, the second register is divided into two groups of storage units. The first column to the j / 2 column of the second register are the first group of storage units, and the (j / 2 + 1) column to the j column of the second register are the second group of storage units. Thus, the selection and calculation of each bit can be implemented in combination with the select instruction, thereby improving efficiency and facilitating implementation.

[0048] Figure 3 shows Figure 2 a schematic flowchart of an example of step S202 in

[0049] As Figure 3 shown, for the second register divided into two groups of storage units, an example of step S202 may include the following steps S301 to S303.

[0050] Step S301: Use a selection instruction to select the bits stored in the s-th column storage units in the first group of storage units and the bits stored in the s-th column storage units in the second group of storage units, where the bits stored in the s-th column storage units in the first group of storage units and the bits stored in the s-th column storage units in the second group of storage units correspond to the elements in the input data located in consecutive i columns in every two consecutive rows.

[0051] For example, in some embodiments of the present disclosure, the format type of the selection instruction is the same as the format type of the input data.

[0052] For example, in some embodiments of the present disclosure, the format type of the selection instruction is BF16, and the format type of the input data is BF16.

[0053] Since the format type of the selection instruction is the same as the format type of the input data, the selection operation can be directly performed on the input data, which is more concise and efficient.

[0054] Figure 4 shows Figure 3 a schematic diagram of an example of step S301 in

[0055] As Figure 4 shown, the storage units of the second register are arranged in 32 rows and 32 columns (the second register has 32 channels (Lane0 to Lane31), and each channel has 32 bits (bit0 to bit32)). The 1st to 16th columns of the second register are the first group of storage units, and the 17th to 32nd columns of the second register are the second group of storage units. Use a selection instruction to select the highest bit positions of the first group of storage units and the second group of storage units, that is, the bits stored in the 16th column storage units in the first group of storage units and the bits stored in the 16th column storage units in the second group of storage units. For example, the bits stored in the 16th column storage units in the first group of storage units and the bits stored in the 16th column storage units in the second group of storage units correspond to the elements in the input data located in the 1st to 32nd columns of the first row and the second row.

[0056] It should be noted that the present disclosure does not limit which column storage units in the first group of storage units and the second group of storage units are selected, as long as the selection is made in units of columns.

[0057] Step S302: Based on the corresponding relationship, perform a multiplication calculation on the bits in the selected storage units and the elements of the corresponding input data to obtain the elements corresponding to the output data that are located in consecutive i columns in every two consecutive rows.

[0058] For example, in some embodiments of the present disclosure, step S302 may include: when the value of the bit in the selected storage unit is 1, use the corresponding element of the input data as the corresponding element in the output data; when the value of the bit in the selected storage unit is 0, use 0 as the corresponding element in the output data.

[0059] Step S303: Shift the first group of storage units and the second group of storage units one bit to the left or one bit to the right to obtain the shifted first group of storage units and the shifted second group of storage units. Use the selection instruction to select the bits stored in the s-th column storage units of the shifted first group of storage units and the bits stored in the s-th column storage units of the shifted second group of storage units, and perform a multiplication calculation on the bits in the selected storage units and the elements of the corresponding input data based on the corresponding relationship. Continue the shift operation and multiplication calculation until the selection and multiplication calculation of all columns in the first group of storage units and the second group of storage units are completed.

[0060] Figure 5 shows Figure 3 a schematic diagram of an example of step S303 in

[0061] As Figure 5 shown, the size of the second register is the same as that of the second register shown in Figure 4 The selection instruction only selects the bits stored in the 16th column (bit15) storage units of the first group of storage units and the bits stored in the 16th column (bit31) storage units of the second group of storage units each time. After the selection operation is completed, step S302 is executed. After step S302 is executed, the first group of storage units and the second group of storage units are shifted one bit to the left. In this way, bit30 will be moved to the position of bit31, bit14 will be moved to the position of bit15, and the remaining bits are shifted one bit to the left in sequence. Then, the multiplication calculation of step 302 is performed until the selection and multiplication calculation of all columns in the first group of storage units and the second group of storage units are completed.

[0062] It should be noted that the first group of storage units and the second group of storage units can also be shifted right by 1 bit, which can be adjusted according to the corresponding relationship between the bits of the elements of the input data and the bits of the masking data. The present disclosure places no restrictions on this. In other embodiments, if a selection instruction is not used but other means are adopted to select the bits in the column direction, the shifting operation can also be omitted as long as the bits of each column can be selected in sequence. The embodiments of the present disclosure place no restrictions on this.

[0063] Return to Figure 3 , for example, in some embodiments of the present disclosure, step S202 may further include step S304: dividing the result obtained from the multiplication calculation by (1 - drop_prob) to obtain the output data, where drop_prob represents the probability that each bit of each element of the masking data is 0.

[0064] For example, each bit of each element of the masking data has a certain probability of being 0. Here, this probability is represented by drop_prob, then the probability that each bit of each element of the masking data is 1 is (1 - drop_prob). Usually, the calculation result is scaled, that is, multiplied by 1 / (1 - drop_prob) or divided by (1 - drop_prob).

[0065] For example, the data processing method provided by the embodiments of the present disclosure may further include: storing the output data in the memory according to the corresponding positions of the input data.

[0066] Since the size and data type of the output data are exactly the same as those of the input data, after obtaining the output data, it is only necessary to store the output data in the memory according to the same coordinates as those of the currently processed input data.

[0067] Next, a specific embodiment is used to illustrate the data processing method provided by the present disclosure.

[0068] For example, in an embodiment of the present disclosure, N = 512, M = 1024, P = 256, i = j = Q / 2 = Z = 32. That is, the size of the input data is 512 * 1024, the size of the masking data is 256 * 64, each element of the masking data is 32 bits, and the storage units of the second register are arranged in 32 rows and 32 columns. The format type of the input data is BF16, the format type of the masking data is FP32, and the format type of the selection instruction is BF16.

[0069] For example, the first 16 columns of the second register are the first group of storage units, and the 17th to 32nd columns of the second register are the second group of storage units

[0070] First, load the input data into the first register and load the masking data into the second register.

[0071] Then, a selection instruction is used to select the bits stored in the 16th column memory cells in the first group of memory cells and the bits stored in the 16th column memory cells in the second group of memory cells, that is, the 16th and 32nd columns in the 32-column memory cells of the second register are selected by using the selection instruction.

[0072] Then, based on the corresponding relationship, the bits in the selected memory cells and the elements of the corresponding input data are multiplied to obtain the elements in the corresponding continuous 32 columns in every two consecutive rows of the output data. The elements in the corresponding continuous i columns in every two consecutive rows of the obtained output data include the elements in the first i columns of the Xth row and the (X + 1)th row in the output data.

[0073] For example, the bits stored in the 16th column memory cells in the first group of memory cells correspond to the elements in the 1st to 32nd columns of the Xth row in the input data, and the bits stored in the 16th column memory cells in the second group of memory cells correspond to the elements in the 1st to 32nd columns of the (X + 1)th row in the input data, where X is a positive integer and X is odd. For example, when X = 1, the bits stored in the 16th column memory cells in the first group of memory cells correspond to the elements in the 1st to 32nd columns of the 1st row in the input data, and the bits stored in the 16th column memory cells in the second group of memory cells correspond to the elements in the 1st to 32nd columns of the 2nd row in the input data. The bits in the selected memory cells and the elements of the corresponding input data are multiplied to obtain the elements in the 1st to 32nd columns of the 1st row and the 2nd row in the output data.

[0074] Then, the first group of memory cells and the second group of memory cells are shifted left by 1 bit to obtain the shifted first group of memory cells and the shifted second group of memory cells. A selection instruction is used to select the bits stored in the 16th column memory cells in the shifted first group of memory cells and the bits stored in the 16th column memory cells in the shifted second group of memory cells, which is equivalent to selecting the 15th and 31st columns in the 32-column memory cells of the second register before shifting. Then, based on the corresponding relationship, the bits in the selected memory cells and the elements of the corresponding input data are multiplied. In a similar manner, shifting, selection, and multiplication are alternately performed until the selection and multiplication of all columns in the first group of memory cells and the second group of memory cells are completed.

[0075] For example, the bits stored in the 16th column storage unit of the first group of shifted storage units correspond to the elements in the input data located in the 33rd to 64th columns of the Xth row, and the bits stored in the 16th column storage unit of the second group of shifted storage units correspond to the elements in the input data located in the 33rd to 64th columns of the (X + 1)th row, where X is a positive integer and X is odd. For example, when X = 1, the bits stored in the 16th column storage unit of the first group of shifted storage units correspond to the elements in the input data located in the 33rd to 64th columns of the 1st row, and the bits stored in the 16th column storage unit of the second group of shifted storage units correspond to the elements in the input data located in the 33rd to 64th columns of the 2nd row. The first group of storage units contains 512 bits, the second group of storage units contains 512 bits, the second register contains 2 * 512 bits, and the 2 * 512 bits correspond to the elements in the 1st to 512th columns of two consecutive rows in the input data. The elements in the 513th to 1024th columns of every two consecutive rows in the input data are corresponded to by another register. For example, the elements in the 513th to 1024th columns of every two consecutive rows in the input data are corresponded to one by one by the 2 * 512 bits stored in the third register. The third register stores masking data that is different from but has the same size as the masking data stored in the second register, and the arrangement manner of the masking data in the third register is the same as that of the masking data in the second register.

[0076] For example, in the initial state, the bits stored in the 16th column storage unit of the first group of storage units correspond to the elements in the input data located in the 1st to 32nd columns of the 1st row, and the bits stored in the 16th column storage unit of the second group of storage units correspond to the elements in the input data located in the 1st to 32nd columns of the 2nd row; shift the first group of storage units and the second group of storage units to the left by 1 bit. At this time, the bits stored in the 16th column storage unit of the first group of storage units correspond to the elements in the input data located in the 33rd to 64th columns of the 1st row, and the bits stored in the 16th column storage unit of the second group of storage units correspond to the elements in the input data located in the 33rd to 64th columns of the 2nd row; then shift the first group of storage units and the second group of storage units to the left by 1 bit again. At this time, the bits stored in the 16th column storage unit of the first group of storage units correspond to the elements in the input data located in the 65th to 96th columns of the 1st row, and the bits stored in the 16th column storage unit of the second group of storage units correspond to the elements in the input data located in the 65th to 96th columns of the 2nd row; and so on until the selection and product calculation of all columns in the first group of storage units and the second group of storage units are completed. The 512 bits in the first group of storage units correspond to the elements in the 1st to 512th columns of the 1st row in the input data, and the 512 bits in the second group of storage units correspond to the elements in the 1st to 512th columns of the 2nd row in the input data.

[0077] Then, divide the result obtained from the product calculation by (1 - drop_prob) to obtain the output data.

[0078] Finally, store the output data in memory at the corresponding positions of the input data.

[0079] The data processing method provided by the embodiments of the present disclosure effectively reduces the number of assembly instructions and greatly improves the operation efficiency. Assuming the input data for four registers, that is, 8 * 32 BF16 - type input data, according to the usual data processing method, it requires approximately 38 assembly instructions, while according to the data processing method provided by the embodiments of the present disclosure, only 6 assembly instructions are needed to complete, and the optimization effect is remarkable.

[0080] For example, in some embodiments of the present disclosure, the data processing method is used for the calculation of the dropout layer of a neural network.

[0081] For a standard neural network, the training process of the neural network is as follows: First, pass the input data through the forward propagation of the neural network, and then backpropagate the loss result to determine how to update the parameters of the neural network so that the neural network can learn. After using the dropout layer, the training process becomes: First, randomly delete half of the hidden neurons in the neural network, and keep the input neurons and output neurons unchanged; then pass the input data through the modified neural network for forward propagation, and then backpropagate the obtained loss result through the modified neural network. After a small batch of training samples complete this process, update the corresponding parameters on the neurons that have not been deleted according to the stochastic gradient descent method; then continue to repeat this process.

[0082] The input part of the dropout layer includes input data and masking data. When the value of the bit of an element in the masking data is 1, the corresponding element of the input data is used as the output of the dropout layer; when the value of the bit of an element in the masking data is 0, the corresponding element of the input data is discarded (such as the deleted neurons).

[0083] As mentioned above, divide the result obtained from the product calculation by (1 - drop_prob) to obtain the output data because the dropout layer needs to perform scaling. During training, some neurons are randomly discarded, but during prediction, it is not possible to randomly discard. If some neurons are discarded, it will cause the problem of unstable results and inaccurate model prediction. One solution is to multiply the weight of each neuron by a probability so that the prediction data and the training data are roughly the same. For example, if the output of a neuron is x, then during training, it has a probability of drop_prob of being discarded and a probability of (1 - drop_prob) of participating in training. Therefore, during prediction, divide the output of this neuron by (1 - drop_prob).

[0084] Figure 6 The schematic block diagram of a data processing device 600 provided by at least one embodiment of the present disclosure is shown. This data processing device can be used to execute Figure 2 the data processing method shown.

[0085] As Figure 6 shown, the data processing device 600 includes a data loading unit 601 and a calculation unit 602.

[0086] The data loading unit 601 is configured to load input data into the first register and load masking data into the second register. The size of the input data is N*M, the size of the masking data is P*Q, each element of the masking data is Z bits, Q*Z = 2*M, each bit of each element of the masking data corresponds to an element of the input data, the storage units of the second register are arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially correspond to i consecutive elements in the same row of the input data. N, M, P, Q, i, j, and Z are all positive integers.

[0087] The calculation unit 602 is configured to perform a multiplication calculation based on the corresponding relationship between each element of the input data and each bit of each element of the masking data to obtain output data.

[0088] The data loading unit 601 can implement Figure 2 step S201 in the data processing method shown, and the calculation unit 602 can implement Figure 2 step S202 in the data processing method shown. For relevant descriptions, reference can be made to the above content and will not be elaborated here. This data processing device 600 and Figure 2 the data processing method shown have the same technical effects and will not be elaborated here.

[0089] For example, the data processing device can be implemented by using hardware, software, firmware, and any feasible combination thereof. The present disclosure does not limit this.

[0090] For example, the data loading unit 601 and the calculation unit 602 can be hardware, software, firmware, and any feasible combination thereof. For example, the data loading unit 601 and the calculation unit 602 can be dedicated or general-purpose circuits, chips, or devices, etc., or can also be a combination of a processor and a memory. Regarding the specific implementation forms of the data loading unit 601 and the calculation unit 602, the embodiments of the present disclosure do not limit this.

[0091] It should be noted that in the embodiments of the present disclosure, each unit of the data processing device 600 corresponds to each step of the foregoing data processing method. For the specific functions of the data processing device 600, reference may be made to the relevant descriptions of the data processing method in the foregoing text, and details are not described herein again. Figure 6 The components and structures of the data processing device 600 shown are exemplary and not restrictive. According to needs, the data processing device 600 may further include other components and structures.

[0092] At least one embodiment of the present disclosure further provides a data processing device, including: a memory for non-temporarily storing computer-executable instructions; and a processor for running the computer-executable instructions, wherein the computer-executable instructions, when run by the processor, execute the data processing method provided by at least one embodiment of the present disclosure.

[0093] Figure 7 A schematic diagram of a data processing device 700 according to an embodiment of the present disclosure is shown. As Figure 7 shown, the data processing device 700 according to an embodiment of the present disclosure may include a processing device 701 and a memory 702, which may be interconnected through a bus 703.

[0094] The processing device 701 may perform various actions and processes according to programs or codes stored in the memory 702. Specifically, the processing device 701 may be an integrated circuit chip with signal processing capabilities. For example, the foregoing processing device may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute various methods, steps, processes, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be an X86 architecture or an ARM architecture, etc.

[0095] The memory 702 stores computer-executable instructions, which, when executed by the processing device 701, implement the data processing method provided by at least one embodiment of the present disclosure. The memory 702 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories.

[0096] At least one embodiment of the present disclosure further provides an electronic device, including the data processing device provided by at least one embodiment of the present disclosure. In one embodiment, the electronic device is, for example, a central processing unit, and the processor is, for example, a single-core or multi-core processor. In one embodiment, the electronic device is a computer system, and the computer system includes one or more processors.

[0097] Figure 8 A schematic diagram of an electronic device 800 according to an embodiment of the present disclosure is shown. As Figure 8 shown, the electronic device 800 according to an embodiment of the present disclosure may include a data processing device 600.

[0098] At least one embodiment of the present disclosure provides a computer-readable storage medium for non-transiently storing computer-executable instructions, which, when executed by a processor, implement the data processing method provided by at least one embodiment of the present disclosure.

[0099] Figure 9 A schematic diagram of a storage medium provided for some embodiments of the present disclosure is shown. As Figure 9 shown, the storage medium 900 is used to store computer-executable instructions 910. For example, when the computer-executable instructions 910 are executed by a computer, one or more steps of the data processing method described above can be executed.

[0100] Similarly, the computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. It should be noted that the memories of the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0101] Embodiments of the present disclosure also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method according to the embodiments of the present disclosure.

[0102] The technical effects of the above data processing device, electronic device, and storage medium are the same as those of Figure 2 the data processing method shown, and will not be elaborated here.

[0103] The following points need to be explained:

[0104] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0105] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0106] As described above, the above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A data processing method, comprising: Loading input data into a first register and loading masking data into a second register, wherein the size of the input data is N*M, the size of the masking data is P*Q, each element of the masking data is Z bits, Q*Z = 2*M, N = 2P, each bit of each element of the masking data corresponds to an element of the input data, the storage units of the second register are arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially correspond to i consecutive elements in the same row of the input data; N, M, P, Q, i, j, Z are all positive integers; Performing a multiplication calculation based on the correspondence between each element of the input data and each bit of each element of the masking data to obtain output data.

2. The data processing method according to claim 1, wherein, The first j / 2 columns of the second register are the first group of storage units, and the (j / 2 + 1)-th column to the j-th column of the second register are the second group of storage units; Performing a multiplication calculation based on the correspondence between each element of the input data and each bit of each element of the masking data to obtain the output data, including: Using a selection instruction to select the bits stored in the s-th column storage unit of the first group of storage units and the bits stored in the s-th column storage unit of the second group of storage units, wherein the bits stored in the s-th column storage unit of the first group of storage units and the bits stored in the s-th column storage unit of the second group of storage units correspond to the elements in i consecutive columns in every two consecutive rows of the input data; Based on the correspondence, multiplying the bits in the selected storage units by the corresponding elements of the input data to obtain the elements in i consecutive columns in every two consecutive rows of the output data; Shifting the first group of storage units and the second group of storage units left by 1 bit or right by 1 bit to obtain the shifted first group of storage units and the shifted second group of storage units, using the selection instruction to select the bits stored in the s-th column storage unit of the shifted first group of storage units and the bits stored in the s-th column storage unit of the shifted second group of storage units, and based on the correspondence, multiplying the bits in the selected storage units by the corresponding elements of the input data, and continuing the shift operation and multiplication calculation until the selection and multiplication calculation of all columns in the first group of storage units and the second group of storage units are completed.

3. The data processing method according to claim 2, wherein, N = 512, M = 1024, P = 256, i = j = Q / 2 = Z = 32.

4. The data processing method according to claim 3, wherein s = 16, the bits stored in the 16th column storage unit of the first group of storage units correspond to the elements in the 1st column to the 32nd column of the X-th row of the input data, and the bits stored in the 16th column storage unit of the second group of storage units correspond to the elements in the 1st column to the 32nd column of the (X + 1)-th row of the input data; The corresponding elements in every two consecutive rows and consecutive i columns in the obtained output data include the elements in the 1st to 32nd columns of the Xth row and the (X + 1)th row in the output data; The bits stored in the 16th column memory cell in the first group of shifted memory cells correspond to the elements in the 33rd to 64th columns of the Xth row in the input data, and the bits stored in the 16th column memory cell in the second group of shifted memory cells correspond to the elements in the 33rd to 64th columns of the (X + 1)th row in the input data, where X is a positive integer and X is odd.

5. The data processing method according to claim 4, wherein, The first group of memory cells contains 512 bits, the second group of memory cells contains 512 bits, the second register contains 2 * 512 bits, and the 2 * 512 bits correspond to the elements in the 1st to 512th columns in every two consecutive rows of the input data.

6. The data processing method according to claim 5, wherein The elements in the 513th to 1024th columns in every two consecutive rows of the input data are in one-to-one correspondence with the 2 * 512 bits stored in the third register. The third register stores masking data that is different from but has the same size as the masking data stored in the second register, and the arrangement of the masking data in the third register is the same as that of the masking data in the second register.

7. The data processing method according to claim 2, wherein Performing a multiplication calculation on the bits in the selected memory cells and the corresponding elements of the input data includes: When the value of the bit in the selected memory cell is 1, using the corresponding element of the input data as the corresponding element in the output data; When the value of the bit in the selected memory cell is 0, using 0 as the corresponding element in the output data.

8. The data processing method according to claim 7, wherein, Based on the corresponding relationship between each element of the input data and each bit of each element of the masking data, performing a multiplication calculation to obtain the output data further includes: Dividing the result obtained from the multiplication calculation by (1 - drop_prob) to obtain the output data, where drop_prob represents the probability that each bit of each element of the masking data is 0.

9. The data processing method according to any one of claims 2 - 8 further includes: Storing the output data in the memory at the corresponding position of the input data.

10. The data processing method according to claim 9, wherein, The format type of the selection instruction is the same as the format type of the input data.

11. The data processing method according to claim 10, wherein, The format type of the selection instruction is BF16, and the format type of the input data is BF16.

12. The data processing method according to claim 11, wherein, The data processing method is used for the calculation of the dropout layer of a neural network. The input part of the dropout layer includes the input data and the masking data. When the value of the bit of the element of the masking data is 1, the corresponding element of the input data is used as the output of the dropout layer; When the value of the bit of the element of the masking data is 0, the corresponding element of the input data is discarded.

13. The data processing method according to claim 1 further includes: Before loading the input data into the first register and loading the masking data into the second register, Perform an alignment operation on the input data and the masking data such that every 2*N elements in the input data correspond to every 1*Z elements in the masking data.

14. A data processing device, comprising: A data loading unit configured to load the input data into a first register and load the masking data into a second register, wherein the size of the input data is N*M, the size of the masking data is P*Q, each element of the masking data is Z bits, Q*Z = 2*M, each bit of each element of the masking data corresponds to an element of the input data, the storage units of the second register are arranged in i rows and j columns, i = j = Q / 2 = Z, and the i bits stored in the same column of the second register sequentially correspond to i consecutive elements in the same row of the input data, where N, M, P, Q, i, j, and Z are all positive integers; A calculation unit configured to perform a multiplication calculation based on the correspondence between each element of the input data and each bit of each element of the masking data to obtain output data.

15. A data processing device, comprising: A processor; And A memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the processor, implement the data processing method according to any one of claims 1-13.

16. An electronic device, comprising the data processing device according to claim 15.

17. A computer-readable storage medium for non-transiently storing computer-executable instructions, Among them, wherein the computer-executable instructions, when executed by a processor, implement the data processing method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Data processing method and device, chip and computer readable storage medium

    CN111428879A

  • Efficient multiplication of small matrices using SIMD registers

    US20040122887A1