Data processing method and device, electronic equipment and storage medium
By splitting the target operator into multiple suboperators and allocating them to multiple processor cores for processing, the performance reduction problem caused by the insertion of layout convert operators in the prior art is solved, and more efficient data processing performance is achieved.
Patent Information
- Application Number
- CN202510361496.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the insertion of the layout convert operator will cause the problem of mismatch in data layout formats, resulting in performance degradation.
By dividing the target operator into multiple suboperators and assigning these suboperators to multiple cores of the neural network processor for parallel processing, the data layout format is directly converted to avoid inserting the layout convert operator.
It reduces the time-consuming process caused by intensive data access by layout convert and improves the performance of neural network processors.
Smart Images

Figure CN120218149A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology. Specifically, the present disclosure relates to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In a deep learning neural network model, for higher performance, a multi-layout mechanism is generally adopted, that is, different operators may support different data layout formats, so there may be a situation where the data layout formats between two adjacent operators do not match.
[0003] In the prior art, a layout convert operator (an operator for adjusting the memory layout of a tensor) is usually inserted between two operators. The layout convert operator is a memory access-intensive operator for pure memory operations. The insertion of the layout convert operator will bring a large amount of time-consuming data access, thus reducing the performance. Summary of the Invention
[0004] Embodiments of the present disclosure provide a data processing method, apparatus, electronic device, and storage medium, which can solve the problem that the insertion of the layout convert operator in the prior art causes performance degradation. The technical solutions provided by the present disclosure are as follows: According to one aspect of the embodiments of the present disclosure, a data processing method is provided. The method includes: Obtain initial data to be processed and a target operator corresponding to the initial data; Use the quotient obtained by dividing the number of output channels of the target operator by a preset number as a first number, and divide the target operator into the first number of sub-operators; Allocate the first number of sub-operators to at least two target cores in a neural network processor; wherein, multiple cores in the neural network processor at least include the at least two target cores; For each sub-operator in each target core, execute the operation of the sub-operator on the initial data through the target core, and use the result obtained by the operation as sub-data corresponding to the sub-operator; Concatenate the sub-data corresponding to each sub-operator respectively, and use the concatenated data as target data corresponding to the initial data; Wherein, the initial data is stored and laid out based on a first memory layout format; the target data is stored and laid out based on a second memory layout format; the second memory layout format is obtained by grouping in the channel dimension based on the first memory layout format with the preset number.
[0005] Optionally, the allocating the first number of sub-operators to at least two target cores in the neural network processor includes: Determining a second number of cores in the neural network processor; Based on the magnitude relationship between the first number and the second number, determining at least one target core in each core and the number of sub-operators to be processed corresponding to each target core; the number of sub-operators to be processed corresponding to the target core is not zero; For each target core, determining the number of sub-operators corresponding to the target core from the first number of sub-operators; where at least one sub-operator corresponding to each target core does not overlap; For each target core, allocating the number of sub-operators corresponding to the target core to the target core.
[0006] Optionally, the based on the magnitude relationship between the first number and the second number, determining at least one target core in each core and the number of sub-operators to be processed corresponding to each target core includes: If the first number is less than the second number, taking the first number of cores in the second number of cores as target cores and setting the number of sub-operators to be processed for each target core to 1; If the first number is not less than the second number, taking each core as a target core and determining the number of sub-operators to be processed corresponding to each target core based on the first number and the second number.
[0007] Optionally, the based on the first number and the second number, determining the number of sub-operators to be processed corresponding to each target core includes: Taking the calculation result of performing integer division of the first number by the second number as a first value; Taking the calculation result of performing a remainder operation of the first number by the second number as a second value; If the second value is zero, taking the first value as the number of sub-operators to be processed corresponding to each target core; If the second value is not zero, determining the second number of first target cores from the target cores, setting the number of sub-operators to be processed corresponding to the first target core to a third value; setting the number of sub-operators to be processed corresponding to each second target core among the target cores to the first value; The second target core is the target core other than the first target core among the target cores; the third value is 1 greater than the first value.
[0008] Optionally, each target core corresponds to a core serial number; Determining the second numerical number of first target cores from the respective target cores includes: Regarding the second numerical number of target cores with the earliest core sequence numbers among the respective target cores as the first target cores.
[0009] Optionally, the concatenating of the respective sub - data corresponding to each sub - operator includes: Determining the order relationship corresponding to each sub - operator; Sequentially concatenating the respective sub - data corresponding to each sub - operator according to the order relationship corresponding to each sub - operator.
[0010] Optionally, there is a connection relationship between the target operator and the adjacent operator of the target operator; the target operator is used to process data stored in a layout based on the first memory layout format, and the adjacent operator is used to process data stored in a layout based on the second memory layout format.
[0011] According to another aspect of the embodiments of the present disclosure, there is provided a data processing apparatus, and the apparatus includes: An acquisition module, configured to acquire initial data to be processed and the target operator corresponding to the initial data; A splitting module, configured to use the quotient obtained by dividing the number of output channels of the target operator by a preset number as the first number, and split the target operator into the first number of sub - operators; An allocation module, configured to allocate the first number of sub - operators to at least two target cores in a neural network processor; where multiple cores in the neural network processor at least include the at least two target cores; An operation module, configured to, for each sub - operator in each target core, execute the operation of the sub - operator on the initial data through the target core, and use the result obtained from the operation as the sub - data corresponding to the sub - operator; A concatenating module, configured to concatenate the respective sub - data corresponding to each sub - operator, and use the concatenated data as the target data corresponding to the initial data; Wherein, the initial data is stored in a layout based on the first memory layout format; the target data is stored in a layout based on the second memory layout format; the second memory layout format is obtained by grouping in the channel dimension based on the first memory layout format with the preset number.
[0012] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, and the electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any of the above - mentioned data processing methods are implemented.
[0013] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above data processing methods are implemented.
[0014] According to one aspect of the embodiments of the present disclosure, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above data processing methods are implemented.
[0015] The beneficial effects brought by the technical solutions provided by the embodiments of the present disclosure are as follows: By splitting the target operator into a first number of sub-operators, each sub-operator operates on the initial data to obtain corresponding sub-data, and the sub-data are spliced together. The spliced data is used as the target data, so that the target data can be directly converted into the second memory layout format without inserting a layout convert operator, reducing the time consumption caused by the layout convert intensive memory access data and improving the performance of the neural network processor.
[0016] By allocating the first number of sub-operators to multiple target cores in the neural network processor, the task of the target operator is divided into sub-tasks of multiple sub-operators. By allocating multiple sub-tasks to multiple cores for parallel processing, the performance of the neural network processor is further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description in the embodiments of the present disclosure.
[0018] Figure 1 It is a schematic diagram of an NHWC memory layout format; Figure 2 It is a schematic diagram of an NC’HWC32 memory layout format; Figure 3 It is a schematic diagram of the calculation process of a neural network processor; Figure 4 It is a schematic diagram of the process of a data processing method provided by the embodiments of the present disclosure; Figure 5 It is a schematic diagram of the mapping relationship between sub-data and sub-operators provided by the embodiments of the present disclosure; Figure 6 It is a schematic diagram of the allocation of sub-operators provided by the embodiments of the present disclosure; Figure 7 It is another schematic diagram of the allocation of sub-operators provided by the embodiments of the present disclosure; Figure 8Schematic diagram of the calculation process of another neural network processor provided by an embodiment of the present disclosure; Figure 9 Schematic diagram of the structure of a data processing device provided by an embodiment of the present disclosure; Figure 10 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0019] The embodiments of the present disclosure will be described below with reference to the accompanying drawings in the present disclosure. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure, and do not constitute limitations on the technical solutions of the embodiments of the present disclosure.
[0020] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" or "A, B" indicates being implemented as "A", or being implemented as "B", or being implemented as "A and B".
[0021] To make the purpose, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be further described in detail below in conjunction with the accompanying drawings.
[0022] With the rapid development of Artificial Intelligence (AI), neural network models are being applied more and more widely in the field of artificial intelligence. In scenarios such as speech recognition and TTS (Text To Speech) in the speech field, object detection, segmentation, human skeleton key point detection, and face recognition in the video field, as well as in the field of large models, text-to-text and text-to-image are two of the most important basic main branches. All of this is based on the DeepLearning model, and for the model to achieve the final effect, it must be executed on hardware. The general approach is to execute on the GPU (Graphics Processing Unit), but with the emergence of the demand for high computing power and low power consumption, a large number of DSAs (Domain Specific Architecture) have started to appear.
[0023] When an AI model runs on specific hardware, it generally needs to store data in the types of Activation (i.e., activation data) and weights (i.e., weight data). In a deep learning framework, data generally has four dimensions, which can be represented as tensors. The memory layout format of data in the memory is generally NHWC or NCHW. Among them, N represents the number of batches, H represents the height of the tensor, W represents the width of the tensor, and C represents the number of input channels.
[0024] Figure 1 It is a schematic diagram of an NHWC memory layout format. Figure 1 In it, layout (i.e., data layout) refers to the arrangement order of multi-dimensional data in the memory. As Figure 1 shown, the storage order of the NHWC format in the memory is to first traverse the Figure 1 dimension as shown in step1 and store the data in continuous memory, then traverse the W dimension according to step2, that is, store the next data in continuous memory, and finally traverse the H dimension according to step3 to store the data in continuous memory.
[0025] In order to reduce the memory data reload of data (reduce power consumption and accelerate calculation) inside the DSA, a specific memory layout format is generally adopted, such as NCHWCx, NHWCCx, etc. Taking NHWCC32 (x = 32) as an example, it can also be expressed as NC’HWC32, where C’ = C / / x. Figure 2 Figure 2 It is a schematic diagram of an NC’HWC32 memory layout format. As shown, NC’HWC32 is in Slice along the dimension into chunks of 32 channels, and each chunk is stored in the NHWC format. Finally, store the next HWC32 chunk according to step 4.
[0026] Assume the convolution Conv has a size of k k s1 d1, and the output tensor is O: [1, , , .
[0027] In the case of the NHWC format layout (i.e., the storage layout is based on the NHWC memory layout format): Then, for each output point, it is necessary to read the input data , read the weight data , then perform multiply-accumulate (MAC), and finally obtain an output point. Traverse the output to get the size of the input data that needs to be read for the entire process is , the output data is , and finally traverse to get the total input data size is , and the output data size is ; In the case of the NC’HWC32 format layout (i.e., the storage layout is based on the NC’HWC32 memory layout format): (1) Read the input data , and read the weight ; (2) Slide the window of the input according to the specification parameters of the conv operator within the range to obtain the intermediate value temp_i; (3) Traverse the next convolution calculation and accumulate the result with the cached output intermediate value temp_i; (4) Until all are traversed to obtain the final output ; (5) Repeat steps (1) to (4), traverse the next until all calculations are completed; Among them, in each loop, when sliding the window , It can be reused. Therefore, without considering the input buffer (only considering the weight buffer and the output buffer), the amount of input data that needs to be read is , but the weight only needs to be read . Therefore, for the entire output process, the amount of input data that needs to be read is k k +k k , that is .
[0028] Comparing with the original data reading method of NHWC, the total change in the amount of input data read is:
[0029] It can be seen that the total amount of data read has almost been reduced by half, mainly due to the reuse of weights; in addition, adding an input buffer can further reduce the amount of input data read.
[0030] In a deep learning neural network model, for higher performance, a multi-layout mechanism is generally adopted, that is, different operators may adopt different data layout formats, so there may be a situation where the data layout formats between two adjacent operators do not match. Figure 3 It is a schematic diagram of the calculation process of a neural network processor, as Figure 3 shown. Operator 1 is connected to Operator 2, and the direction of the arrow represents the data flow direction, that is, the output of Operator 1 is the input of Operator 2. For example, in a neural network model, the activation layer is connected after the convolutional layer (that is, the output of the convolutional layer is the input of the activation layer). Assuming that the convolutional layer corresponds to Operator 1 (i.e., the convolutional operator) and the activation layer corresponds to Operator 2 (i.e., the activation operator), then Operator 1 and Operator 2 are connected. In the above scenario, it is possible that Operator 1 adopts the NHWC memory layout format and Operator 2 adopts the NC’HWC32 memory layout format, where layout-in represents the memory layout format of the input data and layout-out represents the memory layout format of the output data.
[0031] For the scenario where different operators may adopt different memory layout formats, in the existing technology, usually a layout convert operator (an operator for adjusting the memory layout of tensors) is inserted between the two operators. The layout convert operator is a memory access-intensive operator for pure memory operations, and the insertion of the layout convert operator will bring a large amount of time-consuming data access, thus reducing the performance.
[0032] The data processing method, apparatus, electronic device, and storage medium provided by the present disclosure aim to solve the above technical problems in the prior art.
[0033] The technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be described below by describing several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0034] Figure 4 It is a flowchart of a data processing method provided by an embodiment of the present disclosure. As Figure 4 shown, the method includes: Step S110, obtain the initial data to be processed and the target operator corresponding to the initial data.
[0035] Specifically, the data processing method provided by the embodiments of the present disclosure can be applied to an NPU (Neural Network Processing Unit). The NPU is a hardware acceleration unit designed specifically for accelerating artificial intelligence calculations (especially deep learning tasks). Its main goal is to accelerate the calculation process in neural network inference and training, and it is usually optimized for matrix calculations, convolution operations, activation functions, etc. in deep learning.
[0036] The initial data to be processed can be a tensor that needs to be operated on. The target operator corresponding to the initial data can be an operator that operates on the initial data. The target operator can be a convolution operator, an activation operator, a pooling operator, or a normalization operator, etc. The target operator can be specifically set according to the model structure in the actual application. The embodiments of the present disclosure do not limit the specific type of the target operator.
[0037] The operator that has a connection relationship with the target operator is used as the adjacent operator. That is, there is an edge connecting the node A corresponding to the target operator and the node B corresponding to the adjacent operator, and the arrow of the connecting edge points from node A to node B. Among them, the target operator is used to process the data stored in the first memory layout format, and the adjacent operator is used to process the data stored in the second memory layout format. The first memory layout format is different from the second memory layout format, that is, two adjacent operators support different memory layout formats respectively.
[0038] Step S120, use the quotient obtained by dividing the number of output channels of the target operator by a preset number as the first number, and divide the target operator into the first number of sub-operators; Among them, the initial data is stored and laid out based on the first memory layout format; the target data is stored and laid out based on the second memory layout format; the second memory layout format is obtained by grouping in the channel dimension with a preset number on the basis of the first memory layout format.
[0039] Specifically, the target operator is used to perform operations on the initial data. Based on the initial data and the target operator, the target data input to the adjacent operator can be obtained. Among them, the initial data is stored and laid out based on the first memory layout format, and the target data is stored and laid out based on the second memory layout format. That is to say, in the embodiments of the present disclosure, the target data stored and laid out based on the second memory layout format can be output, so that the target data can be directly input to the adjacent operator without inserting a layout convert operator, reducing the time consumption caused by the layout convert intensive memory access data and improving the performance of the neural network processor.
[0040] Among them, the second memory layout format can be obtained by grouping in the channel dimension with a preset number on the basis of the first memory layout format. The preset number can be determined based on the hardware architecture of the neural network processor. For example, the preset number can be set to 32 or 64.
[0041] For example, assuming that the preset number is 32, the first memory layout format can be NHWC, and the second memory layout format can be NC’HWC32, where C’ = C / / 32 (C’ is the quotient obtained by dividing the number of channels by the preset number). Combining Figure 1 and Figure 2 the provided examples, it can be seen that NC’HWC32 is grouped in the channel dimension with a preset number of 32 on the basis of NHWC, that is, 32 channels in NC’HWC32 are used as a group (i.e., Figure 2 C32 in
[0042] After determining the preset number, the quotient obtained by dividing the number of output channels of the target operator by the preset number is used as the first number, and the target operator is split based on the first number to obtain the first number of sub-operators. For example, if the number of output channels of the target operator is 64 and the preset number is 32, then the first number is 2. It should be noted that the output channel data of the target operator is an integer multiple of the preset number.
[0043] Step S130, allocate the first number of sub-operators to at least two target cores in the neural network processor; among them, multiple cores in the neural network processor at least include at least two target cores; Specifically, the neural network processor may include multiple cores (i.e., cores), and multiple cores at least include at least two target cores, and the target core is the core for executing the sub-operator.
[0044] A first quantity of sub-operators can be assigned to at least two target cores in a neural network processor. The task of the target operator is divided into subtasks of multiple sub-operators, and by assigning the multiple subtasks to multiple target cores for parallel processing, the performance of the neural network processor is further improved.
[0045] For each target core, at least one sub-operator assigned to the target core is used as at least one sub-operator corresponding to the target core.
[0046] Step S140: For each sub-operator in each target core, execute the operation of the sub-operator on the initial data through the target core, and use the result obtained from the operation as the sub-data corresponding to the sub-operator. Step S150: Concatenate the sub-data corresponding to each sub-operator respectively, and use the concatenated data as the target data corresponding to the initial data.
[0047] Specifically, for each sub-operator in each target core, read the initial data through the target core, execute the operation of the sub-operator on the initial data, and use the result obtained from the operation as the sub-data corresponding to the sub-operator, that is, the sub-data can be the output result of the initial data passing through the corresponding sub-operator.
[0048] After obtaining the sub-data output by each sub-operator, the sub-data corresponding to each sub-operator respectively can be concatenated, and the concatenated data is used as the target data corresponding to the initial data.
[0049] Optionally, concatenating the sub-data corresponding to each sub-operator respectively includes: Determine the order relationship corresponding to each sub-operator; Concatenate the sub-data corresponding to each sub-operator in sequence according to the order relationship corresponding to each sub-operator.
[0050] Specifically, Figure 5 FIG. is a schematic diagram of a mapping relationship between sub-data and sub-operators provided by an embodiment of the present disclosure. As Figure 5 shown, assume that the target operator is a convolution operator, the number of output channels of the convolution operator is 96, that is, the convolution operator includes 96 convolution kernels, the preset quantity is 32, then the first quantity is 3, the convolution operator is sliced into 3 sub-operators, and one sub-operator includes 32 convolution kernels, then the number of channels of the sub-data output by one sub-operator is 32, that is, Figure 5 C32 in
[0051] In Figure 5In the provided example, there is an order relationship among the three sub-operators after splitting, namely sub-operator 1 (for example, the 1st to 32nd convolutional kernels), sub-operator 2 (for example, the 33rd to 64th convolutional kernels), and sub-operator 3 (for example, the 65th to 96th convolutional kernels). Therefore, by sequentially concatenating the sub-data corresponding to each sub-operator according to the determined order relationship, the memory layout format of the concatenated data is converted to NC’HWC32 (i.e., the second memory layout format).
[0052] It should be noted that steps S110 - S130 can be executed offline by the neural network processor, and steps S140 - S150 can be executed online by the neural network processor.
[0053] In the embodiments of the present disclosure, by splitting the target operator into a first number of sub-operators, each sub-operator operates on the initial data to obtain the corresponding sub-data, and the sub-data are concatenated, and the concatenated data is used as the target data, so that the target data can be directly converted into the second memory layout format without inserting a layout convert operator, reducing the time consumption caused by the layout convert intensive access to data and improving the performance of the neural network processor.
[0054] By allocating the first number of sub-operators to multiple target cores in the neural network processor, the task of the target operator is divided into sub-tasks of multiple sub-operators, and by allocating multiple sub-tasks to multiple cores for parallel processing, the performance of the neural network processor is further improved.
[0055] As an alternative embodiment, allocating the first number of sub-operators to at least two target cores in the neural network processor includes: Determining the second number of multiple cores in the neural network processor; Based on the magnitude relationship between the first number and the second number, determining at least one target core among each core and the number of sub-operators to be processed corresponding to each target core; the number of sub-operators to be processed corresponding to the target core is not zero; For each target core, determining the number of sub-operators to be processed corresponding to the target core from the first number of sub-operators; wherein, there is no overlap among at least one sub-operator corresponding to each target core; For each target core, allocating the number of sub-operators to be processed corresponding to the target core to the target core.
[0056] Specifically, taking the number of multiple sub-operators as the first number and the number of multiple cores in the neural network processor as the second number.
[0057] After determining the first quantity and the second quantity, the magnitude relationship between the first quantity and the second quantity can be judged, the target core for the sub-operator to be executed is determined from multiple cores, and the processing quantity of the sub-operator corresponding to each target core is determined, that is, the number of sub-operators executed by each target core, where the processing quantity of the sub-operator corresponding to the target core is not zero.
[0058] After determining at least one target core, for each target core, the number of sub-operators corresponding to the target core can be selected from the first quantity of sub-operators, and at least one sub-operator corresponding to each target core does not overlap, that is, the sum of the processing quantities corresponding to each target core is the first quantity.
[0059] For each target core, after determining the number of sub-operators corresponding to the target core, the determined number of sub-operators can be assigned to the target core.
[0060] Optionally, any one of the target cores can be used as the core to be assigned corresponding to the current assignment operation, the processing quantity corresponding to the core to be assigned is used as the first processing quantity, and the first processing quantity of sub-operators is selected from the first quantity of sub-operators as the sub-operators corresponding to the core to be assigned. For example, the first processing quantity of sub-operators can be arbitrarily selected from the first quantity of sub-operators, or the first processing quantity of consecutive sub-operators can be selected from the first quantity of sub-operators, where each sub-operator corresponds to the first sub-data, and the order between the sub-operators can be the order between the first sub-data.
[0061] Determine the core to be assigned corresponding to the next assignment operation from the other target cores except the core to be assigned, and determine the sub-operators corresponding to the updated core to be assigned from the remaining sub-operators, and repeat the above assignment operation until all sub-operators are assigned to the target cores.
[0062] As an alternative embodiment, based on the magnitude relationship between the first quantity and the second quantity, determining at least one target core in each core and the processing quantity of the sub-operator corresponding to each target core respectively includes: If the first quantity is less than the second quantity, then the first quantity of cores in the second quantity of cores is used as the target core, and the processing quantity of each target core is set to 1; If the first quantity is not less than the second quantity, then each core is used as the target core, and based on the first quantity and the second quantity, the processing quantity corresponding to each target core is determined.
[0063] Specifically, if the first quantity is less than the second quantity, that is, the number of sub-operators is less than the number of cores, the first quantity of cores among the second quantity of cores can be used as target cores, and the processing quantity of each target core is set to 1. That is to say, when the first quantity is less than the second quantity, there will be idle cores. By setting the processing quantity of each target core to 1, the balanced load of multiple cores is achieved, which is beneficial to improving the performance of the neural network processor.
[0064] In one example, assume that the first quantity is M and the second quantity is N. If M = 3 and N = 4, that is, M is less than N, then 3 sub-operators can be evenly distributed to 3 cores (such as core 1, core 2, and core 3), and core 4 will not be assigned sub-operators for operation, and core 4 is not a target core.
[0065] If the first quantity is not less than the second quantity, that is, the number of sub-operators is not less than the number of cores, each core can be used as a target core, and based on the first quantity and the second quantity, the processing quantity corresponding to each target core is determined. That is to say, when the first quantity is not less than the second quantity, there will be no idle cores, and the resources of the multi-core processor are fully utilized, which is beneficial to improving the performance of the neural network processor.
[0066] Optionally, determining the processing quantity corresponding to each target core based on the first quantity and the second quantity includes: Taking the calculation result of integer division of the first quantity by the second quantity as the first value; Taking the calculation result of the remainder operation of the first quantity by the second quantity as the second value; If the second value is zero, taking the first value as the processing quantity corresponding to each target core; If the second value is not zero, determining the second quantity of first target cores from each target core, setting the processing quantity corresponding to the first target core to the third value; setting the processing quantity corresponding to each second target core among each target core to the first value; The second target core is the target core other than the first target core among each target core; the third value is 1 greater than the first value.
[0067] Specifically, when the first quantity is not less than the second quantity, the calculation result of integer division of the first quantity by the second quantity can be taken as the first value, and the calculation result of the remainder operation of the first quantity by the second quantity can be taken as the second value, where both the first value and the second value are natural numbers.
[0068] Assuming that the first quantity is M and the second quantity is N, the first value is M / / N and the second value is M%N, where “ / / ” represents integer division, and the result of M / / N is the result of rounding down after M is divided by N; “%” represents the remainder operation (also called modulo operation), and M%N calculates the remainder after M is divided by N.
[0069] If the second value is zero, it means that the multiple sub-operators can be evenly distributed to the multiple target cores, and the processing quantity of each target core is set to the first value.
[0070] In an example, if M=4, N=4, that is, M equals N, the first value is 1, and the second value is 0, then the four sub-operators can be evenly distributed to the four cores, and each core processes one sub-operator.
[0071] If the second number is not zero, a second value of first target cores can be determined from multiple target cores, the target cores before the first target core can be used as the second core, the processing quantity of the first target core can be set to the third value (that is, the first value plus 1), and the processing quantity of the second target core can be set to the first value, so as to achieve balanced load on multiple cores as much as possible.
[0072] In one example, if M=5, N=3, that is, M is greater than N, the first value is 1, and the second value is 2, then 2 cores (for example, core 1 and core 3) can be selected from the 3 cores as the first target cores, and the remaining 1 core (for example, core 2) can be selected as the second target core. The processing quantity of the first target core is the first value plus 1, that is, 2, and the processing quantity of the second target core is 1.
[0073] Figure 6 A schematic diagram of sub-operator allocation provided in an embodiment of the present disclosure is shown in FIG. Figure 6 As shown, M=5, N=3, the first value is 1, the second value is 2, core 1 and core 3 are used as the first target cores, and core 2 is used as the second target core. Therefore, sub-operator 1 and sub-operator 2 can be assigned to core 1, sub-operator 3 can be assigned to core 2, and sub-operator 4 and sub-operator 5 can be assigned to core 3.
[0074] In addition, Figure 6 Based on the determined allocation quantity of each core, sub-operator 1 and sub-operator 3 can also be allocated to core 1, sub-operator 5 can be allocated to core 2, and sub-operator 2 and sub-operator 4 can be allocated to core 3. In other words, any allocation method that ensures that the sub-operators corresponding to each core do not overlap and the sum of the processing quantities corresponding to each core is the first quantity can be used as the allocation method of the sub-operators provided in the embodiment of the present disclosure. Figure 6 The examples provided do not constitute a limitation on the manner of assigning sub-operators.
[0075] Optionally, each target core corresponds to a core number respectively; Determining the second numerical number of first target cores from each of the target cores includes: Regarding the second numerical number of target cores with the earliest sorted core numbers among each of the target cores as the first target cores.
[0076] Specifically, each target core corresponds to a core number respectively. Regarding the second numerical number of target cores with the earliest sorted core numbers among each of the target cores as the first target cores, and regarding the remaining target cores as the second target cores.
[0077] Figure 7 The figure is a schematic diagram of a sub-operator allocation provided by an embodiment of the present disclosure. As Figure 7 shown, M = 5, N = 3, the first numerical value is 1, the second numerical value is 2, and the core numbers corresponding to the 3 target cores are Core 1, Core 2, and Core 3 respectively. Then, Core 1 and Core 2 are regarded as the first target cores, and Core 3 is regarded as the second target core. Therefore, sub-operator 1 and sub-operator 2 can be allocated to Core 1, sub-operator 3 and sub-operator 4 can be allocated to Core 2, and sub-operator 5 can be allocated to Core 3.
[0078] In the embodiment of the present disclosure, by determining at least one target core and the processing quantity corresponding to each target core respectively based on the magnitude relationship between the first quantity and the second quantity, the processing quantities corresponding to each target core are kept balanced, realizing the balanced load of the multi-core processor and further improving the performance of the neural network processor.
[0079] Figure 8 The figure is a schematic diagram of a calculation process of a neural network processor provided by an embodiment of the present disclosure. As Figure 8 shown, there are N cores in the NPU, the number of sub-operators is M, Core 1 corresponds to sub-operators 1 to sub-operator S, S is less than M, and S is the processing quantity of Core 1. The output tensor channel of the target operator is aligned with 32, and Batch = 1.
[0080] The specific steps include: (1) Obtaining the output tensor channel number of the target operator (meeting the condition %32 = 0), obtaining the total number of divided parts as M = / / 32, and the shape of each part is [B, H, W, C32_i]; (2) Grouping the M divided sub-operators, and the algorithm process is as follows: a) pre_group = M / / N, pre_group represents the number of sub-operators processed by each core; b) If M < N, then directly set pre_group = 1; c) If M % N == 0, then set pre_group = M / / N; d) If M % N is not equal to 0, let R = M % N, for the first R cores, increment pre_group by 1; for the cores between R and N, the number of groups assigned remains the pre_group value; After multi-core computing is completed, perform the concat operation (at this time, since the memory addresses are continuous, it can be changed to the internal-concat operator. The internal-concat operator can complete the splicing without actual data copying, thus achieving an efficient "zero-cost" operation), axi = 1 (axi = 1 means splicing in the dimension with dimension 1), so the resulting shape becomes [1, / / 32, H, W, C32], where, / / 32 corresponds to dimension 1.
[0081] In the embodiments of the present disclosure, while utilizing multi-core acceleration, the existence of the Layout Convert operator is further reduced, and the number of memory accesses in the DSA hardware is reduced. This not only reduces the overall model time consumption but also reduces the overall power consumption, and can greatly optimize the performance of the neural network processor.
[0082] Figure 9 The following is a schematic structural diagram of a data processing device provided by an embodiment of the present disclosure. As Figure 9 shown, the device in this embodiment may include: An acquisition module 210, configured to acquire the initial data to be processed and the target operator corresponding to the initial data; A splitting module 220, configured to use the quotient obtained by dividing the number of output channels of the target operator by a preset number as the first number, and split the target operator into the first number of sub-operators; An allocation module 230, configured to allocate the first number of sub-operators to at least two target cores in the neural network processor; where multiple cores in the neural network processor at least include the at least two target cores; An operation module 240, configured to, for each sub-operator in each target core, execute the operation of the sub-operator on the initial data through the target core, and use the result of the operation as the sub-data corresponding to the sub-operator; A splicing module 250, configured to splice the sub-data corresponding to each sub-operator respectively, and use the spliced data as the target data corresponding to the initial data; Among them, the initial data is stored and laid out based on the first memory layout format; the target data is stored and laid out based on the second memory layout format; the second memory layout format is obtained by grouping in the channel dimension with the preset quantity on the basis of the first memory layout format.
[0083] As an alternative embodiment, when the allocation module allocates the first quantity of sub-operators to at least two target cores in the neural network processor, it is specifically configured to: Determine the second quantity of multiple cores in the neural network processor; Based on the magnitude relationship between the first quantity and the second quantity, determine at least one target core in each core and the processing quantity of the sub-operators respectively corresponding to each target core; the processing quantity of the sub-operators corresponding to the target core is not zero; For each target core, determine the processing quantity of sub-operators corresponding to the target core from the first quantity of sub-operators; among them, at least one sub-operator corresponding to each target core does not overlap; For each target core, allocate the processing quantity of sub-operators corresponding to the target core to the target core.
[0084] As an alternative embodiment, when the allocation module determines at least one target core in each core and the processing quantity of the sub-operators respectively corresponding to each target core based on the magnitude relationship between the first quantity and the second quantity, it is specifically configured to: If the first quantity is less than the second quantity, use the first quantity of cores among the second quantity of cores as target cores and set the processing quantity of each target core to 1; If the first quantity is not less than the second quantity, use each core as a target core and determine the processing quantity respectively corresponding to each target core based on the first quantity and the second quantity.
[0085] As an alternative embodiment, when the allocation module determines the processing quantity respectively corresponding to each target core based on the first quantity and the second quantity, it is specifically configured to: Use the calculation result obtained by performing integer division of the first quantity by the second quantity as the first value; Use the calculation result obtained by performing the remainder operation of the first quantity by the second quantity as the second value; If the second value is zero, use the first value as the processing quantity corresponding to each target core; If the second value is not zero, determine the second value of first target cores from the respective target cores, and set the processing quantity corresponding to the first target cores to a third value; set the processing quantity corresponding to each second target core among the respective target cores to the first value; The second target core is a target core other than the first target core among the respective target cores; the third value is 1 greater than the first value.
[0086] As an alternative embodiment, each of the target cores corresponds to a core serial number; When the allocation module determines the second value of first target cores from the respective target cores, it is specifically configured to: Use the second value of target cores with the earliest sorted core serial numbers among the respective target cores as the first target cores.
[0087] As an alternative embodiment, when the splicing module splices the respective sub - data corresponding to each sub - operator, it is specifically configured to: Determine the order relationship corresponding to each sub - operator; Splice the respective sub - data corresponding to each sub - operator in sequence according to the order relationship corresponding to each sub - operator.
[0088] As an alternative embodiment, there is a connection relationship between the target operator and the adjacent operator of the target operator; the target operator is used to process data stored in a first memory layout format, and the adjacent operator is used to process data stored in a second memory layout format.
[0089] The device according to the embodiments of the present disclosure can execute the method provided by the embodiments of the present disclosure, and its implementation principle is similar and has corresponding technical effects. The actions performed by each module in the device according to the embodiments of the present disclosure correspond to the steps in the method according to the embodiments of the present disclosure. For the detailed function descriptions of the respective modules of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details are not repeated here.
[0090] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0091] An electronic device is provided in an embodiment of the present disclosure, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method provided in any optional embodiment of the present disclosure. Compared with the prior art, it can be achieved that: by splitting the target operator into a first number of sub-operators, each sub-operator operates on the initial data to obtain corresponding sub-data, and the sub-data are spliced together, and the spliced data is used as the target data, so that the target data can be directly converted into the second memory layout format without inserting a layout convert operator, reducing the time consumption caused by the layout convert intensive access to data and improving the performance of the neural network processor. By allocating the first number of sub-operators to multiple target cores in the neural network processor, the task of the target operator is divided into sub-tasks of multiple sub-operators, and by allocating multiple sub-tasks to multiple cores for parallel processing, the performance of the neural network processor is further improved.
[0092] In an optional embodiment, an electronic device is provided, as Figure 10 shown Figure 10 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present disclosure.
[0093] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the present disclosure. The processor 4001 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0094] The bus 4002 may include a path for transmitting information among the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 it is only represented by a thick line in Figure 10 , but it does not mean that there is only one bus or one type of bus.
[0095] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.
[0096] The memory 4003 is used to store the computer program for implementing the embodiments of the present disclosure and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0097] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0098] The embodiments of the present disclosure provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0099] The embodiments of the present disclosure also provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the foregoing method embodiments when executed by a processor.
[0100] It should be understood that although the flowcharts in the embodiments of the present disclosure indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.
[0101] The foregoing are only optional implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present disclosure, other similar implementation means based on the technical idea of the present disclosure also belong to the protection scope of the embodiments of the present disclosure.
Claims
1. A data processing method, characterized in that: include: Acquire initial data to be processed and a target operator corresponding to the initial data; The quotient obtained by dividing the number of output channels of the target operator by the preset number is taken as the first number, and the target operator is divided into the first number of sub-operators; Allocating the first number of sub-operators to at least two target cores in a neural network processor; wherein the plurality of cores in the neural network processor at least include the at least two target cores; For each sub-operator in each target core, the target core is used to perform an operation of the sub-operator on the initial data, and a result obtained by the operation is used as the sub-data corresponding to the sub-operator; The sub-data corresponding to each sub-operator are spliced together, and the spliced data are used as the target data corresponding to the initial data; Among them, the initial data is stored and laid out based on a first memory layout format; the target data is stored and laid out based on a second memory layout format; and the second memory layout format is obtained by grouping the preset number on the channel dimension based on the first memory layout format.
2. The method according to claim 1, characterized in that: The allocating the first number of sub-operators to at least two target cores in the neural network processor comprises: determining a second number of the plurality of cores in the neural network processor; Based on the size relationship between the first number and the second number, determining at least one target core among the cores and the processing number of the sub-operators corresponding to each target core; the processing number of the sub-operators corresponding to the target core is not zero; For each target core, determine the processing number of sub-operators corresponding to the target core from the first number of sub-operators; wherein at least one sub-operator corresponding to each target core does not overlap; For each target core, the processing number sub-operators corresponding to the target core are allocated to the target core.
3. The method according to claim 2, characterized in that The determining, based on the size relationship between the first number and the second number, at least one target core in each core and the processing number of the sub-operators corresponding to each target core includes: If the first number is less than the second number, taking the first number of cores in the second number of cores as target cores, and setting the processing quantity of each target core to 1; If the first number is not less than the second number, each core is used as a target core, and based on the first number and the second number, the processing quantity corresponding to each target core is determined.
4. The method according to claim 3, characterized in that The determining, based on the first number and the second number, the processing number corresponding to each target core includes: Using a calculation result obtained by integer division of the first number by the second number as a first value; Taking the result of performing a modulo operation on the first number and the second number as a second value; If the second value is zero, the first value is used as the processing quantity corresponding to each target core; If the second value is not zero, determining a second value of first target cores from the target cores, setting the processing quantity corresponding to the first target cores to a third value; and setting the processing quantity corresponding to each second target core in the target cores to the first value; The second target core is a target core other than the first target core among the target cores; and the third value is 1 greater than the first value.
5. The method according to claim 4, characterized in that Each of the target cores corresponds to a core serial number; The determining a second numerical value of first target cores from the target cores comprises: The second target core with the highest core number among the target cores is used as the first target core.
6. The method according to any one of claims 1 to 5, characterized in that The step of splicing the sub-data corresponding to the sub-operators includes: Determine the order relationship corresponding to each sub-operator; The sub-data corresponding to each sub-operator are sequentially spliced according to the order relationship corresponding to each sub-operator.
7. The method according to any one of claims 1 to 5, characterized in that The target operator is connected to an adjacent operator of the target operator; the target operator is used to process data stored in a first memory layout format, and the adjacent operator is used to process data stored in a second memory layout format.
8. A data processing device, characterized in that: include: An acquisition module, used to acquire initial data to be processed and a target operator corresponding to the initial data; A splitting module, configured to take a quotient obtained by dividing the number of output channels of the target operator by a preset number as a first number, and split the target operator into the first number of sub-operators; An allocating module, configured to allocate the first number of sub-operators to at least two target cores in a neural network processor; wherein the plurality of cores in the neural network processor at least include the at least two target cores; An operation module, for each sub-operator in each target core, executing the operation of the sub-operator on the initial data through the target core, and using the result of the operation as the sub-data corresponding to the sub-operator; A splicing module, used for splicing the sub-data corresponding to each sub-operator, and using the spliced data as the target data corresponding to the initial data; Among them, the initial data is stored and laid out based on a first memory layout format; the target data is stored and laid out based on a second memory layout format; and the second memory layout format is obtained by grouping the preset number on the channel dimension based on the first memory layout format.
9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.