Data Processing Method, Apparatus and Storage Medium

By obtaining and using static and dynamic label information to determine the split index in a multi-core processor, the problem of low data splitting and layout efficiency in neural network operations is solved, and more efficient computing speed and efficiency are achieved.

CN114580607BActive Publication Date: 2025-07-22CAMBRICON TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011403593.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-02
Publication Date
2025-07-22
Estimated Expiration
2041-01-20

AI Technical Summary

Technical Problem

In multi-core processors, the speed of neural network computing is limited by the large number of data types, large amount of computing and hardware limitations. How to efficiently split and layout data to improve computing efficiency and speed is an urgent problem.

Method used

By obtaining the static and dynamic tag information of the first data, a split index is determined and added to the dynamic tag information, so that the multi-core processor can split and lay out the data based on these tag information, thereby optimizing neural network operations.

Benefits of technology

It improves the processing efficiency and speed of neural network computing tasks, simplifies the data splitting and layout process, and makes full use of the parallel computing capabilities of multi-core processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580607B_ABST
    Figure CN114580607B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data processing method, apparatus, and storage medium. The board card disclosed therein includes: a storage device, an interface device, a control device, and a chip provided with a data processing device; wherein the data processing device is respectively connected to the storage device, the control device, and the interface device; the storage device is used for storing data; the interface device is used for realizing data transmission between the data processing device and an external device; the control device is used for monitoring the state of the data processing device. The data processing method, apparatus, and storage medium provided by the embodiments of the present disclosure can split and layout data based on the split index in the dynamic label information when performing a neural network operation task, improving the processing efficiency and speed of the neural network operation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a data processing method, apparatus, and storage medium. Background Art

[0002] With the development of computer technology, neural networks have also made significant progress and can perform neural network operations through specific or general-purpose processors. In related technologies, under the influence of factors such as multiple data types, large amounts of operations, and hardware limitations in neural networks, the speed of neural network operations is greatly limited. Summary of the Invention

[0003] In view of this, the present disclosure provides a data processing method, apparatus, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided a data processing method applied to a multi-core processor. The method includes:

[0005] Obtaining first data and label information of the first data, where the first data is used to participate in neural network operations, and the label information includes static label information and dynamic label information;

[0006] Determining a split index required to split the first data into multiple second data according to parameters of an operator corresponding to an operation participated by the determined first data, the number of memory channels in the multi-core processor, and the label information;

[0007] Increasing the split index to the dynamic label information so that the multi-core processor can perform the neural network operation on the first data based on the label information,

[0008] where the static label information is used to characterize information related to the first data participating in the neural network operation, and the dynamic label information is used to characterize information related to the first data and the multi-core processor.

[0009] According to another aspect of the present disclosure, there is provided a data processing apparatus applied to a multi-core processor. The apparatus includes:

[0010] An information acquisition module that acquires first data and label information of the first data, where the first data is used to participate in neural network operations, and the label information includes static label information and dynamic label information;

[0011] An index determination module that determines a split index required to split the first data into multiple second data according to parameters of an operator corresponding to an operation participated by the determined first data, the number of memory channels in the multi-core processor, and the label information;

[0012] A label adding module that adds the split index to the dynamic label information so that the multi-core processor can perform the neural network operation on the first data based on the label information.

[0013] Among them, the static label information is used to represent the information related to the participation of the first data in the neural network operation, and the dynamic label information is used to represent the information related to the first data and the multi-core processor.

[0014] According to another aspect of the present disclosure, a machine learning operation device is provided, and the device includes:

[0015] One or more of the above data processing devices, which are used to obtain data to be operated and control information from other processing devices, perform specified machine learning operations, and transfer the execution results to other processing devices through the I / O interface.

[0016] When the machine learning operation device includes multiple data processing devices, the multiple data processing devices can be connected through a specific structure and transmit data.

[0017] Among them, multiple data processing devices are interconnected and transmit data through the Peripheral Component Interconnect Express (PCIE) bus to support larger-scale machine learning operations; multiple data processing devices share the same control system or have their own control systems; multiple data processing devices share memory or have their own memory; the interconnection method of multiple data processing devices is any interconnection topology.

[0018] According to another aspect of the present disclosure, a combined processing device is provided, and the combined processing device includes:

[0019] The above-mentioned machine learning operation device, a general interconnection interface, and other processing devices;

[0020] The machine learning operation device interacts with the other processing devices to jointly complete the calculation operations specified by the user.

[0021] Among them, the combined processing device further includes: a storage device, which is respectively connected to the machine learning operation device and the other processing devices and is used to store the data of the machine learning operation device and the other processing devices.

[0022] According to another aspect of the present disclosure, a chip is provided, and the chip includes the above combined processing device.

[0023] According to another aspect of the present disclosure, a board is provided, and the board includes: a storage device, an interface device, a control device, and the above chip;

[0024] Wherein, the data processing device is respectively connected to the storage device, the control device, and the interface device;

[0025] The storage device is used for storing data;

[0026] The interface device is used for realizing data transmission between the data processing device and an external device;

[0027] The control device is used for monitoring the state of the data processing device,

[0028] Wherein, the storage device includes: multiple groups of storage units, each group of storage units is connected to the data processing device through a bus, and the storage unit is: DDR SDRAM;

[0029] The data processing device includes: a DDR controller, which is used for controlling data transmission and data storage of each storage unit;

[0030] The interface device is: a standard PCIE interface.

[0031] According to another aspect of the present disclosure, a non - volatile computer - readable storage medium is provided, on which computer program instructions are stored, and characterized in that when the computer program instructions are executed by a processor, the above - mentioned data processing method is implemented.

[0032] The data processing method, device, and storage medium provided by the embodiments of the present disclosure. The method includes: obtaining first data and calculating tag information of the first data of the computing core cluster. The first data of the computing core cluster is used to participate in neural network operations. The tag information of the computing core cluster includes static tag information and dynamic tag information; determining a split index required to split the first data of the computing core cluster into multiple second data according to parameters of an operator corresponding to an operation participated by the determined first data of the computing core cluster, the number of memory channels in the multi - core processor of the computing core cluster, and the tag information of the computing core cluster; adding the split index of the computing core cluster to the dynamic tag information of the computing core cluster, so that the multi - core processor of the computing core cluster can perform neural network operations on the first data of the computing core cluster based on the tag information of the computing core cluster. Among them, the static tag information of the computing core cluster is used to characterize information related to the first data of the computing core cluster participating in neural network operations, and the dynamic tag information of the computing core cluster is used to characterize information related to the first data of the computing core cluster and the multi - core processor of the computing core cluster. This enables data to be split and laid out based on the split index in the dynamic tag information when performing neural network operation tasks, improving the processing efficiency and speed of neural network operation tasks.

[0033] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Brief Description of the Drawings

[0034] The drawings included in and constituting a part of the specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0035] Figure 1a 、 Figure 1b Shows a schematic diagram of a multi-core processor architecture according to an embodiment of the present disclosure.

[0036] Figure 1c 、 Figure 1d Shows a schematic diagram of the architecture of a computing core cluster according to an embodiment of the present disclosure.

[0037] Figure 2 Shows a flowchart of a data processing method according to an embodiment of the present disclosure.

[0038] Figure 3 Shows a schematic diagram of the architecture of a multi-core processor 1 according to an embodiment of the present disclosure.

[0039] Figure 4a Shows a schematic diagram of the calculation task allocation for a multi-core processor 1 to execute an arithmetic task A according to an embodiment of the present disclosure.

[0040] Figure 4b Shows a schematic diagram of the process of a multi-core processor 1 executing an arithmetic task A according to an embodiment of the present disclosure.

[0041] Figure 5 Shows a block diagram of a data processing device according to an embodiment of the present disclosure.

[0042] Figure 6 Is a structural diagram showing a combined processing device 1200 according to an embodiment of the present disclosure.

[0043] Figure 7 Is a schematic structural diagram showing a board 1300 according to an embodiment of the present disclosure. Detailed Description of the Embodiments

[0044] The following will describe in detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0045] The special term "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.

[0046] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0047] With the development of computer technology and big data, multi-core processors have gradually become mainstream, and multi-core processors can be configured with different architectures according to usage requirements. For example, a multi-core processor may include multiple computing core clusters, each computing core cluster may include at least one computing core, and each computing core cluster may be provided with a designated memory. Figure 1a 、 Figure 1b FIG. shows a schematic diagram of a multi-core processor architecture according to an embodiment of the present disclosure. Figure 1c 、 Figure 1d FIG. shows a schematic diagram of the architecture of a computing core cluster according to an embodiment of the present disclosure. The multi-core processor may be a Symmetrical Multi-Processing (SMP) architecture, as Figure 1a shown. Assume that the symmetric multi-processor includes four computing core clusters (cluster) and four memories. The computing cores (core) in the computing core cluster can access the memories through the routing node R. Among them, the 4 memories are respectively the "local memories" of the corresponding 4 computing core clusters. However, due to its symmetric setting, each computing core in each computing core cluster needs to pass through two routing nodes R when accessing the 4 memories. Without congestion, the access rate to any one memory is the same. For example, Figure 1a the "local memory" of computing core cluster x is memory a. During the process of data storage and access, the data corresponding to computing core cluster x is preferentially stored in the "local memory". A corresponding memory channel is provided between the computing core cluster and its corresponding "local memory" so that the computing core cluster can access its own "local memory" through this memory channel. In this way, the probability that the computing cores in different computing core clusters access the same memory can be reduced.

[0048] The multi-core processor may also be a Non Uniform Memory Access Architecture (NUMA) architecture, as Figure 1b shown. Assume that the multi-core processor with non-uniform symmetric memory access includes four computing core clusters and four memories. The computing cores (core) in the computing core cluster can access the memories through the routing node R. Among them, the 4 memories are respectively the "local memories" of the corresponding 4 computing core clusters. However, due to its asymmetric setting, without congestion, each computing core in each computing core cluster accesses its corresponding "local memory" (such as Figure 1bWhen calculating the access of the computing core in the computing core cluster n to the memory s), only one routing node R needs to be passed. For example, when the computing core in the computing core cluster n accesses the memory s), only one routing node R needs to be passed, and the access rate is the fastest; when the computing core in each computing core cluster accesses other non-"local memories", at least two routing nodes R need to be passed, and the access rate is slow.

[0049] Optionally, when multiple computing cores are set in the computing core cluster, the organizational structure in the computing core cluster can also be set as needed. The architecture of the computing core cluster can be a bus competition type, such as Figure 1c As shown, assuming that the computing core cluster includes 4 computing cores, each computing core can access the memory through the routing node R. The architecture of the computing core cluster can also be a shared cache type, such as Figure 1d As shown, assuming that the computing core cluster includes four computing cores and an off-chip cache corresponding to the computing core cluster is set, the four computing cores in the computing core cluster share the off-chip cache. Each computing core in the computing core cluster can directly access the off-chip cache and can access the memory through the routing node R.

[0050] In order to accelerate big data volume tasks such as neural network operation tasks, when using a multi-core processor for operation, data parallelism, model parallelism, and hybrid parallelism are generally used for specific calculation processing. Data parallelism means that different computing cores process different input data of the same neural network, which is equivalent to parallelly calculating multiple batches of input data of the same neural network. Model parallelism means that different computing cores process different parts of the same input data of the same neural network, that is, disassembling each operator and data in the neural network, and each computing core processes a part of the disassembled data, which is equivalent to the entire accelerator parallelly processing a piece of input data. Hybrid parallelism means a combination of data parallelism and model parallelism. For example, model parallelism is used within the computing core cluster, and data parallelism is used between computing core clusters. However, when the multi-core processor executes operation tasks using the above three methods, how to disassemble operators and data, plan the storage location of data to ensure the efficient execution of operation tasks is a technical problem to be solved urgently.

[0051] The present disclosure provides a data processing method. Before actually executing a neural network operation task, by determining the splitting index, target storage space, and target data exchange level of the first data participating in the neural network operation and adding them to the dynamic label information of the first data, the multi-core processor can split, layout, and determine the storage location of the data based on the new dynamic label information when executing the operation task, improving the processing efficiency and speed of the operation task. The data processing method of the present disclosure can be specifically referred to in the following description.

[0052] Figure 2 The flowchart showing the data processing method according to an embodiment of the present disclosure is as follows. As Figure 2As shown, the method is applied to a multi-core processor, and the method includes steps S11 to S13.

[0053] In step S11, first data and label information of the first data are obtained. The first data is used to participate in neural network operations, and the label information includes static label information and dynamic label information. Among them, the static label information is used to characterize information related to the first data participating in the neural network operations, and the dynamic label information is used to characterize information related to the first data and the multi-core processor.

[0054] In this embodiment, the first data may be data of various data categories participating in neural network operations, such as: input neurons, output neurons, hidden neurons, constant neurons, input weights, output weights, constant weights, and auxiliary data. The neural network operations participated by the first data may include operations that appear in neural networks such as convolution operations, fully connected operations, pooling operations, and scaling operations. The present disclosure does not limit this. In actual operation tasks, usually one or more data to be operated are processed to obtain an operation result. The first data may be one or more of the one or more data to be operated and the operation result. The present disclosure does not limit this.

[0055] In this embodiment, the static label information may include information such as the data type, dimension, and dimension value that describe the nature of the first data itself, and also includes information related to the neural network operations participated by the first data. Therefore, the static label information of the same first data may be different in different neural networks. The static label information may be determined after the neural network is established. The static label information of the first data may be applicable to any processor running the neural network, that is, the static label information of the first data is unchanged in different processors.

[0056] Optionally, the static label information of the first data may be automatically detected and determined by the processor during the process of obtaining the first data (the process of the user inputting the first data), or may be determined according to the information input by the user. The present disclosure does not limit this.

[0057] In a possible implementation manner, the static label information may include at least one of the following: data category, static data type, static data dimension, static data dimension order, and dimension value corresponding to each static data dimension.

[0058] In this implementation, the first data represented by the data category indicates what kind of data it belongs to in the neural network, which is determined based on information such as whether it is visible to the user and the operations it participates in within the neural network. The static data type represents the type and number of bits of the first data. For example, the static data type can be a 32-bit floating-point number or an 8-bit fixed-point number, etc. The static data dimension can be one-dimensional, two-dimensional, multi-dimensional, etc., and the static data dimension order can represent the dimension order for storing and / or reading the first data.

[0059] Optionally, the static data dimension can include, but is not limited to, at least one of the following: the channel dimension C, the height dimension H, the width dimension W, the number dimension N, the depth dimension D (for three-dimensional convolution operations), and the time dimension T (for RNN network operations such as LSTM). Recording or representing the static data dimension order can be achieved using the arrangement order of the first 6 letters or other identifiers corresponding to each static data dimension. In the arrangement, the static data dimension with an earlier order or on the left precedes the static data dimension with a later order or on the right. Or, it can also be defaulted that the static data dimension with an earlier order or on the left is higher than the static data dimension with a later order or on the right. For example, if the static data dimension order of the first data is NCHW, it means the first data is a four-dimensional tensor, and its corresponding static data dimension order is the number dimension N, the channel dimension C, the height dimension H, and the width dimension W in sequence, with the highest dimension being the number dimension N and the lowest dimension being the width dimension W. If the static data dimension order of the first data is TNHWC, it represents that the first data is a five-dimensional tensor, and its corresponding static data dimension order is the time dimension T, the number dimension N, the height dimension H, the width dimension W, and the channel dimension C in sequence, with the highest dimension being the time dimension T and the lowest dimension being the channel dimension C. When the actual user input or the multi-core processor retrieves data from the physical memory, the first data will be mapped into a one-dimensional array and stored on the physical memory according to the determined static data dimension order. It may not be easy to determine which dimension a one-dimensional tensor (vector) belongs to. This situation can be determined according to the algorithm definition of the operator. For example, the bias data bias in two-dimensional convolution is one-dimensional. In the algorithm definition, bias needs to be superimposed on the output feature map, and one number is superimposed on each output channel, and the numbers superimposed on different channels are different. Therefore, the dimension of the one-dimensional bias can be regarded as the channel dimension C.

[0060] In this implementation, the dimension value corresponding to each static data dimension represents the length or size of the corresponding static data dimension. For example, if a first data is a matrix, the static data dimension includes rows and columns, and the static data dimension order is row-major. The dimension value of the row is 10 and the dimension value of the column is 4, which means the length of the row is 10 and the length of the column is 4.

[0061] In a possible implementation, the device can determine the data category of the first data according to the out-degree, in-degree of the neural network corresponding to the first data, and the operations in the neural network in which the first data participates.

[0062] In this implementation, the in-degree represents the number of preceding operation nodes that the first data participates in as a data node (the first data is the output of the preceding operation node), and the out-degree represents the number of subsequent operation nodes that the first data participates in as a data node (the first data is the input of the subsequent operation node). For example, if a first data cc is the output of 1 preceding operation node and the input of 3 subsequent operation nodes, then the out-degree of the first data cc is 3 and the in-degree is 1. Different codes can be set for different data categories for differentiation. As shown in Table 1 below, it describes the characteristics and corresponding identifiers of data of different data categories.

[0063] Table 1 Data Categories, Corresponding Identifiers, and Data Characteristics

[0064]

[0065] Among them, the out-degree and in-degree of the instruction are zero, which is used to trigger neural network operations. The out-degree of input neurons, constant neurons, input weights, constant weights, and auxiliary data is greater than 1, and the in-degree is 0. The out-degree of output neurons and output weights is 0, and the in-degree is greater than or equal to 1. The out-degree and in-degree of hidden neurons are both greater than or equal to 1.

[0066] In a possible implementation, the static label information of the first data can be expressed as:

[0067] Static:classification,type1,DIM_A1…An,{x1…xn}

[0068] Among them, static is an identifier indicating that the label information is static label information. classification represents the data category, and type1 represents the static data type. n in DIM_A1…An represents the static data dimension, and A1…An represents the static data dimension order as A1…An. The dimension value of A1 is x1…The dimension value of An is xn. The ",", "{}" are only used to separate different parameters in the static label information in this disclosure, and are not essential contents of the static label information. In actual applications, ",", "{}" may not exist, or may be replaced by other identifiers, and this disclosure does not limit this.

[0069] It should be understood that those skilled in the art can set the static label information, the identifier of the data category, and the positions of the various parameters in the static label information according to actual needs, and this disclosure does not limit this.

[0070] In this embodiment, the dynamic label information is determined based on the static label information, the algorithm characteristics of the neural network, the computing power, performance, etc. of the processor after determining the processor for running the neural network, so that the first data with the dynamic label information can be applicable to the operations of the processor. When the neural network is operated using different processors, the dynamic label information of the first data may be different. When the performance, computing power and other parameters of two processors are the same, the dynamic label information of the first data may be the same.

[0071] In a possible implementation manner, the dynamic label information may include at least one of the following: dynamic data type, dynamic data dimension order, sharding parameter, padding parameter, and data size.

[0072] In this implementation manner, the dynamic data type may be determined according to the type of data that the processor running the neural network can process, the computing power, etc. For example, if a certain processor can process 16-bit floating-point numbers, then when using this processor to run the neural network, the dynamic data type of the data to be processed is 16-bit floating-point numbers. The dynamic data dimension order may be determined according to the requirements of the processor running the neural network for reading or storing data. The sharding parameter may be determined according to the computing power of the processor running the neural network. For example, if a certain processor can perform operations on 8 numbers at a time, the sharding parameter can be set to 8. The padding parameter may be determined according to the dimension value of the static data dimension of the data to be processed and the sharding parameter. The padding parameter may include the length of the data to be padded in different dimensions of the data and / or the padding value required for padding. The data size, also known as the size or amount of data, is the product of the dimension values of each operation dimension and the data bit width of the data in actual operation, determined according to the dimension value of the static data dimension, the sharding parameter, and the padding parameter. For example, if a certain data is a matrix, and the operation dimension values of its two dimensions in the operation process are 4 and 8 respectively, and the data bit width is 4, then the data size of this data is 4×8×4 = 128 bytes.

[0073] In a possible implementation manner, the steps for determining the dynamic label information include: obtaining information about the target processor (i.e., the above-mentioned multi-core processor) running the neural network. The information about the target processor may include information related to the computing power and performance of the target processor, such as the type of data that the target processor can process, the dimension order of reading and storing data by the target processor, the number of data bits (or the number of numbers) processed by the target processor each time, etc. Furthermore, the data processing device can determine the dynamic label information of the first data based on the information of the target processor and the static label information of the first data through at least one of the following operations:

[0074] Determine the dynamic data type according to the type of data that the target processor can process;

[0075] Determine the dynamic data dimension order according to the dimension order of reading and storing data by the target processor;

[0076] Determine the sharding parameter according to the number of bits of data processed by the target processor each time;

[0077] Determine the padding parameter according to the sharding parameter and the dimension value of the static data dimension;

[0078] Determine the data size according to the dimension value of the static data dimension, the sharding parameter, and the padding parameter.

[0079] In a possible implementation, the dynamic label information of the first data can be expressed as:

[0080] dynamic:type2,DIM_B1…Bn,tiling,padding,size

[0081] Wherein, dynamic represents the identifier that the label information is dynamic label information. type2 represents the dynamic data type. DIM_B1…Bn represents that the dynamic data dimension order is B1…Bn. Tiling is the sharding parameter. Padding is the padding parameter, and size is the data size, which represents the size of the storage space occupied by the data after dimension conversion, tiling, or padding. "," is only used to separate different parameters in the dynamic label information in this disclosure, and is not an essential content of the dynamic label information. In actual applications, "," may not exist, or may be replaced by other identifiers, and this disclosure does not limit this.

[0082] In step S12, according to the parameters of the operator corresponding to the operation participated by the determined first data, the number of memory channels in the multi-core processor, and the label information, determine the splitting index required to split the first data into multiple second data.

[0083] In a possible implementation, step S12 may include:

[0084] According to the parameters of the operator corresponding to the operation participated by the determined first data and the label information, determine the splitting strategy corresponding to the first data from the splitting strategy database;

[0085] According to the splitting strategy, the number of memory channels in the multi-core processor, and the label information, determine the splitting index required to split the first data into multiple second data.

[0086] Among them, the splitting strategy includes at least one splittable dimension of the first data and the priority level corresponding to each splittable dimension. The splitting index includes: a target splitting dimension, a starting splitting position and an ending splitting position of each second data on the target splitting dimension of the first data, and the target splitting dimension is the dimension selected from the at least one splittable dimension.

[0087] In this implementation manner, the parameters of the operator are used to represent or describe the types of operations that the operator needs to perform and the parameters required to implement the corresponding operations, such as the parameters of operators for convolution operations, pooling operations, etc.

[0088] In this implementation manner, the splitting strategy can indicate the optional splitting methods corresponding to different multi-core processors, different types of neural network operations, and different data categories in the operation. Accordingly, the data can be further split into multiple data in different ways. The splitting index represents the implementation scheme of splitting the first data into multiple second data corresponding to a splitting method determined from at least one selectable splitting method given in the splitting strategy when the first data participates in the current type of operation as a certain specified data category in the current multi-core processor. The splitting strategy can be preset and determined, and the preset splitting strategy can be stored in the memory in the form of a splitting strategy database. Of course, in other embodiments, the splitting strategy can also be actually determined by the data processing device according to the operations participated by the first data and the characteristics of the arithmetic units of the multi-core processor. The splitting strategy database records the preset splitting strategies, and corresponding splitting strategies are set for data of different data categories in different operators.

[0089] For example, when the first data (such as input neurons, input weights, etc.) participates in the operation as "input", the splitting index can be used to split the first data into multiple second data, thereby implementing the operation. When the first data (such as output neurons, etc.) participates in the operation as "output", the splitting index can be used to indicate or predict that the final output neurons (i.e., the final operation result) obtained due to the operation after the splitting of the "input" will be obtained by performing operations on several intermediate results. What it indicates is actually the splitting relationship between the intermediate results and the output neurons; or due to the size limitation of the output neuron data, it may not be possible to split it, that is, it does not have a corresponding splitting index. In actual operation tasks, usually one or more data to be operated are processed to obtain an operation result. The determined splitting index can be for any one of the one or more data to be operated and the operation result. The present disclosure places no restrictions on this.

[0090] In step S13, the split index is added to the dynamic label information so that the multi-core processor can perform the neural network operation on the first data based on the label information.

[0091] In this embodiment, after adding the split index to the dynamic label information, when the multi-core processor executes an operation task using the first data (the first data is the data to be operated), it can split the first data into multiple second data according to the split index in the dynamic label and then execute the operation task; it can also determine how the relevant calculation process is carried out according to the split index during the execution of the operation task (such as when the first data is the operation result). In this way, with the assistance of the label information, the multi-core processor can simplify the process of splitting and arranging the data involved in the operation task and improve the operation speed of the operation task.

[0092] In a possible implementation manner, according to the parameters of the operator corresponding to the operation participated by the determined first data and the label information, determining the split strategy corresponding to the first data from the split strategy database may include:

[0093] Obtain the computational graph corresponding to the neural network, where the computational graph includes operation nodes, data nodes, and the connection relationships between the data nodes and the operation nodes; among them, the data nodes may include the label information of the first data, and the label information of the first data includes static label information and dynamic label information;

[0094] Obtain the information of the target operation node connected to the data node where the first data is located in the computational graph, and the information of the target operation node includes the parameters of the operator corresponding to the operation implemented by the target operation node;

[0095] Determine the split strategy corresponding to the first data from the split strategy database according to the parameters of the operator and the static label information of the first data.

[0096] In this implementation manner, the static label information and part of the dynamic label information (except the split index, the target storage space, and the target data exchange level) can be first determined according to the computational graph of the neural network, and the determined label information is bound and stored with the corresponding data node.

[0097] In this implementation, the static tag information may further include a data category, and the data category includes any one of the following: instruction, input neuron, output neuron, hidden neuron, constant neuron, input weight, output weight, constant weight, and auxiliary data. Further, according to the parameters of the operator and the data category of the first data, the splitting strategy corresponding to the first data can be determined from the splitting strategy database. For example, referring to Table 2 below, assuming that the operator corresponding to the operation participated by the first data determined based on the above steps is a two-dimensional convolution Conv and its data category is an input weight, then its corresponding splitting strategy is "the splittable dimension is N, the candidate storage space is Cluster, and the candidate data exchange level is No". Assuming that the operator corresponding to the operation participated by the first data determined based on the above steps is a two-dimensional convolution Conv and its data category is an input neuron, then its corresponding splitting strategy is "the splittable dimensions are N, H, W (in decreasing priority order). When the splittable dimension is N, the candidate storage space is Mem and the candidate data exchange level is No; when the splittable dimension is H or W, the candidate storage space is Cluster and the candidate data exchange level is Cluster".

[0098] In a possible implementation, according to the splitting strategy corresponding to the first data, the number of memory channels in the multi-core processor, and the tag information, determining the splitting index for splitting the first data into multiple second data may include:

[0099] Determining a target splitting dimension from the at least one splittable dimension according to the number of memory channels and the operation dimension value of each dimension of the first data in the process of participating in the neural network operation determined according to the tag information, and determining a splitting length corresponding to the target splitting dimension, where the splitting length is the length of the second data in the direction of the target splitting dimension;

[0100] Determining the splitting index for splitting the first data into multiple second data according to the number of memory channels, the operation dimension value of the first data on the target splitting dimension, and the splitting length.

[0101] In this implementation, the operation dimension value refers to the dimension value of each dimension of the first data in the actual operation process.

[0102] In this embodiment, since the purpose of splitting the first data is to utilize the multi-core processor to achieve parallel processing as much as possible. Therefore, before determining the splitting length, the target number of the second data can be determined first, and then the splitting length can be determined according to the dimension value of the target splitting dimension and the target number. First, the size relationship between the operation dimension value of the target splitting dimension and the total number of computing cores of the multi-core processor (i.e., the number of computing cores included in the multi-core processor) and the number of memory channels can be judged, and then the corresponding method can be further adopted according to the size relationship to determine the splitting length, that is, the following Method 1, Method 2, and Method 3.

[0103] Method 1, when the operation dimension value of the target splitting dimension is greater than or equal to the total number of computing cores, the target number of the second data is determined to be the same as the total number of computing cores, and the splitting length is the ratio obtained by dividing the operation dimension value of the first data in the target splitting dimension by the total number of computing cores. When the ratio is not an integer, a preset rounding function (as described below) can be used for rounding processing to obtain the splitting length. Among them, the rounding function can include a floor function (such as the floor function, that is, the Floor function, taking the largest integer not greater than the ratio), a ceiling function (such as the Ceil function, taking the smallest integer greater than or equal to the ratio), a round-off function (such as the roundingfunction, that is, the Round function, obtaining an integer by rounding the ratio according to the specified number of decimal places), etc. In this way, it can be ensured that each computing core participates in the actual operation process, reducing the idle computing power of the multi-core processor and improving the operation efficiency and speed.

[0104] Method 2, when the operation dimension value of the target splitting dimension is less than the total number of computing cores and greater than or equal to the number of memory channels, the number of memory channels is determined to be the target number of the second data, and the splitting length is the ratio obtained by dividing the operation dimension value of the first data in the target splitting dimension by the number of memory channels. When the ratio is not an integer, a preset rounding function can be used for rounding processing to obtain the splitting length.

[0105] Method 3, when the operation dimension value of the target splitting dimension is less than the number of memory channels, the operation dimension value of the target splitting dimension is determined to be the target number of the second data, and the splitting length is the unit length of the operation dimension value of the target splitting dimension.

[0106] By determining the splitting length through the above Method 1, Method 2, and Method 3 for determining the splitting length, it can be ensured that the computing cores in the multi-core processor all participate in the operation task, ensuring the efficient and high-speed execution of the operation task.

[0107] In a possible implementation, the static tag information may include: static data dimensions, and dimension values corresponding to each static data dimension. The dynamic tag information may include a dynamic data type, padding parameters, and data size. Among them, according to the splitting strategy, the number of memory channels in the multi-core processor, and the tag information, the splitting index required to split the first data into multiple second data is determined. It may further include:

[0108] According to the dimension values of each static data dimension of the first data and the padding parameters, determine the operation dimension value of each dimension of the first data during the neural network operation process. The sum of the dimension value of each static data dimension and the padding data length corresponding to this static data dimension in the padding parameters can be determined as the operation dimension value.

[0109] In this implementation, the operation dimension value can also be calculated in other ways based on the existing static tag information and dynamic tag information of the first data. The present disclosure does not limit this.

[0110] In a possible implementation, the operation dimension value can also be pre-calculated and stored in the dynamic tag information. So as to directly obtain the operation dimension value from the dynamic tag information only for determining the splitting length and splitting index, simplify the determination process of the splitting index, and improve efficiency.

[0111] In a possible implementation, according to the number of memory channels and the operation dimension value of each dimension of the first data during the neural network operation process, determine the target splitting dimension from the at least one splittable dimension, and determine the splitting length corresponding to the target splitting dimension. It may include:

[0112] When there are multiple splittable dimensions, sequentially determine whether the splittable dimensions meet the splitting conditions in the order from highest to lowest priority. When it is determined that the current splittable dimension meets the splitting conditions, determine the current splittable dimension as the target splitting dimension.

[0113] Among them, the splitting condition includes: the operation dimension value of the first data on the current splittable dimension is greater than or equal to the number of memory channels.

[0114] For example, assume that the first data has three dimensions "w1, w2, w3", the operation dimension values of each dimension are w1 = 2, w2 = 32, w3 = 64, the splittable dimensions in its corresponding splitting strategy are w1 and w3, and the priority of w1 is higher than that of w3. If the number of channels of the multi-core processor is 4, first determine whether the "splittable dimension with the highest priority, w1" of the first data meets the splitting condition. Since the operation dimension value 2 of w1 is less than the "number of channels 4" and does not meet the splitting condition, then continue to determine whether the next "splittable dimension w3" meets the splitting condition. Since the operation dimension value 32 of w3 meets the splitting condition, w3 of the first data is used as the target splitting dimension.

[0115] In a possible implementation, when all splittable dimensions do not meet the splitting condition, that is, the operation dimension value on each splittable dimension is less than the number of memory channels, the one with the largest operation dimension value on the splittable dimension can be determined as the target splitting dimension. In this way, it can be ensured that more computing cores in the multi-core processor are used to execute the operation task for the first data.

[0116] For example:

[0117] Example 1, assume that the multi-core processor includes 2 memory channels and 2 computing core clusters, and each computing core cluster includes 2 computing cores. If a certain first data includes three dimensions a, b, c, and the dimension values of each dimension are a = 1, b = 4, c = 2. If its splittable dimensions are a, b, and the priority of a is higher than that of b. Then, first determine the splittable dimension a. Since a = 1 is less than the number of memory channels 2 (that is, it does not meet the splitting condition), a cannot be used as the target splitting dimension. Continue to determine the splittable dimension b. Since b = 4 is greater than the number of memory channels, the splittable dimension b is determined as the target splitting dimension. Further, since b = 4 is equal to the total number of computing cores 4, the target number of the second data can be determined to be the same as the total number of computing cores 4, and the splitting length = the dimension value 4 of the target splitting dimension b ÷ the total number of computing cores 4 = 1.

[0118] Example 2, assume that the multi-core processor includes 2 memory channels and 2 computing core clusters, and each computing core cluster includes 2 computing cores. If a certain first data includes three dimensions a, b, c, and the dimension values of each dimension are a = 1, b = 4, c = 2. If its splittable dimensions are a, c, and the priority of a is higher than that of c. Then, first determine the splittable dimension a. Since a = 1 is less than the number of memory channels 2 (that is, it does not meet the splitting condition), a cannot be used as the target splitting dimension. Continue to determine the splittable dimension c. Since c = 2 is less than the total number of computing cores and equal to the number of memory channels, the splittable dimension c can be determined as the target splitting dimension. Further, since c = 2 is less than the total number of computing cores 4 and equal to the number of memory channels, the target number of the second data can be determined to be the same as the number of memory channels 2, and the splitting length = the dimension value 2 of the target splitting dimension c ÷ the number of memory channels 2 = 1.

[0119] Example 2. Assume that the multi-core processor includes 4 memory channels and 2 clusters of computing cores, and each cluster of computing cores includes 2 computing cores. If a first piece of data includes three dimensions a, b, and c, and the dimension values of each dimension are a = 1, b = 2, and c = 2. If its splittable dimensions are a and c, and the priority level of a is higher than that of c. Then, first judge the splittable dimension a. Since a = 1 is less than the number of memory channels 4 (i.e., the splitting condition is not met), a cannot be used as the target splitting dimension. Continue to judge the splittable dimension c. Since c = 2 is still less than the number of memory channels 4 (i.e., the splitting condition is not met). At this time, since c = 2 > a = 1, the splittable dimension c is determined as the target splitting dimension. Further, since c = 2 is less than the number of memory channels 4, the target number of the second data can be determined as the dimension value 2 of the target splitting dimension C, and the splitting length = the dimension value 2 of the target splitting dimension c.

[0120] It should be noted that in the actual operation process, the operation tasks are mostly arithmetic operations between at least two pieces of data to be operated to obtain the operation result. Then, the data to be operated and the operation result can be used as the first data above to determine the splitting index.

[0121] In this embodiment, when the first data participates in the operation task, due to different operators, different target splitting dimensions, and different multi-core processor architectures, in order to facilitate the use of the first data by different computing cores, it is also necessary to set the target storage space of the second data obtained after splitting the first data to ensure that the computing cores that actually use the second data to execute the corresponding operation tasks can read and write conveniently and quickly.

[0122] In a possible implementation manner, step S12 may include: determining at least one splittable dimension corresponding to the first data and the priority level corresponding to each splittable dimension according to the parameters of the operator corresponding to the operation participated by the determined first data, the arithmetic unit characteristics of the multi-core processor, and the tag information (including dynamic data type and data category).

[0123] In this implementation manner, the arithmetic unit characteristics of the multi-core processor may include: the dimensions that data of different data categories can be split when performing operations of different operators, whether splitting different dimensions matches the hardware settings of the multi-core processor itself or the processing capabilities set by the user for the multi-core processor, and the impact of the settings of different splitting dimensions on the operation efficiency and other characteristics related to the hardware settings of the processor itself and the user-defined settings. In this way, while ensuring that the splitting strategy can be determined, the requirements for operation efficiency and speed can be met.

[0124] In this implementation manner, a policy model for determining a splitting strategy can be trained in advance according to the parameters of existing operators, the characteristics of the arithmetic units of the multi-core processor, and the situation of label information; alternatively, a correspondence relationship can be established between the parameters of the operator, the characteristics of the arithmetic units of the multi-core processor, the label information, and the splitting strategy, so as to determine the splitting strategy of the first data in real time.

[0125] In a possible implementation manner, the multi-core processor is provided with a plurality of computing core clusters, and each computing core cluster includes a plurality of computing cores. The method may further include:

[0126] According to the parameters of the operator corresponding to the operation participated by the determined first data and the storage space set in the multi-core processor, determine the candidate storage space corresponding to each splittable dimension;

[0127] According to the target splitting dimension, determine the target storage space corresponding to the first data from the candidate storage spaces, and add the identification information of the target storage space to the dynamic label information.

[0128] Wherein, the storage space respectively includes a plurality of memories (as Figure 1c shown, there is no off-chip cache for the core clusters in the multi-core processor); or the storage space includes a plurality of memories and a plurality of off-chip caches (as Figure 1d shown, the multi-core processor is provided with an off-chip cache for the core clusters). Each computing core can access any one of the plurality of memories, and the plurality of computing cores in each computing core cluster share a corresponding off-chip cache.

[0129] In this implementation manner, a model for determining the candidate storage space can be trained in advance according to the parameters of existing operators and the setting situation of the storage space set in the multi-core processor; alternatively, a correspondence relationship can be established between the parameters of the operator, the storage space set in the multi-core processor, and the candidate storage space, so as to determine the candidate storage space of the first data in real time.

[0130] In this implementation manner, the multi-core processor is provided with a plurality of memories. Each computing core cluster can access through a channel between the corresponding "local memory", or can also access other memories through the channel and the routing node. Depending on the architecture, the off-chip cache in the multi-core processor can be set or not. The candidate storage space corresponding to each splittable dimension can be determined in real time according to the parameters of the operator and the storage space set in the multi-core processor. Different operators will affect the selection of the candidate storage space, and whether the storage space set in the multi-core processor contains an off-chip cache provides an alternative target for the operator.

[0131] In this implementation, the candidate storage spaces can be one or more. During the process of determining the candidate storage spaces, the splittable dimensions corresponding to each candidate storage space can also be determined. Each candidate storage space corresponds to one or more splittable dimensions. Different splittable dimensions can correspond to the same candidate storage space, and a splittable dimension has only one corresponding candidate storage space (that is, a splittable dimension cannot correspond to two different candidate storage spaces). Furthermore, one of the candidate storage spaces whose corresponding splittable dimension includes the target split dimension is determined as the target storage space.

[0132] In this implementation, the identification information of the target storage space can be the number, name, etc. corresponding to the memory (or off-core cache) in which each second data is stored, which can distinguish the identity of the memory from other memories, or it can also be the category of the target storage space, that is, the category identifier of the memory or the category identifier of the off-core cache. The identification information of the target storage space stored in the dynamic tag information can be set according to the split index. For example, assuming that each computing core in a multi-core processor is used for the operation of the first data, that is, each computing core needs to obtain its corresponding second data, then in actual operation, the second data needs to be stored in each memory (assuming the target storage space is a memory). Then, the category identifier of the memory can be used as the identification information of the target storage space and stored in the dynamic tag information, so that the multi-core processor can determine that multiple second data need to be stored in each memory in sequence according to the category identifier of the memory (that is, the identification information of the target storage space); or, if necessary, it can also be set which memory each second data needs to be stored in, then the identity identifier of each memory can be used as the identification information of the target storage space and stored in the dynamic tag information. Those skilled in the art can set the identification information of the target storage space stored in the dynamic tag information according to actual needs, and the present disclosure does not limit this.

[0133] For example, Figure 3 shows a schematic diagram of the architecture of the multi-core processor 1 according to an embodiment of the present disclosure, as Figure 3As shown in the figure, the multi-core processor 1 includes memory 1 and memory 2, computing core cluster 1 (including computing core 1 and computing core 2) and its corresponding off-core cache 1, computing core cluster 2 (including computing core 3 and computing core 4) and its corresponding off-core cache 2. The computing cores access the memory through the routing node R. Among them, the category identifier of the memory is mem, and the category identifier of the off-core cache is cluster. The identification information of memory 1 and memory 2 is mem1 and mem2 respectively, and the identification information of off-core cache 1 and off-core cache 2 is cluster1 and cluster2 respectively. Assume that the first data a is an input neuron, which will be split into two second data a1 and a2 and needs to be stored in the memory. Then, the identifier information "mem" representing the storage space category of the memory (a1 and a2 can be stored in mem1 and mem2 respectively) can be stored in the dynamic label information, or the identification information "mem1" and "mem2" of memory 1 and memory 2 themselves can be stored.

[0134] In a possible implementation manner, the splitting strategy further includes candidate storage spaces, and each splittable dimension is provided with a corresponding candidate storage space. The method further includes: determining a target storage space from the candidate storage spaces according to the target splitting dimension, and adding the identification information of the target storage space to the dynamic label information.

[0135] Among them, the storage space includes multiple memories; or the storage space includes multiple memories and multiple off-core caches. Each computing core can access any one of the multiple memories, and the multiple computing cores in each computing core cluster share a corresponding off-core cache.

[0136] In this implementation manner, the candidate storage space situation of the data can be determined in advance and stored in the splitting strategy database to simplify the process of determining the target storage space and improve the speed and efficiency of data processing.

[0137] In a possible implementation manner, the method further includes:

[0138] Determining the candidate data exchange levels corresponding to each splittable dimension according to the parameters of the operator corresponding to the operation participated by the determined first data, the arithmetic unit characteristics of the multi-core processor, and the label information;

[0139] Determining the target data exchange level corresponding to the first data from the candidate data exchange levels according to the target splitting dimension, and adding the target data exchange level to the dynamic label information.

[0140] In this implementation manner, the target data exchange level can be used to indicate whether the second data needs to be exchanged. This exchange means that in addition to being used by the computing core that calculates the second data (as the data to be operated on or as the result of the operation), whether the second data stored in the off-core cache or memory will also be used by other computing cores in the same computing core cluster or computing cores in other computing core clusters (as the data to be operated on or as the result of the operation). The candidate data exchange levels can include: no exchange, inter-cluster data exchange, and inter-core data exchange. Among them, "no exchange" can mean that the data will not be used by other computing cores, "inter-cluster data exchange" can mean that the data will be used by computing cores in other computing core clusters, and "inter-core data exchange" can mean that the data will be used by computing cores in the same computing core cluster.

[0141] In this implementation manner, a model for determining the candidate data exchange level can be trained based on the parameters of the operator, the characteristics of the arithmetic units of the multi-core processor, and the label information situation; or a correspondence relationship between the parameters of the operator, the characteristics of the arithmetic units of the multi-core processor, the label information, and the candidate data exchange level can also be established, so as to determine the candidate data exchange level of the first data in real time.

[0142] In this implementation manner, the candidate data exchange level can be one or more. During the process of determining the candidate data exchange level, the splittable dimension corresponding to each candidate data exchange level can also be determined. Each candidate data exchange level corresponds to one or more splittable dimensions. Different splittable dimensions can correspond to the same candidate data exchange level, and one splittable dimension has only one corresponding candidate data exchange level (that is, one splittable dimension cannot correspond to two different candidate data exchange levels). Furthermore, among the candidate data exchange levels, the one including the target splittable dimension corresponding to it is determined as the target data exchange level.

[0143] In a possible implementation manner, the splitting strategy further includes the candidate data exchange level corresponding to each splittable dimension, and the method further includes:

[0144] Determine the candidate data exchange level corresponding to the target splittable dimension in at least one candidate data exchange level as the target data exchange level corresponding to the first data, and add the target data exchange level to the dynamic label information.

[0145] In this implementation manner, the candidate data exchange level of the data can be determined in advance and stored in the splitting strategy database to simplify the process of determining the target data exchange level and improve the speed and efficiency of data processing.

[0146] In this embodiment, it can be set in advance according to the splitting strategy for the first data described above. To further illustrate the setting of the splitting strategy, taking the architecture of the multi-core processor shown in Figure 3 and the first data being image data as an example, Table 2 below gives an example of the splitting strategy for the first data in the embodiments of the present disclosure.

[0147] Table 2 Example of Splitting Strategy

[0148]

[0149] Among them, the splittable dimension N represents the quantity dimension, the splittable dimension H represents the height dimension, the splittable dimension W represents the width dimension, and the splittable dimension C represents the channel dimension.

[0150] In the candidate storage space, "Mem" is the category identifier of the memory, and "Cluster" represents the identifier of the off-core cache.

[0151] In the candidate data exchange level, "No" represents no exchange, "Cluster" represents inter-cluster data exchange, and "Core" represents inter-core data exchange. The priority of the splittable dimension that appears earlier in the same data is higher than that of the splittable dimension that appears later. Taking the input neurons in "two-dimensional convolution Conv" as an example, its splittable dimensions are N, H, and W, and the priorities from high to low are N, H, and W in turn.

[0152] In a possible implementation manner, the dynamic label information after adding the splitting index, the identification information of the target storage space, and the target data exchange level can be expressed as:

[0153] dynamic:type2,DIM_B1…Bn,tiling,padding,size,split x[(1_0,1_n-1)…(i_i×n,i_(i+1)n-1)],store,swap

[0154] Among them, split x[(1_0,1_n-1)…(i_i×n,i_(i+1)n-1)] represents the splitting index, where x represents the target splitting dimension, and the splitting length is n. (1_0, 1_n-1) represents that the starting splitting position of the first second data is 0 and the ending splitting position is n-1; (i_i×n, i_(i+1)n-1) represents that the starting splitting position of the i-th second data is i×n and the ending splitting position is (i+1)n-1. store represents the identification information of the target storage space. swap represents the target data exchange level.

[0155] In this embodiment, a model for determining the splitting strategy in real time can be specified in advance according to the correspondence between the reference information related to the splitting strategy (including the splittable dimension, candidate storage spaces, and candidate data exchange dimensions) and the splitting strategy described above. Then, based on the model, the splittable dimension, candidate storage spaces, and candidate data exchange dimensions of the first data can be determined in real time. Furthermore, a splitting index, identification information of the target storage space, and the target data exchange level can be added to the dynamic label information of the first data. Alternatively, a splitting strategy database can be determined in advance according to the reference information, and various required splitting strategies are recorded therein. The present disclosure does not limit this.

[0156] In this embodiment, during the process of a multi-core processor executing an arithmetic task, the arithmetic operations that a computing core can perform on the first data according to the received execution instructions or based on the indication of the label information may include at least one of the following:

[0157] Perform an arithmetic operation on the second data obtained from the target storage space and store the first intermediate result obtained from the arithmetic operation in the target storage space;

[0158] Perform an arithmetic operation on the second data obtained from the target storage space and store the first intermediate result obtained from the arithmetic operation in the off-core cache corresponding to the current computing core;

[0159] Perform an arithmetic operation on the second data obtained from the target storage space and store the first intermediate result obtained from the arithmetic operation in the local memory of the current computing core;

[0160] Obtain the first intermediate results from each memory, perform an arithmetic operation on the multiple first intermediate results to obtain the arithmetic result corresponding to the arithmetic instruction;

[0161] Perform an arithmetic operation on the first intermediate result obtained from the off-core cache of the current computing core and the first intermediate result stored in the local memory of the current computing core to obtain a second intermediate result, and store the second intermediate result in the off-core cache of the current computing core;

[0162] Perform an arithmetic operation on the first intermediate result obtained from the off-core cache of the current computing core and the first intermediate result stored in the local memory of the current computing core to obtain a second intermediate result, and store the second intermediate result in the local memory of the current computing core;

[0163] Perform an arithmetic operation on the second intermediate result obtained from the off-core caches other than the off-core cache of the current computing core among the multiple off-core caches and the second intermediate result stored in the local memory of the current computing core to obtain the arithmetic result corresponding to the arithmetic instruction;

[0164] Store the arithmetic result in the target storage space.

[0165] Application Example

[0166] The following takes "processing operation task A" as an exemplary application scenario to give an application example according to the embodiments of the present disclosure, so as to facilitate understanding of the process of the data processing method. Those skilled in the art should understand that the following application examples are only for the purpose of facilitating understanding of the embodiments of the present disclosure and should not be regarded as a limitation on the embodiments of the present disclosure.

[0167] Assume that the data involved in processing the operation task A includes input neuron I, input weight W, and output neuron O, and the multi-core processor for executing the operation task A is Figure 3 The multi-core processor 1 shown in the figure. The operation task A is a fully connected MLP operation. The input neuron I, input weight W, and output neuron O are data of 1×1024, 1024×4, and 1×4 respectively, and the specific operation is I×W = O. Then, the input neuron I, input weight W, and output neuron O can be classified as "first data", and their corresponding split indexes, target storage spaces, and target data exchange levels can be determined.

[0168] Label information of input neuron I: Static:IN,float32,DIM_NC,{1 1024}dynamic:float16,DIM_NC,C = 256,C = 0,2KB. That is, the data category of I is input neuron, the static data type is float32, the static data dimensions are N and C, the dimension value of dimension N is 1, the dimension value of dimension C is 1024, and the static data dimension order is NC. The dynamic data type is float16, the dynamic data dimension order is NC, the sharding parameter is to shard in the direction of dimension C according to a length of 256, the padding parameter is to pad with "0" in the direction of dimension C, and the data size is 2KB.

[0169] Label information of input weight W: Static:IW,float32,DIM_NC,{1024 4}dynamic:float16,DIM_NC,C = 256,C = 0,8KB. That is, the data category of W is input weight, the static data type is float32, the static data dimensions are N and C, the dimension value of dimension N is 1024, the dimension value of dimension C is 4, and the static data dimension order is NC. The dynamic data type is float16, the dynamic data dimension order is NC, the sharding parameter is to shard in the direction of dimension C according to a length of 256, the padding parameter is to pad with "0" in the direction of dimension C, and the data size is 8KB.

[0170] Label information of output neuron O: Static: ON, float32, DIM_NC, {1 4}; dynamic: float16, DIM_NC, C = 4, C = 0, 8KB. That is, the data category of O is output neuron, the static data type is float32, the static data dimensions are N and C, the dimension value of dimension N is 1, the dimension value of dimension C is 4, and the static data dimension order is NC. The dynamic data type is float16, the dynamic data dimension order is NC, the sharding parameter is to shard in the direction of dimension C with a length of 4, the padding parameter is to pad with "0" in the direction of dimension C, and the data size is 8KB.

[0171] For input neuron I: Assume that according to the pre-determined split strategy database, its splittable dimensions include N and C (the priority of N is higher than that of C), and since input neuron I is 1×1024 and multi-core processor 1 has 2 channels. First, judge for the splittable dimension N. The number of channels 2 is greater than the operation dimension value of input neuron I in dimension N (the determination method refers to the above), so the splittable dimension N cannot be used as the target split dimension; then, judge for the splittable dimension C. The number of channels 2 is less than the operation dimension value of input neuron I in dimension C, so the splittable dimension C is determined as the target split dimension. Further, the split length can be determined as L = 1024÷4 (total number of computing cores) = 512. Then the split index of input neuron I is C[(0,255),(256,511),(512,767),(768,1023)]. Continuing based on the pre-determined split strategy database for the target split dimension C, it can be known that: the target storage space of input neuron I is "Mem" (i.e., memory), and the target data exchange level is "NO", that is, no exchange. Then the label information of input neuron I after addition is: Static: IN, float32, DIM_NC, {1 1024}; dynamic: float16, DIM_NC, C = 256, C = 0, 2KB, C[(0,255),(256,511),(512,767),(768,1023)], Mem, NO.

[0172] Based on the same process as that of the input neuron I, it can be determined that the label information of the input weight W after increase is: Static: IW, float32, DIM_NC, {1024 4} dynamic: float16, DIM_NC, C = 256, C = 0, 8KB, N[(0, 255), (256, 511), (512, 767), (768, 1023)], Mem, NO. The label information of the output neuron O after increase is: Static: ON, float32, DIM_NC, {1 4} dynamic: float16, DIM_NC, C = 4, C = 0, 8KB, C[(0, 3)], Cluster, Core.

[0173] Based on the label information of the above-mentioned input neuron I, input weight W, and output neuron O, when the multi-core processor 1 executes the arithmetic task A, Figure 4a Fig. shows a schematic diagram of the computing task allocation for the multi-core processor 1 to execute the arithmetic task A according to an embodiment of the present disclosure. Figure 4b Fig. shows a schematic diagram of the process for the multi-core processor 1 to execute the arithmetic task A according to an embodiment of the present disclosure. Combining Figure 4a 、 Figure 4b it can be known that the specific process is as follows:

[0174] The multi-core processor 1 or the processor that allocates tasks to the multi-core processor 1, according to the label information of the input neuron I and the input weight W, first splits the input neuron I and the input weight W as Figure 4a shown. The input neuron is split into i1, i2, i3, i4, and the input weight W is split into w1, w2, w3, w4. And as Figure 4b shown, i1, i2, w1, and w2 are respectively stored in the memory 1, and i3, i4, w3, and w4 are respectively stored in the memory 2.

[0175] Then, the computing core 1 fetches i1 and w1 from the memory 1 and performs calculations to obtain the first intermediate result o1, and then stores it in its corresponding local memory NB; the computing core 2 fetches i2 and w2 from the memory 1 and performs calculations to obtain the first intermediate result o2, and stores the first intermediate result o2 in the off-core cache 1; the computing core 3 fetches i3 and w3 from the memory 2 and performs calculations to obtain the first intermediate result o3, and then stores it in its corresponding local memory NB; the computing core 4 fetches i4 and w4 from the memory 2 and performs calculations to obtain the second intermediate result o4, and stores the second intermediate result o4 in the off-core cache 2.

[0176] Next, computing core 1 fetches o2 from off-chip cache 1, adds o2 to o1 to obtain the second intermediate result o5, and then stores it in its corresponding local memory NB; computing core 3 fetches o4 from off-chip cache 2, adds o4 to o3 to obtain the second intermediate result o6, and stores the second intermediate result o6 in off-chip cache 2.

[0177] After that, computing core 1 fetches o6 from off-chip cache 2, adds o5 to o6 to obtain the operation result o7 (i.e., the output neuron O), and stores o7 in memory 1 to complete the operation task.

[0178] Among them, since the target storage space of the output neuron O is the off-chip cache Cluster and the data exchange level is the inter-core data exchange Core, computing cores 2 and 4 will store their operation results o2 and o4 in the corresponding off-chip caches, and computing core 3 will store the obtained o6 in off-chip cache 2. In this way, it can be ensured that computing core 1 and computing core 3 can obtain the required intermediate results or cluster computing results from the off-chip cache.

[0179] It should be noted that although the above embodiments are used as examples to introduce the data processing method as above, those skilled in the art can understand that the present disclosure should not be limited thereto. In fact, users can flexibly set each step and module according to personal preferences and / or actual application scenarios as long as it conforms to the technical solution of the present disclosure.

[0180] Figure 5 The block diagram of a data processing device according to an embodiment of the present disclosure is shown. As Figure 5 shown, the device is applied to a multi-core processor, and the device includes: an information acquisition module 51, an index determination module 52, and a label addition module 53.

[0181] The information acquisition module 51 acquires the first data and the label information of the first data, where the first data is used to participate in neural network operations, and the label information includes static label information and dynamic label information;

[0182] The index determination module 52 determines the splitting index required to split the first data into multiple second data according to the parameters of the operator corresponding to the operation participated by the determined first data, the number of memory channels in the multi-core processor, and the label information;

[0183] The label addition module 53 adds the splitting index to the dynamic label information so that the multi-core processor can perform the neural network operation on the first data based on the label information,

[0184] Among them, the static label information is used to represent information related to the participation of the first data in the neural network operation, and the dynamic label information is used to represent information related to the association between the first data and the multi-core processor.

[0185] In a possible implementation, the index determination module 52 includes:

[0186] A policy determination sub-module, which determines the splitting policy corresponding to the first data from the splitting policy database according to the parameters of the operator corresponding to the operation participated by the determined first data and the label information;

[0187] An index determination sub-module, which determines the splitting index required to split the first data into multiple second data according to the splitting policy, the number of memory channels in the multi-core processor, and the label information,

[0188] Among them, the splitting policy includes at least one splittable dimension of the first data and the priority level corresponding to each splittable dimension,

[0189] The splitting index includes: a target splitting dimension, the starting splitting position and the ending splitting position of each second data on the target splitting dimension of the first data, and the target splitting dimension is the selected dimension among the at least one splittable dimension.

[0190] In a possible implementation, the policy determination sub-module includes:

[0191] A label information acquisition sub-module, which acquires the computational graph corresponding to the neural network. The computational graph includes operation nodes, data nodes, and the connection relationship between the data nodes and the operation nodes; among them, the data nodes record the label information of the data participating in the neural network operation, and the label information includes static label information and dynamic label information;

[0192] A parameter determination sub-module, which obtains the information of the target operation node connected to the data node where the first data is located in the computational graph, and the information of the target operation node includes the parameters of the operator corresponding to the operation for implementing the target operation node;

[0193] A splitting policy determination sub-module, which determines the splitting policy corresponding to the first data from the splitting policy database according to the parameters of the operator and the static label information of the first data;

[0194] Among them, the pre-determined splitting policies are recorded in the splitting policy database.

[0195] In a possible implementation, the index determination sub-module includes:

[0196] A dimension and length determination sub-module determines a target split dimension from the at least one split-able dimension according to the number of memory channels and the operation dimension value of each dimension of the first data determined according to the label information during the neural network operation process, and determines a split length corresponding to the target split dimension, where the split length is the length of the second data in the direction of the target split dimension;

[0197] A split index determination sub-module determines a split index for splitting the first data into a plurality of second data according to the number of memory channels, the operation dimension value of the first data on the target split dimension, and the split length.

[0198] In a possible implementation manner, the static label information includes: static data dimensions, where the index determination sub-module further includes:

[0199] A dimension value determination sub-module determines the operation dimension value of each dimension of the first data during the neural network operation process according to the dimension value of each static data dimension of the first data and the padding parameter.

[0200] In a possible implementation manner, the dynamic label information includes the operation dimension value.

[0201] In a possible implementation manner, determining a target split dimension from the at least one split-able dimension according to the number of memory channels and the operation dimension value of each dimension of the first data determined according to the label information during the neural network operation process, and determining a split length corresponding to the target split dimension, where the split length is the length of the second data in the direction of the target split dimension, includes:

[0202] When there are multiple split-able dimensions, it is sequentially determined whether the split-able dimensions meet the split condition in the order of decreasing priority. When it is determined that the current split-able dimension meets the split condition, the current split-able dimension is determined as the target split dimension,

[0203] where the split condition includes: the operation dimension value of the first data on the current split-able dimension is greater than or equal to the number of memory channels.

[0204] In a possible implementation manner, determining a target split dimension from the at least one split-able dimension according to the number of memory channels and the operation dimension value of each dimension of the first data determined according to the label information during the neural network operation process, and determining a split length corresponding to the target split dimension, where the split length is the length of the second data in the direction of the target split dimension, includes:

[0205] Determine the ratio of the operation dimension value of the first data in the target split dimension to the number of memory channels;

[0206] Use a rounding function to round the ratio to obtain the split length corresponding to the target split dimension,

[0207] wherein, the rounding function includes any one of the following: downward rounding function, upward rounding function, round-off rounding function.

[0208] In a possible implementation manner, the index determination module 52 further includes:

[0209] A level determination sub-module, which determines at least one split dimension corresponding to the first data and the priority level corresponding to each split dimension according to the parameters of the operator corresponding to the operation participated by the determined first data, the arithmetic unit characteristics of the multi-core processor, and the tag information.

[0210] In a possible implementation manner, the multi-core processor is provided with a plurality of computing core clusters, each computing core cluster includes a plurality of computing cores, and the device further includes:

[0211] A candidate storage space determination module, which determines the candidate storage space corresponding to each split dimension according to the parameters of the operator corresponding to the operation participated by the determined first data and the storage space set in the multi-core processor;

[0212] A first dynamic tag information adding module, which determines the target storage space corresponding to the first data from the candidate storage spaces according to the target split dimension, and adds the identification information of the target storage space to the dynamic tag information,

[0213] wherein, the storage space includes a plurality of memories; or the storage space includes a plurality of memories and a plurality of off-core caches, each computing core can access any one of the plurality of memories, and the plurality of computing cores in each computing core cluster share a corresponding off-core cache.

[0214] In a possible implementation manner, the splitting strategy includes the candidate storage space corresponding to each split dimension, and the device further includes:

[0215] A second dynamic tag information adding module, which determines the target storage space from the candidate storage spaces according to the target split dimension, and adds the identification information of the target storage space to the dynamic tag information,

[0216] Wherein, the storage space includes a plurality of memories; or the storage space includes a plurality of memories and a plurality of off-core caches, and each computing core can access any one of the plurality of memories, and a plurality of computing cores in each computing core cluster share a corresponding off-core cache.

[0217] In a possible implementation manner, the apparatus further includes:

[0218] A hierarchical selection module, which determines a candidate data exchange level corresponding to each splittable dimension according to the parameters of the operator corresponding to the operation participated by the determined first data, the arithmetic unit characteristics of the multi-core processor, and the tag information;

[0219] A first level increase module, which determines the target data exchange level corresponding to the first data from the candidate data exchange levels according to the target splitting dimension, and adds the target data exchange level to the dynamic tag information.

[0220] In a possible implementation manner, the splitting strategy further includes a candidate data exchange level corresponding to each splittable dimension, and the apparatus further includes:

[0221] A second level increase module, which determines the candidate data exchange level corresponding to the target splitting dimension in at least one candidate data exchange level as the target data exchange level corresponding to the first data, and adds the target data exchange level to the dynamic tag information.

[0222] The data processing apparatus provided by the embodiments of the present disclosure enables data to be split and laid out based on the splitting index in the dynamic tag information when performing a neural network operation task, improving the processing efficiency and speed of the neural network operation task.

[0223] The apparatus provided by the embodiments of the present disclosure enables the multi-core processor to split, layout, and determine the storage location of data based on the dynamic tag information with a new splitting index, storage space, and data exchange level when performing an operation task, improving the processing efficiency and speed of the operation task.

[0224] The present disclosure provides a machine learning computing device, which may include one or more of the above data processing devices, and is configured to obtain data to be computed and control information from other processing devices and execute specified machine learning computations. The machine learning computing device may obtain neural network computing macro instructions or neural network computing instructions to be executed from other machine learning computing devices or non-machine learning computing devices, and transmit the execution results to peripheral devices (also referred to as other processing devices) through an I / O interface. Peripheral devices such as cameras, displays, mice, keyboards, network cards, Wi-Fi interfaces, and servers. When there are more than one data processing device, the data processing devices can be linked and data can be transmitted through a specific structure. For example, they can be interconnected and data can be transmitted through a PCIE bus to support the computation of larger-scale neural networks. At this time, the same control system can be shared, or each can have its own independent control system; the memory can be shared, or each accelerator can have its own memory. In addition, the interconnection method can be any interconnection topology.

[0225] The machine learning computing device has high compatibility and can be connected to various types of servers through a PCIE interface.

[0226] Figure 6 FIG. is a structural diagram showing a combined processing device 1200 according to an embodiment of the present disclosure. As Figure 6 shown in the figure, the combined processing device 1200 includes a computing processing device 1202, an interface device 1204, other processing devices 1206, and a storage device 1208. According to different application scenarios, the computing processing device may include one or more computing devices 1210. The computing processing device 1202 may be the above machine learning computing device or the above data processing device.

[0227] In different embodiments, the computing processing device of the present disclosure may be configured to execute operations specified by a user. In an exemplary application, the computing processing device may be implemented as a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the computing processing device may be implemented as artificial intelligence processor cores (i.e., the computing cores described above) or partial hardware structures of artificial intelligence processor cores.

[0228] In an exemplary operation, the computing processing device of the present disclosure can interact with other processing devices through an interface device to jointly complete an operation specified by a user. Depending on the implementation, other processing devices of the present disclosure can include one or more types of general-purpose and / or dedicated processors such as a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence processor. These processors can include, but are not limited to, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and the number thereof can be determined according to actual needs. As mentioned above, only with respect to the computing processing device of the present disclosure, it can be regarded as having a single-core structure or a homogeneous multi-core structure. However, when considering the computing processing device and other processing devices together, the two can be regarded as forming a heterogeneous multi-core structure.

[0229] In one or more embodiments, the other processing device can serve as an interface for external data and control for the computing processing device of the present disclosure (which can be specifically implemented as a related computing device for artificial intelligence such as neural network operations), and perform basic controls including but not limited to data transfer, startup and / or stop of the computing device. In other embodiments, the other processing device can also cooperate with the computing processing device to jointly complete an operation task.

[0230] In one or more embodiments, the interface device can be used to transfer data and control instructions between the computing processing device and other processing devices. For example, the computing processing device can obtain input data from other processing devices via the interface device and write it into the storage device (or memory) on the chip of the computing processing device. Further, the computing processing device can obtain control instructions from other processing devices via the interface device and write them into the control cache on the chip of the computing processing device. Alternatively or optionally, the interface device can also read the data in the storage device of the computing processing device and transfer it to other processing devices.

[0231] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is respectively connected to the computing processing device and the other processing device. In one or more embodiments, the storage device may be used to store data of the computing processing device and / or the other processing device. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the computing processing device or other processing device.

[0232] In some embodiments, the present disclosure also discloses a chip (such as Figure 7 the chip 1302 shown in Figure 6 ). In one implementation, the chip is a system-on-chip (SoC) and integrates one or more combined processing devices as shown in Figure 7 . The chip can be connected to other related components through an external interface device (such as the external interface device 1306 shown in Figure 7 ). The related components may be, for example, a camera, a display, a mouse, a keyboard, a network card, or a wifi interface. In some application scenarios, other processing units (such as a video codec) and / or interface modules (such as a DRAM interface) may be integrated on the chip. In some embodiments, the present disclosure also discloses a chip package structure including the above chip. In some embodiments, the present disclosure also discloses a board including the above chip package structure. The following will be described in detail with reference to Figure 7 this board.

[0233] Figure 7 is a schematic structural diagram of a board 1300 according to an embodiment of the present disclosure. As shown in Figure 7 , the board includes a storage device 1304 for storing data, which includes one or more storage units 1310. The storage device can be connected and perform data transmission with a control device 1308 and the above-mentioned chip 1302 through, for example, a bus. Further, the board also includes an external interface device 1306, which is configured to perform data relay or transfer functions between the chip (or the chip in the chip package structure) and an external device 1312 (such as a server or a computer, etc.). For example, the data to be processed can be transmitted from the external device to the chip through the external interface device. Again, for example, the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device may have different interface forms. For example, it may adopt a standard PCIE interface, etc.

[0234] In one or more embodiments, the control device in the disclosed board can be configured to regulate the state of the chip. For this purpose, in one application scenario, the control device can include a microcontroller unit (MCU) for regulating the operating state of the chip.

[0235] According to the above combination Figure 6 and Figure 7 the description of, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which may include one or more of the above boards, one or more of the above chips, and / or one or more of the above combined processing devices.

[0236] According to different application scenarios, the electronic device or apparatus of the present disclosure may include a server, a cloud server, a server cluster, a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a PC device, an Internet of Things terminal, a mobile terminal, a mobile phone, a driving recorder, a navigator, a sensor, a camera, a camera, a video camera, a projector, a watch, a headset, a mobile storage, a wearable device, a vision terminal, an autonomous driving terminal, a vehicle, a household appliance, and / or a medical device. The vehicle includes an airplane, a ship, and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, a rice cooker, a humidifier, a washing machine, a light, a gas stove, an oil fume machine; the medical device includes a nuclear magnetic resonance instrument, a B-ultrasound instrument, and / or an electrocardiogram instrument. The electronic device or apparatus of the present disclosure can also be applied to fields such as the Internet, the Internet of Things, a data center, energy, transportation, public management, manufacturing, education, a power grid, telecommunications, finance, retail, a construction site, and medicine. Further, the electronic device or apparatus of the present disclosure can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing such as the cloud, the edge, and the terminal. In one or more embodiments, the high-computing-power electronic device or apparatus according to the present disclosure can be applied to cloud devices (such as cloud servers), while the low-power-consuming electronic device or apparatus can be applied to terminal devices and / or edge devices (such as smart phones or cameras). In one or more embodiments, the hardware information of the cloud device is compatible with the hardware information of the terminal device and / or the edge device, so that appropriate hardware resources can be matched from the hardware resources of the cloud device according to the hardware information of the terminal device and / or the edge device to simulate the hardware resources of the terminal device and / or the edge device, so as to complete the unified management, scheduling, and collaborative work of the end-cloud integration or cloud-edge-end integration.

[0237] It should be noted that, for the purpose of simplicity, some methods and their embodiments in the present disclosure are expressed as a series of actions and their combinations. However, those skilled in the art can understand that the solutions of the present disclosure are not limited by the order of the described actions. Therefore, according to the disclosure or teachings of the present disclosure, those skilled in the art can understand that some of the steps can be executed in other orders or simultaneously. Further, those skilled in the art can understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved are not necessarily required for the implementation of certain solutions of the present disclosure. In addition, according to different solutions, the descriptions of some embodiments in the present disclosure also have different focuses. In view of this, those skilled in the art can understand that for the parts not detailed in a certain embodiment of the present disclosure, they can also refer to the relevant descriptions of other embodiments.

[0238] In terms of specific implementation, based on the disclosure and teachings of the present disclosure, those skilled in the art can understand that several embodiments disclosed in the present disclosure can also be implemented by other means not disclosed herein. For example, for each unit in the foregoing embodiments of the electronic device or apparatus, based on the consideration of logical functions, it is divided herein, but there may be other division methods in actual implementation. Another example is that multiple units or components can be combined or integrated into another system, or some features or functions of the units or components can be selectively disabled. Regarding the connection relationships between different units or components, the connections discussed in conjunction with the accompanying drawings herein can be direct or indirect couplings between the units or components. In some scenarios, the aforementioned direct or indirect couplings involve communication connections using interfaces, where the communication interfaces can support signal transmission in electrical, optical, acoustic, magnetic, or other forms.

[0239] In the present disclosure, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. The aforementioned components or units may be located at the same position or distributed to multiple network units. In addition, according to actual needs, some or all of the units can be selected to achieve the purpose of the solutions described in the embodiments of the present disclosure. In addition, in some scenarios, multiple units in the embodiments of the present disclosure can be integrated into one unit or each unit exists physically separately.

[0240] In some implementation scenarios, the above-integrated unit can be implemented in the form of a software program module. If it is implemented in the form of a software program module and sold or used as an independent product, the integrated unit can be stored in a computer-readable memory. Based on this, when the solution of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in the memory, and it can include several instructions to enable a computer device (such as a personal computer, a server, or a network device, etc.) to execute some or all of the steps of the method described in the embodiments of the present disclosure. The aforementioned memory can include, but is not limited to, various media that can store program codes such as USB flash drives, flash memory drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs.

[0241] In some other implementation scenarios, the above-integrated unit can also be implemented in the form of hardware, that is, a specific hardware circuit, which can include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit can include, but is not limited to, physical devices, and the physical devices can include, but are not limited to, devices such as transistors or memristors. In view of this, various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs, etc. Further, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high bandwidth memory (HBM), a hybrid memory cube (HMC), a ROM, and a RAM, etc.

[0242] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A data processing method, characterized in that, Applied to a multi-core processor, the method includes: Obtain first data and label information of the first data, where the first data is used to participate in neural network operations, and the label information includes static label information and dynamic label information; Determine a split index required to split the first data into multiple second data according to parameters of an operator corresponding to an operation in which the first data participates, the number of memory channels in the multi-core processor, and the label information; Add the split index to the dynamic label information so that the multi-core processor can perform the neural network operation on the first data based on the label information, wherein the static label information is used to characterize information related to the first data participating in the neural network operation, and the dynamic label information is used to characterize information related to the first data and the multi-core processor; wherein determining a split index required to split the first data into multiple second data according to parameters of an operator corresponding to an operation in which the first data participates, the number of memory channels in the multi-core processor, and the label information includes: Determine a split strategy corresponding to the first data from a split strategy database according to parameters of an operator corresponding to an operation in which the first data participates and the label information; Determine a split index required to split the first data into multiple second data according to the split strategy, the number of memory channels in the multi-core processor, and the label information, wherein the split strategy includes at least one splittable dimension of the first data and a priority level corresponding to each splittable dimension, The split index includes: a target split dimension, a start split position and an end split position of each second data on the target split dimension of the first data, and the target split dimension is a selected dimension among the at least one splittable dimension.

2. The method according to claim 1, wherein Determining a split strategy corresponding to the first data from a split strategy database according to parameters of an operator corresponding to an operation in which the first data participates and the label information includes: Obtain a computational graph corresponding to the neural network, where the computational graph includes operation nodes, data nodes, and connection relationships between data nodes and operation nodes; wherein, the label information of data participating in the neural network operation is recorded in the data nodes, and the label information includes static label information and dynamic label information; Obtain information of a target operation node connected to the data node where the first data is located in the computational graph, and the information of the target operation node includes parameters of an operator corresponding to the operation implemented by the target operation node; Determine a split strategy corresponding to the first data from a split strategy database according to the parameters of the operator and the static label information of the first data; wherein pre-determined split strategies are recorded in the split strategy database.

3. The method according to claim 1, wherein Determining a split index required to split the first data into multiple second data according to the split strategy, the number of memory channels in the multi-core processor, and the label information includes: Determine a target split dimension from the at least one splittable dimension according to the number of memory channels and the operation dimension values of each dimension of the first data determined according to the label information during the neural network operation process, and determine a split length corresponding to the target split dimension, where the split length is the length of the second data in the direction of the target split dimension; Determine a split index for splitting the first data into multiple second data according to the number of memory channels, the operation dimension value of the first data on the target split dimension, and the split length.

4. The method according to claim 3, characterized in that The static label information includes: static data dimensions, Among them, determining the split index required to split the first data into multiple second data according to the split strategy, the number of memory channels in the multi-core processor, and the label information further includes: Determine the operation dimension values of each dimension of the first data during the neural network operation process according to the dimension values of each static data dimension of the first data and the padding parameters.

5. The method according to claim 3, wherein The dynamic label information includes the operation dimension values.

6. The method according to claim 3, characterized in that, Determine a target split dimension from the at least one splittable dimension according to the number of memory channels and the operation dimension values of each dimension of the first data determined according to the label information during the neural network operation process, and determine a split length corresponding to the target split dimension, where the split length is the length of the second data in the direction of the target split dimension, including: When there are multiple splittable dimensions, sequentially determine whether the splittable dimensions meet the split conditions in descending order of priority. When it is determined that the current splittable dimension meets the split conditions, determine the current splittable dimension as the target split dimension, Among them, the split conditions include: the operation dimension value of the first data on the current splittable dimension is greater than or equal to the number of memory channels.

7. The method according to claim 3, wherein Determine a target split dimension from the at least one splittable dimension according to the number of memory channels and the operation dimension values of each dimension of the first data determined according to the label information during the neural network operation process, and determine a split length corresponding to the target split dimension, where the split length is the length of the second data in the direction of the target split dimension, including: Determine the ratio of the operation dimension value of the first data on the target split dimension to the number of memory channels; Use a rounding function to round the ratio to obtain a split length corresponding to the target split dimension, Among them, the rounding function includes any one of the following: floor function, ceiling function, round function.

8. The method according to claim 1, wherein Determine the split index required to split the first data into multiple second data according to the parameters of the operator corresponding to the operation participated by the determined first data, the number of memory channels in the multi-core processor, and the label information further includes: Determine at least one splittable dimension corresponding to the first data and the priority level corresponding to each splittable dimension according to the parameters of the operator corresponding to the operation in which the determined first data participates, the arithmetic unit characteristics of the multi-core processor, and the tag information.

9. The method according to claim 1, characterized in that, The multi-core processor is provided with multiple computing core clusters, and each computing core cluster includes multiple computing cores. The method further includes: Determine the candidate storage space corresponding to each splittable dimension according to the parameters of the operator corresponding to the operation in which the determined first data participates and the storage space provided in the multi-core processor. According to the target splitting dimension, determine the target storage space corresponding to the first data from the candidate storage spaces, and add the identification information of the target storage space to the dynamic tag information. Wherein, the storage space includes multiple memories; or the storage space includes multiple memories and multiple off-core caches. Each computing core can access any one of the multiple memories, and the multiple computing cores in each computing core cluster share a corresponding off-core cache.

10. The method according to claim 1, wherein The splitting strategy includes the candidate storage space corresponding to each splittable dimension. The method further includes: According to the target splitting dimension, determine the target storage space from the candidate storage spaces, and add the identification information of the target storage space to the dynamic tag information. Wherein, the storage space includes multiple memories; or the storage space includes multiple memories and multiple off-core caches. Each computing core can access any one of the multiple memories, and the multiple computing cores in each computing core cluster share a corresponding off-core cache.

11. The method according to claim 1, wherein The method further includes: Determine the candidate data exchange level corresponding to each splittable dimension according to the parameters of the operator corresponding to the operation in which the determined first data participates, the arithmetic unit characteristics of the multi-core processor, and the tag information. According to the target splitting dimension, determine the target data exchange level corresponding to the first data from the candidate data exchange levels, and add the target data exchange level to the dynamic tag information.

12. The method according to claim 1, wherein The splitting strategy further includes the candidate data exchange level corresponding to each splittable dimension. The method further includes: Determine the candidate data exchange level corresponding to the target splitting dimension in at least one candidate data exchange level as the target data exchange level corresponding to the first data, and add the target data exchange level to the dynamic tag information.

13. A non-volatile computer-readable storage medium, characterized in that, Stored thereon are computer program instructions, characterized in that when the computer program instructions are executed by a processor, the data processing method according to any one of claims 1 to 12 is implemented.

14. A data processing device, characterized in that, The device includes a processor and a memory. A computer program is stored in the memory. When the processor executes the computer program, the method according to any one of claims 1-12 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, computer equipment and storage medium

    CN110458285A

  • Data processing method and device, computer equipment and storage medium

    CN110647981A

  • Image data processing method and device, computer equipment and storage medium

    CN111897579A