Tensor processing method and device, electronic equipment and computer readable storage medium
By introducing a slicing unit into the tensor processing device, the number of software slicing times of large-size tensors is reduced, and the problems of poor power consumption control and low utilization in the prior art are solved, thereby achieving more efficient resource utilization.
Patent Information
- Application Number
- CN202510459754.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when the tensor processing device processes large-size tensor data, frequent software slicing scheduling leads to poor power consumption control and a decrease in utilization rate, and cannot efficiently utilize computing resources.
A splitting unit is introduced into the tensor processing device to share the segmentation work of the scheduling module, initially splitting the large-size tensor into intermediate tensors, and further splitting it into basic tensors through the slicing unit, reducing the number of interactions between software and hardware.
The number of software and hardware interactions in the tensor processing device is effectively reduced, power consumption control is improved, and device utilization is improved.
Smart Images

Figure CN120336022A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular, to a tensor processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Currently, in the field of data processing, calculations of two-dimensional or higher-dimensional tensors such as convolution calculations or matrix multiplication operations are involved. In the actual operation process, the scale of some tensors in dimensions such as height or width will exceed the limitations of some hardware in the tensor processing device. For example, the computing units set in the tensor processing device can only fixedly process tensor data with a small size scale. Therefore, in related technologies, software algorithms are used to split the tensor data with a large size scale into a size scale that can be processed by the computing units. However, each split of the input tensor data using the software algorithm requires a data scheduling. If the scale of the data that can be processed by the computing units is small, or the input tensor data is too large, the software algorithm needs to split the input tensor data multiple times, and the number of times of scheduling data is large, resulting in poor power consumption control. Moreover, when the software algorithm is frequently used for splitting and scheduling, the utilization rate of the tensor processing device will also decrease accordingly. Summary of the Invention
[0003] Embodiments of the present disclosure provide a tensor processing method, apparatus, electronic device, and computer-readable storage medium, which can effectively reduce the number of interactions between software and hardware in the tensor processing device, improve power consumption control, and increase the utilization rate of the tensor processing device.
[0004] According to one aspect of the present disclosure, there is provided a tensor processing method for a tensor processing device. The tensor processing device includes a scheduling module, a splitting unit, and a computing unit. The method includes: obtaining a first input tensor through the scheduling module, splitting the first input tensor in at least one dimension to obtain a plurality of intermediate tensors, generating corresponding data processing instructions according to each of the intermediate tensors, and sending the data processing instructions to the splitting unit; parsing, by the splitting unit, the data processing instructions, reading the intermediate tensors according to the parsed data processing instructions, splitting the intermediate tensors in at least one dimension to obtain a plurality of basic tensors, and mapping the plurality of basic tensors to the computing unit; and performing tensor calculations on the plurality of basic tensors in parallel by the computing unit to obtain a tensor calculation result.
[0005] According to another aspect of the present disclosure, a tensor processing device is provided, including a scheduling module, a splitting unit, and a computing unit. Specifically: the scheduling module is configured to obtain a first input tensor, split the first input tensor in at least one dimension to obtain a plurality of intermediate tensors, generate corresponding data processing instructions according to each of the intermediate tensors, and send the data processing instructions to the splitting unit; the splitting unit is configured to parse the data processing instructions, read the intermediate tensors according to the parsed data processing instructions, split the intermediate tensors in at least one dimension to obtain a plurality of basic tensors, and map the plurality of basic tensors to the computing unit; the computing unit is configured to perform tensor calculations on the plurality of basic tensors in parallel to obtain a tensor calculation result.
[0006] Optionally, the splitting unit is further configured to: use the upper limit value of the computable size that can be processed by the computing unit in each dimension as the first splitting granularity in each dimension; perform cyclic splitting on the intermediate tensor in the corresponding dimension according to the first splitting granularity in each dimension to obtain a plurality of basic tensors.
[0007] Optionally, the splitting unit is further configured to: for each dimension, compare the size of the intermediate tensor with the first splitting granularity, and determine the dimension corresponding to the comparison result that the size of the intermediate tensor is greater than the first splitting granularity as the dimension to be split; when there are multiple dimensions to be split, traverse each of the dimensions to be split, and perform cyclic splitting on the tensor obtained by splitting the previous traversed dimension to be split according to the first splitting granularity corresponding to the current dimension to obtain a plurality of basic tensors.
[0008] Optionally, the scheduling unit is further configured to: use the upper limit value of the splitting size that can be processed by the splitting unit in each dimension as the second splitting granularity in each dimension; perform cyclic splitting on the first input tensor in the corresponding dimension according to the second splitting granularity in each dimension to obtain a plurality of intermediate tensors.
[0009] Optionally, the scheduling unit is further configured to: obtain a plurality of first candidate splitting schemes configured in advance, and determine the input tensor size of the first input tensor in each dimension; according to the relationship between the input tensor size in each dimension and the candidate splitting granularities of each of the first candidate splitting schemes in each dimension, determine a first target splitting scheme from the plurality of first candidate splitting schemes; determine the candidate splitting granularities of the first target splitting scheme in each dimension as the third splitting granularity in the corresponding dimension; perform cyclic splitting on the first input tensor in the corresponding dimension according to the third splitting granularity in each dimension to obtain a plurality of intermediate tensors.
[0010] Optionally, the scheduling unit is further configured to: determine the input tensor sizes of the first input tensor in each dimension, and the upper limit of the slicing size that can be processed in a single time by the slicing unit in each dimension; determine the fourth slicing granularity of each dimension according to the ratio relationship between each of the input tensor sizes and each of the size upper limits; perform cyclic slicing on the first input tensor in the corresponding dimension according to the fourth slicing granularity of each dimension to obtain a plurality of intermediate tensors.
[0011] Optionally, the scheduling unit is further configured to: traverse each dimension, and determine the minimum number of slicing times of the first input tensor in each dimension by the scheduling module in the current dimension according to the ratio relationship between the input tensor size of each dimension and the slicing size upper limit of the current dimension; determine the second candidate slicing schemes of the first input tensor in each dimension, and statistically obtain the total number of slicing times of each of the second candidate slicing schemes according to each of the minimum number of slicing times; determine the second candidate slicing scheme with the smallest numerical value of the total number of slicing times as the second target slicing scheme; and determine the fourth slicing granularity of each dimension according to the second target slicing scheme.
[0012] Optionally, the scheduling unit is further configured to: encode each of the intermediate tensors according to a first encoding format to generate corresponding data processing instructions, and add each of the data processing instructions to an instruction queue, where the data processing instruction includes an encoding label for indicating the first encoding format; and sequentially issue each of the data processing instructions to the slicing unit according to the instruction order of the instruction queue.
[0013] Optionally, the slicing unit is further configured to: parse the data processing instruction, read the intermediate tensor according to the parsed data processing instruction, and call the decoding unit to decode the intermediate tensor according to the encoding label to obtain the decoded intermediate tensor.
[0014] Optionally, the scheduling module is further configured to: obtain a second input tensor, and use the second input tensor as an intermediate tensor, where the tensor sizes of the second input tensor in each dimension are less than or equal to the size upper limit that can be processed in a single time by the slicing unit in the corresponding dimension.
[0015] According to another aspect of the present disclosure, an electronic device is provided, characterized in that the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and when the program is run by the processor, the tensor processing method as described above is implemented.
[0016] According to another aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores one or more programs, and the one or more programs can be run by one or more processors to implement the tensor processing method as described above.
[0017] The tensor processing method, device, electronic device, and computer-readable storage medium provided by the present disclosure. The method obtains a first input tensor through a scheduling module, divides the first input tensor in at least one dimension to obtain a plurality of intermediate tensors, then generates corresponding data processing instructions according to each intermediate tensor, and issues the data processing instructions to a splitting unit; then, the splitting unit parses the data processing instructions, reads the intermediate tensors according to the parsed data processing instructions, divides the intermediate tensors in at least one dimension to obtain a plurality of basic tensors, and maps the plurality of basic tensors to a computing unit; then, the computing unit parallelly executes the tensor calculations of the plurality of basic tensors to obtain a tensor calculation result. Therefore, in the embodiment of the present disclosure, a splitting unit for splitting tensor data is added in the tensor processing device to share part of the splitting work originally borne by the scheduling module, so that the scheduling module only needs to perform fewer splitting times, initially divides the first input tensor with too large a size into intermediate tensors with relatively smaller sizes, rather than directly dividing the first input tensor into basic tensors, reducing the number of generated data processing instructions, thereby effectively reducing the interaction times between software and hardware in the tensor processing device, improving power consumption control, and increasing the utilization rate of the tensor processing device.
[0018] Other features and advantages of the present disclosure will be described in the following specification, and part of them will become obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Brief Description of the Drawings
[0019] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0020] Figure 1 It is a system architecture diagram of a tensor processing device to which the tensor processing method of the embodiment of the present disclosure is applied;
[0021] Figure 2 It is an optional flowchart of the tensor processing method provided by the embodiment of the present disclosure;
[0022] Figure 3 It is a schematic diagram of the scalable scale of intermediate tensors provided by the embodiment of the present disclosure;
[0023] Figure 4 It is a schematic diagram of the effect of tensor splitting processing provided by an embodiment of the present disclosure;
[0024] Figure 5 It is a schematic flowchart of the specific splitting method of the splitting unit provided by an embodiment of the present disclosure;
[0025] Figure 6 It is a schematic diagram of the effect of the splitting unit splitting the intermediate tensor provided by an embodiment of the present disclosure;
[0026] Figure 7 It is a schematic flowchart of the specific splitting method of the splitting unit provided by another embodiment of the present disclosure;
[0027] Figure 8 It is a schematic flowchart of the specific splitting method of the scheduling module provided by an embodiment of the present disclosure;
[0028] Figure 9 It is a schematic flowchart of the specific splitting method of the scheduling module provided by another embodiment of the present disclosure;
[0029] Figure 10 It is a schematic diagram of the effect of the first input tensor splitting provided by another embodiment of the present disclosure;
[0030] Figure 11 It is a schematic flowchart of the specific splitting method of the scheduling module provided by another embodiment of the present disclosure;
[0031] Figure 12 It is a schematic flowchart of determining the splitting granularity provided by an embodiment of the present disclosure;
[0032] Figure 13 It is a schematic flowchart of issuing a data processing instruction provided by an embodiment of the present disclosure;
[0033] Figure 14 It is a schematic diagram of the specific process of the tensor processing method provided by a specific example;
[0034] Figure 15 It is a schematic diagram of the structure of the tensor processing device provided by an embodiment of the present disclosure;
[0035] Figure 16 It is a schematic diagram of the structure of the electronic device proposed by an embodiment of the present disclosure. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0037] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:
[0038] Tensor: A multilinear mapping defined on the Cartesian product of some vector spaces and some dual spaces. A tensor can be regarded as a multilinear function, used to represent the linear relationship between multiple vectors, scalars, or other tensors. Among them, a tensor has attributes such as rank and shape. The rank represents the dimension of the tensor. For example, a scalar is a 0th-order tensor, a vector is a 1st-order tensor, and a matrix is a 2nd-order tensor. The shape describes the number of elements in each dimension of the tensor.
[0039] Tensor Engine: A computing system specifically designed to efficiently execute tensor operations. A tensor engine can include hardware or software components. In terms of hardware, a tensor engine can include a calculator specifically for executing tensor operations, such as a Tensor Processing Unit (TPU) or some hardware accelerators that improve tensor operations through parallel computing, such as a graphics processing unit, etc. In terms of software, a tensor engine can include a library or framework that provides tensor operation interfaces and optimization algorithms, such as TensorFlow, PyTorch, etc. Through the mutual cooperation of hardware and software, it can execute tensor-related computing tasks, such as matrix multiplication, convolution, dimensionality reduction, tensor decomposition, etc.
[0040] Graphics Processing Unit (GPU): A microprocessor specifically used for processing graphics and image calculations; it was initially designed to accelerate the rendering of computer graphics, but with the development of technology, the application of GPUs has far exceeded this scope. Among them, GPUs were initially designed to accelerate the rendering of 2D and 3D graphics and improve the performance of games and professional graphics software. GPUs have thousands of cores and can process a large amount of data simultaneously, which makes them very suitable for parallel computing tasks. With the development of technology, GPUs are no longer limited to graphics processing. They are now used in various general computing tasks. GPUs play an important role in the fields of deep learning, machine learning, and artificial intelligence because they can quickly process a large amount of data and accelerate the training and inference of neural networks. In the fields of scientific computing and data analysis, GPUs are used to accelerate complex numerical simulations and data analysis tasks. In cloud computing and data centers, GPUs are used to provide high-performance computing resources to support various compute-intensive applications. In professional applications such as video editing, 3D modeling, and scientific visualization, GPUs can provide real-time high-performance rendering.
[0041] Compute Unit (CU): It is a processing module in electronic devices such as tensor engines. A tensor engine can include multiple compute units, and it contains multiple execution units. A compute unit can be regarded as a processing core in a tensor engine, and each compute unit can independently execute instructions and process data in parallel.
[0042] Currently, in the field of data processing, calculations involving two-dimensional or higher-dimensional tensors such as convolution calculations or matrix multiplication operations are involved. During the actual operation process, the scale of some tensors in dimensions such as height or width will exceed the limitations of some hardware in the tensor processing device. For example, the compute units set in the tensor processing device can only fixedly process tensor data with a small size scale. Therefore, in related technologies, software algorithms are used to split large-size tensor data into a size scale that can be processed by the compute units. However, each split of the input tensor data using software algorithms requires a data scheduling. If the processable data scale of the compute unit is small, or the input tensor data is too large, then the software algorithm needs to split the input tensor data multiple times, and the number of times of scheduling data is large, resulting in poor power consumption control. Moreover, when software algorithms are frequently used for splitting and scheduling, the utilization rate of the tensor processing device will also decrease accordingly.
[0043] Based on this, the present disclosure proposes a tensor processing method, device, electronic device, and computer-readable storage medium. In the embodiments of the present disclosure, a splitting unit for splitting tensor data is added as hardware in the tensor processing device to share part of the splitting work originally undertaken by the scheduling module, so that the scheduling module only needs to perform fewer splitting times, and initially splits the first input tensor with an oversized size into intermediate tensors with a relatively small size, rather than directly splitting the first input tensor into basic tensors, reducing the number of generated data processing instructions, thereby effectively reducing the interaction times between software and hardware in the tensor processing device, improving power consumption control, and increasing the utilization rate of the tensor processing device.
[0044] System architecture description applied in the embodiments of the present disclosure
[0045] Figure 1FIG. 0 is a system architecture diagram of a tensor processing apparatus 100 to which the tensor processing method according to an embodiment of the present disclosure is applied. The tensor processing apparatus 100 may refer to a hardware device for efficiently processing tensor operations, such as a graphics processing unit (GPU), a neural processing unit (NPU), a general-purpose computing on graphics processing units (GPGPU), and other tensor engines or devices including tensor engines. It includes a scheduling module 110, a splitting unit 120, and a computing unit 130. Among them, the tensor processing apparatus 100 may further include a tensor core, and the tensor core may be a hardware unit that efficiently executes tensor operations. The tensor core includes a splitting unit 120 and a computing unit 130. The scheduling module 110 is a part of the tensor processing apparatus 100 responsible for managing and optimizing the execution process of tensor operations. The scheduling module 110 may split tensors according to the current computing resources and task requirements and schedule tasks to the tensor core to achieve optimization of parallel computing.
[0046] A tensor is a multi-dimensional array and is widely used in machine learning and deep learning. The tensor processing apparatus 100 accelerates the training and inference processes of neural networks by optimizing tensor operations such as matrix multiplication and convolution, and is applicable to complex tasks such as image recognition, video processing, and natural language processing. When performing these complex tasks such as image recognition and video processing, the tensor processing apparatus 100 needs to process multi-dimensional tensor data with a large data scale. For example, a color image can be represented by a three-dimensional tensor (height, width, and color channels). However, the hardware used to perform tensor calculations in the tensor processing apparatus 100 is a computing unit 130 with a fixed scale size, which means that the computing unit 130 can only process tensor calculations of a fixed size each time, and the upper limit value of the size that can be processed by the computing unit 130 in a single time is usually less than, or even much less than, the tensor data input to the tensor processing apparatus 100. Therefore, the tensor processing apparatus 100 can use the scheduling module 110 to complete the data chunking and data flow planning of large-size tensor data. For example, after the tensor processing apparatus 100 receives the input tensor data, by calling the scheduling module 110, the input tensor data is split, and at the same time, the transfer process of the split data is specified. However, each split of the input tensor data by calling the scheduling module 110 requires the invocation of multiple software layers and the participation of some hardware. Similarly, the scheduling module 110 also needs to generate an instruction for data scheduling. That is to say, if the scheduling module 110 splits the input tensor data 1024 times, 1024 software and hardware will be invoked, and 1024 instructions will be generated, and each invocation of software and hardware will generate corresponding power consumption.
[0047] After the scheduling module 110 finishes the segmentation of the tensor data, corresponding data processing instructions are generated and sent to the tensor core. The tensor core performs subsequent tensor data processing on the segmented tensor data. The segmentation unit 120 is a hardware arranged in the tensor core. A software program for reading data and segmenting the tensor data can be pre-configured in the segmentation unit 120. When the segmentation unit 120 is called, the segmentation unit 120 can read data and segment the read tensor data according to the established program. Specifically, through the segmentation unit 120, the segmented tensor data can be obtained according to the data processing instructions, and the tensor data segmented by the scheduling module 110 can be segmented again to obtain the basic tensors that the calculation unit 130 can directly calculate. Then, these basic tensors obtained by the secondary segmentation are mapped to the calculation unit 130, and the calculation unit 130 performs parallel tensor calculations on these basic tensors to obtain the tensor calculation results.
[0048] It should be noted that the total number of elements of the tensor data that the segmentation unit 120 can process at one time is greater than the total number of elements of the tensor data that the calculation unit 130 can process at one time, and the total number of elements of the tensor data that the scheduling module 110 can process at one time is greater than the total number of elements of the tensor data that the segmentation unit 120 can process at one time.
[0049] The overall implementation manner of the tensor processing method of the present disclosure embodiment
[0050] The present disclosure embodiment provides a tensor processing method, which can be executed by a tensor processing device. Refer to Figure 2 , Figure 2 which is an optional flow schematic diagram of the tensor processing method provided by the present disclosure embodiment. The tensor processing method includes but is not limited to the following steps 201 to step 203.
[0051] Step 201: Through the scheduling module, obtain the first input tensor, segment the first input tensor in at least one dimension to obtain a plurality of intermediate tensors, generate corresponding data processing instructions according to each intermediate tensor, and send the data processing instructions to the segmentation unit.
[0052] Among them, the scheduling module in the tensor processing device is responsible for slicing and scheduling the input tensor data to optimize the computing efficiency and resource utilization. The first input tensor can refer to the tensor data input to the scheduling module for processing. Specifically, it can refer to the initial tensor data externally input to the tensor processing device, or the intermediate result obtained after other modules or components in the tensor processing device process the externally input tensor data and then pass it to the scheduling module. For example, in application scenarios such as deep learning and machine learning where the tensor processing device is applied, a neural network model is usually pre-deployed in the tensor processing device. The externally input tensor data is preprocessed and feature-extracted using the deployed neural network model, and the obtained intermediate result is passed to the scheduling module as the first input tensor. The first input tensor can be tensor data such as image data and text data. For example, in the scenario of an image recognition task, the first input tensor can be a color image, represented as a three-dimensional tensor, where each dimension of the three-dimensional tensor represents the height, width, and color channels of the image, such as 1000×1000×256. In the scenario of a natural language processing task, the first input tensor can be text data, represented as a two-dimensional or three-dimensional tensor, and each dimension can represent the sequence length, vocabulary size, etc. The intermediate tensor can refer to the small-sized tensor data obtained after the scheduling module slices the input tensor data.
[0053] In a possible implementation, when the size of the first input tensor is relatively large, the scheduling module can split the first input tensor in at least one dimension, which means that the first input tensor can be split in one or more dimensions. For example, taking the first input tensor of the above image data as an example, it can be selected to split in the height dimension, and the first input tensor is split into 10 intermediate tensors of 100×1000×256, or it can be selected to split in both the height and width dimensions, and the first input tensor is split into 10×10 intermediate tensors of 100×100×256; among them, the split dimension can depend on the processing capacity of the split unit, the processing capacity of the computing unit, and the parallelism requirements of the task. Taking the first input tensor as a batch of image data of a deep learning model as an example, the dimension of the first input tensor is a four-dimensional tensor of batch size×height×width×number of channels, and the tensor size of this four-dimensional tensor in each dimension can be specifically 16×224×224×3. First, it can be split in the batch dimension, and 16 image data are divided into multiple small batches. Assuming it is divided into 2 batches, the image data of each batch is 8×224×224×3. Then, these batches of image data are sent as intermediate tensors to the split unit for splitting processing. However, if the width and height dimensions of these image data after splitting exceed the processing capacity of the split unit, the scheduling module can split the first input tensor in multiple dimensions, that is, split in the batch dimension, width dimension, and height dimension. For example, the first input tensor is split into 2×4×4 intermediate tensors of 8×56×56×3, and then sent to the split unit for processing.
[0054] It should be noted that the splitting performed by the scheduling module can refer to splitting the data at the software algorithm level without involving specific physical resource allocation. When splitting the tensor data at the software algorithm level, the goal is how to split according to strategies such as dimension and block size, rather than actually allocating memory or hardware resources. After multiple intermediate tensors are divided, data processing instructions are generated by encapsulating the information corresponding to the intermediate tensors (such as the position information of the intermediate tensors), and the data processing instructions are sent to the split unit. In this process, the scheduling module can not specify which split unit processes the intermediate tensor and does not bind hardware resources. The splitting process of the split unit can refer to splitting the data according to the specific distribution of hardware resources (such as memory address, computing unit layout) and binding to physical hardware resources. The splitting process of the split unit involves actual memory allocation (such as address mapping) and hardware resource binding (such as specifying a computing unit). For example, the split unit maps these data blocks of the base tensor to the local storage of the corresponding computing unit.
[0055] In a possible implementation, the scheduling module can evenly split the first input tensor in each dimension, that is, the tensor sizes of the multiple intermediate tensors after splitting are equal in the same dimension, thereby reducing the number of splitting operations of the scheduling module. However, when evenly splitting the first input tensor, it may occur that the processing capabilities of the splitting units are not fully utilized by all the intermediate tensors after splitting. Therefore, the scheduling module can perform an unequal split of the first input tensor to maximize the utilization rate of the splitting units, that is, the tensor sizes of the multiple intermediate tensors after splitting are not equal in the same dimension.
[0056] In a possible implementation, since the splitting unit can only process one tensor data at a time, and the maximum size of this tensor data is the upper limit of the splitting size that can be processed at one time, where this splitting size upper limit is restricted by the hardware of the splitting unit. Therefore, when the scheduling module obtains a first input tensor with a large size, the scheduling module needs to decompose the first input tensor into smaller parts in one or more dimensions to adapt to the processing capabilities of the splitting units. Since the scheduling module needs to schedule corresponding software and hardware resources for each split, resulting in corresponding power consumption, reducing unnecessary split times can effectively achieve power consumption control and efficiency improvement. That is to say, the first input tensor can refer to a tensor data that exceeds the upper limit of the splitting size that can be processed by the splitting unit at one time, so the scheduling module needs to split the input tensor data and then send it to the splitting unit. In other words, the splitting unit can perform splitting processing on all tensor data with tensor sizes smaller than the upper limit of the splitting size that can be processed at one time, that is, the tensor sizes of the intermediate tensors input to the splitting unit for each split can be the same or different in each dimension, so that the scheduling module can adjust the splitting method of the first input tensor and reduce the number of splits.
[0057] Refer to Figure 3 , Figure 3It is a schematic diagram of scalable intermediate tensor size provided by an embodiment of the present disclosure. Assume that the first input tensor is a feature matrix of text data in a two-dimensional tensor, with dimensions of sequence feature and feature dimension respectively, and the specific tensor size is 256×1024. The upper limit of the split size that the splitting unit can process at one time is 192×1024 respectively. Therefore, the first input tensor exceeds the single-processable capacity of the splitting unit. The scheduling module can split the first input tensor. Specifically, in the first splitting case, the scheduling module can split the two dimensions of the first input tensor respectively to obtain 8 intermediate tensors of 128×256. The tensor sizes of these 8 intermediate tensors are equal in the same dimension, and the size of the obtained intermediate tensors does not exceed the single-processable size of the splitting unit. Therefore, the splitting unit can perform subsequent splitting processing on the intermediate tensors; in the second splitting case, the scheduling module can split the sequence feature dimension of the first input tensor to obtain 2 intermediate tensors, namely an intermediate tensor of 192×1024 and an intermediate tensor of 64×1024. The tensor sizes of these 2 intermediate tensors are not equal in the sequence feature dimension, but the sizes of the two intermediate tensors do not exceed the single-processable size of the splitting unit either. Therefore, the splitting unit can also perform subsequent splitting processing on the intermediate tensors. It can be seen that the size of the intermediate tensors provided by the scheduling module to the splitting unit is scalable, and the size of the intermediate tensors can be adjusted according to the specific task scenario. However, it should be noted that the size of the intermediate tensors is less than or equal to the single-processable size of the splitting unit.
[0058] In a possible implementation manner, the data processing instruction may refer to an instruction generated by the scheduling module according to the splitting result of the first input tensor in the tensor processing device, used to guide the splitting unit and the computing unit on how to process these intermediate tensors. For example, the data processing instruction can be used for the splitting unit to perform data transfer and data splitting, and can also be used to guide the computing unit to perform computing operations and result storage and other steps. Specifically, the data processing instruction may include a data transfer instruction to specify how to transfer the intermediate tensor from the memory to the splitting unit. For example, the data processing instruction may include the storage address of the corresponding intermediate tensor; the data processing instruction may include a computing operation instruction to specify the specific tensor computing operation on the computing unit, such as matrix multiplication, convolution, etc. The data processing instruction may include a result storage instruction to specify the storage location and method of the tensor computing result output by the computing unit. It should be noted that when the data processing instruction is sent to the splitting unit, it can be sent to the computing unit synchronously, or after the splitting unit splits the intermediate tensor, the splitting result and the data processing instruction are sent to the computing unit together.
[0059] Step 202: Parse the data processing instruction through the splitting unit, read the intermediate tensor according to the parsed data processing instruction, split the intermediate tensor in at least one dimension to obtain multiple basic tensors, and map the multiple basic tensors to the computing units.
[0060] In a possible implementation, the splitting unit is used to split the tensor data sent by the scheduling module and split the tensor data to a scale size that the computing unit can directly perform tensor calculation processing, and then output the split tensor data to the computing unit. The basic tensor can refer to the tensor data that can be directly calculated by the computing unit after being split by the splitting unit, and the scale of the basic tensor can be less than or equal to the single processing scale of the computing unit.
[0061] In a possible implementation manner, the splitting unit can be a firmware unit. A processor can be set on the splitting unit. The firmware program running on the processor obtains the corresponding intermediate tensor by parsing the data processing task instruction sent by the scheduling module, then dynamically adjusts the splitting strategy for the intermediate tensor, and allocates the basic tensors obtained by secondary splitting to the corresponding physical storage addresses of the computing units. Specifically, the splitting unit can dynamically adjust the splitting granularity of the intermediate tensor based on the current hardware state (such as the number of computing units, memory bandwidth, etc.). For example, the scheduling module provides M intermediate tensors, the number of intermediate tensors is greater than the maximum number of computing units, and the number that the current computing unit can support for processing is N. Then the splitting unit can dynamically adjust the splitting strategy, split the M intermediate tensors into N groups of basic tensors, with each group containing n basic tensors. Then, the splitting unit can allocate each basic tensor to the corresponding physical storage address of the computing unit according to the storage address of the computing unit, and bind the basic tensor to the corresponding computing unit. It should be noted that the splitting unit can also be a dedicated hardware module. After receiving the data processing instruction from the scheduling module, the splitting unit can parse the data processing instruction, obtain the intermediate tensor according to the data processing instruction, then generate the physical addresses corresponding to each basic tensor according to the pre-burned hardware resource mapping table, then perform physical splitting on the intermediate tensor according to each physical address, and then send the split basic tensors to the target physical address, that is, send them to the corresponding computing units.
[0062] In a possible implementation, the splitting unit is pre-configured firmware or hardware for data splitting. Splitting parameters for data splitting, such as the splitting granularity in each dimension and the upper limit value of the computable size that can be processed by the computing unit in each dimension, can be stored in the splitting unit. Thus, when the splitting unit receives the input tensor data, it can automatically perform the splitting operation according to the pre-configured splitting parameters to generate base tensors that meet the processing requirements of the computing unit, reducing software stack calls and hardware interactions, optimizing the data flow efficiency, reducing system power consumption, and improving the overall computing performance. For example, the splitting unit can be pre-configured with a splitting granularity of 1×1×1. When receiving 8×8×8 tensor data, the splitting unit will automatically split the input tensor into 256 base tensors of 1×1×1 according to the preset splitting parameters; when receiving tensor data of other sizes, such as 4×4×4 tensor data, the splitting unit will automatically split it into 64 base tensors of 1×1×1 according to the preset splitting parameters. That is to say, the splitting unit also performs adaptive splitting based on the pre-configured splitting parameters to ensure that each base tensor directly corresponds to the processing capacity of the computing unit, avoiding multiple software stack calls and hardware interactions.
[0063] In a possible implementation, the scheduling module generates corresponding data processing instructions according to the intermediate tensor and task requirements. The data processing instructions can include the storage location (addressing information) of the corresponding intermediate tensor and specific splitting parameters. After obtaining the data processing instructions, the data processing instructions can be parsed to extract the storage location and splitting parameters of the intermediate tensor. Then, according to the storage location, the corresponding intermediate tensor can be read using the Direct Memory Access (DMA) technology. Next, the splitting unit can split the intermediate tensor according to the splitting parameters in the data instructions to obtain multiple base tensors. Among them, the splitting parameters can refer to the splitting granularity of the intermediate tensor in each dimension or the upper limit value of the computable size that can be processed by the computing unit in each dimension. That is to say, the splitting unit has flexibility and allows the splitting scheme to be dynamically adjusted according to the actual situation in response to the control of the scheduling module during runtime. After the splitting unit splits to obtain multiple base tensors, these splitting results can be stored in a specified memory area, and then the computing unit reads and processes them in sequence. Specifically, the base tensors can be transmitted to the computing unit again through the DMA technology, so as to achieve efficient data flow and computing task allocation, significantly improving the overall system performance. In addition, when the splitting unit performs the splitting operation, it can accurately control the splitting granularity according to the splitting parameters in the data processing instructions to ensure that the size of each base tensor meets the processing requirements of the computing unit, optimizing the computing efficiency and system performance.
[0064] Step 203: Through a computing unit, perform tensor calculations on multiple basic tensors in parallel to obtain a tensor calculation result.
[0065] In a possible implementation manner, the computing unit works in cooperation with the splitting unit. The splitting unit can store the split basic tensors in a memory area directly accessible by the computing unit. Here, this memory area can be a cache shared by the splitting unit and the computing unit. The computing unit can directly read the basic tensors from the shared cache, avoiding frequent data transmission and further reducing latency; or, the splitting unit can directly transmit the basic tensors to the local memory of the computing unit through an internal bus or a direct connection, ensuring the shortest data transmission path. After the computing unit reads these basic tensors, it can perform efficient tensor calculations and finally generate a result tensor.
[0066] In a possible implementation manner, in an application scenario of image recognition, the data input to the tensor processing module can be tensor data representing an image. By calling the scheduling module to perform preliminary splitting on the image tensor data, multiple intermediate image tensors are generated, and then further split into finer-grained image basic tensors by the splitting unit. The computing unit can perform tensor calculations such as matrix multiplication on these image basic tensors and perform accumulation through an accumulation buffer to obtain tensor calculation results such as feature maps and class probabilities; in a natural language processing scenario, the input data can be tensor data representing a word embedding sequence. The computing unit can perform tensor calculations such as matrix multiplication and matrix addition on the split tensor data of the word embedding sequence to obtain tensor calculation results such as word vectors and attention weights.
[0067] Specifically, refer to Figure 4 , Figure 4It is a schematic diagram of the effect of tensor splitting processing provided by an embodiment of the present disclosure. In the related art, a first input tensor is input to a scheduling module, and the scheduling module needs to split the first input tensor into basic tensors that can be directly calculated by a computing unit. For example, if the first input tensor is a three-dimensional tensor of 16×16×16 and the basic tensor is a three-dimensional tensor of 1×1×1, then the first input tensor needs to be split into 4096 three-dimensional tensors. The scheduling module then needs to schedule the corresponding software stacks (such as the application layer, runtime layer, driver layer, etc. in the system) and hardware 4096 times. Each time the corresponding hardware is started for scheduling, there will be corresponding overheads, including initializing the hardware state, configuring registers, loading instructions, etc. The overhead caused by frequently scheduling the hardware startup is large. At the same time, directly splitting the tensor data of a larger size into small-size tensor data that can be directly calculated by the computing unit requires complex algorithms and a large amount of computing resources. Especially when dealing with high-order matrices, the complexity of task decomposition will increase significantly. In addition, for each split tensor data by the scheduling module, corresponding data processing instructions need to be sent to the computing unit, that is, the scheduling module needs to generate 4096 corresponding data processing instructions, which not only increases the operating complexity of the scheduling module but also may cause congestion and delay at the interface of the computing unit. However, in the embodiment of the present disclosure, a new hardware, that is, a splitting unit, is introduced between the scheduling module and the computing unit. The splitting function of the splitting unit is used to assist the scheduling module in splitting the first tensor module. That is to say, the splitting unit can share the splitting work originally undertaken by the scheduling module, as Figure 4 shown, the splitting unit can split the tensor data with a size of 8×8×8 and split it into three-dimensional tensors of 1×1×1. Therefore, the scheduling module can split the first input tensor into 8 intermediate tensors of 8×8×8, and then only generate 8 data processing instructions and send them to the splitting unit. The splitting unit further splits these 8 intermediate tensors of 8×8×8 into 4096 basic tensors of 1×1×1 and then outputs them to the computing unit. It can be seen that during the process of splitting the tensor by the splitting unit, there is no need to call the software stack and hardware required by the scheduling module during splitting. By introducing the splitting unit, the 4096 splitting processes and the generation process of 4096 data processing instructions originally required by the scheduling module are reduced to 8 splitting processes and the generation process of 8 data processing instructions, effectively reducing the number of interactions between software and hardware and reducing power consumption.
[0068] In a possible implementation manner, when the splitting unit splits the intermediate tensor, it can be split according to preset splitting parameters. Specifically, it can refer to Figure 5 , Figure 5 is a schematic flowchart of the specific splitting method of the splitting unit provided by an embodiment of the present disclosure. The tensor processing method further includes but is not limited to the following steps 501 to step 502;
[0069] Step 501: Use the upper limit value of the computable size that can be processed by the computing unit in each dimension as the first segmentation granularity for each dimension;
[0070] Step 502: Perform cyclic segmentation on the intermediate tensor in the corresponding dimension according to the first segmentation granularity of each dimension to obtain multiple basic tensors.
[0071] It can be understood that the computing unit has processing limitations on the tensor data size of each dimension. Therefore, the upper limit value of the computable size that can be processed by the computing unit in each dimension can be used as a segmentation parameter and pre-written into the segmentation unit, and the upper limit value of the computing size of each dimension is used as the corresponding segmentation granularity when the segmentation unit processes the intermediate tensor. The segmentation granularity can refer to the segmentation step length when performing segmentation on one dimension, that is, the size of each dimension slice after segmentation. It should be noted that the first segmentation granularity in each dimension can be the same or different.
[0072] Refer to Figure 6 , Figure 6 FIG. [FIG. ID] is a schematic diagram of the effect of the segmentation unit provided by the embodiment of the present disclosure for segmenting an intermediate tensor. Suppose there is a three-dimensional intermediate tensor with specific dimensions of 100×200×150, and the upper limit value of the computing size of the computing unit in the corresponding dimension is 20×50×30. Therefore, the segmentation unit can simultaneously segment this intermediate tensor in multiple dimensions, and segment it along each dimension into tensor data not exceeding 20×50×30. Specifically, along the first dimension, the intermediate tensor is segmented into 5 parts according to the first segmentation granularity of 20, along the second dimension, the intermediate tensor is segmented into 4 parts according to the first segmentation granularity of 50, and along the third dimension, the intermediate tensor is segmented into 5 parts according to the first segmentation granularity of 30 to obtain 100 basic tensors, and the size of each basic tensor is 20×50×30.
[0073] It should be noted that in the tensor processing device, in order to adapt to different computing requirements and optimization goals, computing units of different scales can be designed to improve the utilization rate of hardware resources and achieve diversified tensor operations. Therefore, the segmentation unit needs to segment the corresponding basic tensors according to the scales of different computing units. That is to say, the sizes of the basic tensors can be different. At this time, the segmentation unit can determine the target computing unit that needs to map the basic tensors from the data processing instructions provided by the scheduling module, and then determine the upper limit values of the computing sizes of the target computing unit in each dimension from the pre-configured segmentation parameters, and then use these upper limit values of the computing sizes as the first segmentation granularity to segment the intermediate tensor; or, the data processing instructions provided by the scheduling module may contain the upper limit values of the computing sizes of the target computing unit in each dimension, so as to guide the segmentation unit to use these upper limit values of the computing sizes as the first segmentation granularity to segment the intermediate tensor.
[0074] In a possible implementation manner, when the slicing unit slices the intermediate tensor in multiple dimensions, it is necessary to ensure that the requirements for the slicing granularity in each dimension are met. Specifically, reference can be made to Figure 7 , Figure 7 FIG. 5 is a schematic flowchart of a specific slicing method of the slicing unit provided in another embodiment of the present disclosure. The tensor processing method further includes, but is not limited to, the following steps 701 to 702;
[0075] Step 701: For each dimension, compare the intermediate tensor size of the intermediate tensor with the first slicing granularity, and determine the dimension corresponding to the comparison result that the intermediate tensor size is greater than the first slicing granularity as the dimension to be sliced;
[0076] Step 702: When there are multiple dimensions to be sliced, traverse each dimension to be sliced, and cyclically slice the tensor obtained by slicing the previous traversed dimension to be sliced according to the first slicing granularity corresponding to the current dimension, to obtain multiple basic tensors.
[0077] It can be understood that since there may be dimensions that do not need to be sliced when slicing the intermediate tensor, therefore, the intermediate tensor size of the intermediate tensor can be compared with the first slicing granularity, and the dimension with the intermediate tensor size greater than the first slicing granularity can be determined as the dimension to be sliced, so as to avoid unnecessary slicing operations and improve processing efficiency. When there are multiple dimensions to be sliced, it is usually necessary to perform slicing operations one by one, and use the result after slicing each dimension as the input for the next slicing. By traversing each dimension to be sliced, it can be ensured that the intermediate tensor of each dimension to be sliced is sliced to meet the requirements of the slicing granularity. Among them, when the dimension to be sliced is the dimension traversed for the first time, the intermediate tensor can be directly sliced according to the first slicing granularity of the current dimension to obtain multiple basic tensors, and then for the subsequent traversed dimensions to be sliced, the basic tensors obtained after slicing the previous traversed dimension to be sliced are sliced according to the first slicing granularity of the current dimension to obtain multiple basic tensors.
[0078] Specifically, assume that the intermediate tensor sizes of an intermediate tensor in dimension 0, dimension 1, and dimension 2 are 8×6×4 respectively, and the first splitting granularity in the corresponding dimensions is 4×4×4. First, compare the intermediate tensor sizes in each dimension with the first splitting granularity to determine the dimensions to be split, that is, dimension 0 and dimension 1 are the dimensions to be split. Traverse in the order of arranging the numerical values of the intermediate tensor sizes in the dimensions to be split from large to small. First, traverse dimension 0, and split the intermediate tensor according to the first splitting granularity of 4 to obtain 2 basic tensors. The tensor size of each tensor in each dimension is 4×6×4; then, traverse to dimension 1, and split the 2 basic tensors output by dimension 0 in the previous traversal according to the first splitting granularity of 4 again to obtain 2 parts, namely a basic tensor with a tensor size of 4×4×4 and a basic tensor with a tensor size of 4×2×4, that is, 4 basic tensors are obtained, 2 basic tensors with a tensor size of 4×4×4 and 2 basic tensors with a tensor size of 4×2×4.
[0079] Since the order of dimension splitting does not affect the set of basic tensors, but has an impact on the intermediate results produced during the splitting process, therefore, according to the task requirements and the processing capabilities of the splitting units, the traversal order of the dimensions to be split can be determined. For example, the numerical values of the intermediate tensor sizes in the dimensions to be split can be arranged from large to small or from small to large, or the numerical values of the first splitting granularity in the dimensions to be split can be arranged from large to small or from small to large, or the number of splitting times required for each dimension to be split can be calculated through the intermediate tensor sizes and the first splitting granularity in the dimensions to be split, and then arranged from large to small or from small to large according to the number of splitting times of the dimensions to be split.
[0080] In a possible implementation manner, when the scheduling module splits the first input tensor in multiple dimensions, the actual processing capabilities of the splitting units need to be considered. Specifically, reference can be made to Figure 8 , Figure 8 is a schematic flowchart of the specific splitting method of the scheduling module provided by an embodiment of the present disclosure. The tensor processing method further includes but is not limited to the following steps 801 to step 802;
[0081] Step 801: Take the upper limit value of the splitting size that can be processed once by the splitting unit in each dimension as the second splitting granularity in each dimension;
[0082] Step 802: Perform cyclic splitting on the first input tensor in the corresponding dimension according to the second splitting granularity in each dimension to obtain multiple intermediate tensors.
[0083] In a possible implementation, when processing tensors, the slicing scheme is crucial. Especially when dealing with large-sized tensors, since the slicing unit can only process one tensor at a time, the maximum size of this tensor is the upper limit of the slicing size that can be processed at a single time. To ensure that each sliced intermediate tensor block can be efficiently processed by the processing unit and avoid being unable to be processed due to hardware resource limitations, therefore, taking the upper limit of the slicing size of each dimension as the second slicing granularity of the corresponding dimension can make full use of the processing capacity of the slicing unit, that is, by maximizing the utilization rate of the slicing unit to reduce the number of slicings of the scheduling module and reduce power consumption. The specific implementation method can be adjusted according to different hardware environments. Next, for each dimension, the first input tensor is sliced according to the corresponding second slicing granularity, and then sliced cyclically in each dimension until the entire first input tensor is sliced into multiple small blocks to obtain multiple intermediate tensors. Suppose there is a first input tensor T with dimensions D1 x D2 x D3, and the upper limits of the slicing sizes of the slicing unit in each dimension are G1, G2, and G3 respectively, then the second slicing granularities also correspond to G1, G2, and G3 respectively. Slicing in the D1 dimension, slicing once every G1, a number of sub-tensors are obtained, and the size of each sub-tensor in the D1 dimension is G1. Among them, if the size in the D1 dimension cannot be divisible by the slicing granularity G1, the size of the sub-tensor in the D1 dimension can be less than G1. Then, slicing in the D2 dimension of each sub-tensor, slicing once every G2, a number of new sub-tensors are obtained, and then slicing in the D3 dimension of each new sub-tensor, slicing once every G3, thus obtaining multiple intermediate tensors. At this time, when the dimension sizes of the first input tensor are different, the slicing process also changes accordingly. If the tensors in each dimension are divisible by the second slicing granularity, then the sizes of the obtained base tensors are the same, but the quantities obtained are different. It should be noted that the second slicing granularity can be equal or unequal in each dimension, which is specifically determined based on the upper limit of the slicing size of the slicing unit in each dimension, and the upper limit of the slicing size of the slicing unit is determined based on the hardware.
[0084] In a possible implementation, when slicing the first input tensor, the input tensor size of the first input tensor can be first compared with the second slicing granularity, and the dimension with the input tensor size greater than the second slicing granularity is determined as the target slicing dimension. When there are multiple target slicing dimensions, generally, the slicing operations need to be performed one by one, and the result after slicing each dimension is used as the input for the next slicing. Among them, when first traversing a target slicing dimension, the first input tensor can be directly sliced according to the second slicing granularity of the current dimension to obtain multiple sub-tensors, and then for the target slicing dimensions traversed subsequently, the sub-tensors obtained after slicing the target slicing dimension traversed last time are sliced according to the second slicing granularity of the current dimension to obtain multiple intermediate tensors.
[0085] It should be noted that when slicing the first input tensor, the slicing order may affect the memory access pattern and cache utilization. Adjust the slicing order of the first input tensor according to the application scenario and hardware characteristics. Specifically, each dimension can be sorted according to the size of the input tensor in each dimension, the second slicing granularity, and the number of slicing times.
[0086] Therefore, using the upper limit value of the slicing size that can be processed in a single time by the slicing unit in each dimension as the slicing granularity in each dimension to slice the first input tensor is a slicing method with the maximum utilization rate of the slicing unit. It slices as much as possible using the slicing unit, sharing part of the slicing work originally borne by the scheduling module, so that the scheduling module only needs to perform fewer slicing times, preliminarily slicing the first input tensor with too large a size into relatively small intermediate tensors, rather than directly slicing the first input tensor into base tensors, reducing the slicing times of the scheduling module, reducing the number of generated data processing instructions, and thus effectively reducing the interaction times between software and hardware in the tensor processing device, improving power consumption control, and increasing the utilization rate of the tensor processing device.
[0087] In a possible implementation manner, since the upper limit value of the slicing size of the slicing unit in each dimension is fixed, the scheduling module can be pre-configured with multiple slicing schemes that match the processing capabilities of the slicing unit. Thus, when the scheduling module receives the first input tensor, it can directly perform slicing processing by calling the corresponding slicing scheme, improving the slicing efficiency. Specifically, it can refer to Figure 9 , Figure 9 which is a schematic flowchart of the specific slicing method of the scheduling module provided by another embodiment of the present disclosure. The tensor processing method further includes but is not limited to the following steps 901 to step 904;
[0088] Step 901: Obtain multiple pre-configured first candidate slicing schemes and determine the input tensor sizes of the first input tensor in each dimension;
[0089] Step 902: Determine a first target slicing scheme from the multiple first candidate slicing schemes according to the relationship between the input tensor sizes in each dimension and the candidate slicing granularities in each dimension of each first candidate slicing scheme;
[0090] Step 903: Determine the candidate slicing granularities in each dimension of the first target slicing scheme as the third slicing granularities in the corresponding dimensions;
[0091] Step 904: Perform cyclic slicing on the first input tensor in the corresponding dimensions according to the third slicing granularities in each dimension to obtain multiple intermediate tensors.
[0092] In a possible implementation, the pre-set first candidate segmentation scheme may include candidate segmentation granularities in each dimension. The candidate segmentation granularities in each first candidate segmentation scheme match the upper limit values of the segmentation sizes of the segmentation unit in each dimension. As a result, when processing tensors, computing resources can be utilized more efficiently, computing performance can be optimized, and the flexibility and adaptability of the system can be ensured. By pre-setting multiple first candidate segmentation schemes, a suitable segmentation method can be selected according to the actual computing task and the size of the first input tensor, unnecessary computing overhead can be reduced, the overall computing efficiency can be improved, and some first candidate segmentation schemes can be designed with specific candidate segmentation granularities for specific task computing requirements to ensure the stability of the intermediate tensors obtained by segmentation.
[0093] In a possible implementation, each first candidate segmentation scheme can be determined based on the upper limit values of the segmentation sizes of the segmentation unit in each dimension. For example, a proportionality coefficient can be obtained and multiplied by the upper limit values of the segmentation sizes in each dimension to obtain the candidate segmentation granularities of the first candidate segmentation scheme in each dimension. Specifically, if the upper limit value of the segmentation size of the segmentation unit in each dimension is 1920 and a proportionality coefficient of 0.75 is obtained, multiplying the proportionality coefficient 0.75 by the segmentation size upper limit value 1920 gives the candidate segmentation granularities of the first candidate segmentation scheme in each dimension as 1440.
[0094] In a possible implementation, the first candidate segmentation scheme can be predefined based on conventional tensor sizes and computing requirements. Therefore, each first candidate segmentation scheme has wide applicability. For example, one first candidate segmentation scheme can be applicable to first input tensors of multiple sizes. At the same time, there are also multiple first candidate segmentation schemes suitable for a first input tensor of the same size. For example, for a three-dimensional first input tensor [8, 6, 4], multiple first candidate segmentation schemes can be adapted. Different first candidate segmentation schemes can segment the first input tensor along different dimensions. Therefore, according to the relationship between the sizes of the input tensors in each dimension and the candidate segmentation granularities of each first candidate segmentation scheme in each dimension, the first target segmentation scheme can be determined from multiple first candidate segmentation schemes. It is expected that using the first target segmentation scheme to segment the first input tensor can minimize the number of segmentations.
[0095] Among them, the relationship between the input tensor size in each dimension and the candidate segmentation granularity of each first candidate segmentation scheme in each dimension can refer to the size relationship between the input tensor size and the candidate segmentation granularity in the same dimension, or it can refer to the ratio relationship between the input tensor size and the candidate segmentation granularity in the same dimension. Taking the size relationship between the two as an example, when the input tensor size in the same dimension is greater than the candidate segmentation granularity, it means that the first input tensor needs to be segmented at least once in this dimension, while when the input tensor size in the same dimension is less than or equal to the candidate segmentation granularity, it can be indicated that there is no need to segment the first input tensor in this dimension. Therefore, by judging the size relationship between the two, the first candidate segmentation scheme with the least number of segmentation times can be selected as the first target segmentation scheme, reducing the number of interactions between software and hardware. Taking the ratio relationship between the two as an example, the ratio between the two can be used to determine whether the first input tensor needs to be segmented in this dimension and the number of times the first input tensor is segmented. For example, if the ratio of the input tensor size to the candidate segmentation granularity in a dimension is 0.8, it means that the input tensor size in this dimension is less than the candidate segmentation granularity and there is no need to segment the first input tensor, while if the ratio of the input tensor size to the candidate segmentation granularity is 3.2, it means that the first input tensor needs to be segmented 4 times in this dimension to obtain four tensors. Similarly, by judging the ratio relationship between the two, the first candidate segmentation scheme with the least number of segmentation times can be selected as the first target segmentation scheme, reducing the number of interactions between software and hardware.
[0096] Refer to Figure 10 , Figure 10It is a schematic diagram of the effect of splitting the first input tensor provided by another embodiment of the present disclosure. Taking the above three-dimensional first input tensor [8, 6, 4] as an example, it is assumed that multiple first candidate splitting schemes are pre-configured in the scheduling module, specifically as follows: First candidate splitting scheme A: The candidate splitting granularity for each dimension is [4, 6, 4]; First candidate splitting scheme B: The candidate splitting granularity for each dimension is [8, 2, 4]; First candidate splitting scheme C: The candidate splitting granularity for each dimension is [8, 6, 3]; First candidate splitting scheme D: The candidate splitting granularity for each dimension is [4, 3, 4]; For each first candidate splitting scheme, by comparing the relationship between the input tensor size of each dimension and the candidate splitting granularity, it can be analyzed that in the first candidate splitting scheme A, the ratio relationship between the input tensor size of dimension 0 and the candidate splitting granularity is 2, and the ratio relationships of the remaining dimensions are 1, that is, the first input tensor is split 2 times along dimension 0, and the remaining dimensions are not split. After splitting, 2 tensors [4, 6, 4] are obtained; In the first candidate splitting scheme B, the ratio relationship between the input tensor size of dimension 1 and the candidate splitting granularity is 3, and the ratio relationships of the remaining dimensions are 1, that is, the first input tensor is split 3 times along dimension 1, and the remaining dimensions are not split. After splitting, 3 tensors [8, 2, 4] are obtained; In the first candidate splitting scheme C, the ratio relationship between the input tensor size of dimension 2 and the candidate splitting granularity is 1.3, and the ratio relationships of the remaining dimensions are 1, that is, the first input tensor is split 2 times along dimension 2, and the remaining dimensions are not split. Among them, the tensor sizes of the intermediate tensors obtained by splitting the first input tensor according to the first candidate splitting scheme C are not equal in dimension 2. After splitting, 1 tensor [8, 6, 3] and 1 tensor [8, 6, 1] are obtained; In the first candidate splitting scheme D, the ratio relationships between the input tensor sizes of dimension 0 and dimension 1 and the candidate splitting granularity are both 2, and the ratio relationship of dimension 2 is 1, that is, the first input tensor is split 2 times along dimension 0 and dimension 1 respectively, and dimension 2 is not split. After splitting, 4 tensors [4, 3, 4] are obtained; Therefore, the first candidate splitting scheme A and the first candidate splitting scheme C have the fewest splitting times. The first candidate splitting scheme A or the first candidate splitting scheme C can be used as the first target splitting scheme. Assuming that the first candidate splitting scheme A is used as the first target splitting scheme, then the third splitting granularity of each dimension is [4, 6, 4]. After splitting the first input tensor [8, 6, 4] according to the third splitting granularity [4, 6, 4], 2 intermediate tensors [4, 6, 4] can be obtained.
[0097] In a possible implementation, the choice of the segmentation granularity may affect the final computing efficiency and memory usage. If the segmentation granularity is too small, it may increase the number of segmentations, resulting in additional overhead; if it is too large, it may exceed the processing capacity of the segmentation unit. Specifically, taking the above three-dimensional first input tensor [8, 6, 4] as an example, at this time, the upper limit of the segmentation size of the segmentation unit in each dimension is [6, 6, 6]. By comparing the above four first candidate segmentation schemes A to D, it can be seen that the intermediate tensor obtained after segmentation according to the first candidate segmentation scheme B is [8, 2, 4], and the intermediate tensors obtained after segmentation according to the first candidate segmentation scheme C are [8, 6, 3] and [8, 6, 1]. The intermediate tensors obtained by segmenting the two first candidate segmentation schemes exceed the processing capacity of the segmentation unit. Therefore, the first candidate segmentation schemes B and C are no longer considered; although the intermediate tensor [4, 3, 4] obtained by segmenting according to the first candidate segmentation scheme D does not exceed the processing capacity of the segmentation unit, due to the too small candidate segmentation granularity and the large number of segmentations, and the intermediate tensor [4, 6, 4] obtained by segmenting according to the first candidate segmentation scheme A also does not exceed the processing capacity of the segmentation unit, and the number of segmentations of the first candidate segmentation scheme A is less than that of the first candidate segmentation scheme D, so, in this case, the first candidate segmentation scheme A can be used as the first target segmentation scheme.
[0098] In a possible implementation, when selecting the first target segmentation scheme from multiple first candidate segmentation schemes, it is also possible to pay attention to whether the ratio between the input tensor size of each dimension and the candidate segmentation granularity is an integer. If the ratio between the input tensor size and the candidate segmentation granularity in any dimension is an integer, it means that the first input tensor can be evenly divided in this dimension according to the candidate segmentation granularity. For example, in the first candidate segmentation scheme A proposed in the above embodiment, the ratio relationship between the input tensor size of dimension 0 and the candidate segmentation granularity is 2, and the ratio relationships of the remaining dimensions are 1, that is, the first input tensor is segmented 2 times along dimension 0, and the remaining dimensions are not segmented. After segmentation, 2 tensors with equal sizes [4, 6, 4] are obtained; if the ratio between the input tensor size and the candidate segmentation granularity in any dimension is a non-integer, it means that the first input tensor cannot be evenly segmented in this dimension according to the candidate segmentation granularity. For example, in the first candidate segmentation scheme C proposed in the above embodiment, the ratio relationship between the input tensor size of dimension 2 and the candidate segmentation granularity is 1.3, and the ratio relationships of the remaining dimensions are 1, that is, the first input tensor is segmented 2 times along dimension 2, and the remaining dimensions are not segmented. Among them, the tensor sizes of the intermediate tensors obtained by segmenting the first input tensor according to the first candidate segmentation scheme C are not equal in dimension 2. Generally speaking, after the non-uniform intermediate tensor is input to the segmentation unit for segmentation, it is easy to have a situation where the tensor sizes of the generated multiple basic tensors are not equal. When facing some computing units that only support tensors of a fixed scale for processing, additional remainder processing is required, such as data splicing or data filling for the remainder part. Therefore, in order to reduce the occurrence of non-uniform basic tensors, when selecting the first target segmentation scheme from multiple first candidate segmentation schemes, the first candidate segmentation scheme with an integer ratio between the input tensor size of each dimension and the candidate segmentation granularity can be preferentially selected as the first target segmentation scheme.
[0099] In a possible implementation, the scheduling module can call a pre-trained neural network model to evaluate each first candidate segmentation scheme. It can respectively extract features from the input tensor sizes of the first input tensor in each dimension, the upper limit values of the segmentation sizes of the segmentation unit in each dimension, and the candidate segmentation granularities of the first candidate segmentation scheme in each dimension, to obtain corresponding tensor size features, segmentation size features, and segmentation granularity features. Then, the tensor size features, segmentation size features, and segmentation granularity features are concatenated, and the concatenated features are input into the pre-trained neural network model for regression prediction to output the execution time and / or segmentation unit utilization rate of the first candidate segmentation scheme. Then, a comparison is made of the execution time and / or segmentation unit utilization rate of each first candidate segmentation scheme, and the first candidate segmentation scheme with the shortest execution time or the highest segmentation unit utilization rate is selected as the first target segmentation scheme. During the training process of the pre-trained neural network model, a sample training set and a sample validation set are obtained. After the features of the sample training set are extracted and concatenated, the concatenated features are input into the neural network model for regression prediction to obtain the output sample metrics (execution time and / or segmentation unit utilization rate). Then, the mean squared error loss function is used to calculate the difference between the sample metrics and the actual metrics of the sample validation set to obtain a loss value. Through optimization algorithms such as gradient descent, the parameters of the neural network model are updated according to the gradient information of the mean squared error loss function to make the loss value as small as possible.
[0100] Therefore, by pre-configuring multiple first candidate segmentation schemes, according to the relationship between the size and segmentation granularity of the first input tensor, and dynamically selecting a suitable first target segmentation scheme according to different criteria (such as the number of segmentations, segmentation unit utilization rate, execution time, whether there is a remainder part, etc.), it is possible to adapt to different hardware environments and computing requirements, and improve computing efficiency and resource utilization.
[0101] In a possible implementation, when the scheduling module segments the first input tensor, it can dynamically adjust the segmentation granularity of each dimension according to the first input tensors of different sizes to reduce the number of segmentations of the scheduling module. Specifically, it can refer to Figure 11 , Figure 11 FIG. 1101 is a schematic flowchart of a specific segmentation method of the scheduling module provided by another embodiment of the present disclosure. The tensor processing method further includes, but is not limited to, the following steps 1101 to step 1103;
[0102] Step 1101: Determine the input tensor sizes of the first input tensor in each dimension, and the upper limit values of the segmentation sizes that can be processed in one time by the segmentation unit in each dimension;
[0103] Step 1102: Determine the fourth segmentation granularity of each dimension according to the ratio relationship between each input tensor size and each size upper limit value;
[0104] Step 1103: Perform cyclic slicing on the first input tensor in the corresponding dimension according to the fourth slicing granularity of each dimension, to obtain multiple intermediate tensors.
[0105] In a possible implementation, the slicing granularity for slicing the first input tensor is determined by the ratio of the input tensor size of the first input tensor to the upper limit value of the size of the slicing unit. This means that the slicing granularity of each dimension can be adaptively and dynamically adjusted, ensuring that each sliced intermediate tensor can be processed by the slicing unit, avoiding exceeding the processing capacity. At the same time, the size of the output intermediate tensor will also change with the size of the first input tensor, so as to be able to slice the first input tensor as evenly as possible, making the intermediate tensor as close as possible to the upper limit value of the processing capacity of the slicing unit, fully utilizing the hardware resources, improving the utilization rate, reducing the scheduling times of the software, and implementing a flexible slicing method.
[0106] In a possible implementation, when the scheduling module slices the first input tensor and dynamically selects the fourth slicing granularity of each dimension, the goal is to maximize the fourth slicing granularity and minimize the number of slicing times. Among them, the fourth slicing granularity needs to be less than or equal to the upper limit value corresponding to this dimension, and the fourth slicing granularity needs to be able to divide the input tensor size without a remainder. Therefore, for each dimension, select the largest possible fourth slicing granularity to reduce the number of times the scheduling module slices the first input tensor.
[0107] In a possible implementation, for each dimension, the ratio relationship between the input tensor size and the upper limit value can represent the initial number of slicing times of the first input tensor in this dimension. Thus, using this ratio relationship, the maximum slicing granularity in this dimension can be mined. This maximum slicing granularity must be less than or equal to the upper limit value of this dimension, and then the fourth slicing granularity is determined. Among them, the fourth slicing granularity can divide the input tensor size of this dimension. Specifically, for each dimension, judge whether the ratio between the input tensor size and the upper limit value is an integer. If this ratio is an integer, the upper limit value of this dimension can be used as the fourth slicing granularity of this dimension; if this ratio is a non-integer, then round up this ratio to obtain the initial number of slicing times, and then start traversing upward from the initial number of slicing times to check whether the value of each number of slicing times can divide the input tensor size. If this number of slicing times can divide the input tensor size, then use the ratio of the input tensor size to this number of slicing times as the fourth slicing granularity; if this number of slicing times cannot divide the input tensor size, then continue to increase the number of slicing times until a number of slicing times that can divide the input tensor size is determined, and then use the ratio of the input tensor size to this number of slicing times as the fourth slicing granularity.
[0108] Specifically, assume that in one dimension, the input tensor size of the first input tensor is 10, and the upper limit value of the size of the splitting unit is 4. According to the ratio between the input tensor size and the upper limit value of the size, which is 2.5, rounding up the ratio 2.5 gives the initial splitting times as 3. Then, starting from the initial splitting times 3 and traversing upward, check whether the splitting times 3 can divide the input tensor size 10. If not, accumulate the splitting times; check whether the accumulated splitting times 4 can divide the input tensor size 10. If still not, continue to accumulate the splitting times; check whether the accumulated splitting times 5 can divide the input tensor size 10. If so, take the ratio of the input tensor size 10 to the splitting times 5, which is 2, as the fourth splitting granularity. That is, for the dimension with the input tensor size 10 of the first input tensor and the upper limit value 4 of the size of the splitting unit, the fourth splitting granularity is 2. Therefore, in this dimension, perform cyclic splitting on the first input tensor according to the fourth splitting granularity 2, and obtain 5 intermediate tensors.
[0109] Specifically, assume that in one dimension, the input tensor size of the first input tensor is 16, and the upper limit value of the size of the splitting unit is 4. According to the ratio between the input tensor size and the upper limit value of the size, which is 4, and the upper limit value of the size can divide the input tensor size. Therefore, the upper limit value of the size of this dimension can be used as the fourth splitting granularity of this dimension. That is, in this dimension, perform cyclic splitting on the first input tensor according to the fourth splitting granularity 4, and obtain 4 intermediate tensors.
[0110] In a possible implementation manner, when determining the splitting granularity of each dimension, it can be considered from the direction of the splitting times of the first input tensor. Specifically, refer to Figure 12 , Figure 12 which is a schematic flowchart of the process for determining the splitting granularity provided by an embodiment of the present disclosure. The tensor processing method further includes but is not limited to the following steps 1201 to step 1204;
[0111] Step 1201: Traverse each dimension, and determine the minimum splitting times of the first input tensor in each dimension by the scheduling module according to the ratio relationship between the input tensor size of each dimension and the upper limit value of the splitting size of the current dimension;
[0112] Step 1202: Determine the second candidate splitting schemes of the first input tensor in each dimension, and count the total splitting times of each second candidate splitting scheme according to each minimum splitting times;
[0113] Step 1203: Determine the second target splitting scheme as the second candidate splitting scheme with the smallest numerical value of the total splitting times;
[0114] Step 1204: Determine the fourth splitting granularity of each dimension according to the second target splitting scheme.
[0115] Among them, the second candidate segmentation scheme may be used to indicate that the scheduling module segments the first input tensor in each dimension. Specifically, the second candidate segmentation scheme may include the number of segmentations of the first input tensor in each dimension and / or the candidate segmentation granularity.
[0116] In a possible implementation manner, for each dimension, calculate the ratio of the size of the input tensor to the upper limit value of the segmentation size to obtain the minimum number of segmentations. That is, use the upper limit value of the segmentation size as the segmentation granularity. Among them, if the ratio is an integer, the ratio can be used as the minimum number of segmentations. Otherwise, round up the ratio, and use the rounded-up value as the initial number of segmentations. Then traverse upward from the initial number of segmentations, and check whether the value of each number of segmentations can divide the size of the input tensor. If the number of segmentations can divide the size of the input tensor, use this number of segmentations as the minimum number of segmentations; if the number of segmentations cannot divide the size of the input tensor, continue to increase the number of segmentations until a number of segmentations that can divide the size of the input tensor is determined, and then use this number of segmentations as the minimum number of segmentations.
[0117] It should be noted that the shape of the first input tensor can be changed to make the splitting in a specific dimension more efficient. For example, the first input tensor is a three-dimensional tensor of 256×512×16, and the single-processable scale of the splitting unit is a three-dimensional tensor of 512×256×16. If the first input tensor is split in each dimension, it is necessary to perform one split on the first input tensor in dimension 1, that is, split it into two three-dimensional tensors of 256×256×16 to meet the single-processable scale of the splitting unit. However, by changing the dimension order of the first input tensor, such as by transposing the first input tensor and swapping dimension 1 and dimension 0, a three-dimensional tensor of 512×256×16 is obtained, so that the single-processable scale of the splitting unit can be satisfied without splitting the first input tensor. It should be noted that changing the dimension order of the tensor will affect the semantics of the tensor data. For example, in image processing, rotating the image data will change the direction of the image. Therefore, whether the first input tensor can change the dimension order needs to be determined according to the specific application scenario and task requirements. Therefore, by calculating the ratio relationship between the input tensor size of each dimension and the upper limit value of the splitting size of the current dimension, the minimum splitting times of the first input tensor in each dimension by the scheduling module in the current dimension are determined, that is, the minimum splitting times of the first input tensor with different shapes in each dimension after dimension transformation are determined. Specifically, taking the dimension of the first input tensor as A×B×C as an example, the minimum splitting times when the input tensor size in dimension A faces the upper limit value of the splitting size in dimension A is a1, that is, the minimum splitting times when the first input tensor in dimension A is restricted by the upper limit value of the splitting size in dimension A is a1, and the minimum splitting times when the input tensor size in dimension B faces the upper limit value of the splitting size in dimension A is b1, that is, the minimum splitting times when the first input tensor in dimension B is restricted by the upper limit value of the splitting size in dimension A is b1. That is to say, by swapping dimension A and dimension B of the first input tensor, the dimension of the first input tensor at this time is B×A×C; in addition, the minimum splitting times when the input tensor size in dimension C faces the upper limit value of the splitting size in dimension A is c1, that is, the minimum splitting times when the first input tensor in dimension C is restricted by the upper limit value of the splitting size in dimension A is c1. That is to say, by swapping dimension A and dimension C of the first input tensor, the dimension of the first input tensor at this time is C×B×A.
[0118] In a possible implementation manner, when allowing dimensional transformation of the first input tensor, by traversing each dimension and selecting different splitting times corresponding to different dimensions for permutation and combination, a plurality of different second candidate splitting schemes are formed. For example, taking the first input tensor as a two-dimensional tensor, the dimension of the first input tensor is A×B, and the corresponding input tensor data is A1×B1. When splitting the first input tensor with respect to dimension A, it can be considered to perform splitting with the upper limit value of the splitting size of dimension A, which is equivalent to splitting the first input tensor of A1×B1, or the upper limit value of the splitting size of dimension B can be considered, which is equivalent to splitting the first input tensor of B1×A1. At this time, by selecting different splitting times corresponding to different dimensions for permutation and combination, 2 different second candidate splitting schemes can be formed; for another example, taking the first input tensor as a three-dimensional tensor, the dimension of the first input tensor is C×D×E, and the corresponding input tensor data is C1×D1×E1. By performing dimensional transformation on the first input tensor, the first input tensor of C1×D1×E1, or the first input tensor of D1×C1×E1, or the first input tensor of C1×E1×D1, or the first input tensor of D1×E1×C1, or the first input tensor of E1×D1×C1, or the first input tensor of E1×C1×D1 can be obtained. By selecting different splitting times corresponding to different dimensions for permutation and combination, 6 different second candidate splitting schemes can be formed.
[0119] After obtaining the second candidate splitting scheme, for each second candidate splitting scheme, the minimum splitting times on each dimension can be accumulated to obtain the total splitting times, and then each second candidate splitting scheme is sorted according to the value of the total splitting times, and the second candidate splitting scheme with the smallest value of the total splitting times is determined as the second target splitting scheme. Through the minimum splitting times on each dimension and the input tensor size, the splitting granularity on each dimension can be calculated, and further, through the second target splitting scheme, the fourth splitting granularity of each dimension can be determined.
[0120] In a possible implementation manner, after the scheduling module splits the first input tensor, a plurality of intermediate tensors are generated, and information such as the storage locations of these intermediate tensors and the sizes of the intermediate tensors on each dimension is encoded into corresponding data processing instructions. Specifically, reference can be made to Figure 13 , reference can be made to Figure 13 , Figure 13 is a schematic flowchart of issuing data processing instructions provided by an embodiment of the present disclosure. The tensor processing method further includes, but is not limited to, the following steps 1301 to step 1302;
[0121] Step 1301: Encoding and processing each intermediate tensor according to the first encoding format to generate corresponding data processing instructions, and adding each data processing instruction to the instruction queue;
[0122] Step 1302: Sequentially issue each data processing instruction to the splitting unit according to the instruction sequence in the instruction queue.
[0123] Among them, the first encoding format refers to converting data information into an instruction format that can be understood and executed by computer hardware. Specifically, encoding the intermediate tensor according to the first encoding format to generate a data processing instruction means compiling the information describing the tensor data and guiding the splitting unit to perform processing (such as splitting granularity, allocation information of the basic tensor, etc.) into an instruction readable by computer hardware. The data processing instructions will be stored in the memory for the hardware to read and execute. For example, the instruction queue can be a memory area for storing instructions to be executed. The instruction queue can be in a first-in, first-out form. The splitting unit will sequentially read and execute according to the order in which the data processing instructions enter the instruction queue. By writing the data processing instructions into the instruction queue, it can ensure that the splitting unit can continuously execute tasks, that is, split multiple intermediate tensors; when the splitting unit completes an instruction task, that is, after splitting an intermediate tensor into multiple basic tensors and mapping these basic tensors to the computing unit, the splitting unit can take out the next data processing instruction from the instruction queue and execute it. Among them, the scheduling module can configure the data processing instructions in the instruction queue to the splitting unit so that the splitting unit executes sequentially. Among them, the data processing instructions include an encoding label for indicating the first encoding format to facilitate the corresponding computer hardware to decode the data processing instructions.
[0124] In a possible implementation manner, after the splitting unit receives the data processing instruction, since the data processing instruction is generated by encoding, it is necessary to decode the data processing instruction and then perform corresponding operations according to the data processing instruction. Among them, the tensor processing device further includes a decoding unit, and the decoding unit can be located in the tensor core. The splitting unit can call the decoding unit to decode the data processing instruction. Specifically, through the splitting unit, the intermediate tensor can be read according to the data processing instruction, and the decoding unit can be called to decode the intermediate tensor according to the encoding label to obtain the decoded intermediate tensor.
[0125] Among them, the data processing instruction may include the address information of the storage location of the intermediate tensor. The splitting unit can, according to this address information, read the encoded data of the intermediate tensor from the memory, and then call the decoding unit to decode according to the encoding label, that is, the first encoding format, so that the splitting unit obtains the decoded data of the intermediate tensor, and then can perform splitting processing on the decoded intermediate tensor.
[0126] It should be noted that the splitting unit can read the encoded data of the intermediate tensor from the memory through the direct memory access (DMA) technology. Specifically, when the scheduling module generates and issues data processing instructions to the splitting unit, it can initialize the DMA controller, including specifying the source address (such as the specific location in the memory) of the intermediate tensor transmission, the target address (such as the splitting unit), the amount of data to be transmitted, and other control information. When the splitting unit receives the data processing instruction, it can send a request to the DMA controller to obtain the corresponding intermediate tensor according to the data processing instruction. The DMA controller responds to the request, reads the corresponding data (intermediate tensor) from the source address, and transmits this data to the target address (i.e., the splitting unit).
[0127] It is worth noting that the data processing instruction can also include calculation processing information for guiding the calculation unit to perform tensor calculations. Therefore, the splitting unit can also issue the data processing instruction to the calculation unit while mapping the base tensor to the calculation unit.
[0128] In a possible implementation, the scheduling unit can first determine the tensor sizes in each dimension of the input tensor, and respectively compare the upper limit of the splitting size that can be processed at one time by the splitting unit in the corresponding dimension, and the upper limit of the calculation size that can be processed at one time by the calculation unit in the corresponding dimension, so as to determine whether the input tensor needs to be split;
[0129] When the tensor size in each dimension of the input tensor is greater than the upper limit of the splitting size that can be processed at one time by the splitting unit in the corresponding dimension, then the input tensor can be determined as the first input tensor, and then the scheduling unit performs splitting processing on the first input tensor. That is to say, the first input tensor can refer to the tensor data whose tensor size in each dimension is greater than the upper limit of the splitting size that can be processed at one time by the splitting unit in the corresponding dimension;
[0130] When the tensor size in each dimension of the input tensor is less than or equal to the upper limit of the splitting size that can be processed at one time by the splitting unit in the corresponding dimension, and greater than the upper limit of the calculation size that can be processed at one time by the calculation unit in the corresponding dimension, then the input tensor can be determined as the second input tensor, and the scheduling module does not need to perform splitting processing on the second input tensor, but can directly use the second input tensor as the intermediate tensor. That is to say, the second input tensor can refer to the tensor whose tensor size in each dimension is less than or equal to the upper limit of the size that can be processed at one time by the splitting unit in the corresponding dimension, and greater than the upper limit of the calculation size that can be processed at one time by the calculation unit in the corresponding dimension;
[0131] When the tensor sizes in all dimensions of the input tensor are less than or equal to the upper limit of the computable size that can be processed by the computing unit in the corresponding dimension, then the input tensor can be determined as the third input tensor. At this time, neither the scheduling module nor the splitting unit needs to split the third input tensor, and the third input tensor can be directly used as the base tensor. That is to say, the third input tensor can refer to the tensor sizes in all dimensions being less than or equal to the upper limit of the computable size that can be processed by the computing unit in the corresponding dimension.
[0132] Therefore, when the tensor obtained by the scheduling module is the first input tensor with a large size, the scheduling module splits the first input tensor in at least one dimension to obtain multiple intermediate tensors with medium sizes, then generates corresponding data processing instructions according to each intermediate tensor, and sends these data processing instructions to the splitting unit; after receiving these data processing instructions, the splitting unit reads the intermediate tensors according to the data processing instructions, and splits the intermediate tensors in at least one dimension to obtain multiple base tensors with smaller sizes, and then maps these base tensors to the computing unit, and the computing unit performs tensor calculations on these base tensors to obtain tensor calculation results.
[0133] When the tensor obtained by the scheduling module is the second input tensor, the scheduling module can directly use the second input tensor as the intermediate tensor without splitting, generate a data processing instruction according to the second input tensor, and send the data processing instruction to the splitting unit; the splitting unit reads the second input tensor according to the data processing instruction, then splits the second input tensor in at least one dimension to obtain multiple base tensors with smaller sizes, and then maps the multiple split base tensors to the computing unit so that the computing unit performs parallel tensor calculations on these base tensors to obtain tensor calculation results.
[0134] When the tensor obtained by the scheduling module is the third input tensor, neither the scheduling module nor the splitting unit needs to split the input tensor. The scheduling module generates a data processing instruction according to the third input tensor and directly sends the data processing instruction to the computing unit instead of the splitting unit. Therefore, the computing unit can read the third input tensor according to the data processing instruction and perform tensor calculations on the third input tensor to obtain tensor calculation results.
[0135] By setting the rules for comparing the tensor sizes in all dimensions of the input tensor data, the scheduling module is only called to perform tensor splitting when the scale of the input tensor data exceeds the upper limit of the processable scale of the splitting unit, so as to maximize the utilization rate of the splitting unit, reduce the number of splits of the scheduling module, and further effectively reduce the number of software and hardware interactions, and improve the power consumption and utilization rate of the tensor processing device.
[0136] Refer toFigure 14 As shown Figure 14 is a schematic diagram of the specific process of the tensor processing method provided by a specific example. As Figure 14 shown, the tensor processing device may include a scheduling module and a tensor core. The tensor core includes a splitting unit and a computing unit. Among them, the scheduling module can be represented as a software module of the tensor processing device, and the tensor core can be represented as a hardware component of the tensor processing device. When the first input tensor is input to the tensor processing device, the tensor processing device will execute steps 1401 to 1408 including but not limited to the following. Specifically, step 1401: Obtain the first input tensor through the scheduling module, and split the first input tensor in at least one dimension to obtain m intermediate tensors; step 1402: Encode the m intermediate tensors respectively according to the first encoding format through the scheduling module to generate corresponding m data processing instructions; step 1403: Add the m data processing instructions to the instruction queue through the scheduling module; step 1404: Read the m intermediate tensors based on the m data processing instructions through the splitting unit; step 1405: Call the decoding unit through the splitting unit to decode the intermediate tensors according to the encoding tags of the data processing instructions to obtain the decoded intermediate tensors; step 1406: Split the m decoded intermediate tensors in at least one dimension to obtain n basic tensors, where m is greater than or equal to n, the scale of the first input tensor is greater than or equal to the scale of the intermediate tensor, and the scale of the intermediate tensor is greater than the scale of the basic tensor; step 1407: Map the n basic tensors to the computing unit through the splitting unit; step 1408: Obtain the n basic tensors through the computing unit, perform parallel tensor calculations on the n basic units, and obtain the tensor calculation result.
[0137] It can be understood that although each step in the above flowcharts is displayed in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0138] Description of the device and equipment in the embodiments of the present disclosure
[0139] Refer to Figure 15 , Figure 15FIG. 0 is an alternative schematic diagram of the tensor processing device provided by the embodiments of the present disclosure. The tensor processing device 1500 includes a scheduling module 1510, a splitting unit 1520, and a computing unit 1530. Specifically:
[0140] The scheduling module 1510 is configured to obtain a first input tensor, split the first input tensor in at least one dimension to obtain a plurality of intermediate tensors, generate corresponding data processing instructions according to each intermediate tensor, and send the data processing instructions to the splitting unit 1520;
[0141] The splitting unit 1520 is configured to parse the data processing instructions, read the intermediate tensors according to the parsed data processing instructions, split the intermediate tensors in at least one dimension to obtain a plurality of basic tensors, and map the plurality of basic tensors to the computing unit 1530;
[0142] The computing unit 1530 is configured to perform tensor calculations on a plurality of basic tensors in parallel to obtain a tensor calculation result.
[0143] In a possible implementation manner, the splitting unit 1520 is further configured to: use the upper limit value of the computable size per dimension of the computing unit 1530 as the first splitting granularity for each dimension; perform cyclic splitting on the intermediate tensors in the corresponding dimension according to the first splitting granularity for each dimension to obtain a plurality of basic tensors.
[0144] In a possible implementation manner, the splitting unit 1520 is further configured to: for each dimension, compare the size of the intermediate tensor with the first splitting granularity, and determine the dimension corresponding to the comparison result that the size of the intermediate tensor is greater than the first splitting granularity as the dimension to be split; when there are multiple dimensions to be split, traverse each dimension to be split, and perform cyclic splitting on the tensor obtained by splitting in the previous traversed dimension to be split according to the first splitting granularity corresponding to the current dimension to obtain a plurality of basic tensors.
[0145] In a possible implementation manner, the scheduling unit is further configured to: use the upper limit value of the splitting size per dimension of the splitting unit 1520 as the second splitting granularity for each dimension; perform cyclic splitting on the first input tensor in the corresponding dimension according to the second splitting granularity for each dimension to obtain a plurality of intermediate tensors.
[0146] In a possible implementation, the scheduling unit is also used to: obtain a plurality of pre-configured first candidate segmentation schemes, and determine the input tensor size of the first input tensor in each dimension; determine a first target segmentation scheme from the plurality of first candidate segmentation schemes according to the relationship between the input tensor size in each dimension and the candidate segmentation granularity of each first candidate segmentation scheme in each dimension; determine the candidate segmentation granularity of the first target segmentation scheme in each dimension as the third segmentation granularity of the corresponding dimension; and perform cyclic segmentation on the first input tensor in the corresponding dimension according to the third segmentation granularity of each dimension to obtain a plurality of intermediate tensors.
[0147] In a possible implementation, the scheduling unit is also used to: determine the input tensor size of the first input tensor in each dimension, and the upper limit of the split size that can be processed at a single time by the splitting unit 1520 in each dimension; determine the fourth split granularity of each dimension based on the ratio relationship between each input tensor size and each size upper limit value; according to the fourth split granularity of each dimension, cyclically split the first input tensor in the corresponding dimension to obtain multiple intermediate tensors.
[0148] In a possible implementation, the scheduling unit is also used to: traverse each dimension, and determine the minimum number of splits of the first input tensor of each dimension in the current dimension by the scheduling module 1510 according to the ratio of the input tensor size of each dimension and the upper limit of the split size of the current dimension; determine the second candidate splitting scheme of the first input tensor in each dimension, and according to each minimum splitting number, obtain the total number of splitting of each second candidate splitting scheme; determine the second candidate splitting scheme with the smallest total number of splitting as the second target splitting scheme; and determine the fourth splitting granularity of each dimension according to the second target splitting scheme.
[0149] In a possible implementation, the scheduling unit is further used to: encode each intermediate tensor according to the first encoding format, generate a corresponding data processing instruction, and add each data processing instruction to the instruction queue, wherein the data processing instruction includes an encoding tag for indicating the first encoding format;
[0150] According to the instruction sequence of the instruction queue, each data processing instruction is sent to the segmentation unit 1520 in sequence.
[0151] In a possible implementation, the splitting unit 1520 is further used to: parse the data processing instructions, read the intermediate tensor according to the parsed data processing instructions, and call the decoding unit to decode the intermediate tensor according to the encoding label to obtain the decoded intermediate tensor.
[0152] In a possible implementation manner, the scheduling module 1510 is further configured to: obtain a second input tensor and use the second input tensor as an intermediate tensor, where the tensor size of the second input tensor in each dimension is less than or equal to the upper limit value of the size that can be processed at one time by the splitting unit 1520 in the corresponding dimension.
[0153] The tensor processing device 1500 of the present disclosure is used to execute the tensor processing method of the above embodiment, and its specific processing process is the same as that of the tensor processing method of the above embodiment, and will not be elaborated here.
[0154] The embodiment of the present disclosure further provides an electronic device 1600, including: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions, and the instructions are executed by the at least one processor so that when the at least one processor executes the instructions, the method of any one of the above embodiments of the present disclosure is implemented.
[0155] Next, in conjunction with Figure 16 The hardware structure of the electronic device will be described in detail. The electronic device includes: a processor 1610, a memory 1620, an input / output interface 1630, a communication interface 1640, and a bus 1650.
[0156] The processor 1610 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure;
[0157] The memory 1620 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1620 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1620, and the processor 1610 is called to execute the tensor processing method of the embodiments of the present disclosure;
[0158] The input / output interface 1630 is used to implement information input and output;
[0159] The communication interface 1640 is used to implement communication and interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.), or can also implement communication through a wireless method (such as mobile network, WIFI, Bluetooth, etc.); and
[0160] The bus 1650 transmits information among various components of the device, such as the processor 1610, the memory 1620, the input / output interface 1630, and the communication interface 1640.
[0161] Among them, the processor 1610, the memory 1620, the input / output interface 1630, and the communication interface 1640 achieve communication connections with each other inside the device through the bus 1650.
[0162] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0163] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0164] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0165] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0166] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0168] It should also be understood that the various embodiments provided by the present disclosure can be combined arbitrarily to achieve different technical effects.
[0169] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A tensor processing method, characterized in that, For a tensor processing device, the tensor processing device includes a scheduling module, a splitting unit, and a computing unit, and the method includes: Through the scheduling module, obtain a first input tensor, split the first input tensor in at least one dimension to obtain a plurality of intermediate tensors, generate corresponding data processing instructions according to each of the intermediate tensors, and send the data processing instructions to the splitting unit; Through the splitting unit, parse the data processing instructions, read the intermediate tensors according to the parsed data processing instructions, split the intermediate tensors in at least one dimension to obtain a plurality of basic tensors, and map the plurality of basic tensors to the computing unit; Through the computing unit, perform tensor calculations on the plurality of basic tensors in parallel to obtain a tensor calculation result.
2. The tensor processing method according to claim 1, wherein The splitting the intermediate tensors in at least one dimension to obtain a plurality of basic tensors includes: Taking the upper limit value of the computable size per single time of the computing unit in each dimension as the first splitting granularity of each dimension; According to the first splitting granularity of each dimension, perform cyclic splitting on the intermediate tensors in the corresponding dimension to obtain a plurality of basic tensors.
3. The tensor processing method according to claim 2, wherein The performing cyclic splitting on the intermediate tensors in the corresponding dimension according to the first splitting granularity of each dimension to obtain a plurality of basic tensors includes: For each dimension, compare the size of the intermediate tensor with the first splitting granularity, and determine the dimension corresponding to the comparison result that the size of the intermediate tensor is greater than the first splitting granularity as the dimension to be split; When there are multiple dimensions to be split, traverse each of the dimensions to be split, and perform cyclic splitting on the tensor obtained by splitting in the dimension to be split in the previous traversal according to the first splitting granularity corresponding to the current dimension to obtain a plurality of basic tensors.
4. The tensor processing method according to claim 1, characterized in that The splitting the first input tensor in at least one dimension to obtain a plurality of intermediate tensors includes: Taking the upper limit value of the splitting size per single time of the splitting unit in each dimension as the second splitting granularity of each dimension; According to the second splitting granularity of each dimension, perform cyclic splitting on the first input tensor in the corresponding dimension to obtain a plurality of intermediate tensors.
5. The tensor processing method according to claim 1, wherein The splitting the first input tensor in at least one dimension to obtain a plurality of intermediate tensors includes: Obtain a plurality of first candidate splitting schemes configured in advance, and determine the input tensor size of the first input tensor in each dimension; According to the relationship between the input tensor size of each dimension and the candidate splitting granularities of each of the first candidate splitting schemes in each dimension, determine a first target splitting scheme from the plurality of first candidate splitting schemes; Determine the candidate splitting granularities of the first target splitting scheme in each dimension as the third splitting granularity of the corresponding dimension; According to the third splitting granularity of each dimension, perform cyclic splitting on the first input tensor in the corresponding dimension to obtain a plurality of intermediate tensors.
6. The tensor processing method according to claim 1, wherein The splitting the first input tensor in at least one dimension to obtain a plurality of intermediate tensors includes: Determine the input tensor sizes of the first input tensor in each dimension, and the upper limit of the slicing size that can be processed in one pass by the slicing unit in each dimension; Determine the fourth slicing granularity of each dimension according to the ratio relationship between each of the input tensor sizes and each of the size upper limits; Perform cyclic slicing on the first input tensor in the corresponding dimension according to the fourth slicing granularity of each dimension to obtain a plurality of intermediate tensors.
7. The tensor processing method according to claim 6, characterized in that, The determining the fourth slicing granularity of each dimension according to the ratio relationship between each of the input tensor sizes and each of the size upper limits includes: Traverse each dimension, and determine the minimum number of slicing times of the first input tensor in each dimension by the scheduling module in the current dimension according to the ratio relationship between the input tensor size of each dimension and the slicing size upper limit of the current dimension; Determine the second candidate slicing schemes of the first input tensor in each dimension, and count the total number of slicing times of each of the second candidate slicing schemes according to each of the minimum number of slicing times; Determine the second candidate slicing scheme with the smallest numerical value of the total number of slicing times as the second target slicing scheme; Determine the fourth slicing granularity of each dimension according to the second target slicing scheme.
8. The tensor processing method according to claim 1, wherein The generating corresponding data processing instructions according to each of the intermediate tensors and sending the data processing instructions to the slicing unit includes: Encode each of the intermediate tensors according to the first encoding format to generate corresponding data processing instructions, and add each of the data processing instructions to an instruction queue, where the data processing instruction includes an encoding label for indicating the first encoding format; Send each of the data processing instructions to the slicing unit in sequence according to the instruction order of the instruction queue.
9. The tensor processing method according to claim 8, wherein The tensor processing device further includes a decoding unit, and the parsing the data processing instruction by the slicing unit and reading the intermediate tensor according to the parsed data processing instruction includes: Parse the data processing instruction by the slicing unit, read the intermediate tensor according to the parsed data processing instruction, and call the decoding unit to decode the intermediate tensor according to the encoding label to obtain the decoded intermediate tensor.
10. The tensor processing method according to claim 1, characterized in that The method further includes: Obtain a second input tensor through the scheduling module and use the second input tensor as an intermediate tensor, where the tensor size of the second input tensor in each dimension is less than or equal to the size upper limit that can be processed in one pass by the slicing unit in the corresponding dimension.
11. A tensor processing device, characterized in that, including: A scheduling module, a slicing unit, and a computing unit, where The scheduling module is configured to obtain a first input tensor, slice the first input tensor in at least one dimension to obtain a plurality of intermediate tensors, generate corresponding data processing instructions according to each of the intermediate tensors, and send the data processing instructions to the slicing unit; The splitting unit is configured to parse the data processing instruction, read the intermediate tensor according to the parsed data processing instruction, split the intermediate tensor in at least one dimension to obtain a plurality of basic tensors, and map the plurality of basic tensors to the computing unit; The computing unit is configured to perform tensor calculations on the plurality of basic tensors in parallel to obtain tensor calculation results.
12. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is run by the processor, it implements the tensor processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be run by one or more processors to implement the tensor processing method according to any one of claims 1 to 10.
Citation Information
Cited By
Method, device and equipment for realizing matrix operation on TPU (Thermoplastic Polyurethane) and medium
CN120596777A
Data processing method, electronic device, storage medium and program product
CN120849771A
Tensor calculation method, electronic equipment and computer readable storage medium
CN121167099A
Weight tensor segmentation method, device, equipment, medium and product
CN121212233A
Operator execution method and device, equipment, storage medium and program product
CN121233489A