Tensor read-write method and device and artificial intelligence chip

By handling memory out-of-bounds access at the hardware level, the problem of slow read and write speed of tensors in large-scale artificial intelligence models is solved, and more efficient data processing performance is achieved.

CN120010788APending Publication Date: 2025-05-16广州壁仞智能科技有限公司 +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510169252.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

As the scale of artificial intelligence models grows, the dimensions and memory requirements of tensors have also increased rapidly, resulting in memory exceeding the limit and memory cannot be allocated, thereby reducing the speed of tensor reading and writing and data processing performance.

Method used

By determining the index coordinates to be read and written, and when the index coordinates exceed the upper limit of the index coordinates of the tensor element, the corresponding physical coordinates are set to a preset value, thereby processing memory out-of-bounds access at the hardware level to avoid reading and writing dirty data.

Benefits of technology

It realizes processing memory out-of-bounds access at the hardware level, avoids dirty data generated by tensor folding, and improves the speed of tensor reading and writing and data processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010788A_ABST
    Figure CN120010788A_ABST
Patent Text Reader

Abstract

The invention provides a tensor read-write method and device and an artificial intelligence chip, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining an index coordinate to be read and written; the index coordinate is used for representing the position of each element in the target tensor; under the condition that the index coordinates to be read and written are greater than the index coordinate upper limit value of each element in the target tensor, setting the corresponding physical coordinates to be read and written in the memory of the index coordinates to be read and written as preset values; the physical coordinate of each element in the target tensor in the memory is determined based on the index coordinate of each element and a folding mode adopted when the target tensor is stored in the memory; and under the condition that the physical coordinates to be read and written are preset values, performing no read-write operation on data corresponding to the physical coordinates to be read and written in the memory. According to the method and the device provided by the invention, the tensor read-write speed is increased, and the data processing performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a tensor reading and writing method, device and artificial intelligence chip. Background Art

[0002] In AI application scenarios, there is a trend that as the model size increases, the tensor dimension increases, especially in large language models, the sequence length and hidden size are getting larger and larger, which is reflected in the increase of tensor dimension. As the tensor dimension and the size of each dimension increase, the memory requirement will also increase rapidly. Exceeding the memory limit will result in the inability to allocate memory for the tensor.

[0003] Related technologies store tensors in memory by folding tensors. Since the folding method is used for storage, the shape of the tensor is inconsistent with the shape of the storage area of ​​the tensor in the memory space. Therefore, when reading and writing tensors in the memory, it is necessary to determine at the software level whether the physical coordinates of the read elements in the memory are within the allocated storage area, that is, to handle memory out-of-bounds access (OOB), thereby avoiding reading dirty data. The above method reduces the speed of tensor reading and writing, resulting in a decrease in data processing performance.

[0004] Therefore, how to increase the speed of tensor reading and writing and improve data processing performance has become a technical problem that needs to be solved urgently in the industry. Summary of the invention

[0005] The present invention provides a tensor reading and writing method, device and artificial intelligence chip, which are used to solve the technical problem of how to increase the speed of tensor reading and writing and improve data processing performance.

[0006] The present invention provides a tensor reading and writing method, comprising: Determine the index coordinates to be read and written; the index coordinates are used to represent the position of each element in the target tensor; In the case where the index coordinates to be read and written are greater than the upper limit value of the index coordinates of each element in the target tensor, the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory are set to preset values; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; When the physical coordinates to be read or written are preset values, no read or write operations are performed on the data corresponding to the physical coordinates to be read or written in the memory.

[0007] In some embodiments, after determining the index coordinates to be read and written, the method further includes: In a case where the index coordinates to be read or written are less than the upper limit of the index coordinates of each element in the target tensor, determining the physical coordinates to be read or written based on the index coordinates to be read or written and the folding method used when the target tensor is stored in the memory; Perform read and write operations on the data corresponding to the physical coordinates to be read and written in the memory.

[0008] In some embodiments, the physical coordinates of each element in the target tensor in the memory are determined based on the following steps: Determine a folding method used when the target tensor is stored in the memory; the data dimension of the target tensor includes a first dimension and a second dimension; Based on the folding manner, determining that the second dimension is folded toward the first dimension, and a folding granularity of the second dimension; Based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension, the physical coordinates of each element in the memory are determined.

[0009] In some embodiments, determining the physical coordinates of each element in the memory based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension includes: The product of the first dimension coordinate value in the index coordinate of each element and the upper limit value of the index coordinate of the second dimension is added to the second dimension coordinate value in the index coordinate of each element to obtain a first calculation result; The first calculation result is divided by the folding granularity of the second dimension to obtain the first dimension coordinate value of each element in the physical coordinate in the memory.

[0010] In some embodiments, determining the physical coordinates of each element in the memory based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension includes: A modulo operation is performed on the first calculation result and the folding granularity of the second dimension to obtain a second dimensional coordinate value of each element in the physical coordinate in the memory.

[0011] The present invention provides a tensor reading and writing device, comprising: A determination unit, used to determine the index coordinates to be read or written; the index coordinates are used to represent the position of each element in the target tensor; A setting unit, configured to set the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory to a preset value when the index coordinates to be read and written are greater than the upper limit value of the index coordinates of each element in the target tensor; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; The read / write unit is used for not performing read / write operations on the data corresponding to the physical coordinates to be read / written in the memory when the physical coordinates to be read / written are preset values.

[0012] The present invention provides an artificial intelligence chip, including a memory management unit; The memory management unit is used to execute the tensor reading and writing method.

[0013] The present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the tensor reading and writing method is implemented when the processor executes the computer program.

[0014] The present invention provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the tensor reading and writing method is implemented.

[0015] The present invention provides a computer program product, comprising a computer program, wherein the computer program implements the tensor reading and writing method when executed by a processor.

[0016] The tensor reading and writing method, device and artificial intelligence chip provided by the present invention determine the index coordinates to be read and written; the index coordinates are used to indicate the position of each element in the target tensor; when the index coordinates to be read and written are greater than the upper limit value of the index coordinates of each element in the target tensor, the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory are set to preset values; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; when the physical coordinates to be read and written are preset values, no read and write operations are performed on the data corresponding to the physical coordinates to be read and written in the memory; since the occurrence of memory out-of-bounds access is judged by the index coordinates, no read and write operations are performed on the data in the memory, thereby realizing the processing of memory out-of-bounds access at the hardware level (memory management) without the need for special processing at the software level (application), avoiding the reading and writing of dirty data generated by tensor folding, improving the speed of tensor reading and writing, and improving data processing performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is a schematic diagram of the tensor provided by the present invention not being folded.

[0020] Figure 2 This is one of the schematic diagrams of tensor folding provided by the present invention.

[0021] Figure 3 This is the second schematic diagram of the tensor folding provided by the present invention.

[0022] Figure 4 It is a flowchart of the tensor reading and writing method provided by the present invention.

[0023] Figure 5 A schematic diagram of the structure of the tensor reading and writing device provided by the present invention.

[0024] Figure 6 It is a schematic diagram of the structure of the artificial intelligence chip provided by the present invention.

[0025] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units or modules is not necessarily limited to those steps or units or modules that are clearly listed, but may include other steps or units or modules that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] In the performance tuning scenarios of artificial intelligence models such as text-based graph models, there are often a large number of memory transfer operations on tensors, such as layout switching, deformation, sequence rearrangement, and deformation + sequence rearrangement. As the demand increases, the shapes of tensors are becoming more diverse.

[0029] Due to the hardware settings of the addressable space in the AI ​​chip, there is an upper limit on the memory size allocation of a single tensor, and each dimension also has an upper limit. For example, for a three-dimensional tensor, it can include three dimensions, namely batch, height, and width, which can be abbreviated as B, H, and W respectively. When storing the tensor in memory, each element can be written into the memory in the order of width, height, and batch.

[0030] The logical tensor can be used to represent the actual shape of a tensor, and the physical tensor can be used to represent the shape of the storage area of ​​the tensor in the memory space.

[0031] For a tensor, if the data of each dimension does not exceed the upper limit of each dimension, the logical tensor can be considered equal to the physical tensor. For example, suppose the upper limit of the H dimension H_LIMIT is 8192. Figure 1 is a schematic diagram of the tensor provided by the present invention without folding, such as Figure 1 As shown in the figure, if the shape of a tensor is 1*8190 (batch*height), which is less than the upper limit of each dimension, the logical tensor can be considered equal to the physical tensor, and the tensor is not folded when stored in memory. Figure 2 is one of the schematic diagrams of tensor folding provided by the present invention, such as Figure 2 As shown in the figure, if the shape of a tensor is 1*8194 (batch*height), the tensor exceeds the upper limit in the H dimension. In this case, the H dimension can be folded to the B dimension, and the folded physical tensor is 33*256 (batch*height). It can be considered that the physical tensor is larger than the logical tensor.

[0032] After the tensor is folded, the memory space actually occupied by the physical tensor is larger than the memory space required by the logical tensor, and the excess space can be considered as a hole. For example, the memory space actually occupied by the physical tensor can store 33*256=8448 elements, while the logical tensor only needs 8194 elements. The memory space occupied by the extra 8448-8194=254 elements is a hole. The data in the hole is dirty data. Dirty data is data that should not be operated when reading and writing tensors, but dirty data occupies the actual memory location. For example Figure 2 The shaded area in the middle indicates a hole in the memory space. The data in the hole is dirty data. The hole appears at the end of the memory.

[0033] Figure 3 This is the second schematic diagram of the tensor folding provided by the present invention, such as Figure 3 As shown in the figure, if the shape of a tensor is 1*11304*11304 (batch*height*width), the folded physical tensor is 9*4096*4096 (batch*height*width). After the tensor is folded, the actual memory space occupied by the physical tensor (12288*12288) is larger than the memory space required by the logical tensor (11304*11304), and the excess part can be considered as holes. These holes may also appear in the middle of the memory. For example Figure 3 The shaded area in the middle represents the hole in the memory space, and the data in the hole is dirty data. It can be seen that due to the introduction of more dimensions, the hole can become discontinuous, and the hole appears in the middle of the memory space.

[0034] The relevant technology needs to determine at the software level whether the physical coordinates of the read element in the memory are within the allocated storage area, that is, to handle memory out-of-bounds access (OOB), so as to avoid reading dirty data. If actual access occurs on the hardware, in order not to affect the correctness of the result of the logical operation in the final result, the developer is also required to write a software program to fill the data in the hole with zeros in advance. In this way, the software program written by the developer contains complex logical branches, and there is a risk of introducing program errors (bugs). The above method reduces the speed of tensor reading and writing, resulting in a decrease in data processing performance.

[0035] In order to solve the shortcomings of related technologies, Figure 4 It is a flowchart of the tensor reading and writing method provided by the present invention, such as Figure 4 As shown, the method includes step 410, step 420 and step 430.

[0036] Step 410: Determine the index coordinates to be read or written; the index coordinates are used to indicate the position of each element in the target tensor.

[0037] Specifically, the subject of executing the tensor reading and writing method provided in the embodiment of the present invention is a tensor reading and writing device. The device can be implemented by a driver in a hardware device, such as a tensor reading and writing program; it can also be implemented by hardware, such as a memory management unit that executes the tensor reading and writing method in an artificial intelligence chip.

[0038] The target tensor is the tensor whose elements need to be accessed. The access methods include read operation and write operation.

[0039] Tensors are multidimensional arrays that can represent various data structures, including scalars (zero-dimensional tensors), vectors (one-dimensional tensors), matrices (two-dimensional tensors), and higher-dimensional arrays. A tensor consists of multiple elements. The specific positions of these elements in the tensor are represented by index coordinates. For example, for a two-dimensional tensor, its shape is 1*8190, which means that the tensor has 1 row and 8190 columns. For any element, the index coordinate (i, j) can be used to access it. Among them, i is the row coordinate value and j is the column coordinate value. Rows can represent batch dimensions, and columns can represent height dimensions. For multidimensional tensors, the number and order of index coordinates are related to the dimension and shape of the tensor.

[0040] Step 420: When the index coordinates to be read or written are greater than the upper limit of the index coordinates of each element in the target tensor, the physical coordinates to be read or written in the memory corresponding to the index coordinates to be read or written are set to preset values; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory.

[0041] Specifically, due to the hardware settings for the addressable space, there is an upper limit on the memory size allocation of a single tensor, and there is also an upper limit for each dimension. In other words, there is an upper limit on the index coordinates of each element in the target tensor. For different dimensions, the upper limit of the index coordinates is different and can be determined based on the hardware settings for the addressable space. For example, for high dimensions (H), the upper limit of the index coordinates H_LIMIT can be 8192; for wide dimensions (W), the upper limit of the index coordinates W_LIMIT can be 8192.

[0042] The index coordinates to be read or written are simply the positions of the elements determined from the actual shape of the tensor. To read or write the elements corresponding to the index coordinates, it is necessary to determine the positions of the elements in memory space. Physical coordinates are used to represent the positions of each element in the target tensor in memory space. Considering that the target tensor is stored in memory after being folded, the physical coordinates of each element in memory can be determined based on the index coordinates of each element and the folding method used when the target tensor is stored in memory.

[0043] When the index coordinates to be read and written are greater than the upper limit of the index coordinates of each element in the target tensor, it means that the element corresponding to the index coordinates to be read and written is located in a hole, and the actual storage location exceeds the limit of the memory space required by the target tensor. The data is not an element in the target tensor, but dirty data. At this time, the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory can be set to a preset value. The preset value can be set as needed. For example, the preset value can be the upper limit of the index coordinate H_LIMIT.

[0044] Step 430: When the physical coordinates to be read or written are preset values, no read or write operation is performed on the data corresponding to the physical coordinates to be read or written in the memory.

[0045] Specifically, when it is determined that the physical coordinates to be read and written are preset values, it can be determined that the data corresponding to the physical coordinates to be read and written is dirty data and is not an element in the target tensor, and the data corresponding to the physical coordinates to be read and written in the memory need not be read or written.

[0046] The tensor reading and writing method provided by the embodiment of the present invention determines the index coordinates to be read and written; the index coordinates are used to indicate the position of each element in the target tensor; when the index coordinates to be read and written are greater than the upper limit value of the index coordinates of each element in the target tensor, the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory are set to preset values; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; when the physical coordinates to be read and written are preset values, no read and write operations are performed on the data corresponding to the physical coordinates to be read and written in the memory; since the occurrence of memory out-of-bounds access is judged by the index coordinates, no read and write operations are performed on the data in the memory, thereby realizing the processing of memory out-of-bounds access at the hardware level (memory management) without the need for special processing at the software level (application), avoiding the reading and writing of dirty data generated by tensor folding, improving the speed of tensor reading and writing, and improving data processing performance.

[0047] It should be noted that each implementation of the present invention can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.

[0048] In some embodiments, after determining the index coordinates to be read or written, the method further includes: When the index coordinates to be read or written are less than the upper limit of the index coordinates of each element in the target tensor, the physical coordinates to be read or written are determined based on the index coordinates to be read or written and the folding method used when the target tensor is stored in the memory; Perform read and write operations on the data corresponding to the physical coordinates to be read and written in the memory.

[0049] Specifically, when the index coordinates to be read or written are less than the upper limit of the index coordinates of each element in the target tensor, it means that the actual storage location in the memory of the element corresponding to the index coordinates to be read or written does not exceed the limit of the memory space required by the target tensor, and the data is an element in the target tensor, not dirty data.

[0050] According to the folding method used when the target tensor is stored in the memory, the index coordinates to be read and written can be mapped to obtain the physical coordinates to be read and written. Then, the data corresponding to the physical coordinates to be read and written in the memory can be read and written.

[0051] The tensor reading and writing method provided in the embodiment of the present invention determines that no out-of-bounds memory access occurs through index coordinates, thereby processing out-of-bounds memory access at the hardware level (memory management) without the need for special processing at the software level (application), avoiding reading and writing dirty data generated by tensor folding, improving the speed of tensor reading and writing, and improving data processing performance.

[0052] In some embodiments, the physical coordinates of each element in the target tensor in memory are determined based on the following steps; Determine the folding method used when the target tensor is stored in the memory; the data dimensions of the target tensor include the first dimension and the second dimension; Based on the folding mode, determining the folding of the second dimension into the first dimension and the folding granularity of the second dimension; Based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension, the physical coordinates of each element in the memory are determined.

[0053] Specifically, the folding mode is used to indicate the folding direction and folding granularity between the dimensions of the target tensor. The folding granularity indicates the dimension size of the sub-tensors after the target tensor is folded into multiple sub-tensors.

[0054] The data dimensions of the target tensor include the first dimension and the second dimension. The second dimension can be folded toward the first dimension. The number of second dimensions can be multiple. For example, for a three-dimensional tensor Batch*Height*Width, it can include three dimensions, namely B dimension (Batch), H dimension (Height) and W dimension (Width). Each dimension has an upper limit of the allocatable space. The upper limit of the B dimension is 1024, the upper limit of the H dimension is 8192, and the upper limit of the W dimension is 8192. When the H dimension and W dimension of the tensor reach the upper limit, it can be folded toward the B dimension. That is, the B dimension is the first dimension, and the H dimension and W dimension are the second dimension. After folding, the target tensor is folded into multiple sub-tensors, and the size of each sub-tensor is subH* subW. subH is the folding granularity of the H dimension, that is, the size of the sub-tensor in the H dimension. num_subH is the number of folds in the H dimension. subW is the folding granularity of the W dimension, that is, the size of the sub-tensor in the W dimension. num_subW is the number of folds in the H dimension. After folding, the three-dimensional tensor can be expressed as (Batch*num_subH*num_subW)*subH*subW. num_subH*num_subW can be considered as an additional Batch dimension. Batch*num_subH*num_subW is the folded Batch, subH is the folded Height, and subW is the folded Width.

[0055] For ease of understanding and explanation, take the two-dimensional tensor Batch*Height as an example, which is folded to (Batch*num_subH)*subH. For ease of expression, newBatch=Batch*num_subH, newH=subH. The two-dimensional tensor is 1*8194 (Batch*Height). The upper limit of the H dimension is H_LIMIT = 8192. According to the folding logic, it is 33*256 (newBatch *newH). Each newBatch is newH elements long. The total size of the physical tensor is newBatch* newH. That is, it contains 33*256=8448 elements, which has exceeded the 8194 elements (data) of the original logical tensor.

[0056] After folding, the physical coordinates of each element in the memory can be determined based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension.

[0057] The tensor reading and writing method provided in an embodiment of the present invention can map the index coordinates of each element to the physical coordinates of each element in the memory according to the folding method used when the target tensor is stored in the memory, and can accurately read and write each element in the target tensor.

[0058] In some embodiments, determining the physical coordinates of each element in the memory based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension includes: The product of the first dimension coordinate value in the index coordinate of each element and the upper limit value of the second dimension index coordinate is added to the second dimension coordinate value in the index coordinate of each element to obtain a first calculation result; Perform integer division operation on the first calculation result and the folding granularity of the second dimension to obtain the first dimension coordinate value of each element in the physical coordinate in the memory; A modulo operation is performed on the first calculation result and the folding granularity of the second dimension to obtain a coordinate value of the second dimension of each element in the physical coordinate in the memory.

[0059] Specifically, take a two-dimensional tensor as an example. The tensor includes a first dimension and a second dimension. The upper limit value of the first dimension is Batch, and the upper limit value of the second dimension is H. The index coordinates of any element in the tensor can be expressed as (b, h). The value range of the first dimension coordinate value b is [0, Batch-1], and the value range of the second dimension coordinate value h is [0, H-1]. If the two-dimensional tensor is 1*8194, the value range of b is [0, 0], and the value range of h is [0, 8193].

[0060] After the two-dimensional tensor is folded, the coordinate value of the first dimension in the physical coordinate can be expressed as newb, and the coordinate value of the second dimension can be expressed as newh. The value range of newb is [0, newBatch-1]. The value range of newh is [0, newH-1]. newBatch=33, newH=256. Then the value range of newb is [0, 32], and the value range of newh is [0, 255].

[0061] The product of the first dimension coordinate value b in the index coordinates of each element and the upper limit value H of the second dimension index coordinates is added to the second dimension coordinate value h in the index coordinates of each element to obtain a first calculation result b*H+h.

[0062] The coordinate mapping relationship between index coordinates and physical coordinates can be expressed as: The first calculation result b*H+h is divided by the folding granularity subH of the second dimension to obtain the first dimension coordinate value newb of each element in the physical coordinate in the memory, which is expressed by the formula: newb=(b*H+h) / subH; where / is the integer division symbol.

[0063] The first calculation result b*H+h is modulo the folding granularity subH of the second dimension to obtain the second dimensional coordinate value newh of each element in the physical coordinate in the memory, which is expressed by the formula: newh=(b*H+h)%subH, where % is the modulo symbol.

[0064] The tensor reading and writing method provided in an embodiment of the present invention determines the physical coordinates of each element in the memory according to the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension, and can accurately read and write each element in the target tensor.

[0065] In some embodiments, for example, an application attempts to access the element corresponding to the index coordinate (0, 8200) in the 2D tensor 1*8194. This coordinate outside the logical boundary range will access the element (newb=32, newh=8) in the physical tensor. Because 32=8200 / 256, 8=8200%256. This element occupies the actual memory location and will be accessed by the user. In fact, the index coordinate should not access the element.

[0066] If the method in the related art is used, actual access to the hardware will occur, but the final result will not affect the correctness of the result of the logical operation. The method in the related art does not require the coordinate processing of (32,8). However, it is necessary to fill the 254 data in the hole with 0 in advance to solve the problem. The logic branch needs to be written at the software level, which is complex and has the risk of introducing program errors.

[0067] In the method provided in the embodiment of the present invention, since the memory out-of-bounds access has been determined through index coordinates at the hardware level, there is no need to perform actual access, generate actual memory reads and writes, and return 0; there is no need for the user to manually add 0. It is meaningless for 0 to participate in the calculation and has no effect on the logical result.

[0068] That is, when the memory is read or written, the hardware will detect whether the accessed coordinates are within the allocated physical tensor range. If they are within the physical range, the hardware reads and writes correctly and returns the data. If they are outside the read boundary, the hardware will not report an error and will directly return 0 to the software. If there is a write operation, the hardware will not operate and will execute nothing. The speed of hardware processing OOB is much faster than software processing.

[0069] The following describes an apparatus provided by an embodiment of the present invention. The apparatus described below and the method described above can refer to each other.

[0070] Figure 5 A schematic diagram of the structure of the tensor reading and writing device provided by the present invention is shown in FIG. Figure 5 As shown, the device comprises: A determination unit 510 is used to determine the index coordinates to be read or written; the index coordinates are used to indicate the position of each element in the target tensor; A setting unit 520, configured to set the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory to a preset value when the index coordinates to be read and written are greater than the upper limit value of the index coordinates of each element in the target tensor; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; The read / write unit 530 is configured to not perform a read / write operation on the data corresponding to the physical coordinates to be read / written in the memory when the physical coordinates to be read / written are preset values.

[0071] The tensor reading and writing device provided by the embodiment of the present invention determines the index coordinates to be read and written; the index coordinates are used to indicate the position of each element in the target tensor; when the index coordinates to be read and written are greater than the upper limit value of the index coordinates of each element in the target tensor, the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory are set to preset values; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; when the physical coordinates to be read and written are preset values, no read and write operations are performed on the data corresponding to the physical coordinates to be read and written in the memory; since the occurrence of memory out-of-bounds access is judged by the index coordinates, no read and write operations are performed on the data in the memory, thereby realizing the processing of memory out-of-bounds access at the hardware level (memory management) without the need for special processing at the software level (application), avoiding the reading and writing of dirty data generated by tensor folding, improving the speed of tensor reading and writing, and improving data processing performance.

[0072] Figure 6 is a schematic diagram of the structure of the artificial intelligence chip provided by the present invention, such as Figure 6 As shown, the artificial intelligence chip 600 includes a memory management unit 610; the memory management unit 610 is used to execute the tensor reading and writing method in the above embodiment.

[0073] Specifically, artificial intelligence chips can be: graphics processing units (GPU), general-purpose computing on graphics processing units (GPGPU), domain-specific architectures (DSA), etc.

[0074] The memory management unit is a key hardware module for processing memory access. It can be responsible for converting index addresses into physical addresses, and can execute the tensor reading and writing methods in the above embodiments to detect whether the access is out of bounds at the hardware level.

[0075] The artificial intelligence chip provided by the embodiment of the present invention realizes the processing of memory out-of-bounds access at the hardware level through the memory management unit, without the need for special processing at the software level (application), avoiding the reading and writing of dirty data caused by tensor folding, improving the speed of tensor reading and writing, and improving data processing performance.

[0076] Figure 7 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 7 As shown, the electronic device may include: a processor (Processor) 710, a communication interface (Communications Interface) 720, a memory (Memory) 730 and a communication bus (Communications Bus) 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic command in the memory 730 to execute the method described in the above embodiment, for example: Determine the index coordinates to be read and written; the index coordinates are used to indicate the position of each element in the target tensor; when the index coordinates to be read and written are greater than the upper limit of the index coordinates of each element in the target tensor, set the physical coordinates to be read and written in the memory corresponding to the index coordinates to be read and written to the preset value; the physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; when the physical coordinates to be read and written are the preset values, do not perform read and write operations on the data corresponding to the physical coordinates to be read and written in the memory.

[0077] In addition, the logic commands in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several commands to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0078] The processor in the electronic device provided in the embodiment of the present invention can call the logic instructions in the memory to implement the above method. Its specific implementation method is consistent with the implementation method of the aforementioned method and can achieve the same beneficial effects, which will not be repeated here.

[0079] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in the above embodiments is implemented.

[0080] Its specific implementation is consistent with the aforementioned method implementation and can achieve the same beneficial effects, so it will not be repeated here.

[0081] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described above is implemented.

[0082] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0083] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A tensor reading and writing method, characterized in that: include: Determine the index coordinates to be read and written; the index coordinates are used to represent the position of each element in the target tensor; When the index coordinates to be read and written are greater than the upper limit of the index coordinates of each element in the target tensor, the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory are set to preset values; The physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; When the physical coordinates to be read or written are preset values, no read or write operations are performed on the data corresponding to the physical coordinates to be read or written in the memory.

2. The tensor reading and writing method according to claim 1, characterized in that: After determining the index coordinates to be read and written, the method further includes: In a case where the index coordinates to be read or written are less than the upper limit of the index coordinates of each element in the target tensor, determining the physical coordinates to be read or written based on the index coordinates to be read or written and the folding method used when the target tensor is stored in the memory; Perform read and write operations on the data corresponding to the physical coordinates to be read and written in the memory.

3. The tensor reading and writing method according to claim 1 or 2, characterized in that: The physical coordinates of each element in the target tensor in the memory are determined based on the following steps: Determine a folding method used when the target tensor is stored in the memory; the data dimension of the target tensor includes a first dimension and a second dimension; Based on the folding manner, determining that the second dimension is folded toward the first dimension, and a folding granularity of the second dimension; Based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension, the physical coordinates of each element in the memory are determined.

4. The tensor reading and writing method according to claim 3, characterized in that: The determining the physical coordinates of each element in the memory based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension includes: The product of the first dimension coordinate value in the index coordinate of each element and the upper limit value of the index coordinate of the second dimension is added to the second dimension coordinate value in the index coordinate of each element to obtain a first calculation result; The first calculation result is divided by the folding granularity of the second dimension to obtain the first dimension coordinate value of each element in the physical coordinate in the memory.

5. The tensor reading and writing method according to claim 4, characterized in that: The determining the physical coordinates of each element in the memory based on the index coordinates of each element, the upper limit value of the index coordinates of the second dimension, and the folding granularity of the second dimension includes: A modulo operation is performed on the first calculation result and the folding granularity of the second dimension to obtain a second dimensional coordinate value of each element in the physical coordinate in the memory.

6. A tensor reading and writing device, characterized in that: include: A determination unit, used to determine the index coordinates to be read or written; the index coordinates are used to represent the position of each element in the target tensor; A setting unit, configured to set the physical coordinates to be read and written corresponding to the index coordinates to be read and written in the memory to a preset value when the index coordinates to be read and written are greater than the upper limit value of the index coordinates of each element in the target tensor; The physical coordinates of each element in the target tensor in the memory are determined based on the index coordinates of each element and the folding method used when the target tensor is stored in the memory; The read / write unit is used for not performing read / write operations on the data corresponding to the physical coordinates to be read / written in the memory when the physical coordinates to be read / written are preset values.

7. An artificial intelligence chip, characterized in that: Includes memory management unit; The memory management unit is used to execute the tensor reading and writing method described in any one of claims 1 to 5.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the tensor reading and writing method according to any one of claims 1 to 5 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the tensor reading and writing method according to any one of claims 1 to 5 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the tensor reading and writing method according to any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Data processing method, data processing device, electronic equipment and storage medium

    CN120744191A