Data access method and device and readable storage medium

By performing broadcast settings and tensor memory layout feature analysis on the input of the target operator, combining and reorganizing memory spaces, the problem of memory access discontinuity in deep learning is solved, and rapid data access and operator computing efficiency is improved.

CN120508412AActive Publication Date: 2025-08-19LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510999389.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-08-19
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

The prior art in deep learning tasks has discontinuity of memory access due to the existence of broadcast dimensions, which affects data acquisition efficiency. Explicit data replication increases data copying time and memory usage, and the method of calculating the offset of data in memory based on the tensor step size still has the problem of inefficient data acquisition.

Method used

By obtaining the input of the target operator for broadcast settings, combining the tensor memory layout features to determine the dimension types of input and output, performing continuous dimension merging, step-by-step dimension sorting and broadcast dimension merging, reorganizing the memory space in segments, and using data access instructions to determine the accessed memory segments, and access tensor data using the adaptive data access method.

Benefits of technology

It realizes rapid data access during operator computing, improves data acquisition efficiency, and optimizes the continuity of memory access and data access speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508412A_ABST
    Figure CN120508412A_ABST
Patent Text Reader

Abstract

The invention discloses a data access method and device and a readable storage medium, relates to the technical field of computers, and performs broadcast setting for input of a target operator. Then, dimension types corresponding to the input and the output respectively are determined in combination with tensor memory layout characteristics of the target operator, so that memory spaces can be assembled in a segmented manner after memories corresponding to the input and the output are combined and rearranged in combination with the dimension types. Therefore, after the data access instruction is received, the memory segment to be accessed can be determined by using the data access instruction, and the tensor data of the target operator can be accessed by adopting the data access mode matched with the memory segment, so that the tensor data can be obtained in an accelerated manner. The method has the advantages that the tensor memory space of the target operator is changed from scattering to segmented assembly memory space, so that the memory arrangement rule for storing the same dimension type is realized, and the corresponding tensor data can be quickly accessed by adopting a data access mode matched with the memory segment during data access.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data access method, device, and readable storage medium. Background Art

[0002] In deep learning tasks, data often requires various preprocessing operations, such as normalization and standardization. These operations often involve performing specific mathematical operations on each element in the dataset. Elementwise broadcast operators allow element-by-element operations on data structures of various shapes.

[0003] The broadcast dimension causes discontinuous memory access, impacting data retrieval efficiency. Currently, implementing broadcast with explicit data copying increases data copying time and memory usage, with little improvement in data retrieval efficiency. Alternatively, the data's in-memory offset is calculated based on the tensor's stride and retrieved using the offset. While this approach avoids unnecessary data copying, it still treats data as discontinuous, resulting in inefficient data retrieval.

[0004] Therefore, how to quickly obtain data from memory during operator calculation is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0005] The present application provides a data access method, device and readable storage medium, which can accelerate the operator calculation process, quickly access data from the memory, and improve the calculation efficiency.

[0006] This application provides a data access method, including: Obtain the input of the target operator, and if the input satisfies broadcast requirements, perform broadcast settings on the input; Obtaining a tensor memory layout feature of the target operator, and using the tensor memory layout feature to determine the dimension types corresponding to the input and output respectively; After performing continuous dimension merging, stride dimension sorting, and broadcast dimension merging on the memory spaces corresponding to the input and the output respectively in combination with the dimension type, the memory space is reorganized in segments; A data access instruction is used to determine an accessed memory segment, and a data access method adapted to the memory segment is used to access the tensor data of the target operator.

[0007] The present application also provides a data access device, comprising: A broadcast setting module is used to obtain the input of the target operator and perform broadcast setting on the input if the input meets the broadcast requirements; A memory layout analysis module, configured to obtain tensor memory layout features of the target operator and determine the dimension types corresponding to the input and output respectively using the tensor memory layout features; a memory layout merging and rearrangement module, configured to perform continuous dimension merging, stride dimension sorting, and broadcast dimension merging on the memory spaces corresponding to the input and the output, respectively, based on the dimension type, and then reorganize the memory spaces in segments; The operator computing module is used to determine the memory segment to be accessed using a data access instruction, and access the tensor data of the target operator using a data access method adapted to the memory segment.

[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data access methods when executing the computer program.

[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data access methods are implemented.

[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned data access methods when executed by a processor.

[0011] This application first sets up broadcasting for the input of the target operator. Then, the dimension types corresponding to the input and output are determined based on the tensor memory layout characteristics of the target operator. In this way, the memory corresponding to the input and output can be merged and rearranged based on the dimension types. After completing the memory merging and rearrangement, the memory space is assembled in segments. In this way, after receiving the data access instruction, the data access instruction can be used to determine the memory segment to be accessed, and the data access method adapted to the memory segment can be used to access the tensor data of the target operator, thereby accelerating the acquisition of tensor data.

[0012] This application has the technical effect of transforming the tensor memory space of the target operator from scattered and disorganized to segmented assembled memory space, so that the memory arrangement pattern of the same dimensional type is stored. When accessing data, a data access method adapted to the memory segment is adopted, which can quickly access the corresponding tensor data and further accelerate the operator's computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0014] Figure 1 A flowchart of a data access method provided in an embodiment of the present application; Figure 2 A schematic diagram of a broadcast setting process provided in an embodiment of the present application; Figure 3 A schematic diagram of a dimension type determination process provided in an embodiment of the present application; Figure 4 A schematic diagram of memory merging and reorganization provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a data access device provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application; Figure 7 A schematic diagram of the specific structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0015] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0016] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0017] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0018] For ease of understanding, the technical terms involved in this application are explained below.

[0019] Element-wise operators: are a type of operation that operates independently on each element in a tensor (or array, matrix, or other data structures).

[0020] Tensor broadcasting is an important concept in deep learning and scientific computing. It allows tensors of different shapes to be automatically expanded to compatible shapes during arithmetic operations without explicitly copying the data. This mechanism can significantly simplify code and improve computational efficiency.

[0021] The basic rules of broadcasting: when two tensors are operated on, frameworks such as NumPy (a basic library for scientific computing) / PyTorch (a dynamic computing framework) / TensorFlow (an industrial-grade application framework) will automatically broadcast according to the following rules: Dimension right alignment: compare the shapes of the two tensors starting from the rightmost dimension; Dimension compatibility condition: the two dimensions are equal; one of the dimensions is 1; automatic expansion: copy the axis with dimension 1 along that dimension to match the corresponding dimension of the other tensor.

[0022] Tensor is a mathematical concept that is a generalization of the concepts of vector and matrix.

[0023] How layout data is stored in memory and the order in which it is calculated.

[0024] CUDA is a parallel computing platform and programming model, or can be regarded as a programming language.

[0025] Kernel: In CUDA (Compute Unified Device Architecture), a kernel refers to a parallel computing function that runs on an NVIDIA GPU (a graphics processing unit).

[0026] CUDA vectorization: CUDA vectorization refers to the process of using CUDA's vector data types and instructions to improve performance. The main purpose of vectorization is to reduce the number of memory accesses and improve instruction throughput.

[0027] Shape,The shape of a tensor is a tuple of integers, indicating the size of the tensor in each dimension, which is called shape in this paper.

[0028] The stride of a tensor is a tuple of integers of the same length as its shape. It indicates the number of elements to skip when crossing a single dimension in memory. In other words, the stride indicates how to calculate the next element's location from the memory address when accessing a specific dimension of the tensor. This article refers to the stride as the step size.

[0029] Ndim, the dimension of the tensor.

[0030] idx of cuda kernel: including threadIdx, blockIdx In CUDA programming, grid refers to thread grid, which is a two-dimensional or three-dimensional layout consisting of multiple thread blocks.

[0031] blockIdx is also a dim3 type variable, which represents the index of the current thread block in the entire thread grid.

[0032] threadIdx is a dim3 type variable that represents the index of the current thread within the thread block to which it belongs.

[0033] vect_num: The number of data to be processed in vectorized batches.

[0034] Please refer to Figure 1 , Figure 1 This is a flow chart of a data access method in an embodiment of the present application, which includes the following steps.

[0035] S101: Obtain the input of the target operator, and if the input meets the broadcast requirements, perform broadcast settings on the input.

[0036] The target operator can be an operator that needs to be broadcast or an operator that does not need to be broadcast. The target operator can be used in scenarios that actually require corresponding operator calculations, such as speech recognition, natural language processing, semantic recognition, and face recognition.

[0037] When the target operator meets the broadcast requirements, you can set the input to broadcast. Broadcast setting means expanding the operator input according to the broadcast calculation rules.

[0038] In a specific embodiment of the present application, when the input satisfies broadcasting, broadcasting is set for the input, including: if the input includes a first input and a second input, determining whether the first input and the second input satisfy broadcasting; if so, setting the stride of the broadcast dimension corresponding to the first input or the second input to 0. That is, determining whether broadcasting is satisfied based on the shape of the input, and if so, setting stride[i] of the broadcast dimension dim_broadcast to 0.

[0039] For the two inputs of the target operator, that is, two tensors, broadcasting is possible when the following conditions are met: 1. Shape alignment: Starting from the last dimension (rightmost), the comparison is carried out dimension by dimension. Each dimension must satisfy the following conditions: the dimension sizes are equal, or the dimension size of one of the tensors is 1, or one of the tensors does not have this dimension (that is, the number of dimensions is small).

[0040] 2. At least one dimension of the tensor needs to be expanded: if the shapes are exactly the same, no broadcasting is required.

[0041] Specifically, such as Figure 2 As shown, first determine the shapes of the two inputs. If the input shapes are unequal and one of them has a shape of 1, then set the stride of that dimension to 0. If the lengths of the two shapes are unequal, then extend the shorter shape to match the longer one, and set the extra shape to 1. For example, if the first input corresponds to shape1 = [2, 3, 4] and the second input corresponds to shape2 = [4], then extend shape2 to shape1 and set the extra shape to 1, so shape2 becomes [1, 1, 4].

[0042] S102: Obtain the tensor memory layout features of the target operator, and use the tensor memory layout features to determine the dimension types corresponding to the input and output respectively.

[0043] The tensor of the target operator includes input and output.

[0044] By obtaining the tensor memory layout characteristics of the target operator, the tensor memory layout characteristics can be used to determine the dimension types corresponding to the input and output respectively.

[0045] Specifically, the memory layout characteristics of the input and output tensors can be analyzed separately, starting from the innermost dimension and analyzing the continuity outward to determine the dimension type.

[0046] In a specific implementation of the present application, the dimension types corresponding to the input and output are determined respectively using the tensor memory layout characteristics, including: using the tensor memory layout characteristics to determine the memory continuity of the input; and using the memory continuity to determine the dimension type of each dimension in the input.

[0047] Among them, memory continuity corresponds to the step size, and memory continuity is used to determine the dimension type of each dimension in the input, including: if the step size of the current dimension is 0, then the dimension type of the current dimension is determined to be a broadcast dimension; if the step size of the current dimension is 1, then the dimension type of the current dimension is determined to be a continuous dimension; if the step size of the current dimension is equal to the product of the size of the next dimension and the step size, then the dimension type of the current dimension is determined to be a continuous dimension; if the step size of the current dimension is not 0, not 1, and not the product of the size of the next dimension and the step size, then the dimension type of the current dimension is determined to be a strided dimension.

[0048] In a specific embodiment of the present application, different step length determination methods are used for different tensors. Specifically, for a contiguous memory tensor (Contiguous Tensor), the step length can be calculated by the following rule: starting from the innermost dimension (the last dimension), the initial step length is 1. Calculate layer by layer outward: s[i]=s[i+1]×shape[i+1]. For a non-contiguous memory tensor (Non-Contiguous Tensor), if the tensor undergoes operations such as transposition and slicing, the step length will be rearranged and may no longer be continuous. The step length determination method can be obtained by executing the following code: Python; copy; C = torch.rand(3, 4).transpose(0, 1) # shape (4, 3); print(C.stride()) # Output (1, 4).

[0049] After obtaining the stride, the memory layout characteristics of the input and output tensors are analyzed separately, starting from the innermost dimension and analyzing the continuity outward. The rules for determining the dimension type include: if the stride is 0, it corresponds to the broadcast dimension (B); if the stride is 1, it corresponds to the continuous dimension (C); if the stride is equal to the size of the next dimension × the stride, it corresponds to the continuous dimension (C); if the stride corresponds to other cases, the corresponding stride dimension (S) is determined.

[0050] Please refer to Figure 3 , the continuity of the tensor memory layout characteristics is analyzed from the innermost dimension outward. Regarding the explanation of the inner and outer layers, the following uses the stride=[20,5,1] as an example to illustrate how to determine the dimension type.

[0051] The stride of the innermost layer is strides[2] = 1; the stride of the middle layer is strides[1] = strides[2] * shape[2] = 5; the stride of the outermost layer is strides[0] = strides[1] * shape[1] = 20. Starting from the highest dimension (i.e. the innermost layer), if the strides[i] of the dimension is 0, the dimension type dim_type is set to B(Broacast); if the strides[i] of the dimension is 1, the dimension type dim_type is set to C(Contiguous); if the stride is equal to the size of the next dimension × the stride, that is, strides[i] == shape[i+1] * strides[i+1], the dimension type is set to dim_type set to C(Contiguous); in other cases, corresponding to the stride (S), the dimension type is set to dim_type set to S(Stride).

[0052] S103 , combining the dimension types to merge the continuous dimensions, sort the stride dimensions, and merge the broadcast dimensions of the memory spaces corresponding to the input and output, and then reorganizing the memory space in segments.

[0053] After determining the dimension types of the input and output dimensions, the memory spaces corresponding to the input and output can be merged and sorted based on the corresponding dimension types. After merging and sorting, the memory spaces are reorganized in segments. This allows fragmented memory spaces to be merged into larger segments, allowing similar vectors to be stored in the same memory for easier data access.

[0054] In a specific implementation of the present application, the memory spaces corresponding to the input and output are respectively merged in continuous dimensions, sorted in stride dimensions, and merged in broadcast dimensions in combination with the dimension type, including: if the dimension type is a continuous dimension, and the correspondence between the input and output is not changed after merging and rearranging, the memory spaces corresponding to the input and output are merged; if the dimension type is a stride dimension, the memory spaces corresponding to the input and output are arranged in ascending order according to the original step size; if the dimension type is a broadcast dimension, and the correspondence between the input and output is not changed after merging and reorganizing, the memory spaces corresponding to the input and output are merged.

[0055] For details, please refer to Figure 4 After memory analysis, the memory layout will have dimension types. Based on the dimension types, all input and output data will be merged and rearranged from the outside to the inside. Specific processing operations include the following.

[0056] If the dimension type is C, continuous dimensions are merged, thereby merging physically continuous dimensions into larger calculation blocks.

[0057] If the dimension type is S, sort across the stride dimension. Sort by stride from smallest to largest to improve memory access locality.

[0058] If the dimension type is B, the broadcast dimensions are merged. By merging adjacent broadcast dimensions, conditional judgments can be reduced.

[0059] It should be noted that during the merging and rearrangement process, in order to ensure that the corresponding relationship between input and output is not changed after the merging and rearrangement, the corresponding relationship between input and output must be checked. If the corresponding relationship is changed, the merging and rearrangement will not be performed.

[0060] For example: First, the dimensions can be grouped according to the dimension type dim_type, and each group has a corresponding dim subscript.

[0061] For example, if vect_C stores the subscripts of dimensions that are accessed continuously in memory, vect_S stores the subscripts of dimensions that are accessed non-contiguously, and vect_B stores the subscripts of dimensions with a shape size of 1 and a stride of 0, such as shape=[5,4,1], stride=[4,1,0], and dim_type is of type [C,C,B], then the value of vect_C is [0,1], vect_S is empty, and vect_B is [2].

[0062] Merge consecutive dimensions, traversing vect_C from the inside out. The merging condition is: adjacent consecutive dimensions satisfy stride[i] == stride[i+1] * size[i+1]. To ensure that the correspondence between input and output is not changed after merging and rearrangement, the three input and output vectors must all satisfy stride[i] == stride[i+1] * size[i+1]. For example, corresponding to a+b=c, a's shape=[3,4,5], stride=[20,5,1], b's shape=[3,4,1], stride=[12,1,0], c's shape=[3,4,5], stride=[20,5,1], a's vect_C value is [0,1], b's vect_C value is [0,1], c's vect_C value is [0,1], then a, b, c satisfy stride[0] == stride[1] * size[1], the dimensions of a, b, c can be merged, if b's shape=[3,4,1], stride=[20,1,0], b's vect_C value is [1], it does not satisfy stride[i] == stride[i+1] * size[i+1], then the dimensions of a, b, c cannot be merged.

[0063] After merging, shape is updated to the product of the shapes corresponding to the vect_C subscripts, stride is the minimum value of the stride corresponding to the vect_C subscripts, and the vect_C subscript is updated to the minimum value of the corresponding subscripts.

[0064] The merged dimensions are shape=[12,5], stride=[5,1] of a, shape=[12,1], stride=[1,0] of b, and shape=[12,5], stride=[5,1] of c.

[0065] For stride dimension sorting, traverse vect_S from the inside out. You can sort in ascending order of the original stride length (small stride length first) to optimize locality.

[0066] For broadcast dimension merging, traverse vect_B from the inside out. Merge all broadcast dimension dimensions. At this time, to ensure that the correspondence between the input and output is not changed after merging and rearrangement, the relationship between the three input and output vectors must also be checked. For example, for a+b=c, a's shape=[3,4,5,6], stride=[120,30,6,1], b's shape=[3,4,1,1], stride=[12,1,0,0], c's shape=[3,4,5,6], stride=[120,30,6,1], a's vect_C value is empty, b's vect_B value is [2,3], and c's vect_C value is []. At this time, a, b, and c satisfy stride[2] == stride[3] * size[2], and the dimensions of a, b, and c can be merged. After merging, update the shape of vect_C to the product of the shapes of the corresponding vect_C subscripts, and the stride to the minimum value of the stride of the corresponding vect_C subscripts. Update the subscript of vect_C to the minimum value of the corresponding subscript. Update the shape of vect_B to the product of the shapes of the corresponding vect_B subscripts, multiply the size of the broadcast dimension (size * = next_size), keep the stride at 0, and update the subscript of vect_B to the minimum value of the corresponding subscript. The merged dimensions are: a's shape = [3, 4, 30], stride = [120, 30, 1], b's shape = [3, 4, 1], stride = [4, 1, 0], and c's shape = [3, 4, 30], stride = [120, 30, 1].

[0067] In one embodiment of the present application, segmented memory space reorganization includes: sorting the memory segments corresponding to the merged continuous dimension, the sorted stride dimension, and the merged broadcast dimension according to priority; after completing the memory segment sorting, the memory segments are assembled. The final dimension order can be assembled in the order of continuous → stride → broadcast. The merged and reordered dimension type becomes (C, S, B), and the corresponding merged shape and stride are obtained.

[0068] That is, the dimensions are assembled sequentially and arranged according to priority, and the resulting new memory space is: [merged continuous dimensions] → [sorted stride dimensions] → [merged broadcast dimensions].

[0069] S104: Determine the memory segment to be accessed using the data access instruction, and access the tensor data of the target operator using a data access method adapted to the memory segment.

[0070] In this application, after operations such as reorganization and sorting, the memory space is segmented into [merged continuous dimension], [sorted stride dimension], and [merged broadcast dimension]. Therefore, the optimal data access method can be set for each memory segment.

[0071] Specifically, when a target operator needs to be called or computed, it is considered to have received a data access instruction. This data access instruction can determine the memory segment to be accessed. By using a data access method that is compatible with this memory segment, the corresponding tensor data can be quickly accessed. Accordingly, data access instructions can be divided into data storage instructions and data read instructions, corresponding to different instructions for writing or reading data.

[0072] In a specific embodiment of the present application, the tensor data of the target operator is accessed using a data access method adapted to the memory segment, including: if the memory segment is continuous memory, the tensor data is accessed in batches from the memory segment in a vectorized manner; if the memory segment is strided memory, the tensor data is accessed from the memory segment in a centralized manner according to an offset manner; if the memory segment is broadcast memory, the tensor data is reused in a data reuse manner.

[0073] The batch accessing of tensor data from the memory segment in a vectorized manner includes: calling a vectorized function to batch access tensor data from the memory segment.

[0074] Specifically, segmented kernel processing is used to calculate operators based on the input and output memory after memory reordering. When the dimension type is C, data is accessed in a vectorized manner. When the dimension type is S, data is accessed by calculating the offset. When the dimension type is B, data is reused.

[0075] That is, for the memory segment of the vector corresponding to the dimensions contained in vect_C, vectorized calculations are performed on both input and output. For the memory segment of the vector corresponding to the dimensions contained in vect_S, data is accessed by calculating offsets and calculations are performed on the data. For the memory segment of the vector corresponding to the dimensions contained in vect_B, data is reused for the input of vect_B that is not empty, and vectorized operations are performed on the other input and output.

[0076] Applying the method provided in the embodiment of the present application, first, a broadcast setting is performed for the input of the target operator. Then, the dimension types corresponding to the input and output are determined in combination with the tensor memory layout characteristics of the target operator. In this way, the memory corresponding to the input and output can be merged and rearranged in combination with the dimension type. After completing the memory merging and rearrangement, the memory space is assembled in segments. In this way, after receiving the data access instruction, the data access instruction can be used to determine the memory segment to be accessed, and the tensor data of the target operator can be accessed using a data access method adapted to the memory segment, thereby accelerating the acquisition of tensor data.

[0077] This application has the technical effect of transforming the tensor memory space of the target operator from scattered and disorganized to segmented assembled memory space, so that the memory arrangement pattern of the same dimensional type is stored. When accessing data, a data access method adapted to the memory segment is adopted, which can quickly access the corresponding tensor data and further accelerate the operator's computing efficiency.

[0078] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0079] Please refer to Figure 5 The embodiment of the present application further provides a data access device, the device comprising: The broadcast setting module 101 is used to obtain the input of the target operator and perform broadcast setting on the input if the input meets the broadcast requirements; A memory layout analysis module 102 is configured to obtain tensor memory layout features of a target operator and determine the dimension types corresponding to the input and output respectively using the tensor memory layout features; The memory layout merging and rearrangement module 103 is used to merge the continuous dimensions, sort the strided dimensions, and merge the broadcast dimensions of the memory spaces corresponding to the input and output respectively according to the dimension type, and then reorganize the memory space in segments; The operator computing module 104 is configured to determine the memory segment to be accessed using a data access instruction, and access the tensor data of the target operator using a data access method adapted to the memory segment.

[0080] When applying the device in this application, first, a broadcast setting is performed for the input of the target operator. Then, the dimension types corresponding to the input and output are determined in combination with the tensor memory layout characteristics of the target operator. In this way, the memory corresponding to the input and output can be merged and rearranged in combination with the dimension type. After completing the memory merging and rearrangement, the memory space is assembled in segments. In this way, after receiving the data access instruction, the data access instruction can be used to determine the memory segment to be accessed, and the tensor data of the target operator can be accessed using a data access method adapted to the memory segment, thereby accelerating the acquisition of tensor data.

[0081] This application has the technical effect of transforming the tensor memory space of the target operator from scattered and disorganized to segmented assembled memory space, so that the memory arrangement pattern of the same dimensional type is stored. When accessing data, a data access method adapted to the memory segment is adopted, which can quickly access the corresponding tensor data and further accelerate the operator's computing efficiency.

[0082] In a specific implementation of the present application, the operator computing module is specifically used to, if the memory segment is continuous memory, batch access tensor data from the memory segment in a vectorized manner; if the memory segment is strided memory, centrally access tensor data from the memory segment in an offset manner; if the memory segment is broadcast memory, reuse tensor data in a data reuse manner.

[0083] In a specific implementation of the present application, the operator computing module is specifically used to call a vectorized function and access tensor data in batches from a memory segment.

[0084] In a specific embodiment of the present application, the broadcast setting module is specifically used to determine whether the first input and the second input meet the broadcast requirements if the input includes a first input and a second input; if so, the step size of the corresponding dimension of the first input or the second input broadcast is set to 0.

[0085] In a specific implementation of the present application, the memory layout analysis module is specifically used to determine the memory continuity of the input using the tensor memory layout characteristics; and to determine the dimension type of each dimension in the input using the memory continuity.

[0086] In a specific implementation of the present application, memory continuity corresponds to the step size, and the memory layout analysis module is specifically used to determine that the dimension type of the current dimension is a broadcast dimension if the step size of the current dimension is 0; if the step size of the current dimension is 1, the dimension type of the current dimension is determined to be a continuous dimension; if the step size of the current dimension is equal to the product of the next dimension size and the step size, the dimension type of the current dimension is determined to be a continuous dimension; if the step size of the current dimension is not 0, not 1 and not the product of the next dimension size and the step size, the dimension type of the current dimension is determined to be a strided dimension.

[0087] In a specific implementation of the present application, the memory layout merging and reordering module is specifically used to sort the memory segments corresponding to the merged continuous dimension, the sorted stride dimension and the merged broadcast dimension according to priority; after completing the memory segment sorting, the memory segments are assembled.

[0088] In a specific implementation of the present application, the memory layout merging and reordering module is specifically used to merge the memory space corresponding to the input and output if the dimension type is a continuous dimension and the corresponding relationship between the input and output is not changed after the merging and reordering; if the dimension type is a strided dimension, the memory space corresponding to the input and output is arranged in ascending order according to the original step size; if the dimension type is a broadcast dimension and the corresponding relationship between the input and output is not changed after the merging and reorganization, the memory space corresponding to the input and output is merged.

[0089] For the description of the features in the embodiment corresponding to the data access device, reference can be made to the relevant description of the embodiment corresponding to the data access method, which will not be repeated here.

[0090] To facilitate those skilled in the art to implement the data access method and device provided in the embodiments of the present application, the data access method and device are described in detail below in conjunction with pseudo code required for actual application.

[0091] In actual applications, corresponding pseudocode can be set for each module to achieve corresponding functions.

[0092] Take the add operator a+b=c as an example, where a and b are inputs and c is the output. First, call broadcast_setup (the function corresponding to the broadcast setup module) to set the input broadcasting, setting the stride of a and b to 0. Then, call analyze_layout (the function corresponding to the memory layout analysis module) to analyze the memory layout of a, b, and c, marking the dimension types: continuous type is marked as C, stride type is marked as S, and broadcast type is marked as B. Then, call reorganize_dims (the function corresponding to the memory layout merging and rearrangement module) to rearrange the memory of a, b, and c while maintaining the corresponding relationship between the elements in a, b, and c. The final memory layout is [continuous, stride, broadcast]. Then, call vectorized_func (the function corresponding to the operator calculation module) to perform the calculation on a, b, and c. The calculation process is segmented according to the memory layout.

[0093] Specifically, if the segment is continuous memory, data is read and written in batches using vectorization, and the vectorized function cal_func_vectorized (another function name corresponding to the operator calculation module) is called to calculate the data. If the segment is strided memory, the offset of each element is calculated, and the values of a and b are obtained based on the offset. The value of a + b is written to c based on the offset. If the segment is broadcast memory, assuming that the memory segment of b is broadcast memory, the elements of b are reused, and the elements of a and c are read and written using vectorization, and the vectorized function cal_func_vectorized is called to calculate the data.

[0094] Corresponding to the above method embodiment, an embodiment of the present application further provides an electronic device. The electronic device described below and the data access method described above can refer to each other.

[0095] See also Figure 6 As shown, the electronic device includes: Memory 332, for storing computer programs; The processor 322 is configured to implement the steps of the data access method of the above method embodiment when executing a computer program.

[0096] For details, please refer to Figure 7 , Figure 7 This is a schematic diagram of the specific structure of an electronic device provided in this embodiment. This electronic device may vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) (for example, one or more processors) and memory 332. The memory 332 stores one or more computer programs 342 or data 344. The memory 332 may be temporary storage or permanent storage. The program stored in the memory 332 may include one or more modules (not shown), each of which may include a series of instruction operations in the data processing device. Furthermore, the processor 322 may be configured to communicate with the memory 332 to execute the series of instruction operations in the memory 332 on the electronic device 301.

[0097] The electronic device 301 may further include one or more power supplies 326 , one or more wired or wireless network interfaces 350 , one or more input / output interfaces 358 , and / or one or more operating systems 341 .

[0098] The steps in the data access method described above can be implemented by the structure of an electronic device.

[0099] Corresponding to the above method embodiments, embodiments of the present application further provide a readable storage medium. The readable storage medium described below and the data access method described above can be referenced in correspondence with each other. Embodiments of the present application further provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above data access method embodiments when executed.

[0100] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0101] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above data access method embodiments are implemented.

[0102] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned data access method embodiments are implemented.

[0103] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0104] Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art may make several improvements and modifications to this application without departing from the principles of this application, and such improvements and modifications also fall within the scope of protection of this application.

Claims

1. A data access method, characterized in that: include: Obtain the input of the target operator, and if the input satisfies broadcast requirements, perform broadcast settings on the input; Obtaining a tensor memory layout feature of the target operator, and using the tensor memory layout feature to determine the dimension types corresponding to the input and output respectively; After performing continuous dimension merging, stride dimension sorting, and broadcast dimension merging on the memory spaces corresponding to the input and the output respectively in combination with the dimension type, the memory space is reorganized in segments; A data access instruction is used to determine an accessed memory segment, and a data access method adapted to the memory segment is used to access the tensor data of the target operator.

2. The method according to claim 1, characterized in that Accessing the tensor data of the target operator using a data access method adapted to the memory segment includes: If the memory segment is continuous memory, accessing the tensor data in batches from the memory segment in a vectorized manner; If the memory segment is a strided memory, the tensor data is centrally accessed from the memory segment according to the offset method; If the memory segment is a broadcast memory, the tensor data is reused in a data multiplexing manner.

3. The method according to claim 2, characterized in that Accessing the tensor data in batches from the memory segment in a vectorized manner includes: Call a vectorized function to access the tensor data in batches from the memory segment.

4. The method according to claim 1, wherein If the input satisfies broadcast requirements, performing broadcast settings on the input includes: If the input includes a first input and a second input, determining whether the first input and the second input satisfy broadcast requirements; If so, the stride of the corresponding dimension of the first input or the second input broadcast is set to 0.

5. The method according to claim 1, wherein The tensor memory layout characteristics are used to determine the dimension types corresponding to the input and output, respectively, including: Determining memory continuity of the input using the tensor memory layout characteristics; The memory continuity is used to determine the dimension type of each dimension in the input.

6. The method according to claim 5, characterized in that The memory continuity corresponds to a step size, and the dimension type of each dimension in the input is determined by using the memory continuity, including: If the step size of the current dimension is 0, the dimension type of the current dimension is determined to be a broadcast dimension; If the step size of the current dimension is 1, the dimension type of the current dimension is determined to be a continuous dimension; If the stride of the current dimension is equal to the product of the size of the next dimension and the stride, the dimension type of the current dimension is determined to be a continuous dimension; If the stride of the current dimension is neither 0 nor 1 and is not the product of the size of the next dimension and the stride, the dimension type of the current dimension is determined to be a strided dimension.

7. The method according to claim 1, characterized in that Segmenting and reorganizing the memory space includes: Sort the memory segments corresponding to the merged continuous dimensions, sorted stride dimensions, and merged broadcast dimensions according to priority; After the memory segments are sorted, the memory segments are assembled.

8. The method according to any one of claims 1 to 7, characterized in that Combining the dimension types, performing continuous dimension merging, stride dimension sorting, and broadcast dimension merging on the memory spaces corresponding to the input and the output, respectively, including: If the dimension type is a continuous dimension and the correspondence between the input and output is not changed after merging and rearranging, then the memory spaces corresponding to the input and the output are merged; If the dimension type is a stride dimension, the memory spaces corresponding to the input and the output are arranged in ascending order according to the original stride length; If the dimension type is a broadcast dimension and the correspondence between the input and output is not changed after merging and reorganizing, the memory spaces corresponding to the input and the output are merged.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data access method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the data access method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Model operation optimization method, product, equipment and medium

    CN118277133A

  • Data processing method, computer program product, equipment and computer medium

    CN119166467A

  • Multi-dimensional data processing method and device, equipment and medium

    CN119271274A

  • Broadcast volume adaptive adjustment method and device, equipment, medium and product

    CN119420442A

  • Large model cross-platform system based on micro operator

    CN119718484A