Cross-border non-aligned memory access staticizing method applied to SIMD architecture

By extracting the dimensional characteristic parameters of the SIMT architecture to perform non-aligned compensation repair and reconstruct the memory access mode, the non-aligned memory access problem between the SIMT architecture and the SIMD architecture is solved, and the data throughput efficiency and computing performance of the SIMD processor are improved.

CN120743849APending Publication Date: 2025-10-03SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510829792.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing technology has the problem of non-aligned memory access when processing memory access between SIMT architecture and SIMD architecture, resulting in low memory access efficiency, especially significant performance loss when processing irregular data structures.

Method used

By extracting multiple dimensional characteristic parameters of the memory access mode under the SIMT architecture, performing modulus and integer division alignment verification, performing non-aligned dimension compensation and repair, reconstructing the memory access pattern, converting non-aligned access into continuous structured data block reading, and using SIMD vectorized loading instructions for optimization.

Benefits of technology

In non-aligned scenarios, the memory access conversion from SIMT architecture to SIMD architecture is realized, which significantly improves the data throughput efficiency of SIMD processor and enhances computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743849A_ABST
    Figure CN120743849A_ABST
Patent Text Reader

Abstract

The invention relates to the field of AI compilers, in particular to a cross-border non-aligned memory access staticizing method applied to an SIMD architecture, which comprises the following steps of: extracting a plurality of dimension characteristic parameters of a memory access mode under the SIMT architecture; carrying out dimension grouping on the extracted multiple dimension characteristic parameters according to the original calculation dimension; traversing each dimension feature parameter in each dimension feature group, and performing modulo alignment verification and integer division alignment verification; when the modulo alignment verification result or the integer division alignment verification result is a non-alignment mode, performing non-alignment dimension compensation repair operation; vectorized memory access under the SIMD architecture is carried out on the memory access matrix, and an over-reading matrix is generated according to the expanded dimension parameters; and carrying out dimension reduction reconstruction on the over-reading matrix, and executing reverse dimension segmentation reduction to obtain an access matrix meeting the SIMD architecture calculation instruction requirements. According to the method and the device, memory access conversion from the SIMT architecture to the SIMD architecture can be realized in a non-aligned scene, the problem of non-aligned memory access in the SIMT architecture is effectively solved, and the memory access efficiency of a computing system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI compilers, and in particular to a method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture. Background Art

[0002] Amidst the accelerating iteration of artificial intelligence technology, the coordinated evolution of hardware architecture and software ecosystems is facing unprecedented challenges. The number of parameters in deep learning models is doubling annually, while algorithm iteration cycles have been shortened to 3-6 months. This exponential technological evolution has exposed the efficiency bottlenecks of traditional, hand-coded heterogeneous computing. Especially during this critical window of hardware architecture transformation, the parallel computing paradigm is undergoing a strategic migration from a SIMT (single instruction, multiple threads)-dominated model to a SIMD (single instruction, multiple data) architecture. This paradigm shift in the underlying architecture places revolutionary demands on compilation technology.

[0003] The compilation dilemma caused by architectural differences is particularly acute in the memory management dimension. The core feature of the SIMT architecture, exemplified by NVIDIA GPUs, is implicit data parallelism achieved through warp scheduling, giving each thread independent memory access paths. While this design provides greater program flexibility, it also results in significant loss of spatial and temporal locality in memory access patterns. In contrast, SIMD architectures (such as the Ascend 910 or AMD CDNA) require strictly aligned memory access patterns in their vectorized execution units, requiring the compiler to explicitly manage data layout to enable batch loading of contiguous memory blocks. The fundamental differences in memory access semantics between these two architectures result in significant performance losses in traditional instruction translation schemes in real-world scenarios, particularly when processing irregular data structures, where memory access efficiency can degrade by orders of magnitude. The MLIR (Multi-Level Intermediate Representation) framework offers a new solution to this dilemma. By constructing a hierarchical and progressive intermediate representation system (from high-level algorithm descriptions to low-level hardware instructions), this framework achieves dimensional decoupling of the compilation optimization process. Its unique dialect extension mechanism allows developers to build an intermediate abstraction layer for specific computing paradigms on top of traditional intermediate representations such as LLVM-IR. This technical feature provides systematic support for cross-architecture instruction conversion by establishing an architecture-independent memory access pattern description in the memory access abstraction layer, and then mapping it to specific hardware features through multi-level progressive optimization. However, existing implementations still have core defects such as insufficient vectorization and excessive data reorganization overhead when dealing with the cross-thread memory interleaving access pattern unique to the SIMT architecture. Therefore, an intelligent conversion method that can deeply understand the memory access characteristics of the two architectures is needed to improve memory access efficiency. Summary of the Invention

[0004] In order to solve the technical problems existing in the prior art, the present invention provides a static method for out-of-bounds non-aligned memory access applied to the SIMD architecture, which can realize the conversion of SIMT architecture to SIMD architecture memory access in non-aligned scenarios, effectively solve the non-aligned memory access problem existing in the SIMT architecture, and improve the memory access efficiency of the computing system.

[0005] The purpose of the present invention can be achieved by taking the following technical solutions:

[0006] A method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture, the method comprising:

[0007] S1. Extract multiple dimensional feature parameters of the memory access mode under the SIMT architecture based on the operators existing in pointer operations. The multiple dimensional feature parameters include offset value, step size, size, number of repeated elements in the dimension, and original calculation dimension;

[0008] S2. The extracted multiple dimensional feature parameters are grouped according to the original calculation dimensions to obtain multiple dimensional feature groups. Within each dimensional feature group, a monotonically increasing sorting is performed according to the number of repeated elements in the dimension.

[0009] S3. Traverse each dimensional feature parameter in each dimensional feature group, perform modulus alignment verification and integer divisibility alignment verification, and obtain modulus alignment verification results and integer divisibility alignment verification results respectively;

[0010] S4. When the modulo alignment verification result or the divisibility alignment verification result is a non-aligned mode, a non-aligned dimension compensation repair operation is performed;

[0011] S5. Perform vectorized memory access on the memory access matrix under the SIMD architecture, generate an excess read matrix according to the expanded dimension parameters, and use SIMD vectorized load instructions to perform stride access optimization;

[0012] S6. Reduce the dimension of the excess read matrix based on the original computing dimension, fold the expanded dimension into a continuous one-dimensional vector, perform reverse dimension splitting and restoration according to the dimension priority, and obtain a memory access matrix that meets the requirements of the SIMD architecture computing instructions.

[0013] Specifically, the operators existing in the pointer operation include: dimension information related operators, shape information related operators, and other operators; dimension information related operators include: vector generation operators, dimension expansion operators, scalar vectorization operators, and broadcast operators; shape information related operators include: modulo, integer division, multiplication, and addition without dimension expansion; other operators include addition operations after integer division / modulo operations.

[0014] Specifically, based on the operators in pointer operations, the dimension characteristic parameters of the memory access mode under the SIMT architecture are extracted, including:

[0015] Initialize the five dimensional feature parameters through the vector generation operator: offset value, step size, size, number of repeated elements in the dimension, and the original calculation dimension. Initialize the step size to 1, the number of repeated elements and the original dimension to 0; initialize the offset value and the size of the current dimension to the offset value and initial size of the vector respectively;

[0016] The offset values ​​of the left and right operands are calculated according to the operator through the modulo, integer division, multiplication, and addition operations without dimension expansion of the shape information related operators; the multiplication and addition operations multiply or add the stride lengths corresponding to the operands on both sides respectively, and the number of repeated elements in the dimension is the second operand in the integer division operation;

[0017] The size is initialized by the length of the vector generation operator and updated by the second operand of the modulo operation.

[0018] The initial default value of the original calculation dimension is 0. When the dimension expansion operator is executed to expand the dimension, if the expanded dimension is less than or equal to the current dimension, the original calculation dimension is incremented by 1.

[0019] Specifically, step S3 includes:

[0020] S31. Perform modulus alignment verification, traverse the feature parameters of each dimension in the dimensional feature group, and sequentially verify the alignment of the modulus dimension size and the reading unit size in each dimension; if the modulus dimension size and the reading unit size are integer multiples, it is determined to be an aligned mode; otherwise, it is considered to be a non-aligned mode, and a modulus alignment verification result is obtained;

[0021] Perform modulo alignment verification, traverse the feature parameters of each dimension in the dimensional feature group, and align the number of low-dimensional repeated elements and the reading unit size in each dimension in turn; when the number of low-dimensional repeated elements and the reading unit size are integer multiples, it is determined to be an alignment mode, otherwise it is considered a non-alignment mode, and the integer alignment verification result is obtained.

[0022] Specifically, the step S4 includes:

[0023] S41. When the modulo dimension size is not aligned with the read unit size, the original offset is recorded, the excess read data is trimmed according to the original offset, a new offset after compensation is calculated, and the starting offset value of this unaligned memory access is adjusted to the new offset after compensation;

[0024] S42. When the divisible dimension representation parameter is not aligned with the reading unit size, the complementary dimension scale is calculated according to the length of a static read and the number of repeated elements in the dimension, and the candidate dimension is generated according to the candidate dimension scale and the modulo dimension size.

[0025] Specifically, step S6 includes:

[0026] S61. Perform dimensionality reduction and reconstruction on the excess read matrix based on the original computing dimension metadata, folding the extended dimension into a continuous one-dimensional vector so that the data arrangement pattern in the one-dimensional dimension is consistent with the expected memory access order in the SIMT architecture.

[0027] S62. Calculate the difference Δoffset between the original offset and the new offset, perform cascade splitting from high dimension to low dimension, use the difference Δoffset as the base offset, perform sliding window splitting along each dimension according to the read unit size, and output a memory access matrix that meets the requirements of the SIMD architecture computing instructions.

[0028] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0029] The present invention proposes a method for staticizing out-of-bounds non-aligned memory access for SIMD architecture. By reconstructing the memory access pattern, discrete non-aligned access requests are integrated into continuous structured data block readings, and the original irregular element-by-element access pattern is converted into batched vector access instructions that conform to SIMD characteristics. It can realize the conversion of SIMT architecture to SIMD architecture memory access in non-aligned scenarios, effectively solve the non-aligned memory access problem existing in SIMT architecture, significantly improve the data throughput efficiency of SIMD processors, and ultimately achieve an improvement in overall computing performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0031] Figure 1 A specific flow chart of a method for statically accessing out-of-bounds non-aligned memory for SIMD architecture according to an embodiment of the present invention

[0032] Figure 2 Schematic diagram of two different high-dimensional lengths with modulo out-of-bounds non-alignment in an embodiment of the present invention;

[0033] Figure 3 Schematic diagram of two different high-dimensional lengths that are divisible out of bounds and non-aligned in an embodiment of the present invention;

[0034] Figure 4 This is a schematic diagram of offset value alignment adjustment in an embodiment of the present invention;

[0035] Figure 5This is a schematic diagram of converting a non-aligned modulo out-of-bounds access into aligned SIMD memory access in an embodiment of the present invention;

[0036] Figure 6 Schematic diagram of converting an out-of-bounds non-aligned integer division into aligned SIMD memory access in an embodiment of the present invention;

[0037] Figure 7 Schematic diagram of the reverse dimensional segmentation and restoration steps in an embodiment of the present invention. DETAILED DESCRIPTION

[0038] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It is obvious that the embodiments described are only some embodiments of the present invention, not all embodiments, and the implementation of the present invention is not limited to these. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] Example 1:

[0040] The present invention discloses a method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture. The method is suitable for vectorized acceleration of out-of-bounds reads under a SIMD architecture and can convert memory accesses in the case of non-aligned out-of-bounds reads into static vectorized memory accesses. Using this method to perform matrix multiplication operations on a heterogeneous parallel computing system of CPU+NPU can improve the overall computing efficiency of the system.

[0041] like Figure 1 As shown, the present invention provides a static method for out-of-bounds non-aligned memory access applied to a SIMD architecture, comprising the following steps:

[0042] S1. Based on the operators existing in pointer operations, the dimension feature parameters of the memory access mode under the SIMT architecture are extracted. The dimension feature parameters include offset value, step size, size, number of repeated elements in the dimension, and original calculation dimension.

[0043] Specifically, the operators in pointer operations include: dimension information related operators, shape information related operators, and other operators. Pointer operations are common operations in SIMT programs, affecting data access patterns and memory access efficiency.

[0044] Dimension information-related operators include: vector generation (make_range), dimension expansion (expand_dim), scalar vectorization (splat), and broadcast. These four types of operators only affect the dimension information of memory access. The vector generation (make_range) operator initializes all eigenvalues, while the other three operations (dimension expansion, scalar vectorization, and broadcast) only modify the dimension information, which remains consistent with the shape of the operator's result. The dimension expansion (expand_dim) operator changes the number of original dimensions and dynamically calculates the original dimension information to which the current dimension belongs based on the position of the inserted dimension.

[0045] Shape information-related operators include modulo (remsi), integer division (divsi), multiplication (mul), and addition without dimension expansion. All four operators change the offset value of the dimension: the previous offset value is used as the left operand, the scalar involved in the operation is used as the right operand, and the operator itself is used as the current operator. The result of the calculation is the new offset value. In addition, multiplication (mul) and addition change the characteristic value of the stride: addition adds the two strides to form the new stride, and multiplication multiplies the stride by another multiplier to form the new stride. Modulo (remsi) and integer division (divsi) may change the size information (size) of the dimension: the right operand of the modulo operation is the new size information of the dimension, and the division operation filters the influence of repeated elements. The original size information is divided by the right operand and rounded up to the integer value as the new size information of the dimension. The number of repeated elements in a dimension (div_num) is bound to the integer division operation (divsi), and its right operand is the number of repeated elements in the dimension.

[0046] Other operators, including addition after integer division or modulo operation, do not adjust the eigenvalues, but directly combine the eigenvalues ​​of the left and right operands to expand a higher-dimensional feature dimension group.

[0047] Specifically, the dimensional characteristic parameters of the memory access mode in the SIMT architecture include: offset (offset), stride (stride), size (size), number of repeated elements in the dimension (div_num), and the original computation dimension (ori_dim). The offset (offset) represents the unit interval between the first memory access unit and the first physical unit in this dimension, identifying the starting position of the memory access. The stride (stride) records the address interval between logically adjacent units in this dimension in physical storage and is used to calculate the physical addresses of subsequent units during vectorized memory access. The size (size) is an implicit upper limit parameter for the dimension size in modulo operations. After dimensionality expansion, this parameter represents the size of the newly added dimension. The number of repeated elements in the dimension (div_num) is a static constant determined at compile time (derived from integer division operations) that represents the number of times a memory address is repeated in this computation. After dimensionality expansion, this parameter represents the number of repetitions of the lower-dimensional address when the higher-dimensional address is fixed. When it is equal to the actual lower-dimensional size, it indicates continuous memory access; when it is greater than the actual lower-dimensional size, it indicates implicit broadcast. The original calculation dimension (ori_dim) identifies the original memory access dimension to which the dimension belongs before expansion. It is used to determine the dimensions that need to be fused during dimensionality reduction processing after dimension expansion.

[0048] Specifically, based on the operators in pointer operations, the offset value, stride, size, number of repeated elements in the dimension, and original calculation dimension of the memory access mode under the SIMT architecture are extracted, including:

[0049] Initialize the five dimensional feature parameters through the vector generation operator: offset value, step size, size, number of repeated elements in the dimension, and the original calculation dimension. Initialize the step size to 1, the number of repeated elements and the original dimension to 0; initialize the offset value and the size of the current dimension to the offset value and initial size of the vector respectively;

[0050] The offset values ​​of the left and right operands are calculated according to the operator through the modulo, integer division, multiplication, and addition operations without dimension expansion of the shape information related operators; the multiplication and addition operations multiply or add the stride lengths corresponding to the operands on both sides respectively, and the number of repeated elements in the dimension is the second operand in the integer division operation;

[0051] The size is initialized by the length of the vector generation operator and updated by the second operand of the modulo operation.

[0052] The initial default value of the original calculation dimension is 0. When the dimension expansion operator is executed to expand the dimension, if the expanded dimension is less than or equal to the current dimension, the original calculation dimension is incremented by 1.

[0053] Table 1 shows the logic for the feature parameter extraction method. Feature extraction is performed for this SIMT memory access based on the operators present in pointer arithmetic. The vector generation operator, make_range, initializes five dimensional feature parameters: offset, stride, size, number of repeated elements in the dimension, and the original computation dimension. The stride is initialized to 1, and the number of repeated elements and the original dimension are initialized to 0. The offset and the dimension size are initialized to the offset and initial size of the vector. The remsi, divsi, mul, and add operations perform the calculations corresponding to the operator on the offset values ​​of the left and right operands. The mul and add operations multiply or add the stride values ​​corresponding to the operands, respectively. The number of repeated elements in the dimension, div_num, is the second operand in the divsi operation. The size is initialized by the length of the vector generation operator, make_range, and then updated by the second operand of the modulo remsi operation. The initial default value of the original calculation dimension ori_dim is 0 (the highest dimension). When the dimension expansion operator expand_dim is executed to expand the dimension, if the expanded dimension is less than or equal to the current dimension, it means that the current dimension is no longer the highest dimension, and the original calculation dimension ori_dim will be incremented by 1.

[0054] Table 1

[0055]

[0056] Taking three cases as an example, the obtained dimensional feature groups [offset, stride, size, div_num, ori_dim] are respectively: The corresponding dimension feature group is [16, 1, 8, 4, 0]. The corresponding dimension feature group is [0, 16, 2, 64, 0]. (index2% 4)×32, the corresponding dimension feature group is [0, 32, 4, 1, 0].

[0057] The integer data type defines the type of pointer operations. For example, all operations mentioned throughout this document are pointer operations. Integers define these operations, while floating-point / integer types define the data pointed to by pointers. Pointers must point to floating-point or integer types, not ptr pointer types, limiting the elements they point to. Matrix feature information includes the current dimension size parameter, lower-dimensional associated dimension parameters, multidimensional step parameters, dimension extension traceability parameters, and dimension offset parameters. The dimension size parameter characterizes the actual space scale allocated to each dimension in the static structured reading mode, which is specifically reflected in the effective data capacity of the corresponding dimension in the memory; the low-dimensional associated size parameter is used to describe the maximum effective size threshold that the preceding dimension can support when eliminating the low-dimensional influence through dimensional reduction; the multi-dimensional step parameter defines the memory address interval between adjacent elements of each dimension, and its value depends on the physical distribution density of the elements of this dimension in the storage space; the dimension expansion traceability parameter records the topological structure information of the original dimension before the dimension expansion operation, which is used to trace the dimensional evolution process of the matrix; the dimension offset parameter quantifies the displacement of the starting position of the data block in the current dimension relative to the base address. This parameter has important characterization significance for non-continuous storage structures. Among them, the dimension size parameter specifically refers to the number of storage units actually occupied by each dimension of the matrix under the static structured memory allocation mechanism; the low-dimensional associated size parameter determines the maximum data scale boundary that the preceding dimension can carry after eliminating the low-dimensional constraints through the dimensionality reduction masking technology.

[0058] S2. The extracted multiple dimensional feature parameters are grouped according to the original calculation dimension to obtain multiple dimensional feature groups. Each dimensional feature group is sorted in a monotonically increasing order according to the number of repeated elements in the dimension.

[0059] Specifically, dimensions are grouped by the original computational dimension ori_dim. Each dimension's feature parameters are divided into different dimension feature groups based on the original dimension. The data in each dimension group does not interfere with each other and is processed independently in the subsequent analysis. Within each dimension group, a strictly monotonically increasing ordering is implemented based on the number of repeated elements (div_num) in the dimension, which serves as the dimension feature value. This ensures that any subsequent division misalignment scenarios are handled with the modulus misalignment analysis of the preceding dimension already completed.

[0060] S3. Traverse each dimensional feature parameter in each dimensional feature group, perform modulo alignment verification and integer alignment verification, and obtain modulo alignment verification results and integer alignment verification results respectively.

[0061] S31. Perform modulus alignment verification, traverse the feature parameters of each dimension in the dimensional feature group, and perform alignment verification on the size of the modulus dimension and the reading unit size in each dimension in turn; when the modulus dimension size and the reading unit size are integer multiples, it is determined to be an alignment mode, otherwise it is regarded as a non-alignment mode, and the modulus alignment verification result is obtained.

[0062] Traverse the feature parameters of each dimension in the dimension feature group, and perform alignment verification on the constants representing the modulus parameter and divisibility parameter in each dimension with the reading unit size of this read. The modulus dimension size is size, and the reading unit size is BLOCK_SIZE. The verification algorithm uses the multiple relationship judgment criterion: if the two values ​​have an integer multiple relationship, the modulus alignment verification result is determined to be aligned mode, otherwise the modulus alignment verification result is considered to be non-aligned mode. Specifically, determine whether the following conditions are all met:

[0063] or

[0064] When this formula is satisfied, it indicates that the current dimension size is a multiple of the read unit size. When size is less than BLOCK_SIZE, the entire access unit is just covered. Otherwise, BLOCK_SIZE is exactly aligned with size. There is no tail block for each read, and no special adjustment is required.

[0065] like Figure 2 As shown, it is a schematic diagram of two different high-dimensional lengths of modulo out-of-bounds non-alignment. In the figure, light gray is aligned data, medium gray is the tail block, and at this time BLOCK_SIZE is 8 and size is 6. There is a dark gray tail block part outside each size length read, which destroys the structure of vectorized memory access. At the same time, the offset value is not actually at 0. The first vector read involves a jump memory access rather than a continuous memory access, so it needs to be repaired and turned into an aligned and non-jump memory access form to improve memory access efficiency. index is a vector with a certain offset value, with a length of 8, which means that this memory access needs to read 8 units. The offset values ​​of these 8 units are shown in the value in index. For Figure 3 The two different index offset values, upper and lower, will result in two different memory access forms even if the length is the same. Both the head and tail of these two memory access forms are in a non-aligned state.

[0066] S32. Perform modulo alignment verification, traverse the feature parameters of each dimension in the dimensional feature group, and perform alignment verification on the number of low-dimensional repeated elements (div_num) and the reading unit size (BLOCK_SIZE) in each dimension in turn; when the number of low-dimensional repeated elements and the reading unit size are integer multiples, it is determined to be an alignment mode, otherwise it is regarded as a non-alignment mode, and the integer alignment verification result is obtained.

[0067] Alignment verification of the number of low-dimensional repeated elements (div_num) and the reading unit size (BLOCK_SIZE);

[0068] Traverse the feature parameters of each dimension in the dimension feature group, and perform alignment verification on the constants representing the modulus parameter and the divisibility parameter in each dimension with the length of this reading. The reading unit size is BLOCK_SIZE, and the divisible dimension representation parameter is div_num. The verification algorithm uses the multiple relationship judgment criterion: if the two values ​​have an integer multiple relationship, the modulus alignment verification result is determined to be aligned mode; otherwise, the modulus alignment verification result is considered to be non-aligned mode. Specifically, determine whether the following conditions are all met:

[0069] or

[0070] When this formula is satisfied, it means that the number of repeated elements and the reading unit size are in a multiple relationship. The repeated elements can just fill the entire reading unit without any tail block, thus satisfying the alignment condition.

[0071] like Figure 3 As shown in the figure, two different high-dimensional length diagrams of integer division out of bounds and non-alignment are shown. Light gray represents aligned data, and medium gray represents the tail block. In this case, BLOCK_SIZE is 8 and divNum is 6, which is a non-aligned case. Since the starting offset value cannot be statically determined at compile time, there are many situations for the size read in the row direction (high dimension). Static vectorized memory access requires a size determined at compile time. It is necessary to obtain the maximum possible size by repairing the size information, thereby staticizing the non-statically compiled value in this memory access.

[0072] S4. When the modulo alignment verification result or the divisibility alignment verification result is a non-aligned mode, a non-aligned dimension compensation repair operation is performed to make the dimension modulo operation meet the alignment requirement.

[0073] S41. When the modulo dimension size is not aligned with the read unit size, the original offset is recorded, the excess read data is trimmed according to the original offset, the new offset after compensation is calculated, and the starting offset value of this non-aligned memory access is adjusted to the new offset after compensation, so that the dimension modulo operation meets the alignment requirements.

[0074] First, record the original offset offset_old, which serves as the data restoration benchmark. Trim the overread data based on the original offset and calculate the new offset after compensation. The formula for calculating the new offset after compensation is as follows:

[0075] offset new =min(offsetold ,BLOCK_SIZE-size)

[0076] Among them, offset new The new offset after compensation represents the calculated new offset value, offset old is the original offset, representing the original offset value; size is the modulo dimension size, representing the size of this dimension; BLOCK_SIZE represents the read unit size. Each memory access reads data of size length. This formula ensures that each read completely covers the size-sized data segment, so that illegal data only appears at the beginning / end of the data block, and the integrity of the core data area is maintained.

[0077] like Figure 4 The figure below shows an offset value alignment adjustment diagram. To ensure that the size of each memory access is a complete unit, when size is not aligned with BLOCK_SIZE, an out-of-bounds misalignment will occur as shown in the bottom of the figure. To ensure that data of size length is read each time, the starting offset value of this misaligned memory access is adjusted to BLOCK_SIZE-size. By over-reading the data on the left side, each vectorized memory access is ensured to be aligned without overflowing tail blocks.

[0078] like Figure 5 As shown, the schematic diagram of the conversion from modulo out-of-bounds non-aligned to aligned SIMD memory access is shown. In the figure, light gray is aligned data, and medium gray is the tail block. For the non-aligned overflow tail block (dark gray part), a part of it will be read in excess so that the number of read units in this memory access just meets the requirement that the number of read units is a positive integer multiple of size (as shown in the second part of the figure). In the result of the modulo operation, the values ​​of the previous row and the next row are completely equivalent. After determining the number of sizes included, the offset value can be adjusted to the alignment (as shown in the third part of the figure). At this time, the read elements include all the units required for this memory access (one 0, 1, 2, two 3, 4, and one 5), and the recorded offset new The order of memory access will be adjusted in subsequent steps. After vectorization over-reads the required data, the offset new The results are rearranged (as shown in the fourth section of the figure) to enable vectorized reads of unaligned units. Information completeness in higher dimensions is ensured by out-of-bounds unaligned division in the next dimension. Therefore, only the order of data within each dimension needs to be rearranged. The dimension data originally read from offset 0 is sliced ​​and rearranged according to the actual offset value, ensuring that the data layout in the lower dimensions is consistent with the expected memory access pattern.

[0079] S42. When the divisible dimension representation parameter div_num is not aligned with the read unit size BLOCK_SIZE, the complementary dimension scale is calculated according to the length of a static read and the number of repeated elements in the dimension, and the candidate dimension is generated according to the candidate dimension scale and the modulo dimension size.

[0080] When the divisible dimension representation parameter div_num is not aligned with BLOCK_SIZE, the number of dimensions read in the high dimension each time is determined by offset. Since offset is a parameter determined at runtime, its exact value cannot be determined at the compilation stage. However, the number of complementary dimension scales is calculated based on the low-dimensional size information and the BLOCK_SIZE size read each time. The number of complementary dimension scales is the maximum number of high dimensions that may be involved in the current situation. The calculation formula for the candidate dimension scale lineNum is as follows:

[0081]

[0082] Among them, rangeNum represents the length of a static read, divNum is the number of repeated elements in the dimension, [] is rounded up, and the dimension scale number lineNum is the maximum number of dimensions that can be read in this memory access when the static read length is rangeNum.

[0083] like Figure 6 The following diagram shows how an out-of-bounds, non-aligned integer division is converted to aligned SIMD memory access. Light gray represents aligned data, medium gray represents the tail block, and dark gray represents over-read data. For the two uncertain cases, the dark black cells in the diagram are over-read, thus statically translating the two different cases into the same one, facilitating subsequent vectorized memory access.

[0084] Perform the minimum operation on the complement dimension scale number lineNum and the modulus dimension size to generate the candidate reading dimension, the candidate reading dimension size candidate The calculation formula is as follows:

[0085] size candidate =min(lineNum,size);

[0086] Among them, size is the modulo dimension size, and the alternate reading dimension size candidate It is the actual size of this dimension finally read.

[0087] The modulo dimension size is the current dimension size. When the maximum number of dimensions accessed is greater than the current dimension size, it indicates that implicit broadcasting has occurred and the subsequent data is a repetition of the previous data. In this case, only the data of the current dimension size length needs to be read. Low-dimensional alignment is achieved through dynamic reorganization of dimension parameters to ensure that the data access granularity meets the requirements of the SIMD architecture. The number of complementary dimension scales is the maximum number of high dimensions that may exist in the low-dimensional out-of-bounds non-aligned scenario. However, considering that for each dimension, there are current dimension constraints (size) and low-dimensional constraints (div_num), the maximum value of the two needs to be taken in actual reading to ensure that all data required for this memory access can be read in the case of over-reading.

[0088] S5. Perform vectorized memory access on the memory access matrix under the SIMD architecture, generate an excess read matrix according to the expanded dimension parameters, and use SIMD vectorized load instructions to perform stride access optimization to obtain an excess read matrix that contains the correct memory access order but contains irrelevant data.

[0089] S51. Generate an excess read matrix according to the expanded dimension parameters, where the data size of the excess read matrix is ​​larger than the actual required size.

[0090] A two-dimensional memory access with a length of 8x256 may become an equivalent high-dimensional form after the above processing. For example, after multiple dimensional expansions, the original 8x256 memory access may become a high-dimensional matrix of 2x4x2x2x2x8x4. The matrix dimension may be much larger than the original dimension, so it is called an over-read matrix. At the same time, due to the possibility of non-aligned scenarios, more data may be introduced when performing compensation and repair operations for non-aligned dimensions. For example, the matrix may become 4x4...x4, and actually read 16x256 data. This is because, to meet the requirements of vectorized memory access, additional padding is required for the non-aligned parts at the head and tail to ensure that there are no non-aligned memory accesses, resulting in more data being read.

[0091] S52. Use SIMD vectorized load instructions to optimize stride access and improve cache line utilization.

[0092] SIMD vector load instructions are specifically memory copy instructions, which copy data from two given structured memory references. The original instruction pattern performs a gather-like memory access operation on each address point by point based on an offset address vector. This memory access operation is compatible with the SIMT architecture but can cause significant performance loss in the SIMD architecture. By changing the gather-like memory access format to one that relies on copying structured memory references based on offset values, stride, and size, the SIMD architecture's memory access performance can be better utilized, ultimately resulting in an excess read matrix with the correct memory access order but containing irrelevant data.

[0093] S6. Reduce the dimension of the excess read matrix based on the original computing dimension, fold the expanded dimension into a continuous one-dimensional vector, perform reverse dimension splitting and restoration according to the dimension priority, and obtain a memory access matrix that meets the requirements of the SIMD architecture computing instructions.

[0094] S61. Perform dimensionality reduction and reconstruction on the excess read matrix based on the original computing dimension metadata, and fold the extended dimension into a continuous one-dimensional vector, so that the data arrangement pattern under the one-dimensional dimension is consistent with the expected memory access order under the SIMT architecture.

[0095] Dimension reconstruction is performed on the over-read data. The read data is a matrix with a dimension higher than the actual calculation dimension. The dimension with the same original calculation dimension ori_dim needs to be reduced to one-dimensional form after reading to be consistent with the legal size of subsequent calculations. The read matrix is ​​reconstructed by dimension reduction according to the original calculation dimension ori_dim, so that each original calculation dimension ori_dim is mapped one-to-one with the original dimension after reconstruction. At this time, the data arrangement pattern under the one-dimensional dimension is consistent with the expected memory access order under the SIMT architecture.

[0096] S62. Calculate the difference Δoffset between the original offset and the new offset, perform cascade splitting from high dimension to low dimension, use Δoffset as the base offset, perform sliding window splitting along each dimension according to the read unit size BLOCK_SIZE, and output a memory access matrix that meets the requirements of the SIMD architecture computing instructions.

[0097] S621. Calculate the difference between the original offset and the new offset:

[0098] Δoffset=offset old -offset new

[0099] S622 , perform cascade splitting from high dimension to low dimension, use Δoffset as the reference offset, perform sliding window splitting along each dimension according to the read unit size BLOCK_SIZE, and output an aligned data layout that meets the requirements of SIMD architecture computing instructions.

[0100] From low dimensions to high dimensions, the one-dimensional vector after dimensionality reduction is split. The starting point of the split is the Δoffset parameter recorded in step 4, and the truncation length is the read unit size BLOCK_SIZE. After truncation, illegal data in excess reading is filtered out, leaving only legal data in the correct arrangement order. Cutting from low dimensions to high dimensions can ensure that the corresponding low dimensions when cutting high dimensions are all legally arranged, and the correctly converted memory access matrix is ​​obtained after splitting.

[0101] like Figure 7 As shown in the figure, the reverse dimension splitting and restoration steps are shown. The light gray and dark gray parts are the correct data required for memory access. After reading through vectorized memory access, a matrix consisting of three parts, light gray, dark gray and black, will be obtained. The data in the black part is illegal. Figure 7 After the split, the illegal black part of the stretched one-dimensional vector is removed, leaving only the legal light gray and dark gray parts. At the same time, the memory layout of this part is consistent with the expected memory layout. This method can change the non-aligned tail block (dark gray part) to an aligned mode by padding the front and back (as shown in the upper right image). The padded memory access does not need to consider the tail block and offset value, and it conforms to the vectorized reading format. The left image before processing is fragmented and can only be split into three blocks for vectorized memory access of the middle block. Scalar memory access or extremely small-grained vector memory access is performed at the head and tail, which has a significant performance loss. The aligned right image only requires a single vector memory access to read all data, which can better utilize the capabilities of the vector square unit. The high-dimensional memory access data in the upper part of the figure is stretched into a one-dimensional vector while maintaining the physical storage mode. As shown in the lower part of the figure, the calculated Δoffset can be used to reversely restore the starting position of the valid data in this memory access. At this time, all data is arranged continuously. Therefore, the result obtained by directly intercepting a window of the read unit size BLOCK_SIZE from Δoffset is the initial valid data.

[0102] This embodiment proposes a method for staticizing out-of-bounds non-aligned memory access applied to SIMD architecture, which can realize the conversion of SIMT architecture to SIMD architecture memory access in non-aligned scenarios, and can effectively solve the non-aligned memory access problem existing in SIMT architecture. By reconstructing the memory access mode, discrete non-aligned access requests are integrated into continuous structured data block readings. Under the premise of maintaining the correctness of the calculation logic, the data throughput efficiency of the SIMD processor is significantly improved. It is suitable for vectorized acceleration of out-of-bounds reading under SIMD architecture, and can convert memory access in the case of non-aligned out-of-bounds reading into static vectorized memory access, which can improve the overall computing efficiency of the system.

[0103] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A static method for out-of-bounds non-aligned memory access applied to SIMD architecture, characterized in that: Including steps: S1. Extract multiple dimensional feature parameters of the memory access mode under the SIMT architecture based on the operators existing in pointer operations. The multiple dimensional feature parameters include offset value, step size, size, number of repeated elements in the dimension, and original calculation dimension; S2. The extracted multiple dimensional feature parameters are grouped according to the original calculation dimensions to obtain multiple dimensional feature groups. Within each dimensional feature group, a monotonically increasing sorting is performed according to the number of repeated elements in the dimension. S3. Traverse each dimensional feature parameter in each dimensional feature group, perform modulus alignment verification and integer divisibility alignment verification, and obtain modulus alignment verification results and integer divisibility alignment verification results respectively; S4. When the modulo alignment verification result or the divisibility alignment verification result is a non-aligned mode, a non-aligned dimension compensation repair operation is performed; S5. Perform vectorized memory access on the memory access matrix under the SIMD architecture, generate an excess read matrix according to the expanded dimension parameters, and use SIMD vectorized load instructions to perform stride access optimization; S6. Reduce the dimension of the excess read matrix based on the original computing dimension, fold the expanded dimension into a continuous one-dimensional vector, perform reverse dimension splitting and restoration according to the dimension priority, and obtain a memory access matrix that meets the requirements of the SIMD architecture computing instructions.

2. The method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture according to claim 1, characterized in that: The operators present in the pointer operation include: dimension information related operators, shape information related operators, and other operators; dimension information related operators include: vector generation operators, dimension expansion operators, scalar vectorization operators, and broadcast operators; shape information related operators include: modulo, integer division, multiplication, and addition without dimension expansion; other operators include addition operations after integer division / modulo operations.

3. The method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture according to claim 2, wherein: Based on the operators in pointer arithmetic, the dimension characteristic parameters of the memory access mode in the SIMT architecture are extracted, including: Initialize the five dimensional feature parameters through the vector generation operator: offset value, step size, size, number of repeated elements in the dimension, and the original calculation dimension. Initialize the step size to 1, the number of repeated elements and the original dimension to 0; initialize the offset value and the size of the current dimension to the offset value and initial size of the vector respectively; The offset values ​​of the left and right operands are calculated according to the operator through the modulo, integer division, multiplication, and addition operations without dimension expansion of the shape information related operators; the multiplication and addition operations multiply or add the stride lengths corresponding to the operands on both sides respectively, and the number of repeated elements in the dimension is the second operand in the integer division operation; The size is initialized by the length of the vector generation operator and updated by the second operand of the modulo operation. The initial default value of the original calculation dimension is 0. When the dimension expansion operator is executed to expand the dimension, if the expanded dimension is less than or equal to the current dimension, the original calculation dimension is incremented by 1.

4. The method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture according to claim 1, wherein: The step S3 comprises: S31. Perform modulus alignment verification, traverse the feature parameters of each dimension in the dimensional feature group, and sequentially verify the alignment of the modulus dimension size and the reading unit size in each dimension; if the modulus dimension size and the reading unit size are integer multiples, it is determined to be an aligned mode; otherwise, it is considered to be a non-aligned mode, and a modulus alignment verification result is obtained; Perform modulo alignment verification, traverse the feature parameters of each dimension in the dimensional feature group, and align the number of low-dimensional repeated elements and the reading unit size in each dimension in turn; when the number of low-dimensional repeated elements and the reading unit size are integer multiples, it is determined to be an alignment mode, otherwise it is considered a non-alignment mode, and the integer alignment verification result is obtained.

5. The method for staticizing out-of-bounds non-aligned memory access applied to SIMD architecture according to claim 4, characterized in that: The step S4 comprises: S41. When the modulo dimension size is not aligned with the read unit size, the original offset is recorded, the excess read data is trimmed according to the original offset, a new offset after compensation is calculated, and the starting offset value of this unaligned memory access is adjusted to the new offset after compensation; S42. When the divisible dimension representation parameter is not aligned with the reading unit size, the complementary dimension scale is calculated according to the length of a static read and the number of repeated elements in the dimension, and the candidate dimension is generated according to the candidate dimension scale and the modulo dimension size.

6. The method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture according to claim 5, characterized in that: The calculation formula of the new offset after compensation is: offset new =min(offset old ,BLOCK_SIZE-size); Among them, offset new The new offset after compensation, offset old is the original offset, size is the modulo dimension, and BLOCK_SIZE represents the read unit size; The calculation formula for the candidate dimension scale is: Among them, lineNum is the scale of the complementary dimension, rangeNum represents the length of a static read, divNum is the number of repeated elements in the dimension, and [] is rounded up; The calculation formula of the candidate reading dimension is: size candidate =min(lineNum,size); Among them, size candidate is the alternate reading dimension, and size is the modulo dimension size.

7. The method for staticizing out-of-bounds non-aligned memory access applied to a SIMD architecture according to claim 1, characterized in that: The step S6 comprises: S61. Perform dimensionality reduction and reconstruction on the excess read matrix based on the original computing dimension metadata, folding the extended dimension into a continuous one-dimensional vector so that the data arrangement pattern in the one-dimensional dimension is consistent with the expected memory access order in the SIMT architecture. S62. Calculate the difference Δoffset between the original offset and the new offset, perform cascade splitting from high dimension to low dimension, use the difference Δoffset as the base offset, perform sliding window splitting along each dimension according to the read unit size, and output a memory access matrix that meets the requirements of the SIMD architecture computing instructions.