Multi-matrix double randomization method and device based on vector processor and electronic equipment

CN122451257BActive Publication Date: 2026-09-08SHANGHAI SUIYUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610903082.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-08
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

[0005]本发明提供了一种基于向量处理器的多矩阵双随机化方法、装置及电子设备,解决了直接将单个待处理矩阵加载至对应的一个向量寄存器内进行双随机化处理的现有技术,无法充分利用向量处理器的硬件算力,致使多矩阵双随机化处理效率降低,开销增大的问题,可以充分利用向量处理器的硬件算力,提升多矩阵双随机化处理效率,降低多矩阵双随机化处理开销

Benefits of technology

[0009]The technical solution of this invention, by determining multiple matrices to be processed corresponding to each vector register when the matrix size of each matrix to be processed is smaller than the register width of each vector register, and the difference between the register width and the matrix size is greater than a preset difference threshold; loading each matrix to be processed into the corresponding vector register, and performing multiple rounds of double randomization processing on each matrix to be processed in each vector register to obtain the target double randomized matrix corresponding to each matrix to be processed, solves the problem of existing technology that directly loads a single matrix to be processed into a corresponding vector register for double randomization processing, which cannot fully utilize the hardware computing power of the vector processor, resulting in reduced efficiency and increased overhead of multi-matrix double randomization processing. It can fully utilize the hardware computing power of the vector processor, improve the efficiency of multi-matrix double randomization processing, and reduce the overhead of multi-matrix double randomization processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451257B_ABST
    Figure CN122451257B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a kind of based on vector processor's multi-matrix double randomization method, device and electronic equipment, comprising: obtaining each to-be-processed matrix in regular floating point tensor, and each vector register in target vector processor;Wherein, the matrix size of each to-be-processed matrix is equal, and the register width of each vector register is equal;When matrix size is less than register width, and the difference between register width and matrix size is greater than the preset difference threshold value, determine the multiple to-be-processed matrices corresponding to each vector register respectively;Each to-be-processed matrix is loaded to corresponding vector register, and each to-be-processed matrix in each vector register is respectively subjected to multi-round double randomization processing, and the target double random matrix corresponding to each to-be-processed matrix is obtained, can make full use of the hardware computing power of vector processor, improve multi-matrix double randomization processing efficiency, reduce multi-matrix double randomization processing overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to a multi-matrix double randomization method, apparatus, and electronic device based on a vector processor. Background Technology

[0002] In the field of deep learning, the Sinkhorn-Knopp (SK) algorithm is widely used to iteratively project any non-negative matrix into a doubly random matrix, i.e., a matrix satisfying that all elements are non-negative, the sum of elements in each row is 1, and the sum of elements in each column is 1. This algorithm has important applications in scenarios such as optimal transport, attention mechanism normalization, and model residual hyperconnection constraints.

[0003] In typical application scenarios, the SK algorithm requires performing double randomization iterations on the residual weight tensor during the forward propagation of each layer of the deep learning model. This algorithm needs to complete 10 to 30 iterations per forward propagation, with each iteration including one row normalization operation and one column normalization operation. Since this computational step is on the critical execution path of the deep learning model, its computational efficiency directly determines the throughput of end-to-end training, significantly impacting model training speed and deployment performance.

[0004] In existing technologies, a single matrix to be processed is often directly loaded into a corresponding vector register for double randomization, resulting in a double randomized matrix corresponding to each matrix to be processed. However, since a single matrix only occupies 3% to 12% of the vector register capacity, using existing technologies to perform double randomization on each matrix to be processed cannot fully utilize the hardware computing power of the vector processor, leading to reduced efficiency and increased overhead in multi-matrix double randomization processing. Summary of the Invention

[0005] This invention provides a method, apparatus, and electronic device for multi-matrix double randomization based on a vector processor. It solves the problem that the existing technology of directly loading a single matrix to be processed into a corresponding vector register for double randomization processing cannot fully utilize the hardware computing power of the vector processor, resulting in reduced efficiency and increased overhead of multi-matrix double randomization processing. It can fully utilize the hardware computing power of the vector processor, improve the efficiency of multi-matrix double randomization processing, and reduce the overhead of multi-matrix double randomization processing.

[0006] In a first aspect, embodiments of the present invention provide a multi-matrix double randomization method based on a vector processor, comprising: obtaining each matrix to be processed in a regular floating-point tensor, and each vector register in a target vector processor; wherein the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal; when the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, determining multiple matrices to be processed corresponding to each vector register respectively; loading each matrix to be processed into the corresponding vector register, and performing multiple rounds of double randomization processing on each matrix to be processed in each vector register respectively, to obtain a target double random matrix corresponding to each matrix to be processed respectively.

[0007] Secondly, embodiments of the present invention also provide a multi-matrix double randomization device based on a vector processor, comprising: a matrix and register acquisition module, used to acquire each matrix to be processed in a regular floating-point tensor, and each vector register in a target vector processor; wherein the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal; a matrix loading position determination module, used to determine multiple matrices to be processed corresponding to each vector register when the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold; and a double randomization processing module, used to load each matrix to be processed into the corresponding vector register, and perform multiple rounds of double randomization processing on each matrix to be processed in each vector register to obtain a target double random matrix corresponding to each matrix to be processed.

[0008] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the vector processor-based multi-matrix double randomization method provided in any embodiment of the present invention.

[0009] The technical solution of this invention, by determining multiple matrices to be processed corresponding to each vector register when the matrix size of each matrix to be processed is smaller than the register width of each vector register, and the difference between the register width and the matrix size is greater than a preset difference threshold; loading each matrix to be processed into the corresponding vector register, and performing multiple rounds of double randomization processing on each matrix to be processed in each vector register to obtain the target double randomized matrix corresponding to each matrix to be processed, solves the problem of existing technology that directly loads a single matrix to be processed into a corresponding vector register for double randomization processing, which cannot fully utilize the hardware computing power of the vector processor, resulting in reduced efficiency and increased overhead of multi-matrix double randomization processing. It can fully utilize the hardware computing power of the vector processor, improve the efficiency of multi-matrix double randomization processing, and reduce the overhead of multi-matrix double randomization processing.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of a multi-matrix double randomization method based on a vector processor according to Embodiment 1 of the present invention.

[0013] Figure 2 This is a flowchart of another multi-matrix double randomization method based on a vector processor provided in Embodiment 2 of the present invention.

[0014] Figure 3 This is a flowchart of another multi-matrix double randomization method based on a vector processor provided in Embodiment 3 of the present invention.

[0015] Figure 4 This is a flowchart of another multi-matrix double randomization method based on a vector processor provided in Embodiment 4 of the present invention.

[0016] Figure 5 This is a schematic diagram of the structure of a multi-matrix double randomization device based on a vector processor according to Embodiment 5 of the present invention.

[0017] Figure 6 This is a schematic diagram of the structure of an electronic device provided in Embodiment Six of the present invention. Detailed Implementation

[0018] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0020] Example 1 Figure 1 This is a flowchart of a multi-matrix double randomization method based on a vector processor according to Embodiment 1 of the present invention.

[0021] This embodiment is applicable to situations where multiple matrices to be processed are subjected to double randomization based on a vector processor. This method can be executed by a vector processor-based multi-matrix double randomization device, which can be implemented in hardware and / or software. This vector processor-based multi-matrix double randomization device can be configured in an electronic device, which can be the chip containing the vector processor or a chip control device capable of controlling the aforementioned chip, as long as it can execute the vector processor-based multi-matrix double randomization method. This embodiment of the invention does not limit the specific device type of the electronic device.

[0022] like Figure 1 As shown, this embodiment discloses a multi-matrix double randomization method based on a vector processor, including S110-S130.

[0023] S110. Obtain each matrix to be processed in the rule floating-point tensor, and each vector register in the target vector processor; wherein, the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal.

[0024] In this embodiment, a regular floating-point tensor can be understood as a multidimensional tensor structure composed of several matrices of identical size to be processed, with all elements within the matrices represented in floating-point format. The matrices to be processed can be understood as square matrices with an equal number of rows and columns. The target vector processor can be understood as a processor used to perform double randomization processing on each matrix to be processed within the regular floating-point tensor. The vector register can be understood as a register used to store the matrices to be processed for use by the target vector processor.

[0025] S120. When the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, determine multiple matrices to be processed corresponding to each vector register.

[0026] In this embodiment, the preset difference threshold can be set according to historical experience or user needs. For example, the preset difference threshold can be set to 1 times, 2 times or other multiples that can represent that the matrix size is much smaller than the register width.

[0027] In this step, specifically, when the matrix size is smaller than the register width, it can be determined whether the difference between the register width and the matrix size is greater than a preset difference threshold. If the difference between the register width and the matrix size is greater than the preset difference threshold, the matrix size is considered to be much smaller than the register width. In this case, the number of matrices to be processed that each vector register can store can be determined based on the matrix size and the register width. Then, based on the number of matrices to be processed that each vector register can store, multiple matrices to be processed corresponding to each vector register can be randomly determined.

[0028] If the difference between the register width and the matrix size is less than or equal to a preset difference threshold, the matrix size can be considered to be similar to the register width. In this case, a matrix to be processed can be randomly determined for each vector register.

[0029] S130. Load each matrix to be processed into the corresponding vector register, and perform multiple rounds of double randomization processing on each matrix to be processed in each vector register to obtain the target double random matrix corresponding to each matrix to be processed.

[0030] In this step, specifically, each matrix to be processed can be loaded into its corresponding vector register in row-major order, that is, the first... The nth matrix to be processed Line 1 Column elements are located at address offset Location. Among them, is the matrix order of the matrix to be processed.

[0031] The matrix size of each matrix to be processed is Taking a vector register width of 128 as an example, it can be determined that in row-major order storage mode, one vector register can completely store 8 matrices to be processed, that is, 32 matrix rows.

[0032] It is worth noting that when the tensor dimension of a regular floating-point tensor is not divisible by the register width of a vector register, it can be resolved using the formula... This determines the number of remaining elements that are not filled in a single vector register. This represents the number of remaining elements that are not filled in a single vector register. For regular floating-point tensors, the tensor dimension is... This is the register width of the vector register. Then, based on the number of remaining elements, the valid element bits in each vector register can be determined, and while the mask covers the valid element bits, the invalid element bits other than the valid elements are masked to 0 or not involved in the operation.

[0033] After loading each matrix to be processed into its corresponding vector register, multiple rounds of double randomization can be performed on each matrix in each vector register to obtain a target double random matrix corresponding to each matrix. Taking a matrix in a vector register as an example, one round of double randomization can be performed on the matrix to be processed to obtain a preliminary double random matrix corresponding to it. It is then determined whether the number of rounds of double randomization for the matrix to be processed exceeds a preset threshold. If not, after using the preliminary double random matrix as a new matrix to be processed, the process returns to performing one round of double randomization on the matrix to be processed to obtain a preliminary double random matrix corresponding to it, until the number of rounds of double randomization exceeds the preset threshold. At this point, the preliminary double random matrix is ​​used as the target double random matrix corresponding to the matrix to be processed.

[0034] The advantage of this setup is that by loading multiple matrices to be processed into a single vector register, a single vector register instruction can perform the same operation on multiple matrices simultaneously, thereby effectively improving the hardware computing power utilization of the vector processor.

[0035] Optionally, multiple rounds of double randomization processing are performed on each matrix to be processed in each vector register to obtain target double random matrices corresponding to each matrix to be processed. This includes: obtaining each computing unit cluster in the target vector processor and evenly distributing each matrix to be processed to each computing unit cluster; and using each computing unit cluster to perform multiple rounds of double randomization processing on each matrix to be processed in each vector register to obtain target double random matrices corresponding to each matrix to be processed.

[0036] In this context, a computing unit cluster can be understood as a collection of multiple computing units.

[0037] Specifically, the total number of matrices to be processed in the regular floating-point tensor can be divided by the number of matrices to be processed that each vector register can store to obtain the number of matrix blocks to be processed. Then, the number of blocks to be processed is rounded up to obtain the number of target matrix blocks corresponding to each computing unit cluster.

[0038] Then, to achieve load balancing, the target matrix blocks can be evenly distributed across the various computing unit clusters. Next, since a computing unit cluster contains multiple computing units, the matrices to be processed within the target matrix blocks can be evenly distributed across the computing units of the corresponding clusters for parallel processing. Finally, since each computing unit corresponds to one sub-thread, the sub-threads corresponding to each computing unit can be used to process the matrices to be processed in parallel. Specifically, there are no data dependencies or synchronization overhead between sub-threads; scheduling between sub-threads is performed at the thread grid and thread block levels; each sub-thread has an independent buffer in its local memory buffer, and these buffers are isolated by address offsets.

[0039] Optionally, a load balancing strategy relying on individual child threads is to use a global thread index for cyclic allocation, rather than allocating to contiguous target matrix blocks. Specifically, the global index for each child thread can be determined using the following formula: .

[0040] in, This represents the global index of the child thread. Indicates the cluster number of the computing unit. Indicates the total number of computing unit clusters. Indicates the total number of calculation units. Indicates the calculation unit number. Indicates the child thread number.

[0041] Then, based on the global index of each sub-thread, the matrices to be processed by each sub-thread can be determined sequentially. Using the global index as... Taking a child thread as an example, it can be determined that the child thread needs to process the first... , ,as well as The system waits to process matrices, thus maintaining good load balancing even when the number of matrices to be processed is not divisible by the degree of parallelism. This represents the total number of child threads. , ,as well as It can also be used to refer to the offset of each matrix to be processed in memory.

[0042] The advantage of this setup is that by evenly distributing the matrices to be processed across the cluster of computing units for multiple rounds of double randomization, the efficiency of multi-matrix double randomization processing can be improved. Secondly, by using the global index of each sub-thread to determine the matrices to be processed sequentially, each matrix can be quickly located, thereby improving the efficiency of multi-matrix double randomization processing.

[0043] Optionally, performing multiple rounds of double randomization on a matrix to be processed in a vector register to obtain a target double random matrix corresponding to the matrix to be processed may include: performing one round of double randomization on a matrix to be processed in the vector register to obtain a preliminary double random matrix corresponding to the matrix to be processed; determining the row normalization error vector corresponding to the preliminary double random matrix based on the row sum vector and the target row sum value; determining the maximum row normalization error corresponding to the preliminary double random matrix based on the row normalization error vector; if the maximum row normalization error is greater than or equal to a preset convergence threshold, then after using the preliminary double random matrix as a new matrix to be processed, returning to the operation of performing one round of double randomization on a matrix to be processed in the vector register to obtain a preliminary double random matrix corresponding to the matrix to be processed, until the maximum row normalization error is less than the preset convergence threshold, or the number of rounds of double randomization is greater than a preset number of rounds threshold, at which point the preliminary double random matrix is ​​used as the target double random matrix corresponding to the matrix to be processed. The preset round threshold can be set based on historical experience or user needs. For example, the preset round threshold can be set to any integer value between 10 and 30.

[0044] Specifically, after performing an initial round of double randomization on the matrix to be processed to obtain a preliminary double random matrix corresponding to the matrix to be processed, it is determined whether the number of rounds of double randomization exceeds a preset round threshold. If not, after calculating multiple row sum vectors corresponding to the preliminary double random matrix, vector subtraction is performed on each row sum vector and its target value, and the absolute value of each vector subtraction result is taken to obtain multiple row normalization error vectors corresponding to the preliminary double random matrix. Then, vector reduction to maximum value operation can be performed on each row normalization error vector to obtain the maximum row normalization error corresponding to the preliminary double random matrix.

[0045] Taking a row and a target value of 1 as an example, the following specific calculation formula can be used to determine the normalized error vector of each row corresponding to the initial double random matrix: .

[0046] in, For row-normalized error vector, For rows and vectors, For rows and target values.

[0047] Then, the maximum row normalization error corresponding to the initial double random matrix can be determined using the following specific calculation formula: .

[0048] in, The maximum value of the row normalization error. This is the row-normalized error vector.

[0049] After obtaining the maximum row normalization error corresponding to the initial double-random matrix, it can be determined whether the maximum row normalization error is less than a preset convergence threshold. If the maximum row normalization error is less than the preset convergence threshold, the matrix to be processed is considered to have converged, and the initial double-random matrix can be used as the target double-random matrix corresponding to the matrix to be processed. If the maximum row normalization error is greater than or equal to the preset convergence threshold, the matrix to be processed is considered not to have converged. In this case, the process can be repeated to perform one round of double randomization on the matrix to be processed to obtain the initial double-random matrix corresponding to the matrix to be processed. This process continues until the maximum row normalization error is less than the preset convergence threshold, or the number of double randomization rounds exceeds the preset round threshold. At this point, the initial double-random matrix is ​​used as the target double-random matrix corresponding to the matrix to be processed.

[0050] The advantage of this setting is that, compared to existing technologies that terminate double randomization processing only when the number of double randomization processing rounds exceeds a preset round threshold for any matrix to be processed, thus failing to adapt to the differences in convergence speed of different matrices, the technical solution of this embodiment terminates the iteration in advance when the maximum value of the row normalization error is less than the preset convergence threshold. This achieves a significant reduction in the number of iterations of the fast-converging matrix with a very small increase in computational load, thereby improving the average throughput.

[0051] Optionally, since a vector register can store multiple matrices to be processed, the vector register can be released in advance when all matrices to be processed in a certain vector register have converged in advance, so as to save storage resources.

[0052] Optionally, a round of double randomization is performed on a matrix to be processed in the vector register to obtain a preliminary double random matrix corresponding to the matrix to be processed. This may include: performing row normalization on the matrix to be processed to obtain a column normalization matrix. Then, since the row-major order storage method is used, the first... Liede The elements of the row are located at an offset from the starting address of the matrix to be processed. At this point, different matrices to be processed are further separated by... Since the arrangement is periodic, clustering / dispersing instructions can be used to normalize the columns of the matrix to be normalized. Specifically, for each group of matrices to be processed loaded into the vector register, the column offset vector can be calculated: assuming the vector register can hold... The nth matrix row, then the offset vector's nth row... The elements are .in, The register width of the vector register. Let be the matrix order of the matrix to be processed. Let the element type be the element type of each element in the matrix to be processed. It can be used as an inline index. It can be used as an in-column index. Then, since the offset vector can be calculated once at startup using a separate pre-computation function and stored in the vector register, it can be directly reused in all subsequent iterations, thus avoiding repeated calculation of the offset vector in each iteration. Therefore, for The matrix to be processed can be collected using collection commands. The matrix to be processed is collected from the local memory buffer based on the offset vector. The elements of a row are mapped into a vector. Where, Indicates the matrix to be processed. All elements of the row, This represents the starting address of the matrix to be processed. This represents the offset vector of the matrix to be processed. Then, each column vector can be element-wise summed to obtain the column sum vector. Finally, each column vector can be divided by the column sum vector to obtain the preliminary double random matrix corresponding to the matrix to be processed.

[0053] The technical solution of this embodiment obtains each matrix to be processed in the regular floating-point tensor and each vector register in the target vector processor. The matrix size of each matrix to be processed is equal, and the register width of each vector register is equal. When the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, multiple matrices corresponding to each vector register are determined. Each matrix to be processed is loaded into its corresponding vector register, and multiple rounds of double randomization processing are performed on each matrix to be processed in each vector register to obtain the target double randomized matrix corresponding to each matrix to be processed. This solves the problem of existing technologies that directly load a single matrix to be processed into a corresponding vector register for double randomization processing, which cannot fully utilize the hardware computing power of the vector processor, resulting in reduced efficiency and increased overhead in multi-matrix double randomization processing. This solution can fully utilize the hardware computing power of the vector processor, improve the efficiency of multi-matrix double randomization processing, and reduce the overhead of multi-matrix double randomization processing.

[0054] Example 2 Figure 2 This is a flowchart of another multi-matrix double randomization method based on a vector processor according to Embodiment 2 of the present invention. This embodiment is a further optimization and extension based on the above embodiments and can be combined with various optional technical solutions in the above embodiments.

[0055] like Figure 2 As shown in the figure, this embodiment discloses a multi-matrix double randomization method based on a vector processor, including S210-S270.

[0056] S210. Obtain each matrix to be processed in the rule floating-point tensor, and each vector register in the target vector processor; wherein, the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal.

[0057] S220. When the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, determine the number of matrices to be processed that each vector register can store based on the matrix size and the register width.

[0058] In this step, specifically, the register width can be divided by the matrix size to obtain the number of matrices to be processed that each vector register can store.

[0059] S230. Based on the relative position of each matrix to be processed in the regular floating-point tensor, the logical index of each vector register, and the number of matrices to be processed that each vector register can store, determine multiple matrices to be processed corresponding to each vector register.

[0060] In this step, specifically, the vector register with the smallest logical index can be used as the target vector register. Based on the number of matrices to be processed that each vector register can store, the number of target matrices with the earliest relative positions are obtained from the regular floating-point tensor. Then, a correspondence between the target vector registers can be established, and after eliminating the influence of the target vector registers and target matrices to be processed, the operation of using the vector register with the smallest logical index as the target vector register is repeated until the number of matrices to be processed corresponding to each vector register is determined.

[0061] S240. Load each matrix to be processed into the corresponding vector register. When each matrix to be processed is arranged in row-major order in the corresponding vector register, perform row normalization on each matrix to be processed in each vector register to obtain multiple row normalized matrices.

[0062] In this step, specifically, since row data from different matrices but with the same row number are interleaved in memory with a specific step size, a step-size loading instruction can be used to collect the data from the same column of all matrices to be processed into a single row vector. Then, each row vector can be element-wise summed to obtain the corresponding row and vector. Finally, each row vector can be divided by the row and vector, and the division results can be split and merged to obtain the row-normalized matrix corresponding to each matrix to be processed.

[0063] For example, suppose a vector register contains only a first matrix to be processed and a second matrix to be processed, and the first matrix to be processed is... The second matrix to be processed is Then you can use This long loading instruction collects the data from the same row of the first and second matrices into a single row vector, resulting in the following four row vectors: , , ,as well as .in, For line numbers, This is the starting address of the first matrix to be processed. Indicates the step size.

[0064] Then, the four row vectors can be summed element-wise to obtain a row sum vector. Each position in the row sum vector corresponds exactly to the sum of elements in a specific row of the matrix to be processed. Finally, each row vector can be divided by the row sum vector to obtain the division result, and then... This long storage instruction splits and merges the division results to obtain row-normalized matrices corresponding to each matrix to be processed.

[0065] The advantage of this setup is that, since the step-size loading instruction and the step-size storage instruction are executed directly by the hardware, no additional address calculation or data rearrangement is required, and a single instruction can complete the data acquisition of the entire vector width. Therefore, the technical solution of this embodiment improves the row normalization efficiency by completing the matrix row normalization based on the step-size loading instruction and the step-size storage instruction.

[0066] S250. Transpose each row of the normalized matrix to obtain multiple transpose matrices.

[0067] In this step, specifically, the data transformation engine built into the target vector processor can be obtained, and the transpose function of the data transformation engine can be used to transpose the normalized matrix of each row to obtain multiple transpose matrices.

[0068] S260. Perform row normalization on each transpose matrix to obtain multiple double-normalized matrices.

[0069] In this step, specifically, a step-size loading instruction can be used to collect the data from the same row of all transpose matrices into a single transpose row vector. Then, each transpose row vector can be element-wise summed to obtain the corresponding transpose row and vector. Finally, each transpose row vector can be divided by its corresponding transpose row and vector, and the division results can be split and merged to obtain the double-normalized matrix corresponding to each transpose matrix.

[0070] S270. Transpose each double normalized matrix to obtain the target double random matrix corresponding to each matrix to be processed.

[0071] In this step, specifically, the transpose function of the data transformation engine can be used to transpose each double normalized matrix to obtain multiple target double random matrices.

[0072] It is worth noting that, since the data transformation engine has asynchronous execution characteristics, a double buffer can be set up. The data transformation engine transposes a matrix in the first buffer, while the computing unit clusters normalize another matrix in the second buffer. This allows the pipeline of transposition and computation to overlap, masking the transposition delay.

[0073] The technical solution of this embodiment, when each matrix to be processed is arranged in row-major order in the corresponding vector register, performs row normalization on each matrix to be processed in each vector register to obtain multiple row normalized matrices; performs transpose on each row normalized matrix to obtain multiple transposed matrices; performs row normalization on each transposed matrix to obtain multiple double normalized matrices; and performs transpose on each double normalized matrix to obtain the target double randomized matrix corresponding to each matrix to be processed. This technique can transform all column normalization operations into step-size loading operations with the same step size as row normalization, unify the memory access mode of row and column normalization, and improve the efficiency of multi-matrix double randomization processing.

[0074] Example 3 Figure 3 This is a flowchart of another multi-matrix double randomization method based on a vector processor provided in Embodiment 3 of the present invention. This embodiment is a further optimization and extension based on the above embodiments and can be combined with various optional technical solutions in the above embodiments.

[0075] like Figure 3 As shown in the figure, this embodiment discloses a multi-matrix double randomization method based on a vector processor, including S310-S350.

[0076] S310. Obtain each matrix to be processed in the rule floating-point tensor, and each vector register in the target vector processor; wherein, the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal.

[0077] S320. When the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, determine multiple matrices to be processed corresponding to each vector register.

[0078] S330. Load each matrix to be processed into the corresponding vector register, and perform exponent calculation on each matrix to be processed in each vector register to obtain the exponent matrix corresponding to each matrix to be processed.

[0079] In this step, specifically, each matrix to be processed in global memory can be loaded into its corresponding vector register. Then, exponentiation can be directly performed on each matrix to obtain an exponent matrix corresponding to each matrix. Alternatively, after subtracting the corresponding maximum value matrix from each matrix to obtain an offset matrix corresponding to each matrix, exponentiation can be performed on each offset matrix to obtain an exponent matrix corresponding to each matrix.

[0080] The advantage of this setup is that by performing exponential calculations on the corresponding matrices to be processed within each vector register, it avoids writing the unprocessed matrices to the local memory buffer and then rereading them for calculation, thereby improving the efficiency of exponential operations.

[0081] Optionally, within each vector register, exponentiation is performed on each corresponding matrix to be processed to obtain an exponent matrix corresponding to each matrix to be processed. This includes: determining the maximum element in each matrix to be processed, and generating a maximum value matrix corresponding to each matrix to be processed based on the maximum element corresponding to each matrix to be processed; subtracting the corresponding maximum value matrix from each matrix to be processed to obtain an offset matrix corresponding to each matrix to be processed; and performing exponentiation on each offset matrix to obtain an exponent matrix corresponding to each matrix to be processed.

[0082] Specifically, the vector reduction maximum value instruction can be used to determine the maximum element in each matrix to be processed, and based on the maximum element corresponding to each matrix, a maximum value matrix corresponding to each matrix to be processed can be generated. The maximum value matrix can be understood as a matrix containing only the maximum element of its corresponding matrix to be processed.

[0083] Then, to ensure that the exponentiation operation does not overflow, the maximum value matrix corresponding to each matrix to be processed can be subtracted from each matrix to obtain an offset matrix corresponding to each matrix to be processed, so that all elements in the offset matrix are less than or equal to zero. Finally, hardware exponentiation instructions can be executed on each offset matrix to obtain an exponent matrix corresponding to each matrix to be processed.

[0084] The advantage of this setup is that, since directly performing exponential operations on the matrices to be processed can cause floating-point overflow for larger element values, the technical solution in this embodiment first subtracts the corresponding maximum value matrix from each matrix to be processed to obtain the offset matrix corresponding to each matrix to be processed, and then performs exponential operations on each offset matrix. This can avoid floating-point overflow during exponential operations, thereby preventing distortion of the subsequent double randomization processing results.

[0085] S340: Obtain the local memory buffer on the chip where the vector processor is located, and write each exponent matrix to the local memory buffer.

[0086] The local memory buffer can be understood as a high-speed storage area exclusively for each thread, with extremely low access latency.

[0087] S350. In the local memory buffer, perform multiple rounds of double randomization on each exponent matrix to obtain the target double random matrix corresponding to each matrix to be processed.

[0088] In this step, specifically, multiple child threads can be acquired, and an independent matrix buffer can be allocated for each child thread within a local memory buffer. These matrix buffers are isolated from each other using offsets to avoid data conflicts; each matrix buffer occupies [a certain amount of space / data]. byte, The maximum matrix order is defined in advance. It is an element type.

[0089] Then, within the local memory buffer, each sub-thread can perform multiple rounds of double randomization on each exponent matrix to obtain the target double random matrix corresponding to each matrix to be processed. The multiple rounds of double randomization do not generate any global memory access; all intermediate results reside in the local memory buffer.

[0090] After obtaining the target double random matrices corresponding to each matrix to be processed, each target double random matrix can be written back to global memory, and the boundary conditions can be handled using vector storage instructions and masks.

[0091] The advantage of this setup is that by performing multiple rounds of double randomization on each exponent matrix within the local memory buffer, only one global memory read and one global memory write can be performed during the entire iteration process, avoiding repeated access to global memory and significantly reducing storage access latency and bandwidth consumption.

[0092] The technical solution of this embodiment improves the efficiency of exponentiation by performing exponentiation calculations on the corresponding matrices to be processed within each vector register, thus avoiding writing unprocessed matrices to the local memory buffer and rereading them for calculation. Secondly, by performing multiple rounds of double randomization processing on each exponent matrix within the local memory buffer, only one global memory read and one global memory write are performed throughout the entire iteration process, avoiding repeated access to global memory and significantly reducing storage access latency and bandwidth consumption.

[0093] Example 4 Figure 4 This is a flowchart of another multi-matrix double randomization method based on a vector processor provided in Embodiment 4 of the present invention. This embodiment is a further optimization and extension based on the above embodiments and can be combined with various optional technical solutions in the above embodiments.

[0094] like Figure 4 As shown in the figure, this embodiment discloses a multi-matrix double randomization method based on a vector processor, including S410-S470.

[0095] S410. Obtain each matrix to be processed in the rule floating-point tensor, and each vector register in the target vector processor; wherein, the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal.

[0096] S420. When the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, determine multiple matrices to be processed corresponding to each vector register.

[0097] S430. Load each matrix to be processed into the corresponding vector register, and obtain the preset iteration round corresponding to each matrix to be processed, as well as the previous iteration round and the subsequent iteration round in each preset iteration round.

[0098] In this embodiment, the preset iteration rounds, the previous set rounds, and the subsequent iteration rounds can be set according to user needs and historical experience. For example, the preset iteration rounds, the previous set rounds, and the subsequent iteration rounds can be set to 30, 15, and 15, respectively.

[0099] S440. Convert each matrix to be processed from single-precision floating-point format to half-precision floating-point format to obtain half-precision format matrices corresponding to each matrix to be processed.

[0100] In this step, specifically, since the initial iteration process of the matrix to be processed is mainly used to determine the approximate structure of the matrix (e.g., which elements are large and which elements are small), the precision requirement is low. Therefore, in order to improve the efficiency of double randomization processing, the hardware type conversion instruction of the target vector processor can be used to convert each matrix to be processed from single-precision floating-point format to half-precision floating-point format. Then, in the pre-set rounds corresponding to each matrix to be processed, double randomization processing is performed on each half-precision format matrix.

[0101] S450. In the pre-defined rounds corresponding to each matrix to be processed, the half-precision format matrix is ​​subjected to double randomization to obtain the processed half-precision format matrix.

[0102] The advantage of this setup is that, since single-precision floating-point format elements occupy 4 bytes and half-precision floating-point format elements occupy 2 bytes, by converting each matrix to be processed from single-precision floating-point format to half-precision floating-point format, and performing double randomization processing on each half-precision format matrix in the pre-defined rounds corresponding to each matrix to be processed, the computational throughput in the early stage of the iteration can be increased by about 2 times without losing computational precision, thereby improving the efficiency of multi-matrix double randomization processing.

[0103] S460. Convert each processed half-precision format matrix from half-precision floating-point format to single-precision floating-point format to obtain single-precision format matrices corresponding to each matrix to be processed.

[0104] In this step, specifically, since the later iterations of the matrix to be processed are used for fine-tuning to make the sum of rows and columns exactly equal to 1, which is sensitive to precision, in order to improve the efficiency of multi-matrix double randomization processing while ensuring the precision of multi-matrix double randomization processing, each processed half-precision format matrix can be converted from half-precision floating-point format to single-precision floating-point format. Then, in subsequent iterations, each single-precision format matrix is ​​double randomized to obtain the target double random matrix corresponding to each matrix to be processed.

[0105] S470. In the subsequent iterations corresponding to each matrix to be processed, the single-precision format matrix is ​​subjected to double randomization to obtain the target double random matrix corresponding to each matrix to be processed.

[0106] The advantage of this setup is that by performing double randomization on the half-precision format matrix in the initial set rounds and double randomization on each single-precision format matrix in subsequent iteration rounds, the efficiency of double randomization on multiple matrices can be improved while ensuring the accuracy of double randomization.

[0107] Optionally, to implement multi-matrix double randomization processing using C++ code, C++ template metaprogramming techniques can be employed to generate optimized code at compile time based on the matrix order of each matrix to be processed. Specifically, step one involves using the matrix order as a compile-time template parameter to determine all buffer sizes, loop counts, and offset calculations related to the matrix order during the compilation phase. This allows the compiler to perform optimizations such as loop unrolling, constant folding, and vector register allocation accordingly. Step two involves distributing templates based on the actual matrix order, calling the corresponding specialized version. Supported matrix order values ​​can include 4, 8, 16, and 32, among others. Step three involves generating the optimal instruction sequence for each specialized version of the matrix order, tailored to specific array sizes and loop structures, to eliminate runtime overhead from conditional branches and dynamic index calculations.

[0108] The technical solution of this embodiment converts each matrix to be processed from single-precision floating-point format to half-precision floating-point format to obtain a half-precision format matrix corresponding to each matrix to be processed. In the predetermined rounds corresponding to each matrix to be processed, the half-precision format matrix is ​​subjected to double randomization to obtain a processed half-precision format matrix. The processed half-precision format matrix is ​​then converted from half-precision floating-point format to single-precision floating-point format to obtain a single-precision format matrix corresponding to each matrix to be processed. In the subsequent iteration rounds corresponding to each matrix to be processed, the single-precision format matrix is ​​subjected to double randomization to obtain a target double random matrix corresponding to each matrix to be processed. This approach can improve the efficiency of double randomization processing of multiple matrices while ensuring the accuracy of the double randomization processing.

[0109] To illustrate the multi-matrix double randomization method and its effects based on a vector processor in this invention, a detailed embodiment is described below: Step 1: Pack the data of multiple small matrices into the same vector register, so that a single vector processor instruction can process the same operation on multiple small matrices simultaneously, maximizing the utilization of the vector processor. Specifically, a regular floating-point tensor can be read from global memory, and each matrix to be processed in the regular floating-point tensor, as well as each vector register in the target vector processor, can be obtained. The matrix sizes of each matrix to be processed are equal, and the register widths of each vector register are equal. Then, when the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, multiple matrices corresponding to each vector register are determined, and each matrix to be processed is loaded into the corresponding vector register in row-major order.

[0110] Step 2: Within each vector register, perform exponentiation calculations on the corresponding matrices to be processed, obtaining exponent matrices corresponding to each matrix. Specifically, this can be achieved by subtracting the maximum value matrix from each matrix to be processed, obtaining the offset matrices corresponding to each matrix. Then, perform exponentiation operations on each offset matrix to obtain the exponent matrices corresponding to each matrix to be processed.

[0111] Step 3: Obtain the local memory buffer on the chip where the vector processor resides, and write each exponent matrix into the local memory buffer. Step 4: Within the local memory buffer, perform multiple rounds of double randomization on each exponent matrix to obtain a double randomized exponent matrix corresponding to each exponent matrix, and write each double randomized exponent matrix back to global memory. Specifically, each exponent matrix can be row-normalized to obtain multiple row-normalized exponent matrices; each row-normalized exponent matrix can be transposed to obtain multiple transposed exponent matrices; each transposed exponent matrix can be row-normalized to obtain multiple double-normalized exponent matrices; and each double-normalized exponent matrix can be transposed to obtain a double randomized exponent matrix corresponding to each exponent matrix.

[0112] It is worth noting that, in order to improve the efficiency and ensure the accuracy of the double randomization process of the exponent matrices during the multiple rounds of double randomization, at least one of the following three optimization methods can be performed during the multiple rounds of double randomization process of the exponent matrices.

[0113] Optimization Method 1: To adaptively terminate the iteration process of the fast convergence matrix early, after performing one round of double randomization on the exponential matrix to obtain the preliminary double random matrix corresponding to it, the maximum row normalization error corresponding to the preliminary double random matrix is ​​determined. If the maximum row normalization error is less than a preset convergence threshold, the exponential matrix is ​​considered to have converged, and the preliminary double random matrix can be used as the double random exponential matrix corresponding to the exponential matrix. If the maximum row normalization error is greater than or equal to the preset convergence threshold, the double randomization operation is repeated until the maximum row normalization error is less than the preset convergence threshold, or the number of double randomization rounds exceeds a preset round threshold.

[0114] Optimization Method 2: To improve the efficiency of double randomization of the exponent matrix while maintaining its accuracy, each exponent matrix can be converted from single-precision floating-point format to half-precision floating-point format. In the pre-defined iterations corresponding to each exponent matrix, double randomization is performed on each half-precision matrix to obtain a processed half-precision matrix. Then, each processed half-precision matrix is ​​converted from half-precision floating-point format to single-precision floating-point format. In subsequent iterations corresponding to each matrix to be processed, double randomization is performed on each single-precision matrix to obtain a double-randomized exponent matrix corresponding to each exponent matrix.

[0115] Optimization Method 3: In order to improve the efficiency of double randomization processing of the exponent matrix and achieve load balancing among the computing unit clusters, we can obtain the computing unit clusters in the target vector processor and distribute each exponent matrix evenly to each computing unit cluster; and use each computing unit cluster to perform multiple rounds of double randomization processing on each exponent matrix to obtain the double random exponent matrix corresponding to each exponent matrix.

[0116] The advantage of this setup is that by loading multiple matrices to be processed into a single vector register, a single vector register instruction can perform the same operations on multiple matrices simultaneously, completely resolving the problem of wasted computational resources caused by the mismatch between small matrix sizes and register widths. Secondly, by employing strategies such as transforming all column normalization operations into loading operations with the same step size as row normalization, performing multi-matrix double randomization processing within a local memory buffer, adaptive early termination of iterations, mixed-precision iterations, and parallel computation across multiple computing units, the efficiency of multi-matrix double randomization processing can be effectively improved while maintaining accuracy, thus saving hardware resources.

[0117] Example 5 Figure 5 This is a schematic diagram of the structure of a multi-matrix double randomization device based on a vector processor according to Embodiment 5 of the present invention.

[0118] This embodiment is applicable to situations where multiple matrices to be processed are subjected to double randomization based on a vector processor. This method can be executed by a vector processor-based multi-matrix double randomization device, which can be implemented in hardware and / or software. This vector processor-based multi-matrix double randomization device can be configured in an electronic device, which can be the chip containing the vector processor or a chip control device capable of controlling the aforementioned chip, as long as it can execute the vector processor-based multi-matrix double randomization method. This embodiment of the invention does not limit the specific device type of the electronic device.

[0119] like Figure 5 As shown, the multi-matrix dual randomization device based on a vector processor disclosed in this embodiment includes a matrix and register acquisition module 51, a matrix loading position determination module 52, and a dual randomization processing module 53.

[0120] The matrix and register acquisition module 51 is used to acquire each matrix to be processed in the regular floating-point tensor and each vector register in the target vector processor; wherein the matrix size of each matrix to be processed is equal and the register width of each vector register is equal.

[0121] The matrix loading position determination module 52 is used to determine multiple matrices to be processed corresponding to each vector register when the matrix size is smaller than the register width and the difference between the register width and the matrix size is greater than a preset difference threshold.

[0122] The double randomization processing module 53 is used to load each matrix to be processed into the corresponding vector register, and to perform multiple rounds of double randomization processing on each matrix to be processed in each vector register to obtain the target double random matrix corresponding to each matrix to be processed.

[0123] The technical solution in this embodiment, through the cooperation of the matrix and register acquisition module 51, the matrix loading position determination module 52, and the dual randomization processing module 53, solves the problem that the existing technology of directly loading a single matrix to be processed into a corresponding vector register for dual randomization processing cannot fully utilize the hardware computing power of the vector processor, resulting in reduced efficiency and increased overhead of multi-matrix dual randomization processing. It can fully utilize the hardware computing power of the vector processor, improve the efficiency of multi-matrix dual randomization processing, and reduce the overhead of multi-matrix dual randomization processing.

[0124] Optionally, the matrix loading position determination module 52 is specifically used to: determine the number of matrices to be processed that can be stored in each vector register based on the matrix size and register width; and determine multiple matrices to be processed corresponding to each vector register based on the relative position of each matrix to be processed in the regular floating-point tensor, the logical index of each vector register, and the number of matrices to be processed that can be stored in each vector register.

[0125] Optionally, the dual randomization processing module 53 includes a first dual randomization processing unit, a second dual randomization processing unit, a third dual randomization processing unit, a fourth dual randomization processing unit, and a fifth dual randomization processing unit.

[0126] The first double randomization processing unit is used to perform row normalization processing on each matrix to be processed in each vector register when each matrix to be processed is arranged in row-major order in the corresponding vector register, to obtain multiple row normalized matrices; to perform transpose processing on each row normalized matrix, to obtain multiple transpose matrices; to perform row normalization processing on each transpose matrix, to obtain multiple double normalized matrices; and to perform transpose processing on each double normalized matrix, to obtain the target double randomized matrix corresponding to each matrix to be processed.

[0127] The second double randomization processing unit is used to perform exponential calculation on each corresponding matrix to be processed in each vector register to obtain the exponential matrix corresponding to each matrix to be processed; obtain the local memory buffer on the chip where the vector processor is located, and write each exponential matrix into the local memory buffer; and perform multiple rounds of double randomization processing on each exponential matrix in the local memory buffer to obtain the target double random matrix corresponding to each matrix to be processed.

[0128] The third double-randomization processing unit is used to perform one round of double-randomization processing on a matrix to be processed in the vector register to obtain a preliminary double-random matrix corresponding to the matrix to be processed; determine the row normalization error vector corresponding to the preliminary double-random matrix based on the row sum vector and the target row sum value; determine the maximum row normalization error corresponding to the preliminary double-random matrix based on the row normalization error vector; if the maximum row normalization error is greater than or equal to a preset convergence threshold, after using the preliminary double-random matrix as a new matrix to be processed, return to the operation of performing one round of double-randomization processing on a matrix to be processed in the vector register to obtain a preliminary double-random matrix corresponding to the matrix to be processed, until the maximum row normalization error is less than the preset convergence threshold, or the number of double-randomization processing rounds is greater than the preset round threshold, then use the preliminary double-random matrix as the target double-random matrix corresponding to the matrix to be processed.

[0129] The fourth double randomization processing unit is used to obtain the preset iteration rounds corresponding to each matrix to be processed, as well as the pre-set rounds and subsequent iteration rounds in each preset iteration round; convert each matrix to be processed from single-precision floating-point format to half-precision floating-point format to obtain half-precision format matrices corresponding to each matrix to be processed; in the pre-set rounds corresponding to each matrix to be processed, double randomization processing is performed on each half-precision format matrix to obtain processed half-precision format matrices; convert each processed half-precision format matrix from half-precision floating-point format to single-precision floating-point format to obtain single-precision format matrices corresponding to each matrix to be processed; in the subsequent iteration rounds corresponding to each matrix to be processed, double randomization processing is performed on each single-precision format matrix to obtain target double random matrices corresponding to each matrix to be processed.

[0130] The fifth dual randomization processing unit is used to obtain each computing unit cluster in the target vector processor and evenly distribute each matrix to be processed to each computing unit cluster; each computing unit cluster is used to perform multiple rounds of dual randomization processing on each matrix to be processed in each vector register to obtain the target dual random matrix corresponding to each matrix to be processed.

[0131] Optionally, the second double randomization processing unit is specifically used to: determine the maximum element in each matrix to be processed, and generate a maximum value matrix corresponding to each matrix to be processed based on the maximum element corresponding to each matrix to be processed; subtract the corresponding maximum value matrix from each matrix to be processed to obtain an offset matrix corresponding to each matrix to be processed; and perform exponential operation on each offset matrix to obtain an exponential matrix corresponding to each matrix to be processed.

[0132] The vector processor-based multi-matrix double randomization apparatus provided in this embodiment of the invention can execute the vector processor-based multi-matrix double randomization method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution. Content not described in detail in this embodiment can be referred to the description in any method embodiment of this application.

[0133] Example 6 Figure 6 A schematic diagram of the structure of an electronic device 10 that can be used to implement embodiments of the present invention is shown. For example... Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or loaded from storage unit 18 into the random access memory 13. The random access memory 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output interface 15 is also connected to the bus 14.

[0134] Multiple components in electronic device 10 are connected to input / output interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0135] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the multi-matrix double randomization method based on a vector processor.

[0136] In some embodiments, the vector processor-based multi-matrix double randomization method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the vector processor-based multi-matrix double randomization method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the vector processor-based multi-matrix double randomization method by any other suitable means (e.g., by means of firmware).

[0137] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0138] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0139] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0141] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0142] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0143] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for multiple matrix double randomization based on a vector processor, characterized in that, The vector processor-based multi-matrix double randomization method includes: Obtain each matrix to be processed in the rule floating-point tensor, and each vector register in the target vector processor; wherein, the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal; When the matrix size is smaller than the register width, and the difference between the register width and the matrix size is greater than a preset difference threshold, multiple matrices to be processed corresponding to each of the vector registers are determined. Each of the matrices to be processed is loaded into the corresponding vector register, and each of the matrices to be processed in each of the vector registers is subjected to multiple rounds of double randomization processing to obtain the target double random matrix corresponding to each of the matrices to be processed.

2. The vector processor based multi-matrix double randomization method according to claim 1, wherein, Determine multiple matrices to be processed corresponding to each of the aforementioned vector registers, including: Based on the matrix size and the register width, determine the number of matrices to be processed that each vector register can store; Based on the relative position of each matrix to be processed in the regular floating-point tensor, the logical index of each vector register, and the number of matrices to be processed that each vector register can store, a plurality of matrices to be processed corresponding to each vector register are determined.

3. The multi-matrix double randomization method based on a vector processor according to claim 1, characterized in that, Each matrix to be processed in each of the vector registers is subjected to multiple rounds of double randomization processing to obtain a target double random matrix corresponding to each matrix to be processed, including: When each of the matrices to be processed is arranged in row-major order in the corresponding vector register, row normalization is performed on each of the matrices to be processed in each of the vector registers to obtain multiple row-normalized matrices. Each row-normalized matrix is ​​transposed to obtain multiple transposed matrices. Each transpose matrix is ​​subjected to row normalization to obtain multiple double-normalized matrices; Each of the double normalized matrices is transposed to obtain the target double random matrix corresponding to each of the matrices to be processed.

4. The multi-matrix double randomization method based on a vector processor according to claim 1, characterized in that, Each matrix to be processed in each of the vector registers is subjected to multiple rounds of double randomization processing to obtain a target double random matrix corresponding to each matrix to be processed, including: Within each of the vector registers, exponent calculations are performed on the corresponding matrices to be processed to obtain exponent matrices corresponding to each of the matrices to be processed. Obtain the local memory buffer on the chip where the vector processor is located, and write each of the exponent matrices into the local memory buffer; Within the local memory buffer, each of the exponent matrices undergoes multiple rounds of double randomization to obtain target double random matrices corresponding to each of the matrices to be processed.

5. The multi-matrix double randomization method based on a vector processor according to claim 4, characterized in that, Within each of the vector registers, exponent calculations are performed on the corresponding matrices to be processed, resulting in exponent matrices corresponding to each of the matrices to be processed, including: The maximum element in each of the matrices to be processed is determined, and a maximum value matrix corresponding to each of the matrices to be processed is generated based on the maximum element corresponding to each of the matrices to be processed. By subtracting the corresponding maximum value matrix from each of the matrices to be processed, the offset matrices corresponding to each of the matrices to be processed are obtained; Perform exponential operations on each of the offset matrices to obtain exponential matrices corresponding to each of the matrices to be processed.

6. The multi-matrix double randomization method based on a vector processor according to claim 1, characterized in that, The matrix to be processed in the vector register is subjected to multiple rounds of double randomization to obtain a target double random matrix corresponding to the matrix to be processed, including: A round of double randomization is performed on a matrix to be processed in the vector register to obtain a preliminary double random matrix corresponding to the matrix to be processed. Based on the row sum vector and row sum target value corresponding to the preliminary double random matrix, determine the row normalization error vector corresponding to the preliminary double random matrix; Based on the row normalization error vector corresponding to the preliminary double random matrix, determine the maximum row normalization error corresponding to the preliminary double random matrix; If the maximum row normalization error is greater than or equal to a preset convergence threshold, then after using the preliminary double random matrix as the new matrix to be processed, the process returns to performing one round of double randomization on a matrix to be processed in the vector register to obtain a preliminary double random matrix corresponding to the matrix to be processed. This process continues until the maximum row normalization error is less than the preset convergence threshold, or the number of rounds of double randomization is greater than the preset round threshold. At this point, the preliminary double random matrix is ​​used as the target double random matrix corresponding to the matrix to be processed.

7. The multi-matrix double randomization method based on a vector processor according to claim 1, characterized in that, Each matrix to be processed in each of the vector registers is subjected to multiple rounds of double randomization processing to obtain a target double random matrix corresponding to each matrix to be processed, including: Obtain the preset iteration rounds corresponding to each of the matrices to be processed, as well as the previous and subsequent iteration rounds in each preset iteration round; Each of the matrices to be processed is converted from single-precision floating-point format to half-precision floating-point format to obtain half-precision format matrices corresponding to each of the matrices to be processed. In the predetermined rounds corresponding to each of the matrices to be processed, the half-precision format matrices are subjected to double randomization to obtain the processed half-precision format matrices. The processed half-precision format matrices are converted from half-precision floating-point format to single-precision floating-point format to obtain single-precision format matrices corresponding to each of the matrices to be processed. In subsequent iterations corresponding to each of the matrices to be processed, the single-precision format matrices are subjected to double randomization to obtain target double random matrices corresponding to each of the matrices to be processed.

8. The multi-matrix double randomization method based on a vector processor according to claim 1, characterized in that, Each matrix to be processed in each of the vector registers is subjected to multiple rounds of double randomization processing to obtain a target double random matrix corresponding to each matrix to be processed, including: Obtain each computing unit cluster in the target vector processor, and evenly distribute each matrix to be processed to each computing unit cluster; Each computing unit cluster is used to perform multiple rounds of double randomization processing on each matrix to be processed in each vector register to obtain a target double random matrix corresponding to each matrix to be processed.

9. A multi-matrix double randomization device based on a vector processor, characterized in that, The vector processor-based multi-matrix double randomization device includes: The matrix and register acquisition module is used to acquire each matrix to be processed in the regular floating-point tensor, and each vector register in the target vector processor; wherein, the matrix size of each matrix to be processed is equal, and the register width of each vector register is equal; The matrix loading position determination module is used to determine multiple matrices to be processed corresponding to each of the vector registers when the size of the matrix is ​​smaller than the width of the register and the difference between the width of the register and the size of the matrix is ​​greater than a preset difference threshold. The double randomization processing module is used to load each of the matrices to be processed into the corresponding vector register, and to perform multiple rounds of double randomization processing on each of the matrices to be processed in each of the vector registers to obtain the target double random matrix corresponding to each of the matrices to be processed.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the vector processor-based multi-matrix double randomization method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method, device and equipment for realizing matrix operation on TPU (Thermoplastic Polyurethane) and medium

    CN120596777A

  • Hardware accelerator facing sparse matrix vector multiplication, equipment and application method

    CN121743654A