Weight matrix acceleration system and method

The weight matrix acceleration system addresses the computational and memory challenges of point cloud Transformer architectures by employing sparse feature extraction and dynamic filtering, enhancing real-time processing and deployment efficiency.

CN120318478AActive Publication Date: 2025-07-15NANJING UNIV

Patent Information

Application Number
CN202510786912.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-15
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing point cloud Transformer architecture needs to intensively calculate the high-weight matrix when generating Q, K, and V matrices, resulting in a sharp increase in computing power demand in large-scale point cloud data, which is difficult to meet real-time requirements. Moreover, the QKV full connection layer parameters are large, making it difficult to deploy on edge computing devices. At the same time, sparse technology will destroy the local geometric structural characteristics of point cloud data, resulting in a decline in key indicators.

Method used

The data preprocessing module is used to map point cloud data to the voxel grid, and filter it through the TopK model, threshold model and zero-value model to generate sparse point cloud data; the features are extracted using 3D sparse convolution operation to generate sparse weight matrix, and the TopK weight in the sparse weight matrix is retained through the TopK weight model, reducing the computational complexity and retaining key features.

Benefits of technology

Significantly reduces computing complexity, improves computing efficiency by more than 50%, maintains lossless accuracy of key indicators, and is suitable for edge computing device deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318478A_ABST
    Figure CN120318478A_ABST
Patent Text Reader

Abstract

The invention provides a weight matrix acceleration system and method, and relates to the technical field of accelerators, and the system comprises a data preprocessing module and a sparse feature extraction module. The data preprocessing module is configured to obtain point cloud data; mapping data points in the point cloud data into a voxel grid; filtering the point cloud data to obtain sparse point cloud data; the sparse feature extraction module is configured to extract features of sparse point cloud data by using sparse convolution operation to obtain data dimensions; obtaining an initialized weight matrix; generating a mask matrix; generating a sparse weight matrix; based on the sparse weight matrix and the data dimension, retaining the TopK weight of each row in the sparse weight matrix to obtain a QKV matrix; the problems that when a QKV matrix is generated by an existing point cloud Transform architecture, dense calculation needs to be carried out on a high-dimensional weight matrix, so that the calculation power requirement in large-scale point cloud data is dramatically increased, and the real-time requirement is difficult to meet are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of accelerator technology, and in particular to a weight matrix acceleration system and method. Background Art

[0002] With the wide application of 3D point cloud data in fields such as autonomous driving, robot perception, and 3D reconstruction, efficient and accurate point cloud feature extraction and understanding have become core requirements. For example, in the autonomous driving scenario, real-time 3D object detection relies on the efficient processing of large-scale point clouds, and it is necessary to reduce the computational latency while ensuring accuracy; in robot navigation, it is necessary to quickly analyze the geometric structure of complex scenes to achieve obstacle avoidance and path planning. However, current point cloud processing algorithms have limitations in feature expression ability and are difficult to capture the long-range dependence and global context information of point cloud data.

[0003] To solve the problem of global feature modeling of point clouds, a point cloud Transformer architecture has been proposed in the prior art. The core idea is to achieve global interaction of point cloud features through the self-attention mechanism. Specifically, this type of method first projects the point cloud features into query (Q), key (K), and value (V) matrices, then calculates the feature weights through dot product attention, and finally fuses the global information. Some methods further introduce local window attention, which divides the point cloud into local regions of a fixed size and calculates the attention only within the regions to reduce the computational complexity. In addition, sparsification techniques (such as weight pruning and low-rank decomposition) are currently also applied to compress the number of parameters in the QKV fully connected layer to reduce memory occupancy.

[0004] However, when generating the Q, K, and V matrices in the current point cloud Transformer architecture, dense calculations need to be performed on high-dimensional weight matrices, resulting in a sharp increase in computing power requirements in large-scale point cloud data (such as in the autonomous driving scenario) and making it difficult to meet the real-time requirements; the number of parameters in the QKV fully connected layer is quadratic with the feature dimension, resulting in a large model volume and making it difficult to deploy on edge computing devices; although the sparsification technique can reduce the amount of calculation, it will destroy the local geometric structure features of the point cloud data, resulting in a significant reduction in key indicators (such as mIoU). Summary of the Invention

[0005] This application provides a weight matrix acceleration system and method to solve the technical problem that when generating the Q, K, and V matrices in the current point cloud Transformer architecture, dense calculations need to be performed on high-dimensional weight matrices, resulting in a sharp increase in computing power requirements in large-scale point cloud data and making it difficult to meet the real-time requirements.

[0006] In the first aspect of this application, a weight matrix acceleration system is provided, including: A data preprocessing module and a sparse feature extraction module; The data preprocessing module is configured to: Obtain point cloud data; Map the data points in the point cloud data to a voxel grid; Based on the point cloud data, use a preset model to filter the point cloud data to obtain sparse point cloud data; the preset model includes: a TopK model, a threshold model, and a zero-value model; The sparse feature extraction module is configured to: Use 3D sparse convolution operations to extract the features of the sparse point cloud data to obtain the data dimension; Obtain an initial weight matrix; Generate a mask matrix based on the initial weight matrix; Generate a sparse weight matrix based on the initial weight matrix and the mask matrix; Based on the sparse weight matrix and the data dimension, use the TopK weight model to retain the TopK weights in each row of the sparse weight matrix to obtain a QKV matrix; the TopK weight model is configured to obtain the TopK weights in each row of the sparse weight matrix using the TopK weight formula; the TopK weight formula is: Q / K / V = ; In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; W sparse is the sparse weight matrix; F emb is the data dimension.

[0007] In some embodiments, after the step of obtaining the point cloud data, it includes: Map the data point coordinates of the point cloud data to the normalized space.

[0008] In some embodiments, the step of mapping the data points in the point cloud data to a voxel grid includes: Determine the voxel grid size according to the maximum data point coordinates and the minimum data point coordinates of the point cloud data; Establish a voxel grid according to the voxel grid size; Calculate the voxel index according to the point cloud data and the voxel grid; the voxel index is: i = ,[[]]END]] j = ,[[]]END]]k = ; In the formula, the data point coordinates of the point cloud data are (x, y, z); v is the boundary length of the voxel grid; Map the data points in the point cloud data to the voxel grid according to the voxel index.

[0009] In some embodiments, the step of filtering the point cloud data based on the point cloud data by using a preset model to obtain sparse point cloud data includes: Repeatedly obtain a preset number of data points in the point cloud data, and sequentially input the data points into the TopK model, the threshold model, and the zero-value model for filtering until all the data points in the point cloud data are filtered, so as to obtain sparse point cloud data.

[0010] In some embodiments, the TopK model is configured as: Obtain the feature vectors of each data point in the point cloud data; Calculate the L2 norm of each data point in the point cloud data by using the feature vectors; Sort each data point in the point cloud data in descending order according to the L2 norm, and retain the data points ranked in the top 30% in the point cloud data.

[0011] In some embodiments, the threshold model is configured as: Obtain the feature vectors of each data point in the point cloud data; Calculate the L2 norm of each data point in the point cloud data by using the feature vectors; Retain the data points in the point cloud data whose L2 norm is greater than 0.1.

[0012] In some embodiments, the zero-value model is configured as: Obtain the feature vectors of each data point in the point cloud data; Judge whether there are data points with the feature vector being 0 in the point cloud data. If so, eliminate the data points with the feature vector being 0.

[0013] In some embodiments, a plurality of windows are set in the QKV matrix, and the sequence lengths of the QKV matrix within the windows are the same; a calculation unit is set within the windows; the calculation unit is configured to perform sparse attention calculation within the window where it is located.

[0014] In some embodiments, the system further includes: An MLP feature enhancement module, and the MLP feature enhancement module is configured as: Determine the original number of channels based on the data dimension; Expand the number of input feature channels of the hidden layer to 4 times the original number of channels; Restore the number of input feature channels of the activation layer to the original number of channels.

[0015] The second aspect of the present application provides a method for accelerating the weight matrix, which is applied to a weight matrix acceleration system described in any one of the above first aspects, and includes: Obtain point cloud data; Map the data points in the point cloud data to a voxel grid; Based on the point cloud data, use a preset model to filter the point cloud data to obtain sparse point cloud data; the preset model includes: TopK model, threshold model, zero value model; Use 3D sparse convolution operation to extract the features of the sparse point cloud data to obtain the data dimension; Obtain an initial weight matrix; Generate a mask matrix based on the initial weight matrix; Generate a sparse weight matrix based on the initial weight matrix and the mask matrix; Based on the sparse weight matrix and the data dimension, use the TopK weight model to retain the TopK weights in each row of the sparse weight matrix to obtain the QKV matrix; the TopK weight model is configured to obtain the TopK weights in each row of the sparse weight matrix using the TopK weight formula; the TopK weight formula is: Q / K / V = ; In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; W sparse is the sparse weight matrix; F emb is the data dimension.

[0016] This application provides a weight matrix acceleration system and method. The system includes a data preprocessing module and a sparse feature extraction module. The data preprocessing module is configured to: obtain point cloud data; map the data points in the point cloud data to a voxel grid; filter the point cloud data based on the point cloud data using a preset model to obtain sparse point cloud data. The preset model includes a TopK model, a threshold model, and a zero value model. The sparse feature extraction module is configured to: extract the features of the sparse point cloud data using 3D sparse convolution operations to obtain the data dimension; obtain an initial weight matrix; generate a mask matrix based on the initial weight matrix; generate a sparse weight matrix based on the initial weight matrix and the mask matrix; retain the TopK weights of each row in the sparse weight matrix based on the sparse weight matrix and the data dimension using the TopK weight model to obtain a QKV matrix. The TopK weight model is configured to obtain the TopK weights of each row in the sparse weight matrix using the TopK weight formula. The TopK weight formula is: Q / K / V = ; In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; W sparse is the sparse weight matrix; F emb is the data dimension, so as to solve the problem that when the current point cloud Transformer architecture generates Q, K, and V matrices, it is necessary to perform intensive calculations on a high-dimensional weight matrix, resulting in a sharp increase in computing power requirements in large-scale point cloud data and making it difficult to meet the real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of this application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic structural diagram of the weight matrix acceleration system in this application; Figure 2 is a flowchart of the data preprocessing module during operation in this application; Figure 3 is a flowchart of the sparse feature extraction module during operation in this application; Figure 4It is a comparison curve graph of the effective value density under the original point cloud Transformer architecture and the point cloud Transformer architecture in this application.

[0019] Explanation of the reference numerals: 1 - Data preprocessing module; 2 - Sparse feature extraction module; 3 - MLP feature enhancement module. Detailed implementation manners

[0020] In order to enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0021] Exemplarily, Point Cloud Transformer (such as PointCloudTransformerV3) performs excellently in tasks such as 3D object detection and point cloud segmentation. However, its self-attention mechanism needs to calculate the global relationship between all points, resulting in the computational complexity of generating the QKV matrix being O(N 2 ) (N is the number of data points), which is difficult to meet the real-time requirements.

[0022] Exemplarily, 1. The computational efficiency of the point cloud Transformer network is low and the computational complexity is high: Currently, when generating query (Q), key (K), and value (V) matrices in the point cloud Transformer (taking PointCloudTransformerV3 as an example), dense calculations need to be performed on high-dimensional weight matrices, resulting in a sharp increase in computing power requirements, especially in large-scale point cloud data (such as 3D object detection in autonomous driving), it is difficult to be deployed in real time. 2. Large memory occupancy: The number of parameters in the QKV fully connected layer is large, resulting in difficulties in deploying on edge devices. 3. Redundancy of invalid features: There are a large number of low-information points (such as ground points, sensor noise) in the original point cloud data, and currently, they cannot be dynamically filtered. 4. Currently, the accuracy loss of sparse methods is significant: Although current weight pruning or low-rank approximation methods can reduce the amount of calculation, they will damage the local geometric structure features of the point cloud data, resulting in a decrease in key indicators such as mIoU.

[0023] To solve this technical problem, this application provides a weight matrix acceleration system and method. The following will explain the weight matrix acceleration system and method: As Figure 1 shown, it is a schematic structural diagram of the weight matrix acceleration system in this application.

[0024] The first aspect of this application provides a weight matrix acceleration system, including: The data preprocessing module 1 and the sparse feature extraction module 2.

[0025] Such as Figure 2 shown, it is the flowchart when the data preprocessing module in this application runs.

[0026] The data preprocessing module 1 is configured to: Obtain point cloud data; the point cloud data P is: R N×3 ; R is the set of real numbers in the point cloud data, and N is the three-dimensional coordinates of the data points in the point cloud data.

[0027] After the step of obtaining the point cloud data, the following steps are included: Map the data point coordinates of the point cloud data to the normalized space.

[0028] Specifically, the normalization operation of the data point coordinates of the point cloud data is as follows: first, translate the coordinates to the origin (subtract the minimum value), and then scale them to the grid unit (divide by grid_size = x max - x min ), and finally round to obtain the discrete grid coordinates.

[0029] The functions of mapping the data point coordinates of the point cloud data to the normalized space are as follows: 1. Eliminate scale differences: The point cloud data may come from different sensors or scenarios, and the coordinate ranges vary greatly. The normalization operation can map all coordinates to a unified scale, enabling the model to maintain a consistent response to data from different sources. 2. Accelerate convergence: Neural networks are easier to learn when processing normalized data because the range of input values is constrained and the gradient updates are more stable, thus accelerating the training process.

[0030] Map the data points in the point cloud data to the voxel grid.

[0031] The step of mapping the data points in the point cloud data to the voxel grid includes the following sub-steps: Determine the voxel grid size according to the maximum data point coordinates and the minimum data point coordinates of the point cloud data; through the maximum data point coordinates x max and the minimum data point coordinates x min of the point cloud data, the boundary of the voxel grid can be determined, that is, the size of the voxel grid.

[0032] Establish a voxel grid according to the voxel grid size; through the size of the voxel grid, the voxel grid size can be established, and the voxel grid is cube-shaped.

[0033] Calculate the voxel index according to the point cloud data and the voxel grid; the voxel index is: i = , j = , k = ; wherein, the data point coordinates of the point cloud data are x, y, z; v is the boundary length of the voxel grid; represents rounding down.

[0034] According to the voxel index, map the data points in the point cloud data to the voxel grid. Voxelization is used for dimensionality reduction and sparse representation of point cloud data: point cloud data is irregular three-dimensional data, and direct processing has a large computational cost; through voxelization, the point cloud data is divided into regular three-dimensional grids, and each voxel is represented by a numerical value, greatly reducing the data volume. Retain the spatial structure: voxelization retains the spatial relationship between points through grid indexing, enabling subsequent convolution operations to effectively extract local features.

[0035] Based on the point cloud data, use a preset model to filter the point cloud data to obtain sparse point cloud data; the preset model includes: TopK model, threshold model, zero-value model. The sparse point cloud data P sparse is: R M×3 ; M is the number of data points in the point cloud data after filtering by the preset model.

[0036] The step of filtering the point cloud data based on the point cloud data using a preset model to obtain sparse point cloud data further includes the following steps: Repeatedly obtain a preset number of data points in the point cloud data, and sequentially input the data points into the TopK model, threshold model, and zero-value model for filtering until all the data points in the point cloud data are filtered, obtaining sparse point cloud data. This application adopts a dynamic filtering mode to filter the data points in the point cloud data. By obtaining a preset number of data points from the point cloud data each time, and then sequentially inputting them into the TopK model, threshold model, and zero-value model for filtering, dynamically select the remaining data points according to the weight importance, avoiding accuracy loss caused by a fixed pruning rate.

[0037] In this embodiment, the TopK model is configured as: Obtain the feature vectors of each data point in the point cloud data; for example, the feature vector of data point P is (x, y, z).

[0038] Using the feature vector, calculate the L2 norm of each data point in the point cloud data; the L2 norm (Euclidean norm) is the square root of the sum of the squares of the vector elements and is used to measure the length or distance of the vector. For example, the L2 norm of the data point P is: = .

[0039] The data points in the point cloud data are sorted from large to small according to the L2 norm, and the top 30% of the data points in the point cloud data are retained.

[0040] Specifically, the implementation of the TopK model: filter_k=0.3 means retaining the top 30% of the important points input to the TopK model. For example, if there are 1,000 data points, then retain the 300 data points with the largest L2 norm. The corresponding code is represented as: point_cloud.sparsify(filter_mode="topk", filter_k=0.3) .

[0041] In this embodiment, the threshold model is configured as: Acquire a feature vector of each data point in the point cloud data; calculate the L2 norm of each data point in the point cloud data using the feature vector; and retain the data points in the point cloud data whose L2 norm is greater than 0.1.

[0042] Specifically, the implementation of the threshold mode: also using the L2 norm, filter_threshold=0.1 means retaining the data points whose L2 norm is greater than 0.1 in the point cloud data. The corresponding code is expressed as: point_cloud.sparsify(filter_mode="threshold", filter_threshold=0.1).

[0043] In this embodiment, the zero value model is configured as: Obtain a feature vector of each data point in the point cloud data; determine whether there is a data point in the point cloud data whose feature vector is 0, and if so, remove the data point whose feature vector is 0.

[0044] Specifically, the zero - value mode is implemented as follows: by checking whether the feature vector of each data point is all zero. If all dimensions are 0, then this point is considered an invalid point. Among them, in the code, a boolean mask is generated through torch.sum(self.feat != 0, dim = 1)==0 to mark the all - zero feature points. In this mode, the filter_k parameter is invalid, and all non - zero feature points are directly retained. The corresponding code is as follows: point_cloud.sparsify(filter_mode = "zero_feat");

[0045] As Figure 3 shown, it is the flowchart when the sparse feature extraction module in this application runs.

[0046] The sparse feature extraction module 2, that is, the top - K weight sparse linear layer, is configured as follows: Using 3D sparse convolution operations to extract the features of the sparse point cloud data, obtaining the data dimension; the data dimension F emb is R M×C ; C is the number of channels of the output features.

[0047] Obtain the initial weight matrix; the initial weight matrix is randomly generated by the system.

[0048] Based on the initial weight matrix, generate a mask matrix; the mask matrix has the same shape as the initial weight matrix, and the weights of the mask matrix are 0 or 1. Among them, 30% of the values in the mask matrix are assigned 1, and the rest are assigned 0. The assignment of 1 and 0 is determined according to the k value of filter_k in the TopK mode.

[0049] Based on the initial weight matrix and the mask matrix, generate a sparse weight matrix; the sparse weight matrix is: W sparse = ; In the formula, W is the initial weight matrix; M is the mask matrix.

[0050] Based on the sparse weight matrix and the data dimension, using the TopK weight model, retain the top - K weights of each row in the sparse weight matrix to obtain the QKV matrix; the TopK weight model is configured to use the TopK weight formula to obtain the top - K weights of each row in the sparse weight matrix; the TopK weight formula is: Q / K / V = ; In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; W sparse is the sparse weight matrix; F emb is the data dimension.

[0051] In this embodiment, a plurality of windows are set in the QKV matrix, and the sequence lengths of the QKV matrix within the windows are the same; a computing unit is provided within the windows; the computing unit is configured to perform sparse attention calculation within the window where it is located. By performing attention calculation through local windows, the computational complexity is reduced from O(M 2 ) to O(MK).

[0052] Specifically, in the local window attention mechanism, the QKV matrix is divided into multiple non-overlapping windows, and attention is calculated independently within each window. The window division is based on the position order rather than the attention intensity, and data points with the same attention are not necessarily divided into the same window. The specific steps are as follows: 1. Window division: Each window contains 48 tokens. If the sequence length is not an integer multiple of the window size, padding will be performed. The sequence is divided into multiple consecutive windows in order, and the tokens within each window calculate attention with each other.

[0053] 2. Attention calculation within the window: Inside each window, the attention calculation is the same as the standard self-attention. The tokens within each window only calculate attention with other tokens within the same window, and there is no direct interaction between windows. The attention calculations of all windows can be processed in parallel to improve efficiency.

[0054] Among them, although the attention is calculated independently within each window, the subsequent layers will mix the information between windows through residual connections and MLP layers. The computational complexity is reduced from O(M 2 ) to O(MK), where M is the sequence length and K is the window size, and K = 48.

[0055] In this embodiment, the system further includes: an MLP (Multi-Layer Perceptron) feature enhancement module, and the MLP feature enhancement module 3 is configured to: Based on the data dimension, determine the original number of channels C; expand the number of input feature channels of the hidden layer to 4 times the original number of channels; restore the number of input feature channels of the activation layer to the original number of channels. Among them, the way of dimension increase first (C→4C) is to expand the feature expression with a higher dimension and refine the non-linear relationship, that is, expand the number of input feature channels of the hidden layer to 4 times the original number of channels; then the way of dimension reduction (4C→C) is to output the features learned in the high-dimensional space in the form of the original number of channels C, retain the information after high-dimensional transformation, and adapt to the dimension requirements of subsequent residual connections and feature fusion, that is, restore the number of input feature channels of the activation layer to the original number of channels.

[0056] Specifically, the MLP feature enhancement module 3 includes two cascaded fully connected layers (the hidden layer dimension is 4C) and the GELU activation function.

[0057] The two cascaded fully connected layers expand the number of channels of the input features by 4 times, increasing the non-linear expression ability. The GELU activation layer restores the number of channels to the original dimension C, completing the compression of feature information, so that after the model completes feature transformation in the high-dimensional space, it returns to the original feature space. Through the setting of the MLP feature enhancement module 3, the residual connections are combined together, and the independent attention information of each QKV window is fused together to achieve inter-level feature transfer.

[0058] As Figure 4 shown, it is a comparison curve graph of the effective value density under the original point cloud Transformer architecture and the point cloud Transformer architecture in this application.

[0059] Exemplarily, by adopting the weight matrix acceleration system provided in this application, the K elements with the largest absolute values in the weight matrix are retained (the rest are set to zero), significantly reducing the effective value density of the model (that is, the proportion of non-zero weights), as Figure 4 shown, Encoder 0 (module layer × 2); Encoder 1 (pooling + module layer × 2); Encoder 2 (pooling + module layer × 2); Encoder 3 (pooling + module layer × 6); Encoder 4 (pooling + module layer × 2); Decoder 3 (unpooling + module layer × 2); Decoder 2 (unpooling + module layer × 2); Decoder 1 (unpooling + module layer × 2); Decoder 0 (unpooling + module layer × 2). This sparsification not only reduces the computational and storage overhead, but also improves the efficiency of hardware deployment.

[0060] Exemplarily, the current optimization solutions for point cloud Transformer focus on: Traditional point cloud Transformer architectures ("Point Transformer V2: Grouped Vector Attention and Partition-based Pooling." NeurIPS, 2021; "PCT: Point Cloud Transformer." Computational Visual Media, 2021) use K-Nearest Neighbor (KNN) for neighborhood search and rely on Relative Position Encoding (RPE) to capture spatial relationships, resulting in high computational complexity and large memory consumption. After the weight matrix acceleration system provided in this application is applied to the point cloud Transformer architecture, KNN neighborhood search is not required, avoiding a large number of distance calculation and sorting operations. By directly sparsifying the weight matrix, unnecessary computational connections are reduced.

[0061] This application provides a weight matrix acceleration system, which has the following beneficial effects: 1. Improved computational efficiency: Through the TopK sparsification strategy, only the TopK non-zero values in the QKV weight matrix are retained (e.g., K = 30%), reducing the computational complexity by more than 50% by skipping zero-value calculations.

[0062] 2. Precision-lossless optimization: Based on the dynamic sparsification strategy and gradient reparameterization technology, at a sparsity of K = 30%, the mIoU metric of the point cloud segmentation task decreases by less than 0.5%.

[0063] The second aspect of this application provides a weight matrix acceleration method, which is applied to a weight matrix acceleration system described in any of the above embodiments, and includes: Obtain point cloud data; Map the data points in the point cloud data to a voxel grid; Based on the point cloud data, use a preset model to filter the point cloud data to obtain sparse point cloud data; the preset model includes: a TopK model, a threshold model, and a zero-value model; Use 3D sparse convolution operations to extract the features of the sparse point cloud data to obtain the data dimension; Obtain an initial weight matrix; Generate a mask matrix based on the initial weight matrix; Generate a sparse weight matrix based on the initial weight matrix and the mask matrix; Based on the sparse weight matrix and the data dimension, using the TopK weight model, retain the TopK weights of each row in the sparse weight matrix to obtain the QKV matrix; the TopK weight model is configured to obtain the TopK weights of each row in the sparse weight matrix using the TopK weight formula; the TopK weight formula is: Q / K / V = ; In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; W sparse is the sparse weight matrix; F emb is the data dimension.

[0064] It should be noted that for the effects of the above method embodiments, reference can be made to the effects of the above system embodiments, which will not be elaborated here.

[0065] The above specific implementation manners further elaborate in detail the purpose, technical solutions, and beneficial effects of the embodiments of the present application. It should be understood that the above are only the specific implementation manners of the embodiments of the present application and are not used to limit the protection scope of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application shall be included within the protection scope of the embodiments of the present application.

Claims

1. A weight matrix acceleration system, characterized in that, include: Data preprocessing module (1) and sparse feature extraction module (2); The data preprocessing module (1) is configured as follows: Get point cloud data; Mapping data points in the point cloud data into a voxel grid; Based on the point cloud data, filtering the point cloud data using a preset model to obtain sparse point cloud data; The preset models include: TopK model, threshold model, and zero value model; The sparse feature extraction module (2) is configured as follows: Using a 3D sparse convolution operation, extracting features of the sparse point cloud data to obtain data dimensions; Get the initialization weight matrix; Based on the initialization weight matrix, generating a mask matrix; Based on the initialization weight matrix and the mask matrix, generating a sparse weight matrix; Based on the sparse weight matrix and the data dimension, the TopK weight of each row in the sparse weight matrix is retained by using the TopK weight model to obtain a QKV matrix; the TopK weight model is configured to obtain the TopK weight of each row in the sparse weight matrix by using a TopK weight formula; the TopK weight formula is: Q / K / V= ; Wherein, Q is the query matrix; K is the key matrix; V is the value matrix; W sparse is the sparse weight matrix; F emb is the data dimension.

2. The weight matrix acceleration system according to claim 1, wherein After the step of obtaining point cloud data, the following steps are included: The data point coordinates of the point cloud data are mapped to a normalized space.

3. The weight matrix acceleration system according to claim 1, wherein The step of mapping the data points in the point cloud data into a voxel grid comprises: Determining a voxel grid size according to a maximum data point coordinate and a minimum data point coordinate of the point cloud data; Establishing a voxel grid according to the voxel grid size; According to the point cloud data and the voxel grid, a voxel index is calculated; the voxel index is: i = , j = , k = ; Where, the data point coordinates of the point cloud data are (x, y, z); v is the boundary length of the voxel grid; Data points in the point cloud data are mapped to the voxel grid according to the voxel index.

4. A weight matrix acceleration system according to claim 1, wherein The step of filtering the point cloud data based on the point cloud data using a preset model to obtain sparse point cloud data includes: A preset number of data points in the point cloud data are repeatedly obtained, and the data points are sequentially input into the TopK model, the threshold model, and the zero value model for filtering until all the data points in the point cloud data are filtered to obtain sparse point cloud data.

5. A weight matrix acceleration system according to claim 1, wherein The TopK model is configured as: Obtaining a feature vector of each data point in the point cloud data; Using the feature vector, calculating the L2 norm of each data point in the point cloud data; The data points in the point cloud data are sorted from large to small according to the L2 norm, and the top 30% of the data points in the point cloud data are retained.

6. The weight matrix acceleration system according to claim 1, wherein The threshold model is configured as: Obtaining a feature vector of each data point in the point cloud data; Using the feature vector, calculating the L2 norm of each data point in the point cloud data; The data points in the point cloud data whose L2 norm is greater than 0.1 are retained.

7. A weight matrix acceleration system according to claim 1, characterized in that, The zero-value model is configured as: Obtaining a feature vector of each data point in the point cloud data; It is determined whether there are data points with the eigenvector being 0 in the point cloud data, and if so, the data points with the eigenvector being 0 are removed.

8. A weight matrix acceleration system according to claim 1, characterized in that A number of windows are set in the QKV matrix, and the sequence lengths of the QKV matrix within the windows are the same; a computing unit is set within the windows; the computing unit is configured to perform sparse attention calculation within the window where it is located.

9. The weight matrix acceleration system according to claim 1, wherein The system further includes: an MLP feature enhancement module (3), and the MLP feature enhancement module (3) is configured to: determine the original number of channels based on the data dimension; expand the number of input feature channels of the hidden layer to 4 times the original number of channels; restore the number of input feature channels of the activation layer to the original number of channels.

10. A weight matrix acceleration method, applied to a weight matrix acceleration system according to any one of claims 1 to 9 above, characterized in that, It includes: acquire point cloud data; map the data points in the point cloud data into a voxel grid; filter the point cloud data based on the point cloud data by using a preset model to obtain sparse point cloud data; the preset model includes: a TopK model, a threshold model, and a zero-value model; extract the features of the sparse point cloud data by using 3D sparse convolution operation to obtain a data dimension; acquire an initial weight matrix; generate a mask matrix based on the initial weight matrix; generate a sparse weight matrix based on the initial weight matrix and the mask matrix; based on the sparse weight matrix and the data dimension, use a TopK weight model to retain the TopK weights of each row in the sparse weight matrix to obtain a QKV matrix; the TopK weight model is configured to obtain the TopK weights of each row in the sparse weight matrix by using a TopK weight formula; the TopK weight formula is: Q / K / V = ; In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; W sparse is the sparse weight matrix; F emb is the data dimension.

Citation Information

Patent Citations

  • Point cloud semantic segmentation method based on joint Transform and sparse convolution

    CN116778161A

  • Sparse on-chip training hardware accelerator architecture and implementation method thereof

    CN118760651A

  • Sparse convolutional neural network accelerator for 3d / 4d point-cloud image recognition

    US20230385982A1

Cited By

  • Universal TopK computing device and method based on insertion merging

    CN122450504A