Data processing method and system based on sparsification and related equipment
Through the weighted data sparse method of diagonal matrix analysis, deep learning models are sparsized and rapidly reconstructed, solving the data transmission bottleneck of deep learning models on heterogeneous accelerators and edge computing devices, and improving data compression efficiency and processing speed.
Patent Information
- Application Number
- CN202510711346.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art lacks effective data sparse methods while ensuring the accuracy of deep learning models, resulting in large data transmission volume and low compression efficiency, especially in heterogeneous accelerators and edge computing devices.
The weight data sparse method based on diagonal matrix analysis is used to sparse the pre-trained depth model, expand it into two-dimensional tensor data, and multiple sparse vectors are generated through preset conversion methods, and packaged and transmitted. The hardware is used to realize data packaging and unpacking, and realize rapid structural reconstruction.
Without reducing the accuracy of the model, data compression efficiency is improved, data transmission bottlenecks in the process of large model inference are alleviated, and data processing speed and transmission efficiency are improved.
Smart Images

Figure CN120471175A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of sparse data processing, and more specifically, to a data processing method, system and related equipment based on sparse data processing. Background Art
[0002] The widespread adoption of large-scale deep learning models (such as Deepseek, LLaMA, and GPT) has placed extremely high demands on computing resources and data transmission bandwidth for inference calculations. In actual deployments, especially on heterogeneous accelerators (such as NPUs, FPGAs, ASICs, etc.) or edge computing devices, data transmission has become a key bottleneck in the deep learning inference process.
[0003] Tensor data such as model weights and activation values is often massive, far exceeding on-chip storage capacity and requiring frequent transfer via high-speed interfaces such as AXI, PCIe, and HBM. Therefore, reducing data transfer while maintaining model accuracy has become a key issue in improving overall system inference performance.
[0004] In existing technologies, data compression is often used to effectively reduce data transmission while ensuring model accuracy. Data compression is a common method for optimizing data transmission. For tensor data with high sparsity, sparse coding methods are often used to compress non-zero elements. However, existing compression methods have limited compression rates, especially for tensor data of arbitrary dimensions. There is a lack of universal data sparsification methods, and most sparse compression algorithms are implemented in software, which has high encoding and decoding overhead, thereby reducing data compression efficiency.
[0005] Therefore, how to improve the efficiency of data compression while ensuring model accuracy is an urgent problem to be solved in this application. Summary of the Invention
[0006] In view of this, the present application discloses a data processing method, system and related equipment based on sparsification, aiming to improve the efficiency of data compression while ensuring model accuracy.
[0007] In order to achieve the above purpose, the disclosed technical solutions are as follows:
[0008] In a first aspect, the present application discloses a data processing method based on sparsification, the method comprising:
[0009] Perform model weight data sparsification on the pre-trained deep model to obtain a sparsified deep learning model;
[0010] Acquire the tensor data of the deep learning model after the sparsification, and expand the tensor data into two-dimensional tensor data;
[0011] The two-dimensional tensor data is transformed by a preset transformation method to obtain a plurality of sparse vectors;
[0012] Packing the multiple vectors after the sparse processing to obtain sparse packed data and performing a data transmission operation;
[0013] When the sparse data transmission is completed, the sparse packed data is unpacked to obtain the unpacked data, and actual calculation is performed on the unpacked data.
[0014] Preferably, the performing of model weight data sparsification on the pre-trained deep model to obtain a sparsified deep learning model includes:
[0015] Calculate the diagonal value of each weight in the pre-trained deep model through the diagonal approximation method;
[0016] A preset sparsification threshold is compared with the diagonal value of each weight to set the diagonal value smaller than the sparsification threshold to 0, thereby obtaining a sparsified deep learning model.
[0017] Preferably, obtaining the tensor data of the sparsified deep learning model and expanding the tensor data into two-dimensional tensor data includes:
[0018] Obtaining the order corresponding to the tensor data of the sparsified deep learning model;
[0019] A tensor dimension is determined according to the order, and the tensor data is expanded into two-dimensional tensor data according to the tensor dimension.
[0020] Preferably, the multiple vectors after the sparsification include at least a numerical range boundary vector, an index associated vector, and an index vector, and the converting of the two-dimensional tensor data by a preset conversion method to obtain the multiple vectors after the sparsification includes:
[0021] Determining the number of elements at the boundary of the numerical range, the number of elements associated with the index, and the number of elements of the index from the two-dimensional tensor data; wherein the number of elements at the boundary of the numerical range represents the number of columns of the two-dimensional tensor; the number of elements associated with the index represents whether the data is actually valid; and the number of elements of the index represents the number of valid data;
[0022] According to the number of elements at the boundary of the numerical range, the number of elements associated with the index and the number of elements of the index, the two-dimensional tensor data is vector-converted to obtain a numerical range boundary vector, an index-associated vector and an index vector.
[0023] Preferably, the step of packing the multiple vectors after the sparse processing to obtain the sparse packed data and performing the data transmission operation includes:
[0024] Extracting data content of the multiple vectors after sparsification; wherein the data content includes at least the number of valid elements, the number of columns after expansion, the space occupied by data elements, the data type, and the dimension size before dimensional expansion; the dimension size before dimensional expansion is determined by the input tensor data;
[0025] Determining the data contents of the plurality of vectors as a metadata portion, and determining the plurality of vectors after the sparse processing as a data portion;
[0026] The data portion is placed after the metadata portion to complete the packing of the multiple sparse vectors to obtain sparse packed data, and the sparse packed data is subjected to a data transmission operation.
[0027] Preferably, when the sparse data transmission is completed, the sparse packed data is unpacked to obtain the unpacked data, and actual calculation is performed on the unpacked data, including:
[0028] Upon completion of the sparse data transmission, metadata information corresponding to a fixed position is obtained from the sparse packed data;
[0029] Obtaining numerical range boundary vector data from the sparse packed data according to the number of expanded matrix columns;
[0030] Obtaining index-associated vector data and index vector data according to the number of valid elements in the numerical range boundary vector data;
[0031] Parsing two-dimensional tensor data through the numerical range boundary vector data, the index-associated vector data, and the index vector data;
[0032] Convert the parsed two-dimensional tensor data into original dimensional data content; wherein the original dimensional data content is the tensor data content changed to a preset dimension;
[0033] The original dimensional data content is determined as the unpacked data and actual calculation is performed.
[0034] A second aspect of the present application discloses a data processing system based on sparsification, the system comprising:
[0035] The sparsification unit is used to sparsify the model weight data of the pre-trained deep model to obtain a sparse deep learning model;
[0036] An acquisition unit, configured to acquire tensor data of the sparsified deep learning model and expand the tensor data into two-dimensional tensor data;
[0037] A conversion unit, configured to convert the two-dimensional tensor data using a preset conversion method to obtain a plurality of sparse vectors;
[0038] a packing operation unit, configured to pack the plurality of vectors after the sparse processing to obtain sparse packed data and perform a data transmission operation;
[0039] The unpacking calculation unit is used to unpack the sparse packed data when the sparse data transmission is completed to obtain the unpacked data and perform actual calculation on the unpacked data.
[0040] Preferably, the thinning unit includes:
[0041] A calculation module, used to calculate the diagonal value of each weight in the pre-trained deep model by using a diagonal approximation method;
[0042] The comparison module is used to compare a preset sparsification threshold with the diagonal value of each weight to set the diagonal value less than the sparsification threshold to 0, thereby obtaining a sparsified deep learning model.
[0043] A third aspect of the present application discloses a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the sparsification-based data processing method as described in any one of the first aspects.
[0044] The fourth aspect of the present application discloses an electronic device comprising a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors as the sparsification-based data processing method described in any one of the first aspects.
[0045] It can be seen from the above technical solution that the present application discloses a data processing method, system and related equipment based on sparsification, which performs model weight data sparsification on a pre-trained deep model to obtain a sparse deep learning model, obtains tensor data of the sparse deep learning model, and expands the tensor data into two-dimensional tensor data. The two-dimensional tensor data is transformed by a preset transformation method to obtain multiple sparse vectors, the multiple sparse vectors are packaged to obtain sparse packaged data and perform data transmission operations. When the sparse data transmission is completed, the sparse packaged data is unpacked to obtain unpacked data, and actual calculations are performed on the unpacked data.
[0046] Based on the above scheme, by introducing a method of weight data sparsification based on diagonal matrix analysis, the sparsity of the model tensor can be improved without reducing the accuracy of the model. The tensor data expansion method of this scheme cooperates with the method of packing multiple vectors after sparseness to improve data sparsity efficiency. And based on the data packing and unpacking implemented by hardware, the tensor data can be quickly restructured before and after transmission, significantly reducing the overall delay. Through the above systematic design, this application can effectively alleviate the data transmission bottleneck in the large model reasoning process, improve data compression efficiency and processing speed, and thus achieve more efficient deep learning model transmission and deployment capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0048] Figure 1 A flowchart of a data processing method based on sparsification disclosed in an embodiment of the present application;
[0049] Figure 2 This is an example diagram of the converted vector disclosed in the embodiments of this application;
[0050] Figure 3 A flowchart of another data processing method based on sparsification disclosed in an embodiment of the present application;
[0051] Figure 4 A schematic diagram of the structure of a data processing system based on sparsification disclosed in an embodiment of the present application;
[0052] Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0054] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0055] As can be seen from the background technology, in the prior art, in order to effectively reduce the amount of data transmission while ensuring the accuracy of the model, data compression is usually used. Data compression is a common means of optimizing data transmission. For tensor data with high sparsity, sparse coding methods are usually used to compress non-zero elements. However, the compression rate of existing compression methods is limited, especially for tensor data of any dimension, there is a lack of universal data sparsification methods, and most sparse compression algorithms are implemented in software, with large encoding and decoding overhead, thereby reducing the efficiency of data compression. Therefore, how to improve the efficiency of data compression while ensuring the accuracy of the model is a problem that needs to be solved urgently in this application.
[0056] In order to solve the above problems, the present application discloses a data processing method, system and related equipment based on sparsification. By introducing a weight data sparsification method based on diagonal matrix analysis, the sparsity of the model tensor can be improved without reducing the accuracy of the model. The tensor data expansion method of this solution cooperates with the method of packing multiple vectors after sparsification to improve data sparsity efficiency. And based on hardware-implemented data packing and unpacking, the tensor data can be quickly restructured before and after transmission, significantly reducing the overall delay. Through the above-mentioned systematic design, the present application can effectively alleviate the data transmission bottleneck in the large model reasoning process, improve data compression efficiency and processing speed, and thus achieve more efficient deep learning model transmission and deployment capabilities. The specific implementation method is specifically described through the following embodiments.
[0057] It should be noted that the data processing method, system and related equipment based on sparsification provided in the present application can be used in sparse data processing, data packaging, data transmission and other fields. The above is only an example and does not limit the application field of the data processing method, system and related equipment based on sparsification provided in the present application.
[0058] refer to Figure 1 As shown in FIG, a data processing method based on sparsification disclosed in an embodiment of the present application mainly includes the following steps:
[0059] S101: Performing sparsification on the model weight data of the pre-trained deep model to obtain a sparsified deep learning model.
[0060] According to the actual application scenario, the trained deep learning model is obtained through pre-training model or re-training.
[0061] In S101, the model weight data is sparsified, that is, the Hessian diagonal value of each weight in the pre-trained deep model is calculated by the diagonal (Hessian) approximation method, and a pre-set sparsification threshold is compared with the Hessian diagonal value of each weight to set the Hessian diagonal value less than the sparsification threshold to 0, thereby obtaining a sparsified deep learning model.
[0062] The Hessian diagonal value represents the importance of the weighted data.
[0063] The thinning threshold is set according to actual conditions and is not specifically limited in this application.
[0064] S102: Obtain the tensor data of the deep learning model after sparsification, and expand the tensor data into two-dimensional tensor data.
[0065] In S102, the order corresponding to the tensor data of the sparse deep learning model is obtained through the dimension expansion module, the tensor dimension is determined according to the order, and the tensor data is expanded into two-dimensional tensor data according to the tensor dimension.
[0066] This application is applicable to tensor data of any dimension. Assume that the tensor dimension is T[d1, d2, ..., d n ], expand the tensor data into a two-dimensional vector T[d1*d2*...*d n-1 , d n ], this expansion method does not change the continuity of data storage. At the same time, the number of rows of the expanded tensor is large and the number of columns is small. It can be combined with the following sparsification method to improve the degree of sparsification.
[0067] S103: Transform the two-dimensional tensor data using a preset transformation method to obtain multiple sparse vectors.
[0068] In S103, the two-dimensional tensor data is transformed using a data compression module and a preset transformation method to obtain multiple sparse vectors. The data compression module is implemented in hardware. The expanded two-dimensional tensor data is transformed using the following method to obtain multiple sparse vectors. The multiple sparse vectors include at least three vectors: a value range boundary (Bound) vector, an index association (Value) vector, and an index (Index) vector.
[0069] In vector search, a bound vector is often used to represent the boundaries of a numerical range. For example, in a two-dimensional vector space, a bound vector can represent the boundaries of a rectangular region, which is used to limit the search range.
[0070] The Value vector is often used to represent the value associated with a specific index. In a database system, the Value vector can store the actual data value associated with each index key.
[0071] The Index vector is used to represent the index information of the data. In a vector database, the Index vector usually contains a pointer or reference to the actual data storage location.
[0072] The specific process of obtaining multiple vectors after sparseness is shown in A1-A2.
[0073] A1: Determine the number of Bound elements, Value elements, and Index elements from the two-dimensional tensor data.
[0074] Among them, the number of elements in Bound indicates the number of columns in the two-dimensional tensor; the number of elements in Value indicates whether it is actual valid data; the number of elements in Index indicates the number of valid data (non-zero data).
[0075] A2: Perform vector conversion on the two-dimensional tensor data based on the number of Bound elements, the number of Value elements, and the number of Index elements to obtain the Bound vector, the Value vector, and the Index vector.
[0076] by Figure 2 Taking the two-dimensional tensor in as an example, the transformed Bound vector, the transformed Value vector, and the transformed Index vector are as follows: Bound: [3, 5, 8]; Value: [A, B, D, G, E, C, H, F]; Index: [0, 2, 4, 2, 5, 0, 1, 3].
[0077] The first vector is the Bound vector. The number of elements in the Bound vector represents the number of columns in the two-dimensional tensor, and each element represents the boundary of each column of data, such as Figure 2 In the example, the first element of the Bound vector, 3, indicates that the subscript range of the first column of data in the Value vector is 0-2, and the data are A, B, and D. The second element, 5, indicates that the subscript range of the second column of data in the Value vector is 3-4, and the data are G and E. The third element, 8, indicates that the subscript range of the third column of data in the Value vector is 5-7, and the data are C, H, and F.
[0078] The second vector is the Value vector. The number of elements in the Value vector is equal to the number of valid data (non-zero data). Each element of the Value vector is the actual valid data and is accessed in column-major order. Please refer to the example above.
[0079] The third vector is the Index vector. The number of elements in the Index vector is equal to the number of valid data (non-zero data). Each element of the Index vector represents the row number of each data element. The first element 0 indicates that the actual data A is in row 0. The second element 2 indicates that the actual data A is in row 2. The fourth element 2 indicates that the actual data G is in row 2.
[0080] The above sparsification method is suitable for two-dimensional tensor data with a large number of rows and a small number of columns. A small number of columns can reduce the number of Bound elements, so it needs to be used together with the above-mentioned dimensional expansion module.
[0081] S104: Pack the multiple sparse vectors to obtain sparse packed data and perform data transmission operations.
[0082] In S104 , the multiple sparse vectors are packaged by a data packaging module to obtain sparse packaged data, where the sparse packaged data includes a data portion and a metadata portion.
[0083] The specific process of packing the multiple sparse vectors, obtaining the sparse packed data and performing the data transmission operation is shown in B1-B3.
[0084] B1: Extract the data content of multiple vectors after sparsification; the data content includes at least the number of valid elements, the number of columns after expansion, the space occupied by the data elements, the data type, and the dimension size before expansion; the dimension size before expansion is determined by the input tensor data.
[0085] The specific data content is meta_header+Bound+Index+Value.
[0086] Among them, meta_header is mainly used to provide metadata information about the page in HTML documents, usually located in the <head> section. meta_header can contain a variety of information, such as page description, keywords, author information, viewport settings, etc.
[0087] The data content is described as follows:
[0088] meta_header=[
[0089] nonzero_count, / / uint32, item 1: number of valid elements
[0090] col_count, / / uint32, item 2: the number of columns in the expanded matrix
[0091] value_offset_bits, / / uint32, item 3: how many bits each value occupies (e.g. 32 for float32)
[0092] data_dtype, / / uint32, item 4: data type (such as 0=float32, 1=int8, 2=float16, etc.)
[0093] dim_count, / / uint32, item 5: number of tensor dimensions (e.g. 3 for a 3-dimensional tensor)
[0094] dim_0_size, / / uint32, item 6: length of dimension 1
[0095] dim_1_size, / / uint32, item 7: length of dimension 2
[0096] ... / / Each subsequent item is the length of a dimension (can be 0 if none)
[0097] dim_31_size, / / uint32, item 31: length of the 32nd dimension (can be 0 if not specified)
[0098] ].
[0099] B2: Determine the data contents of the multiple vectors as the metadata part, and determine the multiple vectors after the sparse processing as the data part.
[0100] B3: Place the data part after the metadata part to complete the packing of the multiple sparse vectors to obtain sparse packed data, and perform data transmission operations on the sparse packed data.
[0101] The data portion consists of three vectors after sparsification (the Bound vector, the Value vector, and the Index vector). These three vectors are placed after the metadata. The metadata is used to extract the data content of the three vectors, including the number of valid elements (non-zero data), the number of columns after expansion, the space occupied by the data elements, the data type, and the dimension size before expansion. Since the dimension size before expansion is determined by the input tensor data, the design supports a maximum of 32 dimensions.
[0102] Perform data transmission of the sparsely packed data. The sparsely packed data can be transmitted using a data transmission unit or other transmission method. This application is applicable to any data transmission unit. The transmission method of this application is not specifically limited.
[0103] S105: When the sparse data transmission is completed, the sparse packed data is unpacked to obtain the unpacked data, and actual calculation is performed on the unpacked data.
[0104] In S105 , when the transmission of the sparse data is completed, the sparse packed data is unpacked by the data unpacking module.
[0105] Specifically, when the sparse data transmission is completed, the sparse packed data is unpacked to obtain the unpacked data, and the unpacked data is actually calculated, as shown in C1-C6.
[0106] C1: When the sparse data transmission is completed, the metadata information corresponding to the fixed position is obtained from the sparse packed data.
[0107] In C1, metadata information is obtained. Since the elements in meta_header are fixed, the corresponding metadata information is directly obtained using a fixed position.
[0108] C2: Get the bound vector data from the sparse packed data according to the number of columns in the expanded matrix.
[0109] In C2, the Bound vector data is obtained based on the number of columns in the expanded matrix. The Index and Value element data are obtained based on the number of valid elements.
[0110] C3: Gets the Value vector data and the Index vector data according to the number of valid elements in the Bound vector data.
[0111] C4: Parse the two-dimensional tensor data through the Bound vector data, Value vector data and Index vector data.
[0112] Based on the Bound, Index and Value vector data, construct the two-dimensional tensor data. This step can be parsed according to the content of the three vector data.
[0113] In order to improve data unpacking efficiency, the above data unpacking steps can be implemented by hardware.
[0114] C5: Convert the parsed two-dimensional tensor data into original dimensional data content; wherein the original dimensional data content is the tensor data content changed to the preset dimension.
[0115] Since the data is still stored continuously before and after the conversion, the dimension information of the tensor, that is, the parsed two-dimensional tensor data is directly changed to the tensor data content of the preset dimension. The tensor data content of the preset dimension is T[d1, d2, ..., d n ].
[0116] C6: Determine the original dimension data content as unpacked data and perform actual calculations.
[0117] In C6, the unpacked data is used for actual calculations. The above-mentioned data transmission method and system are applicable to hardware accelerators, which can increase the data transmission speed in the accelerator, thereby improving the inference efficiency of deep learning models in the accelerator, and can be applied to various deep learning model applications.
[0118] In order to facilitate the understanding of the data processing process based on sparsification, combined with Figure 3 Provide explanation.
[0119] Figure 3 , obtain the trained deep learning model;
[0120] Model weight data sparsification:
[0121] Based on the size of the weight data, the diagonal Hessian approximation method is used to calculate the Hessian diagonal value of each weight, indicating the importance of the weight data. A sparse threshold is set, and weight data with Hessian diagonal values less than the threshold are set to 0 to obtain the sparse model.
[0122] Dimension expansion module:
[0123] The dimension expansion module obtains the order corresponding to the tensor data of the deep learning model after sparseness, determines the tensor dimension according to the order, and expands the tensor data into two-dimensional tensor data according to the tensor dimension.
[0124] Data compression module:
[0125] The data compression module uses a preset conversion method to transform the two-dimensional tensor data to obtain multiple sparse vectors. The data compression module is implemented in hardware. The expanded two-dimensional tensor data is transformed using the following method to obtain multiple sparse vectors. The multiple sparse vectors include at least a Bound vector, a Value vector, and an Index vector.
[0126] Data packaging module:
[0127] The data packing module packs the multiple sparse vectors to obtain sparse packed data, which includes a data part and a metadata part.
[0128] Perform post-sparse data transfer:
[0129] The data part is placed after the metadata part to complete the packing of the multiple sparse vectors to obtain the sparse packed data, and the sparse packed data is then subjected to data transmission operation.
[0130] Data unpacking module:
[0131] Get metadata information. Since the elements in meta_header are fixed, the corresponding metadata information is directly obtained using a fixed position.
[0132] Get the Bound vector data based on the number of columns in the expanded matrix. Get the Index and Value element data based on the number of valid elements.
[0133] According to the number of valid elements in the Bound vector data, obtain the Value vector data and the Index vector data;
[0134] Parse the two-dimensional tensor data through Bound vector data, Value vector data and Index vector data;
[0135] Convert the parsed two-dimensional tensor data into original dimensional data content; wherein the original dimensional data content is the tensor data content changed to a preset dimension;
[0136] The original dimension data content is determined as the unpacked data.
[0137] Use the unpacked data to perform actual calculations:
[0138] The unpacked data is then used for actual calculations. The data transmission method and system described above are applicable to hardware accelerators, increasing the data transmission speed within the accelerator and, in turn, improving the inference efficiency of deep learning models within the accelerator. These methods and systems can be applied to various deep learning model applications.
[0139] This application introduces a weight sparsification method based on Hessian matrix analysis, which can actively improve the sparsity of model tensors without reducing model accuracy. A sparse coding scheme suitable for tensors of arbitrary dimensions is designed. The dimensional expansion module and the data packing module cooperate with each other to improve data sparsity efficiency. In addition, a data packing module and an unpacking module based on hardware implementation are proposed to achieve rapid structural reconstruction of tensor data before and after transmission, significantly reducing overall latency. Through the above-mentioned systematic design, this application can effectively alleviate the data transmission bottleneck in the process of large model inference, improve compression efficiency and processing speed, and thus achieve more efficient deep learning model transmission and deployment capabilities. It is suitable for a variety of application scenarios such as server-side inference, edge smart devices, and dedicated artificial intelligence (AI) chips, and has high practical value and innovative value.
[0140] The beneficial effects of the embodiments of the present application: By introducing a method for sparsifying weight data based on diagonal matrix analysis, the sparsity of the model tensor can be improved without reducing the accuracy of the model. The tensor data expansion method of this solution cooperates with the method of packing multiple vectors after sparsification to improve data sparsity efficiency. And based on hardware-based data packing and unpacking, rapid structural reconstruction of tensor data before and after transmission is achieved, significantly reducing overall latency. Through the above-mentioned systematic design, the present application can effectively alleviate the data transmission bottleneck in the large model inference process, improve data compression efficiency and processing speed, and thus achieve more efficient deep learning model transmission and deployment capabilities.
[0141] Based on the above embodiment Figure 1 The disclosed data processing method based on sparsification is disclosed. The embodiment of the present application also discloses a data processing system based on sparsification, such as Figure 4 As shown, the data processing system based on sparsification includes:
[0142] A sparsification unit 401 is used to sparsify the model weight data of the pre-trained deep model to obtain a sparsified deep learning model;
[0143] An acquisition unit 402 is configured to acquire tensor data of the deep learning model after sparse processing and expand the tensor data into two-dimensional tensor data.
[0144] The conversion unit 403 is used to convert the two-dimensional tensor data using a preset conversion method to obtain multiple sparse vectors;
[0145] A packing operation unit 404 is used to pack the multiple vectors after sparse processing to obtain sparse packed data and perform data transmission operations;
[0146] The unpacking calculation unit 405 is used to unpack the sparse packed data after the sparse data transmission is completed to obtain the unpacked data and perform actual calculation on the unpacked data.
[0147] Furthermore, the sparseness unit 401 includes:
[0148] A calculation module, used to calculate the diagonal value of each weight in the pre-trained deep model by using a diagonal approximation method;
[0149] The comparison module is used to compare the pre-set sparsification threshold with the diagonal value of each weight, so as to set the diagonal value less than the sparsification threshold to 0, thereby obtaining a sparse deep learning model.
[0150] The acquisition unit 402 includes:
[0151] The first acquisition module is used to obtain the order corresponding to the tensor data of the deep learning model after sparsification;
[0152] The first determination module is used to determine the tensor dimension according to the order, and expand the tensor data into two-dimensional tensor data according to the tensor dimension.
[0153] Furthermore, the multiple vectors after sparseness include at least a numerical range boundary vector, an index associated vector, and an index vector. The conversion unit 403 includes:
[0154] A second determination module is configured to determine, from the two-dimensional tensor data, the number of elements at the boundary of the numerical range, the number of elements associated with the index, and the number of elements of the index; wherein the number of elements at the boundary of the numerical range represents the number of columns of the two-dimensional tensor; the number of elements associated with the index represents whether the data is actually valid; and the number of elements of the index represents the number of valid data;
[0155] The vector conversion module is used to perform vector conversion on two-dimensional tensor data according to the number of elements at the numerical range boundary, the number of elements associated with the index and the number of elements of the index, so as to obtain the numerical range boundary vector, the index associated vector and the index vector.
[0156] Furthermore, the packaging operation unit 404 includes:
[0157] An extraction module is configured to extract the data content of the multiple vectors after sparsification, wherein the data content includes at least the number of valid elements, the number of columns after expansion, the space occupied by the data elements, the data type, and the dimension size before dimensional expansion; the dimension size before dimensional expansion is determined by the input tensor data;
[0158] a third determining module, configured to determine the data contents of the plurality of vectors as a metadata portion, and determine the plurality of vectors after the sparse processing as a data portion;
[0159] The packing operation module is used to place the data part after the metadata part to complete the packing of the sparse vectors, obtain the sparse packed data, and perform data transmission operation on the sparse packed data.
[0160] Furthermore, the unpacking calculation unit 405 includes:
[0161] A second acquisition module is used to obtain metadata information corresponding to a fixed position from the sparse packed data when the sparse data transmission is completed;
[0162] A third acquisition module is used to obtain the numerical range boundary vector data from the sparse packed data according to the number of columns of the expanded matrix;
[0163] A fourth acquisition module, configured to acquire index-associated vector data and index vector data according to the number of valid elements in the numerical range boundary vector data;
[0164] A parsing module, configured to parse two-dimensional tensor data through the numerical range boundary vector data, the index associated vector data, and the index vector data;
[0165] A conversion module, configured to convert the parsed two-dimensional tensor data into original dimensional data content; wherein the original dimensional data content is the tensor data content changed to a preset dimension;
[0166] The calculation module is determined to determine the original dimension data content as unpacked data and perform actual calculations.
[0167] The beneficial effects of the embodiments of the present application: By introducing a method for sparsifying weight data based on diagonal matrix analysis, the sparsity of the model tensor can be improved without reducing the accuracy of the model. The tensor data expansion method of this solution cooperates with the method of packing multiple vectors after sparsification to improve data sparsity efficiency. And based on hardware-based data packing and unpacking, rapid structural reconstruction of tensor data before and after transmission is achieved, significantly reducing overall latency. Through the above-mentioned systematic design, the present application can effectively alleviate the data transmission bottleneck in the large model inference process, improve data compression efficiency and processing speed, and thus achieve more efficient deep learning model transmission and deployment capabilities.
[0168] An embodiment of the present application further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned sparsification-based data processing method.
[0169] The present application also provides an electronic device, the structure of which is shown in FIG. Figure 5 As shown, it specifically includes a memory 501 and one or more instructions 502, wherein the one or more instructions 502 are stored in the memory 501 and are configured to be executed by one or more processors 503 to perform the above-mentioned sparsification-based data processing method.
[0170] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0171] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0172] The steps in the methods of the various embodiments of the present application can be adjusted in sequence, combined, and deleted according to actual needs.
[0173] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0174] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
[0175] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data processing method based on sparsification, characterized in that: The method comprises: Perform model weight data sparsification on the pre-trained deep model to obtain a sparsified deep learning model; Acquire the tensor data of the deep learning model after the sparsification, and expand the tensor data into two-dimensional tensor data; The two-dimensional tensor data is transformed by a preset transformation method to obtain a plurality of sparse vectors; Packing the multiple vectors after the sparse processing to obtain sparse packed data and performing a data transmission operation; When the sparse data transmission is completed, the sparse packed data is unpacked to obtain the unpacked data, and actual calculation is performed on the unpacked data.
2. The method according to claim 1, characterized in that The method of performing model weight data sparsification on the pre-trained deep model to obtain a sparsified deep learning model includes: Calculate the diagonal value of each weight in the pre-trained deep model through the diagonal approximation method; A preset sparsification threshold is compared with the diagonal value of each weight to set the diagonal value smaller than the sparsification threshold to 0, thereby obtaining a sparsified deep learning model.
3. The method according to claim 1, characterized in that The obtaining of the tensor data of the deep learning model after the sparsification and expanding the tensor data into two-dimensional tensor data includes: Obtaining the order corresponding to the tensor data of the sparsified deep learning model; A tensor dimension is determined according to the order, and the tensor data is expanded into two-dimensional tensor data according to the tensor dimension.
4. The method according to claim 1, wherein The multiple vectors after the sparseness include at least a numerical range boundary vector, an index associated vector, and an index vector. The two-dimensional tensor data is transformed by a preset transformation method to obtain the multiple vectors after the sparseness, including: Determining the number of elements at the boundary of the numerical range, the number of elements associated with the index, and the number of elements of the index from the two-dimensional tensor data; wherein the number of elements at the boundary of the numerical range represents the number of columns of the two-dimensional tensor; the number of elements associated with the index represents whether the data is actually valid; and the number of elements of the index represents the number of valid data; According to the number of elements at the boundary of the numerical range, the number of elements associated with the index and the number of elements of the index, the two-dimensional tensor data is vector-converted to obtain a numerical range boundary vector, an index-associated vector and an index vector.
5. The method according to claim 1, wherein The packing of the multiple vectors after the sparse processing to obtain the sparse packed data and performing the data transmission operation includes: Extracting data content of the multiple vectors after sparsification; wherein the data content includes at least the number of valid elements, the number of columns after expansion, the space occupied by data elements, the data type, and the dimension size before dimensional expansion; the dimension size before dimensional expansion is determined by the input tensor data; Determining the data contents of the plurality of vectors as a metadata portion, and determining the plurality of vectors after the sparse processing as a data portion; The data portion is placed after the metadata portion to complete the packing of the multiple sparse vectors to obtain sparse packed data, and the sparse packed data is subjected to a data transmission operation.
6. The method according to claim 1, characterized in that When the sparse data transmission is completed, the sparse packed data is unpacked to obtain the unpacked data, and actual calculation is performed on the unpacked data, including: Upon completion of the sparse data transmission, metadata information corresponding to a fixed position is obtained from the sparse packed data; Obtaining numerical range boundary vector data from the sparse packed data according to the number of expanded matrix columns; Obtaining index-associated vector data and index vector data according to the number of valid elements in the numerical range boundary vector data; Parsing two-dimensional tensor data through the numerical range boundary vector data, the index-associated vector data, and the index vector data; Convert the parsed two-dimensional tensor data into original dimensional data content; wherein the original dimensional data content is the tensor data content changed to a preset dimension; The original dimensional data content is determined as the unpacked data and actual calculation is performed.
7. A data processing system based on sparsification, characterized in that: The system comprises: The sparsification unit is used to sparsify the model weight data of the pre-trained deep model to obtain a sparse deep learning model; An acquisition unit, configured to acquire tensor data of the sparsified deep learning model and expand the tensor data into two-dimensional tensor data; A conversion unit, configured to convert the two-dimensional tensor data using a preset conversion method to obtain a plurality of sparse vectors; a packing operation unit, configured to pack the plurality of vectors after the sparse processing to obtain sparse packed data and perform a data transmission operation; The unpacking calculation unit is used to unpack the sparse packed data when the sparse data transmission is completed to obtain the unpacked data and perform actual calculation on the unpacked data.
8. The system according to claim 7, characterized in that The thinning unit includes: A calculation module, used to calculate the diagonal value of each weight in the pre-trained deep model by using a diagonal approximation method; The comparison module is used to compare a preset sparsification threshold with the diagonal value of each weight to set the diagonal value less than the sparsification threshold to 0, thereby obtaining a sparsified deep learning model.
9. A storage medium, characterized in that: The storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the data processing method based on sparsification according to any one of claims 1 to 6.
10. An electronic device, characterized in that: The system comprises a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to perform the sparsification-based data processing method according to any one of claims 1 to 6.