Data processing method, convolution engine, equipment and storage medium
By generating weight blocks and removing zero-value weights to form compressed triples, the problems of high power consumption and high decoding cost of convolution engines in satellite navigation chips are solved, thereby improving data processing speed.
Patent Information
- Application Number
- CN202511440761.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing convolution engines cannot skip zero-weight calculations in satellite navigation chips, resulting in excessive power consumption, large data processing latency, and wasted computing resources. They also lack sparse data formats for convolutional structures, leading to high decoding costs.
By generating weight blocks and removing zero-value weights, compressed triples are formed to achieve '0 multiplication-addition skip', and the Block-CSR compression format is used to optimize data processing and generate binary data segments for calculation.
It reduces the amount of computation in the data processing process, lowers hardware power consumption and decoding costs, and improves data processing speed.
Smart Images

Figure CN120911519A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of neural network models, and in particular to a data processing method, a convolution engine, a device and a storage medium. BACKGROUND
[0002] With the development of satellite navigation technology, in order to enable satellite navigation chips to classify, detect and trend predict data such as spectrum graphs, residual graphs and channel energy graphs, convolutional neural networks (CNN) have been gradually integrated into satellite navigation chips. However, the existing convolution engine has the following problems: 1. Unable to skip 0 weight calculation, resulting in high power consumption and large data processing delay of the convolution engine during inference, and also causing waste of computing resources.
[0003] 2. Lack of sparse data format optimized for convolution structure, high decoding cost. SUMMARY
[0004] The embodiments of the present application provide a data processing method, a convolution engine, a device and a storage medium to solve at least one of the above problems.
[0005] In a first aspect, the embodiments of the present application provide a data processing method, comprising: obtaining effective weight values corresponding to a plurality of output channels of a convolution model; generating a plurality of weight blocks including at least two output channels based on the effective weight values of each output channel; eliminating zero-value weights in each weight block to generate corresponding compressed triplets; wherein the compressed triplets include the effective weight values of the corresponding weight blocks, channel masks and intra-block index addresses.
[0006] The data processing method provided by the embodiments of the present application can realize "0 multiplication and addition skipping" in the data processing process, reduce the amount of calculation in the data processing process, and can generate weight blocks for data processing to speed up the decoding speed, thereby reducing hardware power consumption and decoding cost and improving data processing speed.
[0007] Optionally, the step of generating a plurality of weight blocks including at least two output channels based on the effective weight values of each output channel comprises: dividing the effective weight values corresponding to all output channels into a plurality of initial blocks based on a preset channel number; retaining non-zero blocks in the initial blocks to form a plurality of weight blocks; wherein the preset channel number includes at least two output channels.
[0008] Optionally, the step of eliminating zero-value weights in each weight block to generate corresponding compressed triplets comprises: obtaining all weight values, channel masks and intra-block index addresses in each weight block; eliminating zero-value weights from all weight values of each weight block to obtain effective weight values of each weight block; generating corresponding compressed triplets based on the effective weight values, channel masks and intra-block index addresses.
[0009] Optionally, after the step of generating the corresponding compressed triplets, the data processing method further comprises: obtaining a sparse structure configuration table of the compressed triplets and each convolution layer of the convolution model; based on the sparse structure configuration table, calling the effective weight value, the channel mask, and the intra-block index address in the compressed triplets; and based on the effective weight value and the intra-block index address, performing a calculation operation or a skip operation on the input compressed triplets.
[0010] Optionally, the step of performing the calculation operation or the skip operation on the input compressed triplets based on the effective weight value and the intra-block index address comprises: obtaining a plurality of to-be-calculated values of the input convolution model; based on the intra-block index address, matching the effective weight value corresponding to each to-be-calculated value; based on the to-be-calculated value, determining whether a skip operation needs to be performed, if the to-be-calculated value is a zero value, performing the skip operation, otherwise, based on the to-be-calculated value and the effective weight value corresponding thereto, performing the calculation operation.
[0011] Optionally, the data processing method further comprises: after determining to perform the calculation operation, based on the intra-block index address in each compressed triplet, calling the corresponding output channel in the convolution model and obtaining the corresponding effective weight value; and based on the to-be-calculated value and the effective weight value corresponding to the called output channel, performing the calculation operation to generate an output result.
[0012] Optionally, after the step of generating the corresponding compressed triplets, the data processing method further comprises: packing the compressed triplets into a binary data segment in a binary data format.
[0013] Optionally, the step of obtaining the effective weight values corresponding to the plurality of output channels of the convolution model comprises: obtaining a convolution initial model; performing structure pruning on the convolution initial model to generate a convolution model that retains effective weight values; and obtaining the effective weight values corresponding to the plurality of output channels of the convolution model.
[0014] In a second aspect, an embodiment of the present application provides a convolution engine, which comprises: a compression module configured to obtain effective weight values corresponding to a plurality of output channels of a convolution model; based on the effective weight values of each output channel, generate a plurality of weight blocks comprising at least two output channels; and eliminate zero weights in each weight block to generate corresponding compressed triplets; and a calculation module configured to generate an output result according to the compressed triplets.
[0015] The convolution engine provided by the embodiment of the present application can implement "0 multiplication and addition skip" in the data processing process, reduce the calculation amount in the data processing process, and can generate weight blocks for data processing to accelerate the decoding speed, thereby reducing the hardware power consumption and decoding cost and improving the data processing speed.
[0016] In a third aspect, an embodiment of the present application provides a data processing device, comprising: a processor and a memory, the memory storing instructions; and the processor invoking the instructions in the memory to cause the processor to perform the data processing method of any one of the preceding embodiments of the first aspect of the present application.
[0017] The processor of the data processing device provided by the embodiment of the present application performs the data processing method of any one of the preceding embodiments of the first aspect of the present application by invoking the instructions in the memory, so that the data processing method can realize "0 multiply-add skipping" in the data processing process, reduce the calculation amount in the data processing process, and perform data processing by generating a weight block, thereby accelerating the decoding speed, reducing the hardware power consumption and decoding cost, and improving the data processing speed.
[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium storing instructions, and the instructions being executed by a processor to implement the data processing method of any one of the preceding embodiments of the first aspect of the present application.
[0019] The instructions stored in the computer-readable storage medium provided by the embodiment of the present application can be invoked by a processor and executed to perform the data processing method of any one of the preceding embodiments of the first aspect of the present application, so that the data processing method can realize "0 multiply-add skipping" in the data processing process, reduce the calculation amount in the data processing process, and perform data processing by generating a weight block, thereby accelerating the decoding speed, reducing the hardware power consumption and decoding cost, and improving the data processing speed. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from the structures shown in the drawings without creative labor.
[0021] Figure 1 Flowchart of the first embodiment of the data processing method of the present application; Figure 2 Flowchart of step S110 in the first embodiment of the data processing method of the present application; Figure 3 Flowchart of step S120 in the first embodiment of the data processing method of the present application; Figure 4 Flowchart of step S130 in the first embodiment of the data processing method of the present application; Figure 5 Flowchart of the second embodiment of the data processing method of the present application; Fig. 6(a) is a format diagram of a sparse weight block in the second embodiment of the data processing method of the present application; Fig. 6(b) is a flow diagram of performing a compute operation or a skip operation according to a compressed triple in the second embodiment of the data processing method of the present application; Figure 7 Fig. 7 is a flow diagram of step S260 in the second embodiment of the data processing method of the present application; Figure 8 Fig. 8 is a structural block diagram of an embodiment of the convolution engine of the present application; Figure 9 Fig. 9 is a flow diagram of performing a compute operation or a skip operation in an embodiment of the convolution engine of the present application; Figure 10 Fig. 10 is a structural block diagram of an embodiment of the data processing device of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0023] It should be noted that all directional indications such as up, down, left, right, front, back, etc. in the embodiments of the present application are only used to explain the relative positional relationship, movement condition, etc. between components in a certain posture, such as shown in the drawings. If the certain posture changes, the directional indications also change accordingly.
[0024] In addition, the description of "first", "second", etc. in the present application is only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can realize it. When the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, and is not within the scope of protection required by the present application.
[0025] For the convenience of understanding, the data processing method of the embodiments of the present application will be described below. As shown in Fig. 1, the data processing method in the first embodiment of the present application includes steps S110 to S130. Figure 1
[0026] In step S110, effective weight values corresponding to a plurality of output channels of a convolution model are obtained.
[0027] As Figure 2 shown, in some optional embodiments, step S110 includes steps S111 to S113.
[0028] In step S111, a convolution initial model is acquired.
[0029] In step S112, structural pruning is performed on the convolution initial model to generate a convolution model that retains effective weight values.
[0030] In step S113, effective weight values corresponding to multiple output channels of the convolution model are acquired.
[0031] In the present embodiment, structural pruning is first performed on the pre-acquired convolution initial model, and effective weights are retained for subsequent generation of weight blocks.
[0032] In step S120, based on the effective weight values of each output channel, multiple weight blocks including at least two output channels are generated.
[0033] As Figure 3 shown, in some optional embodiments, step S120 includes steps S121 to S122.
[0034] In step S121, based on a preset number of channels, the effective weight values corresponding to all output channels are divided into multiple initial blocks.
[0035] In step S122, non-zero blocks in the initial blocks are retained to form multiple weight blocks. The preset number of channels includes at least two output channels.
[0036] In the present embodiment, the effective weight values corresponding to all output channels are converted into a block structure form using a Block-CSR (Block Compressed Sparse Row) compression format, obtaining initial blocks, each of which includes a number of output channels corresponding to the preset number of channels. Then, the initial blocks in which the effective weight values corresponding to each output channel are all zero are removed, and the remaining non-zero blocks are retained to form multiple weight blocks. The preset number of channels can be 2, 4, 5, 6, or other numbers, which are not limited by the present application.
[0037] In step S130, zero weights in each weight block are removed to generate corresponding compressed triplets. The compressed triplets include effective weight values of the corresponding weight blocks, channel masks, and intra-block index addresses.
[0038] As Figure 4 shown, in some optional embodiments, step S130 includes steps S131 to S133.
[0039] In step S131, all weight values in each weight block, a channel mask, and an intra-block index address are obtained.
[0040] In step S132, zero values in all weight values of each weight block are removed to obtain valid weight values of each weight block.
[0041] In step S133, a corresponding compressed triple is generated based on the valid weight values, the channel mask, and the intra-block index address.
[0042] In the embodiment, by removing zero value weights in each weight block and only retaining non-zero weights as valid weight values, a corresponding compressed triple is generated according to all weight values in each weight block, a channel mask, and an intra-block index address, and the compressed triple is packaged into a binary data segment in a binary data format, and then the compressed triple packaged into the binary data segment is sent to a convolution unit, so that the convolution model can obtain relevant information of valid weight values, a channel mask, and an intra-block index address in the corresponding compressed triple according to a preset sparse structure configuration table, thereby achieving the purpose of reducing calculation amount and improving data processing efficiency by performing data calculation processing on the obtained navigation information.
[0043] Figure 5 A flowchart of a second embodiment of the data processing method of the present application is shown. The second embodiment has the same structure as the first embodiment, and the differences between the two will be described below, and the same parts will not be described in detail.
[0044] As shown in FIG. 2, in the second embodiment of the present application, the data processing method includes steps S210 to S260. Figure 5
[0045] In step S210, valid weight values corresponding to a plurality of output channels of a convolution model are obtained.
[0046] In step S220, a plurality of weight blocks including at least two output channels are generated based on the valid weight values of each output channel.
[0047] In step S230, zero value weights in each weight block are removed to generate a corresponding compressed triple. The compressed triple includes valid weight values of the corresponding weight block, a channel mask, and an intra-block index address.
[0048] After the corresponding compressed triple is generated, the data processing method performs steps S240 to S260.
[0049] In step S240, a sparse structure configuration table of each convolution layer of the convolution model and the compressed triple are obtained.
[0050] In step S250, based on the sparse structure configuration table, the valid weight value, the channel mask, and the intra-block index address in the compressed triple are called.
[0051] In step S260, based on the valid weight value and the intra-block index address, a calculation operation or a skip operation is performed on the input compressed triple.
[0052] As shown in FIG. 6(a) and FIG. 6(b), in the embodiment, the number of output channels corresponding to each weight block is 4, and the sparse weight format of one of the weight blocks obtained through the sparse structure configuration table is shown in FIG. 6(a), wherein the valid weight value is data[], the intra-block index address is offset[], and the channel mask is mask. By accepting the valid weight value in the compressed triple in the corresponding weight block of each convolution layer of the convolution model and feeding the valid weight value into the subsequent convolution unit for multiplication and addition calculation operation, selecting the enabled output channel according to the channel mask, and adjusting the convolution input alignment mode according to the intra-block index address, data decoding is realized.
[0053] As shown in FIG. 6(a) and FIG. 6(b), in the embodiment, the number of output channels corresponding to each weight block is 4, and the sparse weight format of one of the weight blocks obtained through the sparse structure configuration table is shown in FIG. 6(a), wherein the valid weight value is data[], the intra-block index address is offset[], and the channel mask is mask. By accepting the valid weight value in the compressed triple in the corresponding weight block of each convolution layer of the convolution model and feeding the valid weight value into the subsequent convolution unit for multiplication and addition calculation operation, selecting the enabled output channel according to the channel mask, and adjusting the convolution input alignment mode according to the intra-block index address, data decoding is realized. Figure 7
[0054] In step S261, a plurality of to-be-calculated values of the input convolution model are obtained.
[0055] In step S262, based on the intra-block index address, the valid weight value corresponding to each to-be-calculated value is matched.
[0056] In step S263, based on the to-be-calculated value, it is judged whether a skip operation needs to be performed, if the to-be-calculated value is a zero value, the skip operation is performed, otherwise, based on the to-be-calculated value and the valid weight value corresponding thereto, a calculation operation is performed.
[0057] Specifically, after step S263, step S260 further includes steps S264 to S265.
[0058] In step S264, after it is determined that the calculation operation is performed, based on the intra-block index address in each compressed triple, the corresponding output channel in the convolution model is called, and the corresponding valid weight value is obtained.
[0059] In step S265, based on the to-be-calculated value and the valid weight value corresponding to the called output channel, a calculation operation is performed to generate an output result.
[0060] In the embodiment, when external navigation information is input, a plurality of to-be-calculated values for inputting the convolution model in the navigation information are obtained, and whether to perform a calculation operation or a skip operation on each to-be-calculated value is determined through the compressed triple of the weight block in the application. A large number of zero-value calculation operations can be skipped through the zero-value elimination of each output channel corresponding zero-value weight and the input navigation information, so as to reduce the calculation amount of the input navigation information and improve the data processing efficiency.
[0061] The data processing method provided in the embodiment includes: obtaining effective weight values corresponding to a plurality of output channels of a convolution model; generating a plurality of weight blocks including at least two output channels based on the effective weight values of each output channel; eliminating zero-value weights in each weight block to generate corresponding compressed triples; and wherein the compressed triple includes the effective weight values of the corresponding weight block, a channel mask, and an intra-block index address.
[0062] The data processing method provided in the embodiment can realize "0 multiply-add skip" in the data processing process, reduce the calculation amount in the data processing process, and accelerate the decoding speed by generating weight blocks for data processing, thereby reducing hardware power consumption and decoding cost and improving data processing speed.
[0063] For the above method embodiment, the embodiment of the application also provides a convolution engine 200 as shown in Figure 8 The convolution engine 200 includes a compression module 210 and a calculation module 220.
[0064] The compression module 210 is configured to obtain effective weight values corresponding to a plurality of output channels of a convolution model; generate a plurality of weight blocks including at least two output channels based on the effective weight values of each output channel; and eliminate zero-value weights in each weight block to generate corresponding compressed triples. The calculation module 220 is configured to generate an output result according to the compressed triple.
[0065] In the embodiment, the compression module 210 is configured to perform structural pruning on the convolution model, eliminate zero-value weights, retain effective weight values, and convert the effective weight values corresponding to all output channels into a block structure form in a Block-CSR (Block Compressed Sparse Row) compression format.
[0066] The calculation module 220 in the application includes a sparse decoding controller and a multiply-accumulate unit (MAC-Multiply-Accumulate Unit). As shown in Figure 9 Figure 9 The convolution engine 200 can mask the invalid output channel of the multiplication and addition, skip the invalid calculation path (such as 0 x x, x x 0), avoid wasting decoding resources and computing resources, and reduce the computing power consumption of the hardware.
[0067] The convolution engine 200 provided by the embodiment of the present application can realize the "0 multiplication and addition skipping" in the data processing process, reduce the calculation amount in the data processing process, and can generate the weight block for data processing, so as to accelerate the decoding speed, thereby reducing the hardware power consumption and decoding cost, and improving the data processing speed.
[0068] For the above method embodiment, the embodiment of the present application further provides a data processing device as shown in Figure 10 The processor 301 calls the instructions in the memory 302, so that the processor 301 executes the data processing method of any one of the preceding embodiments of the present application.
[0069] The data processing method provided by the embodiment of the present application comprises: obtaining effective weight values corresponding to a plurality of output channels of a convolution model; generating a plurality of weight blocks comprising at least two output channels based on the effective weight values of each output channel; eliminating zero weights in each weight block to generate corresponding compressed triplets; wherein the compressed triplet comprises the effective weight values of the corresponding weight block, a channel mask, and an intra-block index address.
[0070] The data processing device provided by the embodiment of the present application can realize the "0 multiplication and addition skipping" in the data processing process by implementing the above data processing method, reduce the calculation amount in the data processing process, and can generate the weight block for data processing, so as to accelerate the decoding speed, thereby reducing the hardware power consumption and decoding cost, and improving the data processing speed.
[0071] Further, the data processing device provided by the embodiment of the present application can further comprise a communication interface 303 and a bus 304, and the processor 301, the memory 302 and the communication interface 303 are electrically connected through the bus 304.
[0072] The memory 302 can include a high-speed random access memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 303 (which can be wired or wireless), and the Internet, a wide area network, a local network, a metropolitan area network, etc. can be used. The bus 304 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 Only one bidirectional arrow is used to represent the system network element and at least one other network element, but it does not mean that there is only one bus or one type of bus.
[0073] The processor 301 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 301 or the instructions in the form of software. The processor 301 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiment of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiment of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 302, and the processor 301 reads the information in the memory 302, and combines the hardware to complete the steps of the method of the above embodiment.
[0074] The embodiment of the present application also provides a computer readable storage medium, which can be a nonvolatile computer readable storage medium or a volatile computer readable storage medium, and the computer readable storage medium stores instructions, and the instructions enable a computer to perform the steps of the data processing method when the instructions are executed on the computer.
[0075] The computer readable storage medium provided by the embodiment of the present application stores data and computer executable instructions of the data processing method, and the data processing method comprises the following steps: obtaining effective weight values corresponding to a plurality of output channels of a convolution model; generating a plurality of weight blocks comprising at least two output channels based on the effective weight values of each output channel; removing zero weights in each weight block to generate corresponding compressed triplets; and the compressed triplet comprises the effective weight values of the corresponding weight block, a channel mask, and an intra-block index address.
[0076] The computer readable storage medium provided by the embodiment of the present application can realize the '0 multiply-add skip' in the data processing process, reduce the calculation amount in the data processing process, and accelerate the decoding speed by generating the weight blocks for data processing, thereby reducing the hardware power consumption and decoding cost and improving the data processing speed.
[0077] Those skilled in the art can clearly understand the specific working process of the system, device and unit described above for the convenience and brevity of description, and the corresponding process in the foregoing method embodiments can be referred to, which will not be described herein.
[0078] The integrated unit can be stored in a computer readable storage medium if it is realized in the form of a software function unit and sold or used as an independent product. Based on this understanding, the technical solutions of the present application or the whole or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0079] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalent replacements; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized by, The method comprises: obtaining effective weight values corresponding to a plurality of output channels of a convolution model; generating a plurality of weight blocks comprising at least two output channels based on the effective weight values of each output channel; eliminating zero-value weights in each weight block to generate a corresponding compressed triple; wherein the compressed triple comprises effective weight values of the corresponding weight block, a channel mask, and an intra-block index address.
2. The data processing method according to claim 1, characterized in that, The step of generating a plurality of weight blocks comprising at least two output channels based on the effective weight values of each output channel comprises: dividing the effective weight values corresponding to all output channels into a plurality of initial blocks based on a preset channel number; retaining non-zero blocks in the initial blocks to form a plurality of weight blocks; wherein the preset channel number comprises at least two output channels.
3. The data processing method of claim 1, wherein, The step of eliminating zero-value weights in each weight block to generate a corresponding compressed triple comprises: obtaining all weight values in each weight block, the channel mask, and the intra-block index address; eliminating zero-value weights from all weight values of each weight block to obtain the effective weight values of each weight block; generating a corresponding compressed triple based on the effective weight values, the channel mask, and the intra-block index address.
4. The data processing method according to claim 3, characterized in that, After the step of generating a corresponding compressed triple, the method further comprises: obtaining a compressed triple and a sparse structure configuration table of each convolution layer of the convolution model; based on the sparse structure configuration table, calling the effective weight values, the channel mask, and the intra-block index address in the compressed triple; based on the effective weight values and the intra-block index address, performing a calculation operation or a skip operation on the input compressed triple.
5. The data processing method according to claim 4, characterized in that, The step of performing a calculation operation or a skip operation on the input compressed triple based on the effective weight values and the intra-block index address comprises: obtaining a plurality of to-be-calculated values input into the convolution model; based on the intra-block index address, matching the effective weight values corresponding to each to-be-calculated value; based on the to-be-calculated value, determining whether a skip operation needs to be performed, if the to-be-calculated value is a zero value, performing a skip operation, otherwise, based on the to-be-calculated value and the effective weight values corresponding thereto, performing a calculation operation.
6. The data processing method according to claim 5, characterized in that, The method further comprises: after determining to perform a calculation operation, based on the intra-block index address in each compressed triple, calling the corresponding output channel in the convolution model and obtaining the corresponding effective weight values; based on the to-be-calculated value and the effective weight values of the called output channel, performing a calculation operation to generate an output result.
7. The data processing method of claim 1, wherein, After the step of generating a corresponding compressed triple, the method further comprises: packing the compressed triple into a binary data segment in a binary data format.
8. The data processing method of claim 1, wherein, The step of obtaining effective weight values corresponding to a plurality of output channels of a convolution model comprises: obtaining a convolution initial model; performing structure pruning on the convolution initial model to generate a convolution model retaining the effective weight values; The effective weight values corresponding to a plurality of output channels of the convolution model are obtained.
9. A convolution engine, comprising: The convolution engine comprises: A compression module is configured to obtain the effective weight values corresponding to a plurality of output channels of the convolution model; generate a plurality of weight blocks including at least two output channels based on the effective weight values of each output channel; and eliminate zero weights in each weight block to generate a corresponding compressed triple. A calculation module is configured to generate an output result according to the compressed triple.
10. A data processing device, characterized by The data processing device comprises a processor and a memory, and the memory stores instructions. The processor invokes the instructions in the memory, so that the data processing device implements the data processing method in any one of claims 1 to 8.
11. A computer-readable storage medium having stored thereon instructions, the instructions comprising, The instructions are executed by the processor to implement the data processing method in any one of claims 1 to 8.
Citation Information
Patent Citations
Deep learning accelerator and method for accelerating deep learning operation
CN110322001A
Instructions and logic for vector multiply add with zero skipping
CN113094096A
Block sparse method and device based on convolutional neural network, and processing unit
CN115186802A
Neural network model processing method and device, equipment and storage medium
CN115829020A
Neural network hardware accelerator system with zero-skipping and hierarchical structured pruning methods
US20200401895A1