Data processing method and apparatus, many-core chip, and storage medium
By parallelizing decoding and calculation during large model calculations, the problem of bandwidth occupation by weight data reading in incremental inference of large models is solved, efficient data processing and storage space saving are achieved, and the computing efficiency of end-side devices is improved.
Patent Information
- Application Number
- PCT/CN2025/082088
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-12
- Publication Date
- 2025-10-02
AI Technical Summary
During the incremental inference process of large models, repeated reading of large amounts of weight data occupies a huge amount of analog memory bandwidth, seriously restricting the inference efficiency, especially in the terminal-side devices.
By detecting the large model calculation process, the data that needs to be calculated in the next step is determined, and the corresponding target amount of compressed data is collected from the stored compressed data for parallel decoding and parallel calculation, realizing decoding and calculation at the same time, saving analog memory bandwidth and reducing storage space requirements.
It saves memory bandwidth and storage space without affecting computing efficiency, improves data processing efficiency and reduces costs.
Smart Images

Figure CN2025082088_02102025_PF_FP_ABST
Abstract
Description
Data processing method, device, many-core chip and storage medium Technical Field
[0001] The embodiments of the present disclosure relate to the field of decoding technology, and in particular to a data processing method, device, many-core chip, and storage medium. Background Art
[0002] During incremental inference of large models, a vast amount of weight data must be repeatedly read, consuming significant memory bandwidth and severely limiting inference efficiency. This is particularly evident in on-device inference scenarios. For example, for a 7TB-level model (such as a large language model), a text response of approximately 300 words requires reading over 2TB of weight data. Since the DDR (Double Data Rate) interface bandwidth of typical on-device devices is approximately 100GB / s, loading the weight data alone can take tens of seconds. Summary of the Invention
[0003] Embodiments of the present disclosure provide a data processing method, device, many-core chip, and storage medium.
[0004] In a first aspect, embodiments of the present disclosure provide a data processing method, comprising: detecting a large model calculation process, determining data to be calculated in the next step, and collecting a corresponding target amount of compressed data from compressed data stored off-chip based on the determination result; the target amount of compressed data is determined based on the parallel computing capability of a many-core chip when calculating the large model; the compressed data is obtained by segmenting the data to be compressed and then compressing the multiple segments separately; and parallel decoding the target amount of compressed data to obtain decoded data.
[0005] Data parallel computing is performed according to the decoded data.
[0006] In a second aspect, the present disclosure provides a data processing device, the device comprising: a scheduling module, a decoding module, and a calculation module;
[0007] The scheduling module is used to detect the large model calculation process, determine the data to be calculated in the next step, and collect a corresponding target amount of compressed data from the compressed data stored outside the data processing device based on the determination result; the target amount is determined based on the parallel computing capacity during the large model calculation; the compressed data is obtained by segmenting the data to be compressed and compressing the multiple segments separately;
[0008] The decoding module is used to decode the target amount of compressed data in parallel to obtain decoded data;
[0009] The computing module is configured to perform data parallel computing based on the decoded data.
[0010] In a third aspect, the present disclosure provides a many-core chip, including:
[0011] multiple processing cores; and
[0012] An on-chip network is configured to exchange data between the multiple processing cores and external data; wherein one or more instructions are stored in one or more of the processing cores, and one or more of the instructions are executed by one or more of the processing cores to enable one or more of the processing cores to perform the data processing method.
[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned data processing method when executed by a processor.
[0014] In a fifth aspect, the present disclosure provides a computer program product comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned data processing method.
[0015] The embodiment provided by the present disclosure detects the large model calculation process to determine the data that needs to be calculated in the next step, collects the corresponding target number of compressed data from the compressed data stored outside the chip according to the determination result for parallel decoding, and performs parallel calculation based on the decoded data, so that what is input into the chip is compressed data, thereby saving analog memory bandwidth; and realizes decoding and calculation at the same time, so that there is no need to prepare additional storage space to store decoded data, thereby saving space, reducing costs, and not affecting computing efficiency, and both the decoding process and the calculation process are parallel schemes, which improves data processing efficiency; in addition, the target number in the parallel decoding process is determined according to the parallel computing capability, and the compressed data of the target number are decoded in parallel, so that the parallelism of decoding is coordinated with the parallel computing capability, thereby realizing the coupling of decoding and calculation.
[0016] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings. In the accompanying drawings:
[0018] FIG1 is a flow chart of a data processing method provided by an embodiment of the present disclosure;
[0019] FIG2 is a block diagram of a data processing device provided by an embodiment of the present disclosure;
[0020] FIG3a is a schematic diagram of a data processing device for decoding a row or a column of compressed data provided by an embodiment of the present disclosure;
[0021] FIG3 b is a schematic diagram of a data processing device for decoding multiple rows and columns of compressed data according to an embodiment of the present disclosure;
[0022] FIG4 is a block diagram of a many-core chip provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] To enable those skilled in the art to better understand the technical solutions of the present disclosure, exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.
[0025] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0026] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. Similar words such as "connected" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0027] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.
[0028] During incremental inference of large models, a vast amount of weight data must be repeatedly read, consuming significant memory bandwidth and severely limiting inference efficiency. This is particularly evident in on-device inference scenarios. For example, for a 7TB-level model (such as a large language model), a text response of approximately 300 words requires reading over 2TB of weight data. Since the DDR (double-data-rate synchronous dynamic random access memory) interface bandwidth of typical on-device devices is approximately 100GB / s, loading the weight data alone can take tens of seconds.
[0029] Since the amount of weight data in large models is too large, the weight data is usually compressed. During inference calculations, the compressed weight data needs to be decompressed as a whole to obtain the complete weight data before calculation. This process requires the weight data to be decompressed as a whole before calculation, and the decompressed weight data needs to be stored additionally, which wastes storage space and increases costs.
[0030] The embodiment provided by the present disclosure determines the data to be calculated in the next step by detecting the data calculation process, and collects the corresponding target number of compressed data from the stored compressed data according to the determination result for parallel decoding, and performs parallel calculation based on the decoded data, so that what is input into the chip is compressed data, thereby saving analog memory bandwidth; and realizes decoding-while-computing, so that there is no need to prepare additional storage space to store decoded data, thereby saving space, reducing costs, and not affecting computing efficiency, and both the decoding process and the calculation process are parallel schemes, which improves data processing efficiency; in addition, the target number in the parallel decoding process is determined according to the parallel computing capability, and the compressed data of the target number are decoded in parallel, so that the parallelism of decoding is coordinated with the parallel computing capability, thereby realizing the coupling of decoding and computing.
[0031] The data processing method according to the embodiment of the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in a memory, or the method can be executed by a server.
[0032] FIG1 is a flow chart of a data processing method provided by an embodiment of the present disclosure. Referring to FIG1 , the method may include: looping through steps S11-S13 until all compressed data stored outside the many-core chip is calculated:
[0033] S11. Detect the large model calculation process, determine the data that needs to be calculated in the next step, and collect the corresponding target number of compressed data from the compressed data stored outside the chip based on the determination result; the target number is determined based on the parallel computing capability of the multi-core chip when calculating the large model.
[0034] In the embodiment of the present disclosure, the data may include but is not limited to large model weight data. Any large amount of data calculation process is applicable to the embodiment of the present disclosure. The embodiment of the present disclosure is explained below using large model weight data as an example.
[0035] In the embodiment of the present disclosure, in the large model calculation, the main calculation is the matrix multiplication calculation located in the feedforward network, and the large amount of weight data loaded is mainly used for this matrix multiplication calculation. In order to facilitate the transmission of a large amount of weight data, it is necessary to design a compression scheme to compress the large amount of weight data and store the compressed data of the large amount of weight data in DDR. The data transmitted to the large model calculation chip is also compressed data, thereby saving analog memory bandwidth.
[0036] In the embodiment of the present disclosure, the compressed data is obtained by segmenting the data to be compressed (such as the large model weight data) and then compressing the multiple segments separately. The large model weight data can be segmented and compressed in advance and stored in the DDR memory. The compression method during segmented compression can be determined according to the data calculation method. For example, for the matrix multiplication operation, the specific operation process is the multiplication of the vector and the matrix. For each input vector data, there is a row of a specific weight matrix corresponding to it. In order to achieve parallel operation, the weight matrix can be compressed by row.
[0037] In the embodiment of the present disclosure, after the compressed data is transmitted to the many-core chip, it is necessary to decode the compressed data to obtain decoded data so as to perform calculations based on the decoded data.
[0038] In the embodiment of the present disclosure, since the amount of large model weight data is huge, if the compressed data of the large model weight data is decoded and then stored, it will inevitably take up a lot of storage space and affect the data processing efficiency. Therefore, decoding can be performed while calculation is performed to save storage space for the decoded data and save costs. The decoding speed can be set based on the calculation speed so as not to affect the calculation efficiency.
[0039] In the embodiment of the present disclosure, in order to achieve decoding and calculation at the same time, the data calculation process can be detected to determine the data that needs to be calculated in the next step (i.e., the next set of calculation data) so that the next set of calculation data can be sent to the calculation module of the computing chip before the current calculation is completed.
[0040] In the embodiment of the present disclosure, determining the data to be calculated in the next step may include:
[0041] According to the preset calculation cycle, the starting row and / or starting column for data collection in the next step is determined.
[0042] In an embodiment of the present disclosure, for example, at the start of each preset calculation cycle, a starting row and / or starting column for collecting data in the next step may be determined.
[0043] In the embodiment of the present disclosure, the data to be calculated in the next step can be determined based on the parallel computing capability of the hardware of the computing module (for example, the number of parallel computing rows, the number of parallel computing columns, etc.). For example, a preset number of data (such as the data obtained by multiplying the number of parallel computing rows by the number of parallel computing columns) can be input each time when performing the calculation, and only one data calculation can be completed in each computing cycle. The preset number of data can be calculated in parallel, and then the data required for calculation in the next computing cycle can be prepared at the beginning of each computing cycle, that is, the next group of compressed data can be collected for decoding, so that the previous group of decoded data can be calculated just after the decoding of this group of compressed data is completed. Before collecting, it is necessary to confirm the starting row and / or starting column of the next group of compressed data.
[0044] In the disclosed embodiment, there is no limit on the calculation period and it can be defined as needed. The duration of the calculation period needs to ensure that the target amount of compressed data collected can be fully decompressed within the period and that the calculation process does not wait as much as possible. That is, within the calculation period, the target amount of compressed data is decoded and the calculation of the previous set of decoded data is completed, so as not to affect the calculation efficiency.
[0045] In embodiments of the present disclosure, when collecting a target amount of compressed data, data collection can be performed based on a preset data collection window. For example, a starting window position for data collection can be determined. Based on this starting window position, the data collection window is used to collect the target amount of compressed data from the matrix corresponding to the compressed data. The data collection window includes the number of rows and columns of the collected data; the number of data contained within these rows and columns is equal to the target amount of data.
[0046] In the embodiment of the present disclosure, the target number may include a target number of rows and / or a target number of columns;
[0047] Collecting a corresponding target amount of compressed data from the stored compressed data according to the determination result, including:
[0048] Collecting the target number of rows of compressed data starting from the starting row in the matrix corresponding to the compressed data according to the order of rows; and / or,
[0049] According to the arrangement order of the columns, compressed data of a target number of columns are collected starting from the starting column in the matrix corresponding to the compressed data.
[0050] In the embodiment of the present disclosure, parallel decoding and calculation can be performed from the perspective of rows alone, or from the perspective of columns alone. In subsequent embodiments, the parallel decoding and calculation scheme of the embodiment of the present disclosure is explained from the perspective of parallel rows and columns.
[0051] In the disclosed embodiments, to achieve parallel decoding and computation of compressed data from both row and column perspectives, it is necessary to first select a target number of rows and columns of compressed data from the matrix corresponding to the compressed data. The target number of rows and columns are the number of rows and columns that can be computed in parallel by the computing chip during parallel computation.
[0052] In the embodiment of the present disclosure, for example, if a compressed data contains N rows and M columns (both M and N are positive integers), if the target number of rows is n1 and the target number of columns is m1, then before each calculation, the compressed data of the 0th row to the n1th row and the 0th column to the m1th column is first collected for decoding, and the decoded data obtained after decoding is used for calculation, wherein n1 and m1 are both the degree of parallelism. During the calculation of the decoded data of the 0th row to the n1th row and the 0th column to the m1th column, the compressed data of the m1+1th column to the 2m1th column of the 0th row to the n1th row is collected for decoding, and the decoded data obtained after decoding is used for calculation. And so on, until all the data of the 0th row to the n1th row are sent for calculation, then the compressed data of the 0th column to the Mth column of the n1+1th row to the 2n1th row is read, and the above process is repeated until all the compressed data calculations are completed.
[0053] In the embodiment of the present disclosure, the above scheme is used to transmit compressed data to the computing chip one by one according to the computing parallelism, thereby providing a technical basis for stably realizing decoding and computing at the same time without affecting computing efficiency.
[0054] In the embodiment of the present disclosure, the scheme of step S11 can be implemented by a preset scheduling module, which can be set between the computing module of the large model computing chip and the memory for storing compressed data, and is used to determine which part of the weight data the computing module needs to calculate, so that the corresponding compressed data can be retrieved from the memory for decoding.
[0055] S12. Perform parallel decoding on the target amount of compressed data to obtain decoded data.
[0056] In the embodiment of the present disclosure, after the calculation is tested through the solution of step S11 and the target number of compressed data is sequentially obtained and input into the many-core chip, each group of compressed data needs to be decoded in parallel.
[0057] In an embodiment of the present disclosure, decoding a target amount of compressed data in parallel to obtain decoded data may include:
[0058] The compressed data is cached, aligned, and decoded in parallel with a target number of rows and / or a target number of columns as the parallelism.
[0059] In the embodiments of the present disclosure, the compressed data of the target number of rows are cached, aligned and decoded in parallel with the target number of rows as the parallelism; and / or the compressed data of the target number of columns are cached, aligned and decoded in parallel with the target number of columns as the parallelism.
[0060] In an embodiment of the present disclosure, if the target amount of compressed data collected is only multiple rows and one column of data (i.e., a column vector), the multiple rows and one column of data are decoded in parallel from the row dimension. For example, for a group of n (n is a positive integer) rows and 1 column of target data, n cache decoding units can be set, and each cache decoding unit caches, aligns and decodes 1 row and 1 column of data (i.e., 1 data) respectively.
[0061] In an embodiment of the present disclosure, if the target amount of compressed data collected is only 1 row and multiple columns of data (i.e., a row vector), the 1 row and multiple columns of data are decoded in parallel from the column dimension. For example, for a group of 1 row and m (m is a positive integer) columns of target data, m cache decoding units can be set, and each cache decoding unit caches, aligns and decodes 1 row and 1 column of data (i.e., 1 data) respectively.
[0062] In an embodiment of the present disclosure, if the target number of compressed data collected is multi-row and multi-column data (i.e., a matrix), the multi-row and multi-column data can be decoded in parallel from both row and column dimensions. For example, in the case of parallel decoding with a target number of rows (such as row a) and a target number of columns (such as column b), the collected row a compressed data can first be assigned row by row to a group of cache decoding units corresponding to the corresponding row. In the group of cache decoding units corresponding to each row, b cache decoding units can be included respectively. The b cache decoding units are cache decoding units corresponding to each compressed data on the b columns of the corresponding row. Through the b cache decoding units corresponding to each row, the compressed data of the b columns of the row can be cached, aligned, and decoded in parallel. Through this embodiment, the collected multi-row and multi-column compressed data can be decoded in parallel from both row and column dimensions.
[0063] In an embodiment of the present disclosure, for compressed data of multiple rows and columns, when setting a group of cache decoding units corresponding to each row and a cache decoding unit corresponding to each column of the group of cache decoding units, the identifier of the group of cache decoding units corresponding to each row and the identifier of the cache decoding unit corresponding to each column in the group of cache decoding units corresponding to each row can be recorded, so that for a group of cache decoding units corresponding to any row, the same row of compressed data in the collected compressed data (matrix) is always cache-decoded, and for any cache decoding unit corresponding to each column in the group of cache decoding units corresponding to any row, the same column of compressed data in any row of compressed data is cache-decoded.
[0064] In the embodiment of the present disclosure, for example, if, based on the parallel processing capability of the computing module, a decoded data set of 10 rows and 5 columns can be parallelized at a time, then the target amount of compressed data collected from the memory each time is also compressed data of 10 rows and 5 columns. Accordingly, the cache decoding unit can include 10 rows of cache decoding units, each row of cache decoding units including 5 columns of cache decoding queues, thus including a total of 50 cache decoding units, which can cache and decode one compressed data in each of the 10 rows and 5 columns (50) of compressed data, thereby achieving parallel storage and decoding of the compressed data (matrix). Wherein, if NiMj is used to represent a cache decoding unit, Ni represents the i-th row, Mj represents the j-th column, i and j are positive integers, i is less than or equal to the maximum number of rows, and j is less than or equal to the maximum number of columns. For any cache decoding unit NiMj of the i-th row and j-th column, the compressed data of the i-th row and j-th column in the compressed data (matrix) are always cached and decoded. For example, for the 50 cache decoding units mentioned above, the cache decoding unit in the 2nd row and 5th column always caches and decodes the compressed data in the 2nd row and 5th column of the 10 rows and 5 columns of compressed data (50), reducing the possibility of data confusion.
[0065] In the embodiment of the present disclosure, the cache decoding unit may include a cache unit and a decoding unit. The cache unit is used to cache and align the processed compressed data, and the decoding unit is used to decode the aligned compressed data.
[0066] In the embodiment of the present disclosure, the compressed data may include but is not limited to compressed data obtained through Huffman coding. The cache unit can read the large model weight data compressed through Huffman coding from the memory when the update signal is valid. After the cache unit completes the caching and receives the ready signal from the decoding unit, the cached data is sent to the decoding unit. The decoding unit can decode one weight data per clock cycle.
[0067] In the embodiment of the present disclosure, the cache unit can align the compressed data input to the memory, thereby aligning the output data of multiple decoding units, facilitating format conversion and subsequent data calculations (such as matrix multiplication), and realizing parallel decompression and parallel calculation of data calculations.
[0068] In the disclosed embodiment, data can be read in bytes, with each byte being 8 bits long. However, the maximum length of compressed data is 12 bits, and the read bit width does not match. Therefore, data alignment can be achieved by setting a cache unit.
[0069] In the disclosed embodiment, the cache unit can be configured with a two-level cache. The cached data can be 48 bits (needing to be a common multiple of 8 and 12, 6 bytes). The data cache output must be 12 bits (i.e., the input to the decoding unit is 12 bits). The cache unit data must be output four times before new memory data can be read. Therefore, a 2-bit counter can be used for data bit selection to implement this function.
[0070] In the disclosed embodiment, a cache unit can use a state machine to indicate the state of cached data. The state machine can include five states: a waiting state, three data loading states, and a working state. Because memory data is read and the address is changed only when the update signal is valid, three data loading states can be set to replace the initial reset data 0.
[0071] In the embodiment of the present disclosure, the ports of the cache unit can be configured as follows, as shown in Table 1:
[0072] Table 1
[0073] In the embodiment of the present disclosure, an arbitrator is set between the memory and multiple cache decoding units (the multiple cache decoding units can be regarded as a whole cache decoding part), mainly between the memory and the cache units.
[0074] In the embodiment of the present disclosure, since there is a bus between the memory and the cache decoding part, there is a bus conflict between multiple cache decoding units, so an arbiter can be set between the memory bus and the cache unit. After the arbiter reads the compressed data from the memory, it can be distributed to each cache decoding unit in sequence (for example, according to the order of arrangement of the rows).
[0075] In the embodiment disclosed herein, this is related to the degree of parallelism, and is also related to how much memory data can be read at one time (i.e., the aforementioned target amount of compressed data). In theory, the amount of compressed data read from the memory by the arbitrator should just match the degree of parallelism, so that each time the compressed data is read from the memory, it can be distributed to each cache decoding unit, and there will be no situation where the cache decoding unit does not work.
[0076] In the embodiment of the present disclosure, the decoding module may include a counter and a decoder with priority, and a two-level cache may also be set (based on the embodiment of the aforementioned cache unit outputting 12 bits, the decoding unit may cache two 12 bits of compressed data) to ensure the timing stability of the compressed data input and the decoded data output, and to solve the problem of compressed data being truncated due to bit width limitations during the decoding process. After resetting, the data is loaded level by level, and the decoding is judged bit by bit from high to low.
[0077] In the disclosed embodiment, a pointer can be set to mark the progress of the current decoding. If the compressed data is in INT4 format (int4 is a 4-byte signed integer data format, the sign occupies 1 bit, and the remaining 31 binary bits represent the value, and the maximum positive number is 0x7fffffff), after Huffman coding, the minimum encoding bit width of the compressed data is 2 bits and the maximum encoding bit width is 12 bits. Therefore, when the two-level storage of the decoding unit is set, the compressed data input of each level of storage is 12 bits, so that each level of storage can ensure the bit width of storing one compressed data (because some compressed data is less than 12 bits, when storing, multiple compressed data less than 12 bits may be stored in the same level, and a compressed data may have part of the data stored in the previous level storage and part of the data stored in the next level storage, that is, the truncated situation mentioned above. In this case, if there is only one level of storage, the truncated compressed data cannot be decoded in time, which may affect the subsequent computing efficiency), the output of the decoding unit is 4 bits.
[0078] In the embodiment of the present disclosure, the compressed data may be obtained based on prefix coding, for example, based on Huffman coding.
[0079] In an embodiment of the present disclosure, parallel caching, alignment, and decoding of compressed data include:
[0080] The compressed data is decoded in order of priority from high to low; wherein the priority is determined according to the encoding bit width of the input compressed data, and the smaller the encoding bit width, the higher the priority.
[0081] In the disclosed embodiments, prefix encoding refers to encoding a character set in such a way that the encoding of any character in the character set is not a prefix of the encoding of any other character. Huffman encoding is a prefix encoding method. Therefore, for compressed data based on Huffman encoding, a decoder with a priority level can be used in the decoding unit. The priority level is determined by the encoding bit width from smallest to largest.
[0082] In the embodiment of the present disclosure, the ports of the decoding unit can be set as follows, as shown in Table 2:
[0083] Table 2
[0084] In the embodiment of the present disclosure, the workflow of the cache unit and the decoding unit in the cache decoding unit is introduced below:
[0085] First, after being reset, the cache unit pulls the update signal high, reading in new compressed data to occupy the first-level cache and simultaneously transitioning to a new state. The same operation is repeated over the next two clock cycles, filling both levels of cache with valid compressed data and transitioning to a working state. It then receives a ready signal from the decoding unit and, when the ready signal is valid, transmits the cached compressed data to the decoding unit. Whenever the ready signal is valid and compressed data is transmitted to the decoding unit, a counter in the cache unit increments by one. Each counter value corresponds to one of the four 12-bit segments of the 48-bit cache data, divided from the highest bit to the lowest bit.
[0086] The decoding unit (which may include a counter and a priority encoder) also has a two-level cache to ensure the timing stability of compressed data input and decoded data output and to prevent truncation of encoded data due to bit width limitations during the decoding process. After reset, the two-level cache of the decoding unit is filled level by level, and decoding is performed bit by bit from high to low. If the compressed weight data is in INT4 format, the maximum encoding bit width after Huffman encoding is 12 bits, so the compressed data input of the decoding unit is 12 bits and the decoded data output is 4 bits.
[0087] The counter of the decoding unit starts counting after being reset. It needs to wait for the data of the cache unit to be loaded and the two-level cache inside the decoding unit to be loaded. A total of 6 clock cycles are required. At the same time, the counter of the decoding unit is kept unchanged during the subsequent time. After the two-level cache of the decoding unit is loaded, the decoding stage is entered.
[0088] The decoding unit sets a pointer to mark the progress of the current decoding, and pieces together the two-level cache into a 24-bit compressed data. After reset, the pointer is reset to the highest bit. The pointer position is changed after each decoding and recognition operation. If the bit width index pointed to by the pointer after movement is less than or equal to 11, it means that the 12-bit data in the higher-level cache has been fully decoded and new data needs to be loaded. At this time, the ready signal ready is pulled high, and the pointer is moved to the new position after the compressed data is loaded. Since the compressed data input is a complete data stream, a standard design of decoding one can be used in one clock cycle (which can be regarded as the calculation cycle of the embodiment of the present disclosure). The valid signal of the compressed data is pulled high when the counter of the decoding unit is greater than 6. After that, the compressed data output in each clock cycle is valid.
[0089] S13. Perform data parallel calculation based on the decoded data.
[0090] In the embodiment of the present disclosure, after decoding a group of compressed data, a group of decoded data can be obtained, and the group of decoded data is sent to the calculation module for parallel calculation.
[0091] In the embodiment of the present disclosure, performing data parallel computing based on the decoded data may include:
[0092] Calculate each row of data in the matrix corresponding to the decoded data in parallel with the first input data; and calculate each column of data in each row of data in parallel with the data of the corresponding column in the first input data; or
[0093] Each column of data in the matrix corresponding to the decoded data is calculated in parallel with the second input data; and each row of data in each column of data is calculated in parallel with the data of the corresponding row in the second input data.
[0094] In an embodiment of the present disclosure, the compressed data may include large model weight compressed data, and the decoded data may include: large model weight decoded data;
[0095] Perform data-parallel computations based on the decoded data, which may include:
[0096] During the incremental inference process of the large model, matrix multiplication calculations are performed in the feedforward neural network layer of the large model based on the decoded data of the large model weights.
[0097] In the embodiment of the present disclosure, for example, for large model weight data, parallel computing mainly involves the multiplication and addition of input data (such as the first input data, which can be a row vector) and decoded data (matrix). For example, when an input row vector is operated with the decoded weight data matrix, the row vector can be multiplied in parallel with each column of the weight data matrix, first realizing the parallelism of inter-column calculations, in the process of multiplying the row vector with each column of the weight data matrix in parallel, each data in the row vector and each corresponding data in each column can be multiplied in parallel, realizing the parallelism of multiplication of data in each column, after each column of weight data is multiplied with the input row vector, multiple multiplication results will be obtained, and these multiplication results can be added to obtain the calculation result corresponding to the column weight data, and the addition process of multiple multiplication results corresponding to each column weight data can also be carried out in parallel, thereby realizing another level of parallel computing.
[0098] In the embodiment of the present disclosure, the above scheme is explained using the input data as a row vector (i.e., the first input data) as an example. When the input data is a column vector (i.e., the second input data), the scheme of the embodiment of the present disclosure is still used. The computational parallel scheme is similar to the above scheme. The difference is that when the column vector is multiplied with the weight data (matrix), it needs to be multiplied in parallel with each row of the weight data, which will not be repeated here.
[0099] In the embodiment of the present disclosure, the parallel computing solution can be implemented by a preset computing module, which can be a matrix multiplication acceleration unit in a Feedforward Neural Network (FFN).
[0100] In an embodiment of the present disclosure, before the decoded data is input into the calculation module, the format of the decoded data can first be converted by a format conversion module. For example, the format of the input decoded data is converted into FP16 (half-precision floating point) format, and then a target number (such as 1024) of FP16 format decoded data are accepted at one time to perform multiplication and addition calculations with the input data (such as vector x).
[0101] In the disclosed embodiment, the format conversion module can be implemented using combinational logic. It can accept decoded data input in INT4 and INT8 (int8 is an 8-byte signed integer data, the sign occupies 1 bit, the remaining 63 binary bits represent the value, and the maximum positive number is 0x7fffffffffffffff) format, and convert it to FP16 format output. The conversion ideas of INT4 and INT8 are the same. For example, the sign bit can be extracted first, and then the first non-zero bit is found from the highest bit to the lowest bit to determine the exponent bit, and the decimal place is determined by shifting the bits, and finally the sign bit, exponent bit, and decimal place are pieced together according to the FP16 data format to obtain the final result.
[0102] In the embodiment of the present disclosure, the ports of the format conversion module can be configured as follows, as shown in Table 3:
[0103] Table 3
[0104] In an embodiment of the present disclosure, the computing module may mainly involve hardware implementation of multiplication and addition operations in FP16 data format.
[0105] In the embodiment of the present disclosure, first, the multiplication operation can be implemented by a half-precision floating-point multiplier. The functional part adopts combinational logic and the output part adopts sequential logic to ensure that each clock cycle (which can be regarded as the calculation cycle of the embodiment of the present disclosure) is calculated once. The functional part mainly includes the processing of the sign bit, exponent borrowing, multiplication calculation and out-of-range decimals. Accept two half-precision floating-point inputs (input decoded data and input data to be calculated with the decoded data), and output a half-precision floating-point number. Special cases include any one input (decoded data or any one of the input data) or two inputs (decoded data and input data) are 0. In this case, the output is 0. The sign bit judgment can be implemented using an exclusive OR gate. In the calculation part, according to the data format of the half-precision floating-point number, the input is split into a sign part, an exponent part and a decimal part, and stored in separate registers for subsequent calculations. The multiplication part directly adopts behavioral-level description (behavioral-level description describes the function or mathematical model of the circuit, establishes the relationship between the input and output of the circuit or system, has a high degree of abstraction, and has nothing to do with hardware implementation; in the behavioral-level description, the source program can use a large number of arithmetic operations, relational operations, inertial delays, transmission delays, and other ultra-high-speed integrated circuit hardware description language VHDL description statements that are difficult or impossible to perform logic synthesis). The exponent matching part uses if-else and nested if forms to find the first non-zero number and then correct the exponent. Finally, the three split parts are reassembled and output.
[0106] Secondly, a half-precision floating-point adder can be used to perform addition operations. Similar to the multiplier, this uses a combination of combinational and sequential logic, performing calculations once per clock cycle. However, the functionality differs, primarily in that the order of the two values to be added must be balanced before performing addition or subtraction operations based on the sign. Therefore, compared to the multiplier, the adder can define a separate shift register to calculate the order balance during addition. The calculation phase begins with order balancing, followed by sorting by sign, determining the mantissa and exponent, and finally combining the results for output.
[0107] Finally, the above multipliers and adders can be instantiated multiple times to construct multiplier arrays and adder arrays. These arrays can be combined to form a fully connected layer matrix multiplication acceleration module to achieve parallel computing; a counter can be used to determine when the matrix multiplication operation requires column switching.
[0108] In the disclosed embodiment, the decoded data for parallel operation comes from the decoding unit, and the input data (such as row vector or column vector) can be given to the input end of the computing module (mainly the multiplier array) by the vector memory controlled by the address generation unit.
[0109] In the embodiment of the present disclosure, an activation unit may be further provided, the input of which is connected to the output of the calculation module; the activation unit is used to perform nonlinear combination on all calculation results of the calculation module.
[0110] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0111] In addition, the present disclosure also provides a data processing device, a multi-core chip, and a computer-readable storage medium, all of which can be used to implement any data processing method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.
[0112] Figure 2 is a block diagram of a data processing device provided in an embodiment of the present disclosure. Figures 3a and 3b are schematic diagrams of two data processing devices provided in an embodiment of the present disclosure. Figure 3a is a schematic diagram of a data processing device provided in an embodiment of the present disclosure for decoding a row or column of compressed data (the number of cache decoding units is determined based on the number of columns in a row or the number of rows in a column), and Figure 3b is a schematic diagram of a data processing device provided in an embodiment of the present disclosure for decoding multiple rows and columns of compressed data.
[0113] 2 , 3 a and 3 b , an embodiment of the present disclosure provides a data processing device 300 , which includes: a scheduling module 301 , a decoding module 302 and a calculation module 303 ;
[0114] The scheduling module 301 is used to detect the large model calculation process, determine the data that needs to be calculated in the next step, and collect the corresponding target amount of compressed data from the compressed data stored outside the data processing device based on the determination result; the target amount is determined based on the parallel computing capability during the large model calculation; the compressed data is obtained by segmenting the data to be compressed and then compressing multiple segments separately.
[0115] The decoding module 302 is configured to decode the target amount of compressed data in parallel to obtain decoded data.
[0116] The calculation module 303 is used to perform data parallel calculation according to the decoded data.
[0117] In the embodiment of the present disclosure, the decoding module 302 may include: a plurality of cache decoding units 3021; the number of cache decoding units 3021 is determined based on the parallel computing capability during large model calculations, and a target number is determined based on the number of cache decoding units 3021, which is the same as the number of cache decoding units 3021;
[0118] The cache decoding unit 3021 is used to decode the target amount of compressed data in parallel to obtain decoded data.
[0119] In the disclosed embodiment, the plurality of cache decoding units 3021 are arranged in a plurality of rows and columns;
[0120] Each row cache decoding unit 3021 is used to decode the compressed data of the corresponding row in the matrix corresponding to the compressed data in parallel;
[0121] Each column cache decoding unit 3021 in each row cache decoding unit 3021 is used to decode in parallel the compressed data in the corresponding column of the compressed data in the corresponding row in the compressed data corresponding matrix.
[0122] In the embodiment of the present disclosure, the cache decoding unit 3021 includes: a cache unit 30211 and a decoding unit 30212;
[0123] a cache unit 30211, configured to cache and align the obtained compressed data;
[0124] The decoding unit 30212 is used to decode the aligned compressed data.
[0125] In the embodiment of the present disclosure, the decoding module may further include: an arbitrator 3022;
[0126] The arbiter 3022 is configured to distribute a target amount of compressed data collected by the scheduling module to the multiple cache decoding units 3021 .
[0127] In the embodiment of the present disclosure, the computing module 303 includes: a first computing array; the first computing array includes a plurality of computing subunits; the computing subunits are used to:
[0128] Calculate in parallel the data of each column in each row of the matrix corresponding to the decoded data and the data of the corresponding column in the first input data; or,
[0129] The decoded data corresponds to each row of data in each column of the matrix and the data in the corresponding row of the second input data and is calculated in parallel.
[0130] In the embodiment of the present disclosure, the calculation subunit may include a multiplier. The calculation module may also include an adder.
[0131] In the embodiment of the present disclosure, the data processing apparatus may further include a format conversion module 304 and an activation unit 305;
[0132] The format conversion module 304 is connected between the decoding module 302 and the calculation module 303 and is used to convert the format of the decoded data 302;
[0133] The input of the activation unit 305 is connected to the output of the calculation module 303 ; the activation unit 305 is used to perform nonlinear combination on all calculation results of the calculation module 303 .
[0134] In the disclosed embodiment, the data processing device, the scheduling module 301, the decoding module 302, and the computing module 303 operate in a pipelined and parallel manner, which hardly affects the working efficiency of the computing module 303. Since the computing module 303 performs block calculations, in order to save the cache between the computing module 303 and the decoding module 302 and ensure the timely supply of data required by the computing module, the decoding process of the decoding module 302 and the data storage form of the DDR memory need to be coordinated with the operation process of the computing module 303 to form a coupling device. In order to solve the problem of a single DDR memory emulation channel supplying multiple cache decoding units in the decoding module 302, an arbitrator needs to be added. In order to solve the problem of the memory output compressed data being determined and the input rate of the compressed data input to the computing chip being uncertain, a cache decoding unit is added, and the cache decoding unit is used to achieve parallel decoding of multiple rows and / or columns of compressed data. In addition to the static structure, the data processing device of the embodiment of the present disclosure is also dynamically sent to the calculation module 303 at a fixed speed (for example, 1 number per clock). When the calculation module 303 calculates the decoded data of the i-th column, the parallel cache decoding unit decodes the compressed data of the i+1-th column, thereby realizing pipeline parallelism of data processing.
[0135] FIG4 is a block diagram of a many-core chip provided by an embodiment of the present disclosure.
[0136] 4 , an embodiment of the present disclosure provides a many-core chip, the many-core chip including:
[0137] multiple processing cores 401; and
[0138] The on-chip network 402 is configured to exchange data between multiple processing cores and external data; wherein one or more processing cores 401 store one or more instructions, and one or more instructions are executed by one or more processing cores 401 to enable one or more processing cores 401 to execute the data processing method.
[0139] In some embodiments, the electronic device may be a brain-inspired chip. Because brain-inspired chips can use vectorized computing and require external memory, such as Double Data Rate (DDR) synchronous dynamic random access memory, to load parameters such as weight information of the neural network model, the disclosed embodiments utilize batch processing for higher computational efficiency.
[0140] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned data processing method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0141] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned data processing method.
[0142] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium).
[0143] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically contains computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0144] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0145] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0146] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0147] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0148] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0149] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0150] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0151] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A data processing method, wherein: The method comprises: Detecting the large model calculation process, determining the data to be calculated in the next step, and collecting a corresponding target amount of compressed data from the compressed data stored off-chip based on the determination result; the target amount is determined based on the parallel computing capability of the many-core chip when calculating the large model; the compressed data is obtained by segmenting the data to be compressed and compressing the multiple segments separately; Parallel decoding of the target amount of compressed data to obtain decoded data; Data parallel computing is performed according to the decoded data.
2. The data processing method according to claim 1, wherein: The compressed data includes large model weight compressed data; The decoded data includes: large model weight decoded data; The performing data parallel calculation according to the decoded data includes: During the large model incremental inference process, matrix multiplication calculations are performed in the feedforward neural network layer of the large model based on the large model weight decoding data.
3. The data processing method according to claim 1, wherein: The step of determining the data to be calculated in the next step includes: According to the preset calculation cycle, the starting row and / or starting column for data collection in the next step is determined.
4. The data processing method according to claim 3, wherein: The target number includes a target number of rows and / or a target number of columns; The collecting of a corresponding target amount of compressed data from the stored compressed data according to the determination result includes: Collecting the target number of rows of compressed data starting from the starting row in the matrix corresponding to the compressed data according to the arrangement order of the rows; and / or, According to the arrangement order of the columns, the compressed data of the target number of columns are collected starting from the starting column in the matrix corresponding to the compressed data.
5. The data processing method according to claim 4, wherein: The parallel decoding of the target amount of compressed data to obtain decoded data includes: The compressed data is cached, aligned, and decoded in parallel with the target number of rows and / or the target number of columns as a degree of parallelism.
6. The data processing method according to claim 5, wherein: The compressed data is obtained based on prefix encoding; The parallel caching, aligning and decoding of the compressed data includes: The compressed data is decoded in order of priority from high to low; wherein the priority is determined according to the encoding bit width of the input compressed data, and the smaller the encoding bit width, the higher the priority.
7. The data processing method according to claim 1, wherein: The performing data parallel calculation according to the decoded data includes: Calculate each row of data in the matrix corresponding to the decoded data in parallel with the first input data; and calculate each column of data in each row of data in parallel with the data of the corresponding column in the first input data; or Each column of data in the matrix corresponding to the decoded data is calculated in parallel with the second input data; and each row of data in each column of data is calculated in parallel with the data of the corresponding row in the second input data.
8. A data processing device, wherein: The device includes: a scheduling module, a decoding module and a calculation module; The scheduling module is used to detect the large model calculation process, determine the data to be calculated in the next step, and collect a corresponding target amount of compressed data from the compressed data stored outside the data processing device based on the determination result; the target amount is determined based on the parallel computing capacity during the large model calculation; the compressed data is obtained by segmenting the data to be compressed and compressing the multiple segments separately; The decoding module is used to decode the target amount of compressed data in parallel to obtain decoded data; The computing module is configured to perform data parallel computing based on the decoded data.
9. The data processing apparatus according to claim 8, wherein: The decoding module includes: a plurality of cache decoding units; the number of the cache decoding units is determined according to the parallel computing capability during large model calculation, the target number is determined according to the number of the cache decoding units, and the target number is the same as the number of the cache decoding units; The cache decoding unit is used to decode the target amount of compressed data in parallel to obtain decoded data.
10. The data processing apparatus according to claim 9, wherein: The plurality of cache decoding units are arranged in a plurality of rows and a plurality of columns; The cache decoding unit in each row is used to decode the compressed data in the corresponding row in the matrix corresponding to the compressed data in parallel; Each column of the cache decoding units in each row is used for decoding in parallel the compressed data in the corresponding column of the compressed data in the corresponding row in the compressed data corresponding matrix.
11. The data processing apparatus according to claim 9, wherein: The cache decoding unit includes: a cache unit and a decoding unit; The cache unit is used to cache and align the obtained compressed data; The decoding unit is used to decode the aligned compressed data.
12. The data processing apparatus according to claim 8, wherein: The decoding module further includes: an arbitrator; The arbiter is configured to distribute the target amount of compressed data collected by the scheduling module to the multiple cache decoding units.
13. The data processing apparatus according to claim 8, wherein: The computing module includes: a first computing array; the first computing array includes a plurality of computing subunits; the computing subunits are configured to: Parallel calculation is performed on each column of data in each row of the matrix corresponding to the decoded data and the data in the corresponding column of the first input data; or, The data of each row in each column of the matrix corresponding to the decoded data and the data of the corresponding row in the second input data are calculated in parallel.
14. The data processing apparatus according to claim 10, wherein: The device also includes a format conversion module and an activation unit; The format conversion module is connected between the decoding module and the calculation module, and is used to convert the format of the decoded data; The input of the activation unit is connected to the output of the calculation module; the activation unit is used to perform nonlinear combination on all calculation results of the calculation module.
15. A many-core chip, wherein: include: Multiple processing cores; as well as An on-chip network is configured to exchange data between multiple processing cores and external data; wherein one or more instructions are stored in one or more of the processing cores, and one or more of the instructions are executed by one or more of the processing cores to enable one or more of the processing cores to execute the data processing method as described in any one of claims 1 to 7.
16. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the computer program implements the data processing method according to any one of claims 1 to 7.
17. A computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein: When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device implements the data processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Model hyper-parameter valuing method and device, processing core, equipment, chip and medium
CN116542286A
Data transmission method and device, electronic equipment and computer readable storage medium
CN117319373A
Data processing method and device, equipment and storage medium
CN117519996A
Data processing method and device based on many-core chip, chip, system and equipment
CN117634571A
Data processing method and device, many-core chip and storage medium
CN118249816A