Task processing method, device and equipment based on adaptive block floating point data

Through the adaptive block floating-point data processing method, the block size and shared index are dynamically configured, which solves the problem of imbalance between precision and efficiency in block floating-point data representation, achieves more efficient data compression and processing performance, adapts to different data characteristics, and improves the flexibility and configurability of task processing.

CN120704639APending Publication Date: 2025-09-26SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510810520.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing block floating-point data representation methods cannot balance processing accuracy and efficiency when faced with scenarios with variable data distribution characteristics, especially when the numerical differences within the block are extremely large or contain a large number of zero values ​​or special values, resulting in loss of accuracy and low efficiency.

Method used

An adaptive block floating-point data processing method is adopted. By dynamically configuring the block size and shared index determination strategy, the block floating-point parameters are dynamically adjusted according to the data characteristics, including block size, shared index and special value processing mode, and the mantissa representation is optimized to adapt to different data characteristics.

Benefits of technology

It achieves the goal of improving compression efficiency and processing performance while ensuring data processing accuracy, adapting to different data characteristics, improving flexibility and configurability, and ensuring a balance between task processing accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704639A_ABST
    Figure CN120704639A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of task processing, and discloses a task processing method, device and equipment based on adaptive block floating point data, and the method comprises the steps: receiving a group of floating point data corresponding to a target task as a current data block; determining a block floating point parameter of the current data block according to a preset rule or at least one data characteristic of the current data block; calculating a sharing index for the current data block according to a sharing index determination strategy; processing each floating point data in the current data block based on the sharing index to obtain respective corresponding mantissas; combining the sharing index and each mantissa into a block floating point representation of the current data block; and processing the target task based on the block floating point representation of the current data block. According to the task processing method and device, the problem that the processing precision and the processing efficiency are unbalanced when task processing is carried out based on block floating point data representation in the prior art is solved, and the balance of the task processing precision and the task processing efficiency can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of task processing technology, for example, to a task processing method, apparatus, and device based on adaptive block floating-point data. Background Art

[0002] Floating-point numbers are a way of representing real numbers in computers and are widely used in fields such as scientific computing and artificial intelligence. Common floating-point formats, such as IEEE 754 standard single-precision (FP32), half-precision (FP16), and bfloat16 (BF16), represent data with a large dynamic range but also incur significant storage and computational overhead.

[0003] Block Floating Point (BFP) is a data representation method designed to balance storage efficiency and computational accuracy. Its core concept is to have a group of data (a "block") share a common exponent, with only the mantissa of each data point within the block stored. Compared to traditional floating-point formats that store exponents for each data point separately, BFP can significantly save storage space and potentially improve the computational efficiency of certain operations (such as matrix multiplication and FFT). The BFP format is commonly used in scenarios such as neural network acceleration and digital signal processing, and is particularly popular in hardware optimization for applications such as FPGAs and application-specific integrated circuits (ASICs).

[0004] Existing BFP implementations usually use a fixed block size and a strategy for determining a shared exponent based on the maximum absolute value within the block. Although this approach is simple and effective, when faced with scenarios where data distribution characteristics vary, a fixed block size and a single exponent determination strategy may not optimally balance accuracy and compression rate. For example, when the numerical differences within a block are extremely large, smaller values ​​may be quantized to zero due to the shared exponent, resulting in a loss of accuracy; when the numerical distribution within the block is flat, smaller blocks or finer exponent control may be more beneficial. In addition, for data blocks containing a large number of zero values ​​or special values ​​(such as NaN, Infinity), the existing BFP processing method is not efficient enough. In summary, when processing the corresponding target task based on floating-point data, the existing scheme has the problem of unbalanced processing accuracy and processing efficiency.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application. Summary of the Invention

[0006] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0007] The method, apparatus, device and storage medium for task processing based on adaptive block floating-point data provided by the embodiments of the present disclosure solve the problem of imbalance between processing accuracy and processing efficiency when performing task processing based on traditional block floating-point data representation.

[0008] The task processing method based on adaptive block floating-point data in the embodiment of the present disclosure includes:

[0009] Receive a set of floating-point data corresponding to the target task as the current data block;

[0010] Determining block floating-point parameters of the current data block according to a preset rule or at least one data characteristic of the current data block, the block floating-point parameters including at least a block size, a shared index determination strategy, and a special value processing mode;

[0011] Determine the strategy based on the sharing index and calculate a sharing index for the current data block;

[0012] Based on the shared index, each floating point data in the current data block is processed to obtain the corresponding mantissa;

[0013] The shared index and each tail array are combined into a block floating point representation of the current data block;

[0014] The target task is processed based on the block floating-point representation of the current data block.

[0015] In some embodiments, the block size is a variable integer selected from a predefined set, the set including at least two different block size values.

[0016] In some embodiments, the sharing index determination strategy includes:

[0017] The index of the floating-point data with the largest absolute value in the current data block is used as the shared index;

[0018] and / or, using an index corresponding to a specific statistic of the absolute value of floating-point data in the current data block as a shared index;

[0019] And / or, the sharing index of the current data block is determined by differential coding or smoothing filtering in combination with the sharing index of one or more previous historical data blocks.

[0020] In some embodiments, the above method further comprises:

[0021] Check whether there is a special value in the current data block;

[0022] If a special value is detected, a special value marker is added to the block floating point representation of the current data block.

[0023] In some embodiments, if the current data block is detected as an all-zero block, the mantissa portion in the corresponding block floating-point representation is padded with a predefined all-zero pattern or omitted.

[0024] The task processing device based on adaptive block floating-point data in the embodiment of the present disclosure includes:

[0025] A receiving module is used to receive a set of floating-point data corresponding to the target task as the current data block;

[0026] a parameter configuration module, configured to determine block floating-point parameters of the current data block according to a preset rule or at least one data characteristic of the current data block, the block floating-point parameters including at least a block size, a shared index determination strategy, and a special value processing mode;

[0027] The index calculation module is used to determine the strategy based on the sharing index and calculate a sharing index for the current data block;

[0028] The mantissa processing module is used to process each floating point data in the current data block based on the shared exponent to obtain the corresponding mantissa;

[0029] The encapsulation module is used to synthesize the shared index and each tail array into a block floating point representation of the current data block;

[0030] The processing module is used for processing the target task based on the block floating point representation of the current data block.

[0031] In some embodiments, the block size is a variable integer selected from a predefined set, the set including at least two different block size values.

[0032] In some embodiments, the sharing index determination strategy includes:

[0033] The index of the floating-point data with the largest absolute value in the current data block is used as the shared index;

[0034] and / or, using an index corresponding to a specific statistic of the absolute value of floating-point data in the current data block as a shared index;

[0035] And / or, the sharing index of the current data block is determined by differential coding or smoothing filtering in combination with the sharing index of one or more previous historical data blocks.

[0036] An electronic device provided by an embodiment of the present disclosure includes at least one processor;

[0037] and a memory communicatively coupled to the at least one processor;

[0038] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned task processing method based on adaptive block floating-point data.

[0039] The storage medium provided by the embodiment of the present disclosure stores program instructions, which, when run, execute the above-mentioned task processing method based on adaptive block floating-point data.

[0040] The method, apparatus, device, and storage medium for processing tasks based on adaptive block floating-point data provided by the embodiments of the present disclosure can achieve the following technical effects:

[0041] When processing the corresponding target task based on floating-point data, the present disclosure can better adapt to the statistical characteristics of different data by dynamically configuring the block size and index determination strategy, and achieve a better balance between accuracy and compression rate. Therefore, when processing tasks, the compression efficiency and processing performance can be improved while ensuring data processing accuracy. It has higher flexibility and configurability, thereby ensuring a balance between task processing accuracy and efficiency.

[0042] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition,

[0044] Figure 1 1 is a flowchart of a task processing method based on adaptive block floating-point data provided by an embodiment of the present disclosure;

[0045] Figure 2 1 is a flow chart of another task processing method based on adaptive block floating-point data provided by an embodiment of the present disclosure;

[0046] Figure 3 1 is a schematic structural diagram of a task processing device based on adaptive block floating-point data provided by an embodiment of the present disclosure;

[0047] Figure 4 1 is a structural diagram of a task processing device based on adaptive block floating-point data provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0048] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The accompanying drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0049] The terms "first," "second," and the like in the embodiments of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to facilitate the description of the embodiments of the present disclosure herein. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.

[0050] Unless otherwise stated, the term "plurality" means two or more.

[0051] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.

[0052] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0053] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0054] The following describes the task processing method, apparatus, device, and storage medium based on adaptive block floating-point data provided by the embodiments of the present disclosure in conjunction with the accompanying drawings.

[0055] Figure 1 This is a flowchart of a task processing method based on adaptive block floating-point data provided by an embodiment of the present disclosure.

[0056] like Figure 1 As shown, the task processing method based on adaptive block floating-point data may include:

[0057] S101, receiving a set of floating-point data corresponding to a target task as a current data block;

[0058] S102, determining block floating-point parameters of the current data block according to a preset rule or at least one data characteristic of the current data block, where the block floating-point parameters include at least a block size, a sharing index determination strategy, and a special value processing mode;

[0059] S103, determining a strategy based on the sharing index, and calculating a sharing index for the current data block;

[0060] S104, processing each floating point data in the current data block based on the shared index to obtain the corresponding mantissa;

[0061] S105, combining the shared index and each tail array into a block floating point representation of the current data block;

[0062] S106 , processing the target task based on the block floating-point representation of the current data block.

[0063] In some embodiments, the block size is a variable integer selected from a predefined set, the set including at least two different block size values.

[0064] In some embodiments, the sharing index determination strategy includes:

[0065] The index of the floating-point data with the largest absolute value in the current data block is used as the shared index;

[0066] and / or, using an index corresponding to a specific statistic of the absolute value of floating-point data in the current data block as a shared index;

[0067] And / or, the sharing index of the current data block is determined by differential coding or smoothing filtering in combination with the sharing index of one or more previous historical data blocks.

[0068] In some embodiments, Figure 1 The methods also include:

[0069] Check whether there is a special value in the current data block;

[0070] If a special value is detected, a special value marker is added to the block floating point representation of the current data block.

[0071] In some embodiments, if the current data block is detected as an all-zero block, the mantissa portion in the corresponding block floating-point representation is padded with a predefined all-zero pattern or omitted.

[0072] Figure 2 This is a flow chart of another task processing method based on adaptive block floating point data provided by an embodiment of the present disclosure, combined with Figure 2 ,right Figure 1 The task processing method based on adaptive block floating-point data is further described in .

[0073] Specifically, 1. Data reception and pre-analysis:

[0074] 1) Data Reception: The system receives floating-point data corresponding to the target task from external storage (such as DRAM), on-chip cache (such as SRAM), or the output of the previous processing unit. This floating-point data can be a single tensor, a shard / tile of a tensor, or a continuous data stream. The input floating-point data format can be standard single-precision (FP32), half-precision (FP16), bfloat16 (BF16), or other floating-point representations.

[0075] 2) Pre-analysis (optional): Perform online or near-online statistical analysis of the received data blocks (whose size may be related to the subsequently determined BFP block size N, or a preset analysis window). The purpose of this step is to extract key data features that can guide the subsequent adaptive configuration of BFP parameters. Pre-analysis may include, but is not limited to, one or more of the following:

[0076] a. Dynamic range assessment: Calculate the maximum absolute value, minimum value, and overall amplitude range (max-min) of the values ​​within the data block.

[0077] b. Numerical distribution characteristics: Analyze the mean, variance, and median of the data, or more carefully estimate the skewness and kurtosis of the data distribution to determine the degree of concentration or dispersion of the data and whether there is a long tail.

[0078] c. Special value statistics: Statistics on the frequency of occurrence of zero values ​​(Zeros), non-numbers (NaNs), and infinities (Infinities) in a data block and their distribution within the block.

[0079] d. Change and correlation analysis: For time series data or spatially adjacent data blocks, the rate of change of values ​​over time / space, local smoothness, or correlation between data blocks can be analyzed.

[0080] This pre-analysis process can be performed by dedicated lightweight hardware logic (such as comparators, counters, accumulator arrays), firmware on an embedded controller, or at specific checkpoints along the data path. The analysis results will be used as quantitative indicators to input into the parameter configuration step.

[0081] 2. Adaptive block parameter configuration:

[0082] The core of this step is to dynamically select or calculate a set of optimized block floating point (BFP) conversion parameters for the current data block (or a series of subsequent data blocks) to be processed based on the statistical characteristics of the data block output by the pre-analysis module and / or based on instructions from external systems (such as compiler optimization passes, runtime decision engines, and configuration parameters passed in when the user API is called). These parameters include at least:

[0083] 1) Block size (N) determination: N defines the number of data elements in a group that share the same index.

[0084] aN selection logic: For example, if pre-analysis indicates that the current data block has a small dynamic range and a relatively flat distribution of values, the parameter configuration module can select a larger N value (such as 32, 64, or even larger, depending on the upper limit supported by the hardware and the specific application scenario) to maximize compression efficiency and potential SIMD processing benefits. Conversely, if the data block exhibits a high dynamic range, contains many important high-frequency details, or has drastic value changes, a smaller N value (such as 4, 8, or 16) should be selected to more flexibly adapt to local data characteristics and reduce quantization errors caused by shared exponents.

[0085] Sources and constraints of bN: N can be selected from a predefined set of optional values ​​supported by the hardware (e.g., {8, 16, 32}), or calculated in real time (within the hardware constraints) by a small algorithm based on pre-analyzed metrics.

[0086] 2) Sharing index determination strategy selection: Select a specific algorithm or rule for calculating the sharing index of the N data elements.

[0087] a. Based on the maximum absolute value within a block: traverse N elements within a block and find the element with the largest absolute value. Its index (after appropriate bias adjustment and saturation processing) is used as the shared index. This is a traditional BFP strategy and is simple to implement.

[0088] b. Indices based on robust statistics within the block: For example, instead of directly using the maximum index, an equivalent index corresponding to a statistic of the absolute values ​​of all elements in the block (such as the geometric mean, the p-norm, or the kth percentile of the sorted values, such as the 99th percentile) is used. This strategy aims to reduce the dominant influence of extreme outliers on the shared index, thereby better protecting the representation accuracy of the majority of non-extreme values ​​within the block.

[0089] c. Index based on historical information or differential encoding: For consecutive BFP blocks in a data stream, their sharing indices may have local correlations in time or space. This strategy exploits this correlation. For example, the sharing index of the current block can be expressed as the difference with the sharing index of the previous (or multiple) blocks, or generated through a simple prediction model (such as linear prediction using the previous index). This helps reduce the total number of bits required to store the sharing index, further improving the compression ratio.

[0090] d. Strategy selection logic: It can be a lookup table or decision tree based on pre-analysis results (such as outlier detection results, smoothness evaluation of inter-block exponential changes), or specified by an external control signal.

[0091] 3) Special value handling mode configuration: defines how to identify, encode and decode special floating point values ​​(zero, NaN, Infinity) within a data block.

[0092] a. All-Zero Block Optimization Mode: This mode is enabled if pre-analysis detects that the current block of N elements is all zero. The shared exponent can be set to a predefined special "all-zero indication" code, and all N mantissas can be uniformly represented as zero or physically omitted, with only a flag bit or the special exponent code indicating this all-zero state.

[0093] b. NaN / Infinity Propagation and Encoding Mode: Configure how the system handles blocks containing NaNs or Infinity values. For example, a global "contains special values" flag could be set for blocks containing such values. Alternatively, a specific bit pattern could be used in the mantissa encoding of each NaN / Infinity value to uniquely identify it. The goal is to preserve these special semantics while minimizing storage overhead.

[0094] c. Subnormal number handling strategy: determines how to treat subnormal numbers in the original floating-point number during BFP conversion (whether to flush them to zero first), and whether or how to represent the "subnormal" mantissa relative to the shared exponent in BFP representation.

[0095] 3. Sharing index calculation and coding:

[0096] 1) Calculation: Based on the sharing index determination strategy selected for the current N-element data block in step 2, the value of the sharing index is accurately calculated.

[0097] If the Max-Abs strategy is selected, the hardware or software logic will iterate over the N data in the block, find the one with the largest absolute value, and extract its exponent. This exponent will then need to be adjusted based on the target BFP format. For example, if the original data is FP32 and the target format is something like BFP8, the IEEE 754 8-bit exponent (with a bias of 127) may need to be converted to a BFP (e.g.) 5-bit exponent, and any necessary saturation (clamping) will be performed to ensure it fits within the BFP exponent representation range.

[0098] b. If a statistically robust strategy is chosen: first calculate the selected statistic (such as the logarithmic mean, or find the percentile value), then derive the equivalent index from the statistic, and perform the same adjustments and saturation processing as above.

[0099] c. If the differential encoding strategy is selected: the difference between the original sharing index of the current block (obtained by one of the above methods) and the reference index (e.g., the decoded sharing index of the previous block) is calculated. This difference will be used as the encoded "index".

[0100] 2) Encoding: The calculated and adjusted shared exponent (or exponent difference) is encoded. The number of bits is determined by the BFP variant used (for example, BFP8 may have a specific number of bits for the shared exponent). If a special exponent encoding is used (such as a predefined exponent value for an all-zero block), that special encoding is used directly.

[0101] 4. Mantissa calculation and encoding:

[0102] 1) Processing object: Process the N floating-point data elements in the current data block one by one, but may skip those elements that have been specially marked or whose information has been captured by the special value processing module according to the special value processing mode.

[0103] 2) Mantissa extraction and alignment: For each floating-point data element that needs to be encoded normally:

[0104] a. Get the original sign bit (Sign), actual exponent (actual exponent, which has been debiased), and mantissa bits (including the implicit leading '1' that may exist in IEEE 754).

[0105] b. Compare the actual exponent of the element with the shared exponent of the block (which should also be unbiased or in the same biased basis) calculated in step 3. The difference between the two, exp_diff, determines the number of logical shifts required for the mantissa of the element (including the implicit bit).

[0106] c. If exp_diff is positive, meaning the element's actual exponent is greater than the shared exponent, the mantissa is left-shifted by exp_diff bits (which may overflow the BFP mantissa, requiring saturation). If it is negative, the mantissa is right-shifted by the absolute value of exp_diff. If it is zero, the mantissa is not shifted. This process effectively unifies the scale of all data elements to a common base defined by the shared exponent.

[0107] 3) Quantization and rounding: After the alignment and shift operations, the mantissa obtained may still have more bits than the mantissa specified by the target BFP format (for example, in BFP8, it may be a 7-bit mantissa, excluding the sign bit). Therefore, the aligned mantissa needs to be quantized:

[0108] a. Truncation: The simplest way is to directly discard the low-order part that exceeds the target mantissa.

[0109] b. Rounding: A more precise method, such as using the rounding mode recommended by IEEE 754 such as "round-to-nearest-even", can reduce the cumulative deviation.

[0110] 4) Encoding: The sign bit of the original floating-point number is combined with the BFP mantissa bits obtained after quantization and rounding to form the "mantissa" part of the target BFP format. This "mantissa" part actually contains the sign information.

[0111] 5) Special value mantissa processing: If the special value processing mode requires a specific encoding of the mantissa part of NaN or Infinity (for example, to retain part of the mantissa part of the original NaN to distinguish different types of NaN, or to quickly identify them in hardware), the above-mentioned normal mantissa processing is not performed, and the mantissa encoding is generated according to the special mode.

[0112] 5. Data encapsulation:

[0113] 1) Structural Design: The shared exponent encoded in step 3, the optional special value indicator generated by the special value processing module (these indicators may be one or more independent flags or implicitly expressed through a specific encoding pattern of the shared exponent or mantissa itself), and the N (signed) mantissas generated for the N elements in step 4 are arranged and combined according to a predefined, hardware-friendly BFP data unit format.

[0114] 2) Order and format: A typical packing format might be: the shared exponent appears first, followed by the special value flag field (if any), and then the N mantissas are packed tightly together. The order in which the mantissas are stored within the pack can be their logical order in the original block, or some reordering for hardware access convenience.

[0115] 3) Alignment and padding: The total length of the entire BFP data unit may need to be aligned to a boundary that is easy for the hardware to handle (e.g., 32 bits, 64 bits, or cache line size), which may involve adding a small amount of padding bits at the end of the data unit.

[0116] 4) Output: The encapsulated BFP data unit is output as a compact bit stream or byte sequence to the target storage location (such as L1 / L2 cache, DRAM) or directly sent to the subsequent computing unit that supports the BFP format.

[0117] 6. Task processing:

[0118] The target task is processed based on the block floating point representation of the current data block.

[0119] In some specific embodiments, for example, embodiment 1: BFP encoding with adaptive block size and index strategy adjustment.

[0120] Assume that the input data is a tensor in FP32 format.

[0121] Step 1: Data reception and pre-analysis. Receive an FP32 tensor. The pre-analysis module calculates statistical metrics such as the dynamic range, mean, variance, and zero-value ratio of the tensor (or its sub-blocks).

[0122] Step 2: Adaptive block parameter configuration. Based on pre-analysis results:

[0123] 1) If the overall dynamic range of the tensor is small and the proportion of zero values ​​is low, the parameter configuration module can choose a larger block size N (such as 32 or 64) and adopt the traditional "maximum absolute value within the block" as the index determination strategy.

[0124] 2) If the tensor contains many wildly changing values ​​or important low-amplitude signals, a smaller block size N (such as 8 or 16) can be selected, and the "99th percentile of the absolute value within the block" can be used as the index determination strategy to avoid a few extreme large values ​​"polluting" the shared index, thereby protecting the accuracy of most data.

[0125] 3) If the exponents between consecutive data blocks change smoothly, the "exponent differential encoding" strategy can be enabled, that is, the first block stores the complete exponent, and subsequent blocks store the difference from the previous block's exponent.

[0126] Step 3: Sharing Index Calculation and Coding

[0127] For example, with a block size of N = 16 and an exponent determination strategy of "99th percentile absolute value within the block," the exponent calculation module finds the FP32 number corresponding to the 99th percentile of the absolute value of the current 16 FP32 values ​​and extracts its exponent as the shared exponent exp_shared. This exp_shared is encoded and stored.

[0128] Step 4: Mantissa calculation and encoding

[0129] For each FP32 value val_fp32 within the block:

[0130] 1) Convert it to a BFP mantissa. This process may involve shifting the mantissa of val_fp32 by the difference between exp_shared and the exponent of val_fp32 itself, and then truncating or rounding it to the number of mantissa bits allowed by the BFP format (for example, 7 bits for BFP8).

[0131] 2) If val_fp32 is zero or NaN / Inf, the special value handling module kicks in. If the special value flag is configured, only a special flag may be stored, and the mantissa may be a predefined value or not stored.

[0132] Step 5: Data Encapsulation

[0133] Combine the encoded exp_shared, possible special value flags, and 16 encoded mantissa arrays into a BFP data unit, with the exponent first and the mantissa data next.

[0134] Embodiment 2: Optimization processing for all-zero blocks.

[0135] In neural networks, activation values ​​may contain a large number of zeros.

[0136] Steps 1 and 2: The pre-analysis module detects that all FP32 values ​​in a data block (e.g., N=16) are zero. The parameter configuration module enables the special value processing mode of "all-zero block optimization".

[0137] Step 3: The exponent calculation module can set the shared exponent to a predefined zero-marked exponent, or still calculate an exponent depending on the context (for example, if the hardware expects a valid exponent).

[0138] Step 4: The mantissa processing module does not actually process or store the mantissa, or encodes all mantissas as zero, according to the "all-zero block optimization" mode.

[0139] Step 5: The encapsulation module sets an "all-zero block" flag in the BFP data unit. During decoding, the decoder reads this flag and directly outputs N zeros without further processing the exponent and mantissa, thus saving storage and decoding time.

[0140] Example 3: BFP decoding process.

[0141] A BFP data unit encoded according to the method of the present invention is received.

[0142] Step 1: Data unpacking: Separate the shared exponent exp_shared_coded, special value flag (if any) and each coded mantissa mant_coded_i.

[0143] Step 2: Shared exponent decoding: Decode exp_shared_coded to get exp_shared. If exponential differential coding is used, the exponent of the previous block needs to be combined to reconstruct the exponent of the current block.

[0144] Step 3: Special value check: Check the special value flag.

[0145] 1) If it is an "all zero block" mark, directly output N zero values.

[0146] 2) If it is marked as "Contains NaN / Inf", the NaN / Inf value is reconstructed according to predefined rules or passed to subsequent processing units.

[0147] Step 4: Mantissa decoding and FP value reconstruction: For non-special value marked blocks or non-special value elements within a block, exp_shared and mant_coded_i are used to reconstruct the FP32 (or other target floating point format) value. This process combines and adjusts the shared exponent and mantissa (including the sign bit) to form the inverse process of the floating point number. Specifically, it involves converting exp_shared back to a standard floating point exponent according to the BFP convention (such as the bias associated with is_exp_a), and shifting mant_coded_i left to fill the necessary number of bits to combine the sign bit, exponent, and mantissa.

[0148] The adaptive block floating-point data-based task processing method provided by the disclosed embodiments, when processing corresponding target tasks based on floating-point data, dynamically configures the block size and exponent determination strategy to better adapt to the statistical characteristics of different data, achieving a better balance between accuracy and compression rate. This allows for improved compression efficiency and processing performance while ensuring data processing accuracy during task processing. This method offers greater flexibility and configurability, ensuring a balance between task processing accuracy and efficiency. For example, for data with a small dynamic range, larger blocks can be used to improve compression rate; for data with a large dynamic range, smaller blocks or a more robust exponent determination strategy can be used to ensure accuracy. Furthermore, an optimized exponent determination strategy (e.g., based on statistics rather than a simple maximum) can reduce the impact of extreme values ​​on the shared exponent, thereby protecting the accuracy of other values ​​within the block and alleviating the inherent problem of BFP values ​​being flushed to zero. Finally, by specifically processing and encoding special values ​​(e.g., all-zero blocks), storage usage and unnecessary computation can be further reduced, thereby achieving a balance between task processing accuracy and efficiency. For example, if a block is marked as all zeros, its mantissa does not need to be stored and processed.

[0149] and Figure 1 Corresponding to the task processing method based on adaptive block floating point data, the present disclosure also provides a task processing device based on adaptive block floating point data, such as Figure 3 As shown, the device may specifically include:

[0150] A receiving module 301 is configured to receive a set of floating-point data corresponding to a target task as a current data block;

[0151] A parameter configuration module 302 is configured to determine block floating-point parameters of the current data block according to a preset rule or at least one data characteristic of the current data block, wherein the block floating-point parameters include at least a block size, a sharing index determination strategy, and a special value processing mode;

[0152] An index calculation module 303 is configured to calculate a sharing index for the current data block according to the sharing index determination strategy;

[0153] A mantissa processing module 304 is configured to process each floating-point data in the current data block based on the shared index to obtain a corresponding mantissa.

[0154] An encapsulation module 305 is configured to combine the shared index and each tail array into a block floating point representation of the current data block;

[0155] The processing module 306 is configured to process the target task based on the block floating-point representation of the current data block.

[0156] In some embodiments, the block size is a variable integer selected from a predefined set, the set including at least two different block size values.

[0157] In some embodiments, the sharing index determination strategy includes:

[0158] Using the exponent of the floating-point data with the largest absolute value in the current data block as the shared index;

[0159] and / or, using an index corresponding to a specific statistic of absolute values ​​of floating-point data in the current data block as a sharing index;

[0160] And / or, the sharing index of the current data block is determined by differential coding or smoothing filtering in combination with the sharing index of one or more previous historical data blocks.

[0161] In some embodiments, the task processing device based on adaptive block floating-point data may further detect whether there is a special value in the current data block;

[0162] If a special value is detected, a special value mark is added to the block floating point representation of the current data block.

[0163] In some embodiments, if the current data block is detected as an all-zero block, the mantissa portion in the corresponding block floating-point representation is padded with a predefined all-zero pattern or omitted.

[0164] Specifically, the receiving module in the task processing device based on adaptive block floating-point data can actually be a data input interface:

[0165] 1) Function: Responsible for receiving floating-point data to be converted from external systems (such as the load / store unit of the processor core, DMA controller, endpoint of the on-chip network NoC, or the previous-level pipeline computing unit).

[0166] 2) Features: Supports one or more standard bus protocols (such as AXI, AHB) or proprietary internal interconnect protocols. Able to handle data transmission of different bit widths (such as 32-bit, 64-bit, etc.).

[0167] In addition, the task processing device based on adaptive block floating-point data may further include a pre-analysis module:

[0168] 1) Function: If enabled, this module performs real-time or near-real-time statistical analysis on data blocks flowing in through the data input interface. It uses built-in computational logic (e.g., comparators for finding maximum / minimum values, counters for counting zero / special values, and accumulators for calculating sums or square sums to aid in variance estimation) to extract the aforementioned data characteristics.

[0169] 2) Output: The quantitative statistical indicators obtained from the analysis (e.g., maximum absolute value, zero value count, dynamic range indicator, etc.) are passed to the parameter configuration module.

[0170] 3) Implementation: This can be a dedicated small hardware data path or an embedded microcontroller executing a predefined analysis program. Its design should focus on low latency and low power consumption.

[0171] The above parameters configure the module:

[0172] 1) Input: Receives statistical metrics from the pre-analysis module and / or direct parameter setting requests from external control sources (such as instructions written by the CPU through configuration registers, metadata embedded in the instruction stream by the compiler, or control signals dynamically issued by the runtime system).

[0173] 2) Function: As the core of the adaptive logic, this module dynamically determines the optimal or most appropriate BFP parameter combination for the current data block to be processed based on the received information and internally stored configuration rules (which may exist in the form of a lookup table (LUT), a small rule engine, or a lightweight decision model such as a pre-trained neural network): block size N, shared index determination strategy code, and special value processing mode code.

[0174] 3) Output: Broadcast or unicast the selected / calculated set of BFP parameters (e.g., the value of N, an enumeration value representing the exponent strategy, a bit mask representing the special value mode, etc.) to the subsequent related processing modules (exponent calculation module, mantissa processing module, special value processing module).

[0175] 4) State management: This may include a set of configuration registers or a small block of SRAM to store the currently active BFP parameter set and possibly multiple predefined parameter profiles for fast switching.

[0176] The above index calculation module:

[0177] 1) Input: Receive the current data block (consisting of N original floating-point data elements, where N is specified by the parameter configuration module) and the shared index determination strategy code specified by the parameter configuration module.

[0178] 2) Function: Contains hardware logic to implement multiple index calculation algorithms. Based on the received strategy code, it activates the corresponding calculation path:

[0179] a. If it is the Max-Abs strategy: start the maximum value search unit and perform index extraction, bias adjustment and saturation logic on the result.

[0180] b. If it is a statistically robust strategy: it may include logarithmic operations, sorting (partial), or multiply-accumulate units to calculate the required statistics and derive the exponent from them.

[0181] c. If it is a history / differential strategy: it is necessary to access a temporary register or a small FIFO that stores the shared index of the previous (or several) blocks and perform a subtraction or prediction operation.

[0182] 3) Output: The calculated shared index. This index may be an encoded bit string that can be directly used for encapsulation, or an intermediate value that has not yet been encoded and is ultimately encoded by the encapsulation module or its subunits.

[0183] The above-mentioned mantissa processing module:

[0184] 1) Input: Receives the N raw floating-point data elements of the current data block, the (adjusted) shared exponent value provided by the exponent calculation module, and the block size N specified by the parameter configuration module. The configuration of the special value handling mode also indirectly affects this module (for example, the mantissa of an all-zero block may be directly set to zero without complex processing).

[0185] 2) Function: Perform parallel or pipelined mantissa conversion operations on each floating-point data element in the block (which requires regular BFP conversion). The core logic includes:

[0186] a. Sign bit transfer.

[0187] b. Perform an exact logical left / right shift on the element's mantissa (including the implicit bit) based on the difference between the shared exponent and the element's own exponent.

[0188] c. Perform quantization on the shifted mantissa (truncate or round to the target BFP mantissa number of bits).

[0189] 3) Output: N encoded BFP mantissas (including the sign bit at this time).

[0190] The task processing device based on adaptive block floating-point data may further include a special value processing module:

[0191] 1) Input: Receives the N original floating-point data elements of the current data block and the special value processing mode code specified by the parameter configuration module.

[0192] 2) Function:

[0193] a. Test each element in the block in parallel whether it is zero, NaN or Infinity.

[0194] b. Depending on the selected special value processing mode:

[0195] i. If configured for all-zero block optimization, when an all-zero block is detected, an "all-zero block" signal / flag is generated and the normal operation of the mantissa processing module on this block may be suppressed.

[0196] ii. If configured to handle NaN / Infinity, generate specific markers for these values ​​or instruct the mantissa processing module to generate special encoding patterns for their mantissa parts.

[0197] 3) Output: A set of special value flag signals (e.g., a bit indicating an all-zero block, a bit indicating a block containing NaN / Inf, or a more fine-grained special state for each element) that are sent to the wrapper module. This module is typically tightly integrated with the mantissa processing module, or its output is used as one of the control inputs of the mantissa processing module.

[0198] The above encapsulated modules:

[0199] 1) Input: Receives the encoded shared exponent from the exponent calculation module, N encoded mantissas from the mantissa processing module, and (if enabled) various special value indication flags from the special value processing module.

[0200] 2) Function: Strictly following the predefined BFP data unit target format, all the above input components (shared exponent, special flags, N mantissas) are assembled and concatenated in an orderly manner at the bit level to form the final, compact BFP representation data packet. This process may also include necessary bit alignment operations or the addition of padding bits to meet the alignment requirements of the output interface or storage unit.

[0201] 3) Output: One or more encapsulated BFP data units.

[0202] The task processing device based on adaptive block floating-point data may further include a data output interface:

[0203] 1) Function: Responsible for sending the BFP data units generated by the encapsulation module to the external system (such as the storage controller that writes back to the main memory, the injection port of the on-chip network NoC, or directly to the next-level hardware processing unit configured to receive this BFP format).

[0204] 2) Features: Similar to the data input interface, it supports the corresponding bus protocol or internal interconnection standard.

[0205] The processing module is used to process the target task based on the BFP representation data packet.

[0206] Combine Figure 4 As shown, an embodiment of the present disclosure further provides a task processing device 400 based on adaptive block floating-point data, comprising a processor 404 and a memory 401. Optionally, the system may further comprise a communication interface 402 and a bus 403. The processor 404, the communication interface 402, and the memory 401 may communicate with each other via the bus 403. The communication interface 402 may be used for information transmission. The processor 404 may invoke logic instructions in the memory 401 to execute the task processing method based on adaptive block floating-point data of the above embodiment.

[0207] In addition, the logic instructions in the memory 401 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0208] Memory 401, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. Processor 404 executes the program instructions / modules stored in memory 401 to perform functional applications and data processing, thereby implementing the task processing method based on adaptive block floating-point data in the above-mentioned embodiments.

[0209] The memory 401 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Furthermore, the memory 401 may include high-speed random access memory and non-volatile memory.

[0210] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute a task processing method based on adaptive block floating-point data.

[0211] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0212] The technical solution of the embodiments of the present disclosure may be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code, or a transient storage medium.

[0213] The above description and accompanying drawings sufficiently illustrate the embodiments of the present disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments represent only possible variations. Unless expressly required, individual components and functions are optional, and the order of operations may vary. Portions and features of some embodiments may be included in or replaced with portions and features of other embodiments. As used in the description of the embodiments, unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" are intended to include the plural forms as well. Similarly, the term "and / or" as used in this application means including any and all possible combinations of one or more of the associated listed items. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof. In the absence of further limitations, the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, or device comprising the elements. In this document, each embodiment may focus on the differences from other embodiments, and similar parts between the embodiments can be referenced to each other. For methods, products, etc. disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, then the relevant parts can be referenced to the description of the method section.

[0214] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. Technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. Technicians can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0215] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, the functional units in the embodiments of the present disclosure may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0216] The flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.

[0217] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0218] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0219] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0220] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0221] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0222] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0223] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of this disclosure can be achieved, and this document is not limited here.

[0224] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A task processing method based on adaptive block floating point data, characterized in that: The method comprises: Receive a set of floating-point data corresponding to the target task as the current data block; Determining block floating-point parameters of the current data block according to a preset rule or at least one data characteristic of the current data block, the block floating-point parameters including at least a block size, a sharing index determination strategy, and a special value processing mode; Calculating a sharing index for the current data block according to the sharing index determination strategy; Based on the shared index, each floating-point data in the current data block is processed to obtain a corresponding mantissa; Combining the shared index and each tail number into a block floating point representation of the current data block; The target task is processed based on the block floating point representation of the current data block.

2. The method according to claim 1, characterized in that The block size is a variable integer selected from a predefined set, the set including at least two different block size values.

3. The method according to claim 1, characterized in that The sharing index determination strategy includes: Using the exponent of the floating-point data with the largest absolute value in the current data block as the shared index; and / or, using an index corresponding to a specific statistic of absolute values ​​of floating-point data in the current data block as a sharing index; And / or, the sharing index of the current data block is determined by differential coding or smoothing filtering in combination with the sharing index of one or more previous historical data blocks.

4. The method according to claim 1, wherein The method further comprises: Detecting whether there is a special value in the current data block; If a special value is detected, a special value mark is added to the block floating point representation of the current data block.

5. The method according to claim 4, characterized in that If the current data block is detected as an all-zero block, the mantissa part in the floating-point representation of the corresponding block is filled with a predefined all-zero pattern or omitted.

6. A task processing device based on adaptive block floating point data, characterized in that: The method comprises: A receiving module is used to receive a set of floating-point data corresponding to the target task as the current data block; a parameter configuration module, configured to determine block floating-point parameters of the current data block according to a preset rule or at least one data characteristic of the current data block, the block floating-point parameters including at least a block size, a sharing index determination strategy, and a special value processing mode; An index calculation module, configured to determine a sharing index for the current data block according to the sharing index determination strategy; a mantissa processing module, configured to process each floating-point data in the current data block based on the shared exponent to obtain a mantissa corresponding to each of the floating-point data; An encapsulation module, configured to combine the shared index and each tail array into a block floating point representation of the current data block; A processing module is configured to process the target task based on the block floating-point representation of the current data block.

7. The device according to claim 6, characterized in that The block size is a variable integer selected from a predefined set, the set including at least two different block size values.

8. The device according to claim 6, characterized in that The sharing index determination strategy includes: Using the exponent of the floating-point data with the largest absolute value in the current data block as the shared index; and / or, using an index corresponding to a specific statistic of absolute values ​​of floating-point data in the current data block as a sharing index; And / or, the sharing index of the current data block is determined by differential coding or smoothing filtering in combination with the sharing index of one or more previous historical data blocks.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.

Citation Information

Cited By

  • Large model optimization method and device based on data quantization

    CN121235129A