System and method for micromachine learning using block floating points
By optimizing the computing process of IoT devices through block floating-point algorithms and a micro machine learning platform, the problems of high computing and memory requirements of floating-point operations are solved, and efficient AI reasoning and low-power AI applications are achieved.
Patent Information
- Application Number
- CN202580000628.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-12
AI Technical Summary
Existing floating-point operations have high computational and memory requirements on IoT devices, resulting in low efficiency, and traditional block floating-point algorithms have challenges in hardware efficiency and compatibility with deep learning architectures.
Using block floating-point algorithm, through floating-point to block floating-point converter, 8-bit integer multiplier, adder and accumulator, the calculation process of deep learning model is optimized, including converting floating-point data into block floating-point format, and using micro machine learning platform for calculation acceleration.
It reduces computational complexity and memory bandwidth requirements, improves computational efficiency, is suitable for low-power AI application scenarios, reduces dependence on cloud services, and enhances privacy protection.
Smart Images

Figure CN120641912A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a system and method for efficiently executing deep learning models on resource-constrained hardware. Background Art
[0002] In the current development of smart cities and smart homes, Internet of Things (IoT) devices are key to enabling real-time artificial intelligence (AI) and driving automation. However, most AI applications require extensive hardware resources, including high computing power, high data communication bandwidth, and large memory. Due to the inherent resource constraints of IoT devices, many AI applications rely on cloud services to handle compute-intensive tasks. This reliance on cloud services has serious drawbacks, such as increased latency, which hinders real-time applications, and the potential privacy concerns associated with transmitting sensitive data over the network.
[0003] To address these limitations, efficient AI acceleration technology is crucial for enabling on-device inference in IoT systems. The challenge lies in reducing the computational and memory costs of inferencing machine learning or deep learning models, a challenge particularly acute in resource-constrained environments. For example, traditional floating-point operations have high computational and memory requirements, making them inefficient in IoT applications.
[0004] The Block Floating Point (BFP) algorithm offers a solution by reducing computational complexity and memory bandwidth requirements. However, existing implementations can still face challenges related to hardware efficiency and compatibility with modern deep learning architectures. Therefore, there is a need for an optimized system and method that can effectively integrate the BFP algorithm into machine learning frameworks. Summary of the Invention
[0005] According to a first aspect of the present invention, a system for performing deep learning model calculations is provided, characterized in that it includes: a first floating-point to block floating-point (FP to BFP) converter, a second FP to BFP converter, an 8-bit integer multiplier, an adder, an accumulator, and a block floating-point to floating-point (BFP-to-FP) converter. The first FP to BFP converter is used to receive pixel data in a 32-bit floating-point format and convert the pixel data by reducing the mantissa of the pixel data to 8 bits, thereby obtaining a block floating-point (BFP) format with an 8-bit mantissa. The second FP to BFP converter is used to receive filter data in a 32-bit floating-point format and convert the filter data by reducing the mantissa of the filter data to 8 bits, thereby obtaining a BFP format with an 8-bit mantissa. The 8-bit integer multiplier is used to receive BFP-represented data from the first FP to BFP converter and the second FP to BFP converter, and perform a multiplication-accumulation operation to generate 16-bit product data. The adder is configured to receive the 16-bit product data from the 8-bit integer multiplier and accumulate the data of multiple 16-bit products to generate 64-bit cumulative sum data. The accumulator is configured to receive the 64-bit cumulative sum data from the adder and perform cumulative sum aggregation. The block floating-point to floating-point converter is configured to receive the 64-bit cumulative sum data from the accumulator and convert the 64-bit cumulative sum data into output data represented by 32-bit FP.
[0006] According to a second aspect of the present invention, a method is provided for converting floating-point (FP) data into a block floating-point (BFP) format, the method comprising the following steps: receiving data in a 32-bit floating-point format through an FP to BFP converter; determining a shared exponent of a block of floating-point values through the FP to BFP converter, the shared exponent being the maximum exponent among all the floating-point values in the block; calculating a right shift amount for each floating-point value based on a difference between an original exponent of each floating-point value and the determined shared exponent through the FP to BFP converter; right-shifting the mantissa of each floating-point value through the FP to BFP converter; and truncating the 24-bit mantissa of each floating-point value to 8 bits through the FP to BFP converter to obtain a BFP format.
[0007] According to a third aspect of the present invention, a tiny machine learning (tiny ML) platform is provided, characterized in that it includes: a tiny machine learning operation accelerator, a single instruction multiple data multiply accumulate (SIMDMAC) accelerator, a microcontroller unit (MCU) core, a memory, and a multiplexer (MUX). The tiny machine learning operation accelerator is used to perform convolutional neural network (CNN) operations using approximate computing technology, and includes: a first FP to BFP converter, a second FP to BFP converter, and an 8-bit integer multiplier. The first FP to BFP converter is used to receive pixel data in a 32-bit floating point format and convert the pixel data by reducing the mantissa of the pixel data to 8 bits, thereby obtaining a block floating point (BFP) format with an 8-bit mantissa. The second FP to BFP converter is used to receive filter data in a 32-bit floating point format and convert the filter data by reducing the mantissa of the filter data to 8 bits, thereby obtaining a BFP format with an 8-bit mantissa. An 8-bit integer multiplier is used to receive BFP-represented data from the first FP-to-BFP converter and the second FP-to-BFP converter and perform a multiplication-accumulation operation to generate 16-bit product data. A single-instruction multiple-data multiplication-accumulation accelerator is coupled to the micro machine learning operation accelerator and is used to perform signal feature extraction and enhance vectorized execution of CNN calculations. A microcontroller unit core is used to manage execution control, memory access, and coordination between the micro machine learning operation accelerator and the single-instruction multiple-data multiplication-accumulation accelerator. A memory is coupled to the microcontroller unit core, the micro machine learning operation accelerator, and the single-instruction multiple-data multiplication-accumulation accelerator. A multiplexer is used to dynamically route data between the microcontroller unit core, the memory, the micro machine learning operation accelerator, and the single-instruction multiple-data multiplication-accumulation accelerator. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings, in which:
[0009] Figure 1A An architectural diagram showing a system for executing deep learning model calculations according to an embodiment of the present invention;
[0010] Figure 1B A diagram showing how to convert one-dimensional input data into a two-dimensional numerical matrix;
[0011] Figure 2A and Figure 2B A schematic diagram illustrating a conversion process of reducing a mantissa to 8 bits using an FP to BFP converter according to an embodiment of the present invention is shown;
[0012] Figure 3A diagram showing the processing flow of the main arithmetic operations performed in the convolutional layer when converting the basic data type FP32 to BFP; and
[0013] Figure 4 A schematic diagram of the architecture of a micro machine learning platform according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0014] In the following description, preferred examples will be used to illustrate systems and methods for performing micro-machine learning using block floating point. Those skilled in the art will appreciate that various modifications, including additions and / or substitutions, may be made without departing from the scope and spirit of the present invention. Specific details may be omitted to avoid obscuring the present invention; however, this disclosure is intended to enable those skilled in the art to practice the teachings herein without undue experimentation.
[0015] The following description will be used as a reference Figure 1A . Figure 1A The illustrated architecture can be used to optimize the execution of deep learning / machine learning models (i.e., convolutional neural networks (CNNs)) that involve a large number of floating-point multiplication operations. To improve computational efficiency, system 100 employs a block floating point (BFP) algorithm that reduces the complexity of multiplication by having multiple data points within a block share a common exponent. The architecture employed reduces computational cost and memory bandwidth requirements, making it particularly suitable for low-power AI applications such as Tiny Machine Learning (Tiny ML) and edge-based neural network inference.
[0016] The system 100 includes a plurality of processing units 110, an accumulator (ACC) 102, and a block floating-point to floating-point converter (BFP to FP converter; BFP-to-FP (floating-point) converter) 104. The plurality of processing units 110 are configured to execute in parallel, wherein each processing unit 110 is responsible for processing a different data block (i.e., data 1 to data N). All processing units 110 are coupled to the ACC 102, and the ACC 102 is configured to perform at least one computational task, such as a BFP multiply-accumulate (MAC) operation, data accumulation, and feature extraction, thereby optimizing the execution of a deep learning model. The ACC 102 is coupled to the BFP to FP converter 104, which is configured to convert 64-bit BFP data into a 32-bit floating-point (FP32) format to make it compatible with subsequent processing stages.
[0017] The processing unit 110 for executing “Data 1 ” includes a first floating point to block floating point converter (FP to BFP converter) 112 , a second FP to BFP converter 114 , an 8-bit integer multiplier 116 , and an adder 118 .
[0018] The first FP to BFP converter 112 is configured to receive pixel data in a 32-bit floating point format and convert the pixel data by reducing the mantissa of the pixel data to 8 bits (i.e., reducing it to an 8-bit representation), thereby obtaining a BFP format with an 8-bit mantissa. The second FP to BFP converter 114 is configured to receive filter data in a 32-bit floating point format and convert the filter data by reducing the mantissa of the filter data to 8 bits (i.e., reducing it to an 8-bit representation), thereby obtaining a BFP format with an 8-bit mantissa.
[0019] In one embodiment, pixel data in 32-bit floating point format may represent the input pixel data of the CNN, and each pixel may be represented in a 32-bit floating point format. Specifically, the input pixel data may be converted into a two-dimensional numerical matrix and used as the input source of the CNN. For example, Figure 1B A schematic diagram showing how one-dimensional input data is converted into a two-dimensional numerical matrix. The input data may be generated by an audio signal, which is initially represented as a one-dimensional sequence. During the conversion process, the input one-dimensional sequence is reorganized into a two-dimensional numerical matrix, where each element in the matrix corresponds to a pixel data point derived from the original sequence. In one embodiment, the filter data in 32-bit floating point format corresponds to the filter weights used in the convolutional layer of the CNN. For example, the processing unit 110 can use a filter (or kernel) implemented in matrix form to perform a convolution operation on the input pixel data to extract meaningful features, such as edges, textures, or patterns.
[0020] Figure 2A and Figure 2B A schematic diagram of a conversion process of reducing a mantissa to 8 bits using an FP to BFP converter according to an embodiment of the present invention is shown.
[0021] exist Figure 2AIn the FP to BFP conversion, the first step involves determining the shared exponent for the floating-point value block. Each input value in the 32-bit floating-point format consists of a sign bit, an 8-bit exponent, and a 23-bit mantissa. Since the BFP format requires multiple values to share a single exponent, the FP to BFP converter first scans all input values within the block and determines the largest exponent. Once the largest exponent is determined, it is designated as the shared exponent for the entire block (common exponent). In the first step of the FP to BFP conversion, all values within the block are allowed to align with each other under a common exponent, reducing storage requirements and simplifying calculations. In this example, the shared maximum exponent selected from 0x79, 0x7B, 0x7A, and 0x7C is 0x7C.
[0022] Then, in Figure 2B In the second step of the FP to BFP conversion, after determining the shared exponent for the block, the mantissa of each floating-point value is adjusted to align with the selected shared exponent. Since the original 32-bit floating-point values have independent exponents, they are all normalized to the shared exponent when converted to BFP format.
[0023] The FP-to-BFP converter can align the mantissa by right-shifting it. If the original exponent of the input value is smaller than the shared exponent, its mantissa must be right-shifted by the difference between the two exponents. This shift operation allows values to be properly scaled within the block while maintaining numerical integrity. For example, if the shared exponent is 0x7C, but the input value originally had an exponent of 0x79, its mantissa must be right-shifted by (0x7C - 0x79) = 3 bits to align with the shared exponent.
[0024] Additionally, the FP to BFP converter incorporates a hidden leading bit (bit 24) into the 23-bit mantissa. In the FP32 configuration, the mantissa consists of 23 explicit bits, but there is an implicit leading bit (bit 24) that is always assumed to be "1" for normalized numbers. The hidden leading 1 bit is explicitly added to the mantissa before shifting and truncating, effectively making it a 24-bit value instead of just a 23-bit value.
[0025] Next, the FP-to-BFP converter performs mantissa truncation and precision adjustment. Because the BFP format limits the mantissa to 8 bits, the adjusted mantissa must be truncated from its expanded 24-bit representation. This process involves selecting the highest-order 8 bits of the shifted mantissa while discarding the lower-order bits. In one embodiment, the discarded bits can be rounded to reduce quantization error, thereby minimizing loss of numerical precision.
[0026] In this way, the FP to BFP converter constructs the final BFP representation. In one embodiment, the generated BFP format consists of the following parts: a 1-bit sign (the same as the original floating-point value); an 8-bit shared exponent (determined in the first step); and an 8-bit truncated mantissa (obtained by right-shifting and truncating the original mantissa). Therefore, each converted value is stored in a compact 17-bit format (1-bit sign / 8-bit exponent / 8-bit mantissa), which reduces storage and computation requirements compared to a 32-bit floating-point representation. The generated BFP format can support the use of 8-bit integer multipliers for calculations, which helps in efficient processing and reduces hardware complexity.
[0027] Reference again Figure 1A . The 8-bit integer multiplier 116 is configured to receive the BFP representation with an 8-bit mantissa from the first FP to BFP converter 112 and the second FP to BFP converter 114 and perform a multiply-accumulate (MAC) operation. In an embodiment involving CNN calculations, basic operations are implemented through matrix multiplication, and element-by-element multiplication is performed using pixel data and filter weights. The 8-bit integer multiplier 116 simplifies this operation by multiplying the BFP mantissa values of the pixel data and the filter data, and produces an intermediate result, which is then accumulated to generate the convolution output. By utilizing 8-bit integer operations, the system 100 can reduce computational complexity and power consumption compared to 32-bit floating point multipliers.
[0028] The 8-bit integer multiplier 116 can generate a 16-bit output data representing the product of two 8-bit BFP mantissa values. Specifically, the 8-bit integer multiplier 116 performs a multiplication on two 8-bit BFP mantissa values, one of which comes from the first FP-to-BFP converter 112 (pixel data) and the other comes from the second FP-to-BFP converter 114 (filter data). Since each operand is 8 bits, the product result is 16 bits.
[0029] Adder 118 is configured to receive the 16-bit output of 8-bit integer multiplier 116 and perform an accumulation operation. Specifically, adder 118 can be used to sum the multiple 16-bit products generated by successive multiplications of BFP mantissa values during the MAC process. Adder 118 can thus output a 64-bit cumulative sum representing the intermediate convolution result before further processing, such as exponent adjustment and activation function.
[0030] The processing unit 110 described above is a module for performing operations on "data 1"; and the architecture and processing flow of this module are also applicable to instance scenarios of other processing units 110 that process different data blocks (such as data N).
[0031] All processing units 110 have their own adders 118, which are connected to the ACC 102. Each adder 118 can accumulate 16-bit multiplication results and output 64-bit accumulated sum data to the ACC 102 for further processing. The ACC 102 is configured to receive and aggregate these accumulated sum data from multiple processing units 110, thereby achieving efficient parallel computation. The ACC 102 is also configured to provide a 64-bit feedback signal to the adder 118. The feedback provided by the ACC 102 allows the adder 118 to continue accumulating over multiple cycles, thereby retaining partial sums from previous operations and incorporating them into subsequent computations. By utilizing the feedback loop from the ACC 102, the system 100 can support iterative accumulation, thereby being able to handle large-scale matrix multiplications in CNN operations while maintaining numerical accuracy.
[0032] The BFP-to-FP converter 104 is configured to receive the 64-bit accumulated sum data from the ACC 102 and convert it to a 32-bit floating-point representation. The BFP-to-FP converter 104 conversion process includes extracting the shared exponent from the BFP format, adjusting the accumulated mantissa accordingly, and reconstructing the final FP32 output data. By performing this conversion / transformation, the BFP-to-FP converter 104 is compatible with subsequent processing stages that operate in standard floating-point precision.
[0033] In one embodiment, the system 100 can be applied to software on mobile platforms with model deployment on mobile devices, microcontrollers, and other edge devices. Figure 3 A diagram showing the software processing flow of the main arithmetic operations performed in the convolutional layer when converting the basic data type FP32 to BFP. The flow is divided into three stages: data preparation (stage A), BFP conversion and operations (stage B), and output (stage C).
[0034] In phase A, the system 100 processes input data blocks (i.e., pixel data) and filter data blocks (i.e., weight data) in FP32 format. A first FP to BFP converter 112 is configured to receive pixel data, while a second FP to BFP converter 114 receives filter data. Each FP to BFP converter identifies the maximum exponent in a block of values. For input data, the obtained maximum exponent is called "max_input_exp", and for filter data, the obtained maximum exponent is called "max_filter_exp". In phase A, all data points in the block are allowed to share a shared exponent. This step is necessary for BFP conversion and can therefore reduce memory bandwidth and computational complexity.
[0035] In stage B, the system 100 converts the FP32 input data and filter data into BFP format using the shared exponent determined in stage A. The first FP to BFP converter 112 can convert the pixel data into a BFP representation, and the second FP to BFP converter 114 performs the same operation on the filter data. Once the data is converted into BFP format, the system 100 performs a BFP multiply-add-accumulate operation, and an 8-bit integer multiplier 116 can perform element-by-element multiplication of the 8-bit BFP mantissa from the pixel data and the filter data to produce a 16-bit product result. The product result is then passed to the adder 118, which can perform iterative accumulation and generate a 64-bit cumulative sum. The accumulated result is called "bfp_total", which will be further processed in stage C.
[0036] In stage C, the system 100 completes the computation by processing "bfp_total" and converting the accumulated result back to floating point format. The adder 118 may continue to perform BFP multiply-add-accumulate and transmit the 64-bit accumulated sum (i.e., "bfp_total") from the multiple processing units 110 to the ACC 102, which is configured to aggregate and manage accumulated sums from parallel computations.
[0037] As described above, the ACC 102 facilitates efficient processing of large-scale CNN matrix multiplications by coordinating accumulations across multiple processing units 110. After accumulation by the ACC 102, "bfp_total" is passed to the BFP to FP converter 104, which converts it to FP32 format. This conversion process includes extracting the shared exponent, adjusting the mantissa, and reconstructing the final FP32 format output to maintain compatibility with subsequent processing stages, such as activation functions (i.e., ReLU) and pooling layer processing in CNN calculations.
[0038] Figure 4The present invention shows an architectural diagram of a Tiny ML platform 200 according to an embodiment of the present invention. The configuration of the system 100 can be applied to the Tiny ML platform 200. The Tiny ML platform 200 can be used to perform lightweight machine learning tasks and optimize neural network inference using hardware accelerators. The Tiny ML platform 200 includes a microcontroller unit (MCU) core 202, a memory 204, a multiplexer (MUX) 206, a Tiny ML operation accelerator 208, and a single instruction multiple data multiply-accumulate (SIMD MAC) accelerator 210. These components can interact with each other through an Advanced Extensible Interface (AXI) bus for data transmission and a Rocket Custom Coprocessor (RoCC) interface for control signals.
[0039] The MCU core 202 serves as the central processing unit, managing execution control, memory access, and coordination between hardware accelerators. It interacts with memory 204, which stores model parameters, intermediate feature maps, and computation results. The MCU core 202 communicates with the micro machine learning computation accelerator 208 and SIMD MAC accelerator 210 using the RoCC interface, sending control instructions to direct the micro machine learning operations.
[0040] Memory 204 serves as a storage unit, providing the model weights, input data, feature maps, and computational results required by micro machine learning during inference. Memory 204 is connected to MCU core 202 via an AXI bus and is also connected to other hardware accelerators (i.e., micro machine learning operation accelerator 208 and SIMD MAC accelerator 210).
[0041] MUX 206 serves as a data routing component, controlling data flow between the MCU core 202, memory 204, and hardware accelerators. Micro machine learning accelerator 208 and SIMD MAC accelerator 210 are designed to provide deep learning model (i.e., convolutional neural network) computation and signal processing, and MUX 206 dynamically routes data to the corresponding processing units to improve parallel execution efficiency.
[0042] In the past Figure 1AIn the system 100 , the first and second FP-to-BFP converters 112 / 114 , integer multiplier 116 , and adder 118 of the processing unit 110 may be applied to the micro machine learning operation accelerator 208 and the SIMD MAC accelerator 210 .
[0043] For example, the micro machine learning operation accelerator 208 is configured to perform deep learning model operations (i.e., CNN operations) using approximate computing techniques. The configuration of the micro machine learning operation accelerator 208 can be the same or similar to the architecture established by the first and second FP-to-BFP converters 112 / 114 of the processing unit 110. The micro machine learning operation accelerator 208 can provide BFP computations for CNN workloads, including matrix multiplication, element-wise operations, and activation functions required for CNN inference. By applying the configuration of the processing unit 110 to the micro machine learning operation accelerator 208, the micro machine learning operation accelerator 208 can convert floating-point inputs to a BFP format with an 8-bit mantissa. The SIMD MAC accelerator 210 is used to enhance vectorized execution through SIMD-based MAC operations. The configuration of the SIMD MAC accelerator 210 can be the same or similar to the 8-bit integer multiplier 116 and adder 118 of the processing unit 110. Therefore, the SIMD MAC accelerator 210 can collaborate with the micro machine learning operation accelerator 208 to facilitate model execution and network computations as described above. The SIMD MAC accelerator 210 can be further configured to perform accumulation operations similar to the ACC 102, summing the partial results generated by the element-wise multiplications performed by the micro machine learning operation accelerator 208. After the accumulation is completed, these results will be stored in the memory 204 for further calculation or final output processing.
[0044] The BFP to FP conversion process can be performed by the MCU core 202, which manages execution control and data processing, or by dedicated logic implemented in the micro machine learning operation accelerator 208. After conversion, the results in FP32 format are stored in the memory 204, and the MCU core 202 can access these results to perform activation functions, pooling operations, or further post-processing.
[0045] The device platform software layer 220 is coupled to the micro machine learning platform 200. In one embodiment, the device platform software layer 220 integrates TensorFlow Lite for MCUs, a mobile machine learning library optimized for microcontrollers and edge devices. The learning library can be modified to utilize hardware accelerators, enabling efficient execution of CNN inference on constrained hardware. The application software layer 222 is coupled to the device platform software layer 220 and includes a machine learning demonstration application that can run on the micro machine learning platform 200 to demonstrate practical use cases for micro machine learning inference.
[0046] The micro-machine learning platform 200 can be combined with the device platform software layer 220 and the application software layer 222 to provide micro-machine learning inference tasks for everyday use, such as voice command processing in IoT devices. In the IoT environment, the micro-machine learning platform 200 can support real-time voice command recognition and response, making it suitable for application scenarios such as robotic cleaners, wearable devices, smart sensors, and electric vehicles.
[0047] For example, when an IoT device receives a voice command from a user (or from an ultrasonic source), it can initiate a task. The audio signal can be captured by the IoT device's microphone and converted into a digital representation. The MCU core 202 first processes the raw voice data and transfers it to the memory 204 for temporary storage before sending it out for feature extraction. The SIMD MAC accelerator 210 can perform signal feature extraction and convert the voice data into a form suitable for the micro machine learning computation accelerator 208 to perform micro machine learning inference.
[0048] After extracting the speech features, they are passed to the micro machine learning accelerator 208, which performs CNN operations using approximate computing techniques. During this process, an FP-to-BFP converter is used to convert the speech features from FP32 format to BFP format, which can be applied to input data and filter weights. The converted values are then subjected to a BFP multiplication-accumulation operation, using an 8-bit integer multiplier and adder to perform matrix multiplication. The accumulator collects and accumulates the calculation results, which are then passed to the BFP-to-FP converter, which converts the final output to FP32 format.
[0049] In one embodiment, the final classification output determines the corresponding IoT device action based on the recognized voice command. For example, if the detected command is "Start cleaning," the robot cleaner will receive a control signal and start the vacuuming process. If the command is "Check your heart rate," the wearable device will retrieve real-time health data and display the heart rate on the screen. If the command is "Turn off the lights," the smart sensor will send a wireless signal to control the smart lighting. If the command is "Start automatic parking," the electric car will connect to the autonomous driving module and execute the parking maneuver.
[0050] Edge-based processing enables speech recognition to be performed locally on IoT devices, eliminating reliance on cloud computing, reducing latency, and enhancing privacy. The micro machine learning platform optimizes energy-efficient inference, making it ideal for low-power IoT scenarios. Furthermore, the platform reduces the memory usage and computational complexity of block data dot product calculations, enabling a low-cost MCU core. The proposed solution simultaneously improves processing speed and energy efficiency.
[0051] In this disclosure, matrix operations referred to or referred to include the dot product of two vectors. For example, if two vectors a = [a1, a2, ..., a n ] and b=[b1,b2,...,b n ], then the dot product can be defined as:
[0052] The functional units and modules of the apparatus and method according to the embodiments of the present invention can be implemented using computing devices, computer processors or electronic circuits, including but not limited to application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, and other programmable logic devices configured or programmed according to the teachings of this disclosure; for example, they can be implemented as FPGA-based micro machine learning platforms, integrated circuit-based micro machine learning platforms, or other forms of micro machine learning platforms. Those skilled in the art of software or electronics can, based on the teachings of this disclosure, smoothly write computer instructions or software codes that are executed in computing devices, computer processors, or programmable logic devices.
[0053] All or part of the method according to an embodiment of the present invention may be executed in one or more computing devices, including server computers, personal computers, laptop computers, and mobile computing devices (such as smartphones and tablet computers).
[0054] Embodiments may include computer storage media, transient and non-transitory memory devices having stored thereon computer instructions or software code that may be used to program or configure a computing device, computer processor, or electronic circuit to perform any of the processes of the present invention. Storage media, transient and non-transitory memory devices may include, but are not limited to, floppy disks, optical disks, Blu-ray disks, DVDs, CD-ROMs, magneto-optical disks, ROMs, RAMs, flash memory devices, or any other type of medium or device suitable for storing instructions, code, and / or data.
[0055] Each functional unit and module according to various embodiments may also be implemented in a distributed computing environment and / or a cloud computing environment, where all or part of the machine instructions are executed in a distributed manner by one or more processing devices interconnected by a communication network, such as an intranet, a wide area network (WAN), a local area network (LAN), the Internet, and other forms of data transmission media.
[0056] The foregoing description of the present invention has been presented for purposes of illustration and description only. It is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Numerous modifications and variations will occur to those skilled in the art.
[0057] The embodiments were chosen and described in order to best explain the principles of the invention and its practical applications, thereby enabling others skilled in the art to understand the invention and understand its various embodiments and modifications as are suited to the particular use contemplated.
Claims
1. A system for performing deep learning model calculations, characterized in that: include: a first floating-point to block floating-point (FP to BFP) converter configured to receive pixel data in a 32-bit floating-point format and convert the pixel data by reducing a mantissa of the pixel data to 8 bits, thereby obtaining a block floating-point (BFP) format having an 8-bit mantissa; a second FP to BFP converter for receiving filter data in a 32-bit floating point format and converting the filter data by reducing a mantissa of the filter data to 8 bits, thereby obtaining a BFP format with an 8-bit mantissa; an 8-bit integer multiplier, configured to receive the BFP data from the first FP-to-BFP converter and the second FP-to-BFP converter, and perform a multiplication-accumulation operation to generate 16-bit product data; an adder, configured to receive the 16-bit product data from the 8-bit integer multiplier and accumulate a plurality of the 16-bit product data to generate 64-bit accumulated sum data; an accumulator, configured to receive the 64-bit accumulated sum data from the adder and perform accumulation and aggregation; and A block floating point to floating point (BFP-to-FP) converter is configured to receive the 64-bit accumulated sum data from the accumulator and convert the 64-bit accumulated sum data into output data represented by 32-bit FP.
2. The system according to claim 1, wherein: The first FP to BFP converter and the second FP to BFP converter are both further configured to determine a shared exponent for a block of floating point values before converting the mantissa to 8 bits.
3. The system according to claim 2, characterized in that The shared exponent is determined as the largest exponent among all floating point values in the block of floating point values.
4. The system according to claim 3, wherein: The first FP to BFP converter and the second FP to BFP converter are further configured as follows: calculating a right shift amount for each of the floating-point values by subtracting the original exponent of each of the floating-point values from the determined maximum exponent; and The calculated right shift amount is applied to the mantissa of each of the floating-point values to align the floating-point values within the block under the shared exponent, and then the mantissa is truncated to 8 bits.
5. The system according to claim 4, wherein: The first FP-to-BFP converter and the second FP-to-BFP converter are each further configured to incorporate a hidden leading bit into the mantissa before performing the right shift and truncation.
6. The system according to claim 5, characterized in that The right shift and the truncation operations of the first FP-to-BFP converter or the second FP-to-BFP converter include selecting the most significant 8 bits from the 24-bit extended mantissa after the right shift.
7. The system according to claim 1, wherein: The BFP format with 8-bit mantissa contains a 17-bit representation, including a 1-bit sign, an 8-bit shared exponent, and an 8-bit mantissa obtained from the first FP to BFP converter or the second FP to BFP converter.
8. The system according to claim 1, wherein: The adder is configured to receive data of the multiple 16-bit products generated by successive multiplication-accumulation operations performed by the 8-bit integer multiplier.
9. The system according to claim 1, wherein: The BFP to FP converter is used to reconstruct the output data of the 32-bit FP representation by extracting the shared exponent and adjusting the accumulated mantissa accordingly.
10. The system according to claim 1, wherein: The 8-bit integer multiplier performs element-wise multiplication on the pixel data and the filter data in a convolutional neural network (CNN).
11. A method for converting floating-point (FP) data into block floating-point (BFP) format, characterized in that: include: Receive data in 32-bit floating-point format through a floating-point to block floating-point (FP to BFP) converter; determining, by the FP to BFP converter, a shared index for a block of floating-point values, the shared index being a maximum index among all the floating-point values in the block; calculating, by the FP to BFP converter, a right shift amount for each of the floating-point values based on a difference between an original exponent of each of the floating-point values and the determined shared exponent; Shifting the mantissa of each floating-point value right by the FP to BFP converter; as well as The 24-bit mantissa of each of the floating-point values is truncated to 8 bits by the FP to BFP converter to obtain a BFP format.
12. The method according to claim 11, characterized in that The data in the 32-bit floating point format is pixel data or filter data used for convolutional neural network (CNN) calculations.
13. The method according to claim 11, characterized in that Also includes: Before performing the right shift and truncation, the hidden leading bit is incorporated into the mantissa.
14. The method according to claim 13, characterized in that The right shift and the truncation involve selecting the most significant 8 bits from the 24-bit extended mantissa after the right shift as the significant bits.
15. The method according to claim 11, characterized in that The BFP format contains a 17-bit representation, including a 1-bit sign, an 8-bit shared exponent, and an 8-bit mantissa.
16. A tiny machine learning (tiny ML) platform, characterized in that: include: A micro machine learning accelerator for executing convolutional neural networks (CNNs) using approximate computing techniques. CNN) operations, and includes: a first floating-point to block floating-point (FP to BFP) converter configured to receive pixel data in a 32-bit floating-point format and convert the pixel data by reducing a mantissa of the pixel data to 8 bits, thereby obtaining a block floating-point (BFP) format having an 8-bit mantissa; a second FP to BFP converter for receiving filter data in a 32-bit floating point format and converting the filter data by reducing a mantissa of the filter data to 8 bits, thereby obtaining a BFP format having an 8-bit mantissa; and an 8-bit integer multiplier configured to receive the BFP data from the first FP-to-BFP converter and the second FP-to-BFP converter and perform a multiplication-accumulation operation to generate 16-bit product data; a single instruction multiple data multiply-accumulate (SIMD MAC) accelerator coupled to the micro machine learning computation accelerator to perform signal feature extraction and enhance vectorized execution of CNN computations; A microcontroller unit (MCU) core for managing execution control, memory access, and coordination between the micro machine learning accelerator and the single instruction multiple data multiply-accumulate accelerator; a memory coupled to the microcontroller unit core, the micro machine learning operation accelerator, and the single instruction multiple data multiply-accumulate accelerator; and A multiplexer (MUX) for dynamically routing data between the microcontroller unit core, the memory, the micro machine learning operation accelerator, and the single instruction multiple data multiply-accumulate accelerator.
17. The micro machine learning platform according to claim 16, characterized in that The SIMD multiply-accumulate accelerator is also used to perform signal feature extraction by processing incoming voice data prior to micro machine learning inference.
18. The micro machine learning platform according to claim 16, characterized in that The micro machine learning platform is coupled to a device platform software layer, wherein the software layer is used to optimize inference execution on microcontrollers and edge devices.
19. The micro machine learning platform according to claim 16, wherein: The micro machine learning platform is coupled to a microphone and configured to receive an audio signal captured by the microphone, thereby converting the audio signal into digitally represented data.
20. The micro machine learning platform according to claim 19, characterized in that The micro machine learning platform is used to determine a classification output based on the voice command recognized from the audio signal, and to generate a control signal based on the classification output to perform an operation on the IoT device.