Systems and methods for matrix multiplication instructions with specified deviation floating point operands

By designing a scalable node engine with multiple matrix processors and configurable data formats, and optimizing data formats and data paths, the high demand problem of data and computing resources in machine learning model training in the prior art is solved, and efficient matrix operation and performance improvement is achieved.

CN120085915APending Publication Date: 2025-06-03TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129658.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-05-23
Filing Date
2020-03-02
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has high demand for data and computing resources in the training of machine learning models, and the data format and data pipelines of traditional GPUs are not suitable for training models, resulting in inefficiency.

Method used

A scalable node engine with multiple matrix processors and configurable data formats is designed to improve the efficiency of matrix operations by optimizing data formats and data paths. Specific measures include using low-bit floating-point format to store matrix operands and using high-bit floating-point formats in intermediate and final results to improve data bandwidth and retain results accuracy.

Benefits of technology

It significantly improves the efficiency and performance bandwidth of machine learning model training, reduces the delay between matrix multiplication results, and can effectively process large amounts of training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085915A_ABST
    Figure CN120085915A_ABST
Patent Text Reader

Abstract

A microprocessor system includes a matrix calculation unit and a control unit. The matrix calculation unit includes a plurality of processing elements. The control unit is configured to provide matrix processor instructions to the matrix calculation unit. The matrix processor instruction specifies a floating point operand formatted using a first floating point representation format. The matrix calculation unit accumulates the intermediate result values calculated using the floating point operands. The intermediate result value is represented in a second floating point format.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with the filing date of March 2, 2020, application number 202080033150.6, and invention name "Systems and Methods for Matrix Multiplication Instructions with Specified Deviation Floating Point Operands".

[0002] Cross - Reference to Related Applications

[0003] This application claims the priority of U.S. Patent Application No. 16 / 403,083, entitled "Data Path of a Scalable Matrix Node Engine with Hybrid Data Formats", filed on May 3, 2019, and also claims the priority of U.S. Patent Application No. 16 / 421,225, entitled "Scalable Matrix Nodes with Configurable Data Formats", filed on May 23, 2019, the disclosures of which are incorporated herein by reference in their entireties. Background Art

[0004] Machine learning training is a data - and compute - intensive operation. This process is both tedious and time - consuming, requiring large amounts of relevant training data and computing resources to process it. Moreover, the data and computing resources only increase with the complexity of the problem being solved. To train machine learning models, high - power CPUs use training data to perform complex matrix operations to determine appropriate weights. To increase the training speed, Graphics Processing Units (GPUs) are used as an alternative or supplement to traditional CPUs. GPUs allow some of the training in training to be parallelized and help optimize certain mathematical operations. However, traditionally, GPUs are designed for processing graphics problems, such as rendering a three - dimensional world onto a two - dimensional display. When applied to machine learning, GPUs may require a large amount of computing power to meet the computing power they provide. Moreover, the data formats and data pipelines used by GPUs are designed for graphics processing, not for training machine learning models. Therefore, there is a need for a powerful, computationally strong, and energy - efficient machine learning training system. Such a system should support high data bandwidth to significantly increase the amount of training data that can be processed. Moreover, the data formats and data pipelines should be optimized for training data and the resulting machine learning models. Brief Description of the Drawings

[0005] Various embodiments of the present invention are disclosed in the following detailed description and the drawings.

[0006] Figure 1 is a flowchart illustrating an embodiment of a process for training a machine learning model.

[0007] Figure 2 is a block diagram illustrating an embodiment of a system for training a machine learning model.

[0008] Figure 3 is a block diagram illustrating an embodiment of a node engine for performing matrix calculations.

[0009] Figure 4 is a block diagram illustrating an embodiment of an 8-bit floating-point format.

[0010] Figure 5 is a block diagram illustrating an embodiment of a 21-bit floating-point format.

[0011] Figure 6 is a flowchart illustrating an embodiment of a process for performing matrix calculations.

[0012] Figure 7 is a flowchart illustrating an embodiment of a process for performing matrix calculations.

[0013] Figure 8 is a flowchart illustrating an embodiment of a process for performing multiple interleaved matrix calculations. DETAILED DESCRIPTION

[0014] The present invention may be implemented in many ways, including as a process, apparatus, system, composition of matter, computer program product embodied on a computer-readable storage medium, and / or a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor. In this specification, these implementations or any other form that the present invention may take may be referred to as techniques. Generally, the order of the steps of the disclosed processes may be altered within the scope of the present invention. Unless otherwise stated, components such as a processor or a memory described as being configured to perform a task may be implemented as a general component temporarily configured to perform the task at a given time or a specific component manufactured to perform the task. As used herein, the term 'processor' refers to one or more devices, circuits, and / or processing cores configured to process data such as computer program instructions.

[0015] The following provides a detailed description of one or more embodiments of the present invention and the accompanying drawings that illustrate the principles of the present invention. The present invention is described in connection with these embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is defined only by the claims, and the present invention encompasses many alternatives, modifications, and equivalents. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. These details are provided for purposes of illustration, and the present invention may be practiced without some or all of these specific details in accordance with the claims. For clarity, technical material known in the technical field related to the present invention has not been described in detail so as not to unnecessarily obscure the present invention.

[0016] A scalable node engine with multiple matrix processors and configurable data formats is disclosed. As a core component of a training platform for machine learning models, the node engine can be arranged in a network to perform training on machine learning models. As the computational and data requirements increase, the number of node engines in the network can be increased to handle the additional requirements. Compared with traditional CPUs and GPUs that delegate tasks for similar workloads, the disclosed node engine is very efficient in terms of performance per watt per square millimeter. The node engine architecture achieves this performance improvement in part by optimizing the data formats and data paths for machine learning workloads. For example, the node engine includes multiple matrix processors, and each matrix processor can interleave multiple matrix operations. A node engine with a group of eight matrix processors can compute the result of matrix multiplication in each cycle. When waiting for data for a first set of related matrix operations to stop, each matrix processor can interleave a second set of related matrix operations to utilize computational resources that would otherwise be idle. In some embodiments, matrix operands are stored using a low-precision floating-point format and intermediate and final results are computed using a high-precision floating-point format. The low format increases the read data bandwidth of the matrix processor, while the high format preserves the accuracy and precision of the matrix results, for example, by preventing loss of accuracy in the quantized results. Different configurable data formats can be selected to specify different data format configurations, for example, to change the number of bits allocated to the mantissa field and the exponent field. This allows optimization of the data format based on the specific matrix operations for a particular machine learning task. Additionally, the data format can include a configurable bias for the biased exponent. This improves the range of the exponent and allows utilization of a larger range.

[0017] In some embodiments, the node engines are arranged in a mesh network. Each node engine includes a control unit, a memory, registers, multiple matrix processors, and a post-processing unit such as a vector computing unit. The control unit can process custom instructions, which include matrix computation instructions involving one of the multiple matrix processors and are used to synchronize results between different matrix processors and the node engine. Matrix results can be stored in a register file and processed using vector operations by the post-processing unit. The software running the node engine is capable of performing large matrix operations and subdividing problems. Different sub-parts of the problem may be distributed to different node engines and different matrix processors of each node engine. For example, two large matrices can be sliced such that each slice is optimized to the matrix size of the matrix processor. Then, the slices can be assigned to different matrix processors of different node engines, where matrix multiplication is performed on the slices. The results of each matrix multiplication can be combined to compute the multiplication result of the original larger matrix.

[0018] In some embodiments, a microprocessor system includes a matrix calculation unit and a control unit. The matrix calculation unit includes one or more processing elements. For example, the matrix calculation unit includes a matrix of calculation units for determining a calculation result of two elements from two operands. An 8×8 matrix calculation unit includes 64 calculation units. Similarly, an M×N matrix calculation unit includes M×N calculation units. The matrix calculation unit is part of a matrix processor controlled via the control unit. In some embodiments, the control unit is configured to provide matrix processor instructions to the matrix calculation unit. For example, the control unit provides a matrix multiplication instruction to the matrix processor for execution by the matrix calculation unit. The matrix processor instructions specify floating-point operands that are formatted using an exponent that has been biased using a specified configurable bias. For example, the matrix multiplication instruction specifies two floating-point matrix operands. Each element of the matrix operands is formatted using a specific floating-point format and a configurable exponent bias. Along with the matrix operands, the matrix processor instructions specify the floating-point format used by the matrix elements (such as a format that allocates 1 bit for a sign bit, 4 bits for an exponent, and 3 bits for a mantissa) and a specific exponent bias. In various embodiments, the bias can be configured by specifying a value corresponding to the exponent bias. In some embodiments, the bias is reconfigurable. For example, a matrix instruction can specify a new bias for reconfiguring the configurable bias. In some embodiments, the floating-point format supports denormal numbers to increase the number of values that can be represented.

[0019] In some embodiments, matrix processor instructions specify floating-point operands formatted using a first floating-point representation format. For example, the instruction specifies an 8-bit floating-point format that allocates 4 bits for the exponent, 3 bits for the mantissa, and a single sign bit. The specified format is used for the elements of the matrix operand. The format can be selected to increase the data bandwidth of the matrix calculation unit in the matrix processor. The matrix calculation unit accumulates intermediate result values calculated using the floating-point operands, and the intermediate result values are in a second floating-point representation format. For example, the intermediate result uses a different floating-point format, such as a 21-bit floating-point format. As another example, the intermediate result can use a different floating-point format, such as 27 bits or another suitable floating-point format. The number of bits dedicated to the intermediate result can be selected to prevent loss of accuracy when quantifying the result. A format using a large number of bits to represent the intermediate result can be selected to prevent overflow errors and / or underflow errors that may occur when using the first floating-point format. The matrix calculation unit outputs the accumulated intermediate result as an output formatted in a third floating-point representation format. For example, multiple accumulated intermediate results can be moved out of the matrix processor as matrix results. The result can be output using a third format compatible with the bus to which the matrix processor is connected. For example, the node engine can utilize an internal bus with a width of 64 bytes. The intermediate accumulated result can be output from the matrix calculation unit as a 16-bit floating-point value, allowing 32 elements to be moved from the matrix processor for each move instruction. An accumulated result with 64 elements can be moved from the matrix processor to the register file of the node engine using two move instructions, where each instruction moves 32 elements. An upper move instruction can be used to move the upper 32 elements (e.g., elements 32 to 63), and a lower move instruction can be used to move the lower 32 elements (e.g., elements 0 to 31). In some embodiments, the move instruction is a non-destructive instruction, and when moving a value from the source accumulator of the matrix processor to a memory location outside the matrix processor (such as an output array or register), the content of the source accumulator is not cleared.

[0020] Figure 1 is a flowchart illustrating an embodiment of a process for training a machine learning model. For example, Figure 1The process can be used to train models for autonomous driving or driver - assisted driving. When the vehicle is driven, such as by a human driver autonomously or by a combination of manual driving and assisted driving, driving data can be captured. The captured data is prepared as training data and used to train a new machine - learning model to improve the driving experience. The new driving experience can be improved in areas such as safety, efficiency (power, time, etc.), comfort, performance, convenience, and the like. Once the new model is trained and validated, the newly trained model is deployed to a vehicle where it is used by one or more machine - learning networks to implement improved driving features and functions. The new features can include autonomous driving features or assisted - driving features, such as autonomous lane changes, autonomous lane merges onto highways, autonomous exits from highways, improved detection of obstacles and road scenes, and driving based on autonomous navigation. In various embodiments, the machine - learning model can be trained on a training platform that utilizes multiple node engines and where each node engine includes multiple matrix processors and a configurable data format.

[0021] At 101, data is captured for machine - learning training. In some embodiments, when the vehicle is driven by a human, an autonomous - driving system, or both, data corresponding to the vehicle's driving is captured. The data captured on the vehicle's driving conditions can include image - sensor data, vehicle operation parameters (e.g., speed, steering, etc.), vehicle - type information (e.g., left - hand drive, right - hand drive, vehicle model, etc.), whether autonomous driving is enabled, the time of the last disengagement of autonomous driving, the detected obstacles, driving conditions, and the like. The data can be captured passively without interfering with the vehicle's driving and without driver assistance.

[0022] In various embodiments, a vehicle may be equipped with sensors arranged in different ways to capture different forms of data. In some embodiments, the sensor data may be visual data, ultrasonic data, LiDAR data, or other suitable sensor data. For example, images are captured from a high-dynamic range forward camera. As another example, ultrasonic data is captured from a lateral ultrasonic sensor. In some embodiments, the vehicle is attached with multiple sensors for capturing data. For example, in some embodiments, eight surround cameras are attached to the vehicle and provide 360-degree visibility around the vehicle, with a range up to 250 meters. The camera sensors arranged in different ways may include a wide forward camera, a narrow forward camera, a rear camera, a front side camera, and / or a rear side camera. In some embodiments, additional ultrasonic sensors and / or radar sensors are used to capture surrounding details. For example, twelve ultrasonic sensors may be attached to the vehicle to detect both hard and soft objects. Additional forward radars may also be used to capture data on the surrounding environment. In various embodiments, despite heavy rain, fog, dust, and other vehicles, the radar sensors are still able to capture surrounding details. Each sensor is used to capture the environment around the vehicle, and the captured data is stored for consideration as training data for a deep learning network.

[0023] Once captured, the data captured from one or more vehicles is transmitted to a machine learning training platform. For example, a vehicle having wireless connectivity (such as cellular or WiFi connectivity) may transmit the data wirelessly to the machine learning training platform. As another option, when a technician is repairing the vehicle, the captured data may be downloaded from the vehicle. In various embodiments, the data captured from multiple vehicles (such as a fleet of vehicles) is aggregated at the machine learning training platform and used as at least one training data source for the training data source.

[0024] At 103, the captured data is prepared for training a machine learning model. At 101, the data captured from the vehicle is prepared as training data. In some scenarios, the data is divided into training data and validation data. Preparing the data may include selecting (or excluding) the captured data to identify particularly good training data. In some embodiments, the data is annotated to identify features for training. For example, lane markings, traffic lights, traffic signs, vehicles, pedestrians, etc. may be annotated to enhance the usefulness of the training data as part of data preparation. As another example, the data may be converted to a different format or preprocessed as part of the preparation process. In some embodiments, the data may be converted from the source data format to a format compatible with a matrix processor. For example, data captured as fixed-point data may be converted to floating-point data to improve precision.

[0025] At 105, a machine learning model is trained. Using the training data prepared at 103, one or more machine learning models are trained. The training can utilize both the training dataset and the validation dataset. In some embodiments, the training utilizes a machine learning platform composed of multiple node engines, where each node engine includes multiple matrix processors. By utilizing multiple node engines organized, for example, in a grid or another suitable architecture, complex machine learning training problems can be parallelized and executed faster and more efficiently. Also, since each node engine includes multiple matrix processors, each node can perform multiple matrix operations in parallel. In some embodiments, by operating multiple matrix processors in parallel, the node engine can output the result of matrix multiplication in each clock cycle. The latency for waiting for data reads is significantly reduced, the latency between matrix multiplication results is significantly reduced, and the performance bandwidth is significantly increased.

[0026] The result of the training is one or more trained machine learning models. In some embodiments, multiple models are trained, each for a potentially different neural network. For example, one machine learning model can be trained to utilize sensor data from a front-facing camera as input, while another model can be trained to utilize sensor data from a side ultrasonic sensor as input.

[0027] At 107, the trained machine learning model is distributed. For example, the trained model is distributed to a vehicle and installed on that vehicle. The model can be installed via an over-the-air update, by a technician when servicing the vehicle, or in other ways. In some cases, the model is packaged in a data format for easy installation on the vehicle. For example, the model can be compressed to minimize the time and bandwidth required to transfer the model to the vehicle. In some embodiments, multiple models (e.g., each for a different neural network engine operating on the vehicle) can be packaged together and transferred to the vehicle as a single package.

[0028] At 109, the trained machine learning model is applied. For example, a convolutional neural network on a vehicle utilizes the new model to process sensor data and implement autonomous driving or driver assistance features. In some embodiments, more than one model is applied and / or more than one neural network is utilized. For example, on some vehicles, multiple neural networks are utilized to process different data from different sensors. Once the new model is utilized, data reflecting the performance of the new model can be captured, and this data is used for future training. Figure 1 The process can be used to continuously improve the performance of the machine learning network. In this way, the process loops back to 101 where data is captured. The data can be analyzed to identify difficult use cases of the currently deployed model, and the corresponding captured data can be utilized for future training.

[0029] Figure 2is a block diagram illustrating an embodiment of a system for training a machine learning model. Using Figure 2 's training system, a machine learning model can be trained to implement autonomous driving and / or driver assistance driving functions. In some embodiments, Figure 2 's training system is used to perform Figure 1 's process. In the example shown, the training system utilizes certain training-related subsystems of the vehicle subsystem 201 located on the vehicle. The training-related subsystems communicate with the server side of the training system located in one or more training data centers 221. The vehicle subsystem 201 includes sensors 203, a deep learning network 205, an AI processor 207, a vehicle control module 209, a network interface 211, a vehicle data capture system 213, and a captured data storage device 215. For example, additional vehicle subsystems may exist to perform other functions, but are not shown. One or more training data centers 221 include a training platform 223, a training data storage device 227, and a model data storage device 229. The training platform 223 includes at least one or more node engines 225. The node engines are connected (e.g., in a mesh network) to perform parallel processing for machine learning training. In some embodiments, the training platform 223, the training data storage device 227, and the model data storage device 229 are located in a single data center, but may also be distributed or replicated across multiple data centers.

[0030] In some embodiments, a vehicle (not shown) includes a vehicle subsystem 201 to implement autonomous and driver assistance functions and capture a number of data that can be used to train one or more machine learning models to implement and / or improve functions and / or new features. In various embodiments, different vehicle subsystems may be communicatively connected. For example, sensor data from a sensor 203 is fed to a vehicle data capture system 213 for storage in a captured data storage device 215. The captured data is sent to a training platform 223 via a network interface 211. As another example, sensor data from a sensor 203 is fed to a deep learning network 205 running on an AI processor 207. The output of the deep learning network 205 running on the AI processor 207 is fed to a vehicle control module 209. In various embodiments, the network interface 211 is a wireless network interface, such as a wireless network interface including WiFi and / or cellular network connectivity. The network interface 211 is used to communicate with a remote server, make phone calls, send and / or receive text messages, transmit sensor data to the training platform 223, etc. In some embodiments, depending on the situation, the vehicle subsystem 201 may include additional or fewer subsystems. For example, in some embodiments, an image preprocessor (not shown) is used to preprocess the captured sensor data. As another example, in some embodiments, a post-processing component (not shown) is used to perform post-processing on the output of the deep learning network 205 before providing the output to the vehicle control module 209. In some embodiments, a trigger classifier component (not shown) is used to identify driving data as potential training data.

[0031] In some embodiments, the sensor 203 includes one or more sensors. The sensor 203 may be attached to the vehicle at different locations of the vehicle and / or oriented in one or more different directions. For example, the sensor 203 may be attached to the front, side, rear, and / or roof of the vehicle in the forward, backward, lateral, etc. directions. In some embodiments, the sensor 203 may be an image sensor such as a high dynamic range camera. In some embodiments, the sensor 203 includes non-visual sensors. The sensor 203 may include radar, LiDAR, and / or ultrasonic sensors, etc. In certain embodiments, the sensor 203 is not installed on the vehicle with the vehicle control module 209. For example, the sensor 203 may be installed on an adjacent vehicle and / or attached to the road or environment and included as part of a system for capturing sensor data.

[0032] In some embodiments, the deep learning network 205 is a deep learning network for implementing autonomous vehicle control. For example, the deep learning network 205 can be an artificial neural network such as a convolutional neural network (CNN) that is trained using sensor data and whose output is provided to the vehicle control module 209. The machine learning model used by the deep learning network 205 can be trained using Figure 2 the system of

[0033] In some embodiments, the artificial intelligence (AI) processor 207 is a hardware processor for running the deep learning network 205. In some embodiments, the AI processor 207 is a dedicated AI processor for performing inference on sensor data using a convolutional neural network (CNN). The AI processor 207 can be optimized for the bit depth of the sensor data and / or for deep learning operations such as neural network operations including convolution, dot product, vector, and / or matrix operations. In some embodiments, the AI processor 207 is implemented using a graphics processing unit (GPU). In various embodiments, the AI processor 207 is coupled to a memory that is configured to provide instructions to the AI processor that, when executed, cause the AI processor to perform a deep learning analysis on the received input sensor data and determine a machine learning result for at least partially autonomously operating the vehicle.

[0034] In some embodiments, the vehicle control module 209 is used to process the output of the artificial intelligence (AI) processor 207 and transform the output into vehicle control operations. In some embodiments, the vehicle control module 209 is used to control the vehicle for autonomous driving and can adjust the speed and / or steering of the vehicle. For example, the vehicle control module 209 can be used to control the vehicle by braking, steering, changing lanes, accelerating, and merging into another lane, etc. In some embodiments, the vehicle control module 209 is used to control vehicle lighting, such as brake lights, turn signals, headlights, etc. In some embodiments, the vehicle control module 209 is used to control vehicle audio conditions, such as the vehicle's audio system, playing audio alerts, enabling microphones, enabling speakers, etc. In some embodiments, the vehicle control module 209 is used to control a notification system including a warning system to notify the driver and / or passengers of driving events, such as potential collisions or approaching an expected destination. In some embodiments, the vehicle control module 209 is used to adjust sensors such as the vehicle's sensor 203. For example, the vehicle control module 209 can be used to change the parameters of one or more sensors, such as modifying orientation, changing output resolution and / or format type, increasing or decreasing the capture rate, adjusting the captured dynamic range, adjusting the focus of a camera, enabling and / or disabling sensors, etc. In various embodiments, the vehicle control module 209 is used to implement autonomous driving and / or driver assistance control of the vehicle.

[0035] In some embodiments, network interface 211 is a communication interface for sending and / or receiving data including the captured sensor data. In various embodiments, network interface 211 includes a cellular or wireless interface for interfacing with a remote server such as training platform 223 to make and receive voice calls, send and / or receive text messages, transmit sensor data, receive updates to the autonomous driving system including newly trained machine learning models, etc. For example, network interface 211 can be used to receive updates to instructions and / or operating parameters for sensor 203, deep learning network 205, AI processor 207, vehicle control module 209, and / or vehicle data capture system 213. For example, the machine learning model of deep learning network 205 can be updated using network interface 211. As another example, network interface 211 can be used to update the firmware of sensor 203 and / or the operating parameters of vehicle data capture system 213, such as filters and / or parameters for determining the type and quantity of data to be captured.

[0036] In some embodiments, vehicle data capture system 213 and captured data storage device 215 are used to capture and store data associated with vehicle driving conditions. The data captured by vehicle data capture system 213 is stored in captured data storage device 215. Captured data storage device 215 can be implemented using any suitable data storage device such as a hard disk drive, non-volatile memory, etc. In some embodiments, captured data storage device 215 is implemented using a database, file system, or another data organization means. The captured data of vehicle driving conditions can include image sensor data, vehicle operating parameters (e.g., speed, steering, etc.), vehicle type information (e.g., left-hand drive, right-hand drive, vehicle model, etc.), whether autonomous driving is enabled, the time of the last disengagement of autonomous driving, detected obstacles, driving conditions, etc. Data can be captured passively without disturbing vehicle driving and without requiring driver assistance. The data captured by vehicle data capture system 213 includes the data captured from sensor 203.

[0037] In some embodiments, vehicle data capture system 213 communicates with training platform 223 via network interface 211. Network interface 211 can be a wireless network such as WiFi and / or a cellular network. Vehicle data capture system 213 uses network interface 211 to transmit the captured data stored in captured data storage device 215 to training platform 223. In some embodiments, network interface 211 is used to download a trained machine learning model for installation in deep learning network 205 running on the vehicle.

[0038] In Figure 2In the example, the server - side components of the training system are located in one or more data centers of one or more training data centers 221 and include a training platform 223, a training data storage device 227, and a model data storage device 229. The training platform 223 includes one or more computer servers for receiving the captured data from the vehicle data acquisition system 213. The training platform 223 is communicatively connected to the vehicle data acquisition system 213 via a wireless network interface 211 through a computer network such as a limited network or an optical network of one or more training data centers 221. The training platform 223 also includes one or more node engines 225. For example, multiple node engines 225 can be connected in a mesh network. The training platform 223 receives the captured data from the vehicle data capture system 213, processes the data into usable training (and validation) data, and uses the node engines 225 to train one or more new machine - learning models. The training data storage device 227 is used to store the captured data received from one or more vehicles. In some embodiments, the processed captured data, including annotated data, used as training data is stored in the training data storage device 227. Once the training is completed, the model data storage device 229 is used to store the trained machine - learning models. For example, different versions of the trained machine - learning models can be stored in the model data storage device 229 and used to determine the relative functionality of different models and identify areas for improvement. In some embodiments, one or more data storage devices are used to implement the training data storage device 227 and the model data storage device 229.

[0039] In some embodiments, the node engine 225 includes a plurality of connection nodes that can be used to parallelize computing tasks. Each connection node includes at least one matrix processor, and possibly more than one matrix processor. For example, a single node can include eight matrix processors, each of which is capable of determining at least one matrix multiplication result. In some embodiments, the matrix multiplication result requires a single matrix processor at least a minimum number of clock cycles to compute. By expanding each node to include multiple matrix processors, after an initial latency corresponding to the minimum number of clock cycles for computing the matrix multiplication, the node can output the result of a matrix multiplication for each clock cycle. For example, if the matrix multiplication takes eight clock cycles to complete, after an initial latency of seven clock cycles, a node with eight matrix processors can determine the result of the matrix multiplication in each clock cycle. In various embodiments, the throughput is also determined by memory access, including the latency for accessing matrix operands. In various embodiments, the node engine is capable of performing matrix calculations using a variety of digital formats. For example, the node can utilize fixed-point formats and floating-point formats. For the floating-point format, the node can be configured to operate in a variety of formats such as 8-bit format, 16-bit format, and 32-bit format. For each bit depth, one or more different formats can be selected. Depending on the computing objective, different formats can be used to represent numerical values. A format can be selected to allocate higher precision to the mantissa of the floating point number, and another format can be selected to allocate higher precision to the exponent of the floating point number. In some embodiments, the floating-point format utilizes a configurable bias to further customize the computing operation. The configurability of the digital format allows the training system to target different machine learning operations, for example, based on expected input values, intermediate values, and output values. In various embodiments, the configurability of the node, including support for a variety of floating-point formats and floating-point formats with configurable biases, significantly improves the bandwidth and performance of matrix calculation operations without sacrificing precision and accuracy. Similarly, power consumption and efficiency are also significantly improved.

[0040] Figure 3FIG. is a block diagram illustrating an embodiment of a node engine for performing matrix computations. In the example shown, node engine 300 includes a control unit 301, a memory 303, a load register 305, a post-processing unit register file 307, multiplexers 309 and 311, matrix processors 313 and 351-357, an output array 315, and a post-processing unit 317. In various embodiments, the node engine may include multiple matrix processors to compute multiple matrix operations in parallel. In the example shown, node engine 300 includes eight matrix processors 313 and 351-357. Each matrix processor includes a data input array, a weight input array, multiple output accumulators, and a matrix calculation unit. In the example shown, matrix processor 313 includes a data input array 321, a weight input array 323, and two output accumulators 329 and 331. The data input array and the weight input array feed inputs to the matrix calculation unit 325. For example, data in the input arrays (e.g., data input array 321 and / or weight input array 323) is shifted a certain number of bytes (e.g., eight bytes) over multiple cycles (e.g., eight consecutive cycles) to feed the matrix calculation unit 325. In some embodiments, each matrix processor includes a single data input array and a single weight input array. The matrix calculation unit 325 includes a matrix of computing units such as computing unit 327. The M×N dimensional matrix calculation unit includes M×N computing units. The size of each input array is designed to fit the entire input matrix, and the size of each output accumulator is designed to fit the entire matrix result. In some embodiments, the node engine supports multiple floating-point formats, which include Figure 4 8-bit floating-point formats 400 and 410 of Figure 5 and Figure 1 , Figure 6 , Figure 7 and / or Figure 8 of the process.

[0041] In some embodiments, the node engine 300 may include additional components and additional control lines not shown. For example, the node engine 300 may include additional registers such as scalar registers, one or more memory caches, a data formatter for formatting values for the matrix processor, and additional control lines from the control unit 301 to sub-components such as multiplexers 309 and 311 and matrix processors 351 to 357, as several examples. In some embodiments, certain registers (not shown) are dedicated to storing configurable parameters such as digital formats and configurable biases for floating-point numbers. In some embodiments, the buses connecting the different components of the node engine 300 are wide data buses. The size of the bus can be selected to optimize the transfer of matrix values. For example, the width of the bus can all be 64 bytes. This allows an 8×8 matrix of 64 one-byte elements to be transferred from memory to registers, matrix processors, etc. as a single unit.

[0042] In the example shown, the control unit 301 is communicatively coupled to one or more components of the node engine 300, the one or more components including a memory 303, a matrix processor 313, an output array 315, and a post-processing unit 317. Although not shown, the control unit 301 is also communicatively coupled to each of the remaining matrix processors 351 to 357. In various embodiments, the control unit 301 is used to synchronize the processing of computational operations, which include matrix operations and post-processing operations (such as vector operations) and / or accesses to memory and registers. For example, the control unit 301 sends signals to the matrix processor 313 to schedule matrix calculation instructions and may monitor ready signals from the matrix processor 313 to indicate when new instructions can be received and / or when matrix operations are completed and when matrix results are ready.

[0043] In some embodiments, the memory 303 is a memory module for storing the input operands and output results of matrix calculations and post - processing calculations. The memory 303 may include one or more caches (not shown). In the example shown, the memory 303 is connected to the load register 305, multiplexers 309 and 311, and the post - processing unit register file 307. Additional connections or fewer connections are possible, depending on the flexibility required to store data in and retrieve data from the memory. As shown, data can be read from the memory into the load register 305 and the post - processing unit register file 307 and read from the load register 305 and the post - processing unit register file 307 into the memory. The connection to the registers allows data values to be quickly stored in the registers, for example, as arguments for matrix or vector calculations. The memory 303 is also connected to the multiplexers 309 and 311 so that the input matrices can be retrieved from the memory. In some embodiments, the memory access to the memory 303 is controlled by a memory arbiter (not shown) to optimize memory requests, for example, by queuing memory requests and prioritizing certain memory reads over others. In some embodiments, the memory 303 is a static random - access memory (SRAM).

[0044] In some embodiments, the node engine 300 includes registers such as the load register 305 and the post - processing unit register file 307. These registers can be used to optimize memory access. As several examples, the registers can be used to store values retrieved from the memory 303, to store these values before writing them to the memory 303, to store the input and output values of the matrix processor, and to store the input and output values of the post - processing unit. In some embodiments, the post - processing unit register file 307 is a register file for the post - processing unit 317 and is compatible with different lane configurations (e.g., 64 - lane configuration, 32 - lane configuration, and / or 16 - lane configuration) of the post - processing unit 317. For example, the registers of the post - processing unit register file 307 can be addressed using various byte formats such as 1 - byte values, 2 - byte values, and 4 - byte values. In some embodiments, the size of each register is 64 bytes and the register can store 64 1 - byte elements, 32 2 - byte elements, or 16 4 - byte elements. In various embodiments, the data format can be configured and includes various 8 - bit floating - point formats, 16 - bit floating - point formats, and 32 - bit floating - point formats.

[0045] In some embodiments, multiplexers are used to select the input operand sources to the matrix processors. In the example shown, multiplexers 309 and 311 are used to select the source for the data input matrix and the weight input matrix for matrix processor 313. Depending on the control signals received at each multiplexer, data can be sourced from memory 303 or the post-processing unit register file 307. In some embodiments, data sourced from memory 303 is retrieved via the registers of load register 305. In some embodiments, multiplexers 309 and 311 are also used to select the data input matrices and weight input matrices for matrix processors 351 to 357. By offsetting the processing of multiple matrix processors of the node engine, a pair of multiplexers is used to select the inputs for all matrix processors of the node engine. In various embodiments, multiplexers 309 and 311 are used to control which matrix processor receives which matrix operand. Depending on the configuration, a single matrix processor, a subset of all matrix processors, or all matrix processors receive the selected matrix operands. In various alternative embodiments, node engine 300 includes additional multiplexers (not shown) dedicated to each of matrix processors 351 to 357.

[0046] In some embodiments, matrix processor 313 receives matrix operation instructions and performs matrix calculations such as matrix multiplication. For each matrix instruction, matrix processor 313 stores one or more matrix operands in one or more input arrays. For example, a data matrix is stored in a data input array such as data input array 321, while a weight matrix is stored in a weight input array such as weight input array 323. In various embodiments, the matrix operands are a pair of data and weight matrices, a pair of data and gradient matrices, a pair of weight and gradient matrices, or another pair of suitable matrix operands. In various embodiments, matrix processor 313 is used to calculate multiple related matrix calculations as part of the process of calculating the matrix multiplication of matrices that are too large to fit in input arrays 321 and 323 of matrix processor 313. The results of the related matrix calculations are combined as part of the process of calculating the matrix multiplication of the larger matrix. In various embodiments, matrix processor 313 interleaves multiple matrix operations (related or unrelated). For example, matrix processor 313 can interleave performing one or more related matrix operations on a first pair of matrices with performing one or more related matrix operations on a second pair of matrices. For example, matrix processor 313 can perform matrix multiplication on matrices W A and D A that are part (e.g., slices) of larger matrices W 1 and D 1 respectively, and then perform matrix multiplication on matrices W B and G BA matrix W of a portion (e.g., a slice thereof) 2 and G 2 perform matrix multiplication. The matrix multiplication result of matrix W 1 and D 1 is a partial result for calculating the matrix multiplication of the larger matrix W A and D A The matrix multiplication result of matrix W 2 and G 2 is a partial result for calculating the matrix multiplication of the larger matrix W 2 and G 2 The input matrix W 1 and D 1 as well as the input matrix W 2 and G 2 are stored in a pair of weight and data input arrays, such as arrays 321 and 323. In some embodiments, separate output accumulators 329 and 331 are used to accumulate intermediate and / or final results W 1 *D 1 and intermediate and / or final results W 2 *G 2 . For example, output accumulator 329 is used to accumulate intermediate and / or final results of the matrix multiplication associated with matrix W 1 and D 1 and output accumulator 331 is used to accumulate intermediate and / or final results of the matrix multiplication associated with matrix W 2 and G 2 .

[0047] In some embodiments, the sizes of the data input array and the weight input array are designed to fit the entire matrix in a linearized form. For example, a matrix processor capable of performing matrix multiplication on two matrices of size M×N and N×O has an input array of size M×N elements for receiving the corresponding M×N and N×O input matrices and another input array of size N×O elements. In some embodiments, the matrix processor performs calculations on two 8×8 matrices, and the sizes of the weight input array and the data input array are each designed to receive 64 elements. Similarly, the size of the output accumulator is designed to store the entire result matrix. The size of the output accumulator for storing the matrix multiplication result between two matrices of size M×N and N×O is designed to receive M×O elements. In some embodiments, the matrix processor performs calculations on two 8×8 matrices and stores the intermediate and final matrix results in an accumulator whose size is designed to fit 64 elements corresponding to the 8×8 result matrix.

[0048] In the example shown, the input array feeds a matrix calculation unit 325. The matrix calculation unit 325 consists of a matrix of calculation units (such as calculation unit 327). Each calculation unit is a processing element that can receive two operands, one from each input matrix, and perform a calculation such as multiplication on the two input operands. In some embodiments, the calculation is multiplication and addition. For example, two input elements are multiplied, and the result is added to the current result in the accumulator and stored back in the accumulator. In some embodiments, each calculation unit such as calculation unit 327 includes an arithmetic logic unit for performing arithmetic logic operations such as multiplication, division, addition, or subtraction operations. In some embodiments, multiple operations can be performed in the same clock cycle, such as the multiplication and addition operations required to perform a partial dot product. Each calculation unit can include an adder, a multiplier, and / or one or more accumulators corresponding to one or more pairs of data and weight input arrays. In some embodiments, each calculation unit such as calculation unit 327 includes a floating-point multiplier and one or more accumulators. Although in Figure 3 the output accumulators 329 and 331 are depicted as separate from the calculation unit 327, in some embodiments, the corresponding portions of the output accumulators 329 and 331 are integrated into their respective calculation units. For example, the accumulators of each calculation unit together form the output accumulators 329 and 331.

[0049] In various embodiments, the calculation units of the matrix calculation unit 325 support floating-point operations, such as floating-point multiplication and addition. In various embodiments, each calculation unit includes a multiplier and one or more accumulators to perform multiplication and addition operations in a single cycle. Before each matrix calculation begins, the specified accumulator can be cleared. During the process of performing the matrix calculation, the specified accumulator is used to accumulate and store intermediate results. In some embodiments, the matrix processor 313 is an 8×8 matrix processor, and the matrix calculation unit 325 includes 64 calculation units. Each cycle, 128 elements can be loaded into the matrix calculation unit 325, and two input elements serve as operands for each of the 64 calculation units. Each calculation unit can also access the accumulator value stored in the specified accumulator.

[0050] In some embodiments, matrix multiplication takes multiple clock cycles to complete. For each clock cycle, a single row and a single column are retrieved from the input operands. For example, a row is retrieved from the matrix stored in the data input array, while a column is retrieved from the matrix stored in the weight input array. In some embodiments, data is retrieved by shifting entire rows or entire columns of data in the input arrays. Each row and each column are vectors, and each vector is replicated across the entire computing unit. Each row is replicated "down" the rows of the matrix computing unit 325, and each column is replicated "across" the columns of the matrix computing unit 325. For an 8×8 matrix processor, each column of the weight input matrix is 8 elements and each row of the data input matrix is 8 elements. For each pass, a single weight column is replicated for each of the eight columns of the matrix computing unit 325, and a single data row is replicated for each of the eight rows of the matrix computing unit 325. By replicating data across and down one row and one column at a time, an 8×8 matrix processor can complete matrix multiplication in 8 cycles. During each cycle, the intermediate results of the multiplication and accumulation are stored in a designated accumulator. By the eighth and final cycle, the final matrix result is stored in the designated accumulator. Matrix processors with different dimensions (e.g., 4×4 or 16×16 matrices) can be used with input arrays, accumulators, and computing units of corresponding sizes.

[0051] In some embodiments, the input data elements are 8-bit floating-point values. By leveraging 8-bit values, the bandwidth performance of the matrix processor is significantly improved. By leveraging configurable floating-point values and configurable biases, the precision and accuracy required for machine learning training are retained, and the bandwidth is increased. Using the 8-bit format, a 64-byte × 64-byte matrix processor can compute the matrix multiplication of two 8×8 matrices (a total of 128 elements). In contrast, using the 32-bit format, a 64-byte × 64-byte matrix processor can compute the matrix multiplication of two 4×4 matrices (a total of only 32 elements). By optimizing the matrix elements using the configurable 8-bit floating-point format, the bandwidth for loading matrix elements into the matrix processor is significantly increased. The power consumption per unit area is also significantly improved. To prevent overflow and underflow errors, the intermediate and final results stored in the designated accumulator utilize a larger bit format, such as 21-bit, 27-bit, or other appropriate floating-point formats. Using 8-bit elements as input elements and using the 21-bit format to store intermediate results retains the precision and accuracy required for training, while also maintaining the high input bandwidth of the matrix processor. In various embodiments, each output accumulator stores each element of the result matrix using a 21-bit floating-point number, such as Figure 5Format 500. In some embodiments, matrix processor 313 is an 8×8 matrix processor that performs matrix operations using 8-bit floating-point input values and calculates intermediate and final matrix results using 21-bit floating-point values. The input array is 64 bytes (64 8-bit elements), and the output accumulator is 168 bytes (64 21-bit elements). In various embodiments, the output accumulator is specified by matrix calculation instructions. Similarly, the 8-bit floating-point format and exponent bias can be configured by matrix calculation instructions and / or one or more register parameters.

[0052] In some embodiments, matrix processor 313 supports multiple different 8-bit floating-point formats. For example, different formats 400 and 410 are supported and can be selected based on the computational task. Each format allocates a different number of bits to represent the exponent and mantissa of the floating-point number. Depending on the use case, one format or the other is selected. In cases where highly accurate numbers are needed, more bits can be allocated to the mantissa, and a format such as format 400, which has more mantissa bits than format 410, can be selected. A format with more mantissa bits can be selected to perform gradient descent, where very small deltas are needed to preserve accuracy. As another example, a format with more mantissa bits can be selected to perform forward propagation to calculate the cost function. As another optimization, each floating-point format utilizes a configurable bias. The configurable bias is used to shift the exponent range. For example, in the absence of an exponent bias, an exponent represented by 3 bits can specify an exponent value between 2 0 and 2 7 (inclusive). A bias of 5 shifts the range of the exponent to have an exponent value between 2 -5 and 2 +2 (inclusive). As another example, 4 bits are used to represent the exponent, and a bias of 15 shifts the range of the exponent from 2 0 and 2 31 (inclusive) to between 2 -15 and 2 +16 (inclusive). In various embodiments, by optimizing the number of bits in the exponent field and the number of bits in the bias, the range expressed using the exponent and the number coverage of the floating-point numbers can be optimized to preserve the accuracy and precision of the expected inputs and results.

[0053] In some embodiments, the floating-point format supports denormal numbers. For example, an exponent field with a zero value does not require a normalized mantissa without leading zeros. By supporting denormal numbers, the exponent range and the number of values that can be represented can be increased. In various embodiments, each computational unit, such as computational unit 327, includes support for performing floating-point operations using one or more denormal operands.

[0054] In some embodiments, the value of a configurable bias is limited by the number of bits used to represent the configurable bias. For example, a 3-bit configurable bias can have eight different values (0 through 7, inclusive). In some embodiments, as an optimization, the values represented by the configurable bias are not consecutive. For example, the eight values represented by a 3-bit configurable bias are not limited to the values 0 through 7. Instead, the bias can be selected from eight different values. For example, the configurable bias can be selected from eight predetermined values: 1, 3, 5, 7, 9, 11, 15, and 17. In some embodiments, the predetermined values are determined based on the most useful biases. The predetermined values can be selected at least in part to maximize the range and minimize the overlap between the ranges of different biases. In some embodiments, the configurable bias is specified by a matrix processor instruction and / or stored in a register (not shown). In some embodiments, the configurable bias can be reconfigured. For example, after performing an arithmetic operation, the configurable bias can be reconfigured to adjust to the new range of the result. In some embodiments, the reconfiguration is specified as part of a compute instruction. For example, the instruction can specify a new bias for reconfiguring the configurable bias.

[0055] In some embodiments, the compute units of a matrix compute unit can be grouped to also support matrix operations for larger input number formats. For example, the compute units of each 8×8 matrix compute unit that operates on 8-bit floating-point matrix elements as input can be grouped to perform 4×4 matrix operations using 16-bit floating-point matrix elements as input. In some embodiments, the size of the output accumulator is designed to prevent loss of accuracy of the quantized result. For example, a 16-bit floating-point format that uses 1 bit for the sign bit, 8 bits for the exponent, 7 bits for the mantissa, and a non-configurable exponent bias uses a 27-bit intermediate floating-point format for the floating-point result. The 27-bit floating-point format can allocate 1 bit for the sign bit, 9 bits for the exponent, and 17 bits for the mantissa. Support for the grouped operation mode makes the matrix compute unit more versatile in part by supporting more operand formats.

[0056] In various embodiments, the grouped operation mode performs matrix operations by splitting an input operand into multiple components and providing each split component to a different compute unit of the group. Each split component is represented as a floating-point number, and when added together, the different split components sum to the original operand. For example, the input operand is split into the most significant bits (i.e., the high component) and the least significant bits (i.e., the low component). In various embodiments, the exponent of the high component uses the same exponent value as the input operand, while the exponent of the low component is adjusted to account for subtracting the most significant bits from the input operand. In some embodiments, the least significant bit component is normalized. In some embodiments, the compute units support denormal numbers and the component can be represented as a denormal number.

[0057] In various embodiments, when performing a multiplication on two input operands using an operand number format that is twice the size of the compute unit format (e.g., 16-bit floating-point operands instead of 8-bit floating-point operands), four compute units are grouped together, and each input operand has a corresponding high and low component. By pairing the high-high, high-low, low-high, and low-low components and providing the different pairs to different compute units of the group, the high and low components of each input operand are provided to the processing elements. At each compute unit of the group, a matrix multiplication is performed and the result is stored in an output accumulator associated with the compute unit. In some embodiments, the output accumulator utilizes a floating-point format with a higher number of bits than the original input operands. For example, for 16-bit input operands without a configurable exponent bias, the output accumulator may utilize 27 bits. When the output results of the grouped units are added together, the result is the matrix multiplication of the original input operands. In some embodiments, the result is shifted out of the matrix compute unit and added together using a post-processing unit such as a vector compute unit. For example, a floating-point addition instruction is used to add the component results to determine the multiplication result. A floating-point vector add instruction may be used to add the components of the result vector. In various embodiments, the matrix compute unit is Figure 3 matrix compute unit 325 of Figure 3 and the post-processing unit is

[0058] post-processing unit 317 of

[0059] In some embodiments, a matrix instruction specifies a particular matrix operation, a particular matrix processor, an accumulator for storing the matrix result, and the location of the matrix operands. The location of the matrix operands can be specified using register values or memory addresses. For example, a matrix instruction can specify matrix multiplication, matrix multiplication processor 313, output accumulator 329, a register of the post-processing unit register file 307, and a memory address of memory 303. In some embodiments, control unit 301 issues the matrix instruction. In some embodiments, the operations include matrix multiplication, matrix addition, dot product, matrix inversion, etc. In some configurations, the output accumulator of each matrix processor uniquely identifies the matrix processor. By specifying a particular output accumulator as part of the matrix instruction, the matrix processor is selected in an inherent manner. For example, using the A0 - A11 naming scheme for accumulators, the first output accumulator and the second output accumulator (e.g., A0 and A1) are mapped to matrix processor 313, the third output accumulator and the fourth output accumulator (e.g., A2 and A3) are mapped to matrix processor 351, the fifth output accumulator and the sixth output accumulator (e.g., A4 and A5) are mapped to matrix processor 352, and so on. In this example, accumulators 329 and 331 are referred to as A0 and A1, respectively. Since only matrix processor 313 can store the result into accumulator A1, a matrix multiplication instruction specifying accumulator A1 is issued to matrix processor 313.

[0060] In some embodiments, output array 315 is used to retrieve the results of one or more matrix processors. In some embodiments, output array 315 includes multiplexers to determine from which matrix processor to load the results into the output array. In some embodiments, the output array is a 64 - byte array and two move instructions are required to move the matrix result from the matrix processor to the output array. For example, a matrix result using 21 - bit floating - point values requires 168 bytes. Each 21 - bit floating - point value is converted to a 16 - bit floating - point value during the move command. With only two move instructions, a result matrix of 64 elements is converted from 64 21 - bit floating - point values to 64 16 - bit floating - point values. For example, the up - move instruction moves the highest 32 elements into the output array, and the down - move instruction moves the remaining lowest 32 elements into the output array. In various embodiments, the output array is 64 bytes so that the result of the first move is first stored in a register (such as a register of the post - processing unit register file 307) before the second move is executed. In various embodiments, the output array is a temporary output array until the values are moved to memory or a register. In some embodiments, the move instructions are non - destructive and do not clear the matrix result from the matrix processor, e.g., by clearing the source accumulator.

[0061] In some embodiments, the post - processing unit 317 is used to perform post - processing such as normalization, expansion, activation functions, pooling, and the like. In some embodiments, the post - processing unit 317 is a vector - computing engine that operates on each element of a vector. The post - processing unit can utilize different numerical formats including floating - point formats, such as 1 - byte numerical format, 2 - byte numerical format, and 4 - byte numerical format. In some embodiments, the number of channels of the post - processing unit 317 can be configured. For example, the post - processing unit 317 with a 64 - byte vector can operate on 64 1 - byte elements, 32 2 - byte elements, or 16 4 - byte elements corresponding to 64 - channel configuration, 32 - channel configuration, and 16 - channel configuration. In the illustrated example, the post - processing unit 317 utilizes the post - processing unit register file 307 to retrieve data for input and to store the post - processing results. In some embodiments, additional post - processing units (not shown) can be included in the node engine as needed to perform additional machine - learning functions.

[0062] Figure 4 is a block diagram illustrating an embodiment of an 8 - bit floating - point format. In the illustrated example, the 8 - bit floating - point formats 400 and 410 are different 8 - bit floating - point formats for representing floating - point numbers using a sign, a mantissa, and an exponent. In some embodiments, node engines such as the node engine 300 and matrix processors such as Figure 3 the matrix processor 313 utilize the 8 - bit floating - point formats 400 and 410 for matrix operations. By performing matrix operations using 8 - bit floating - point formats such as formats 400 and 410 instead of 16 - bit floating - point format, 32 - bit floating - point format, or other floating - point formats, the bandwidth of the matrix processor is significantly increased. In some embodiments, formats 400 and 410 support configurable biases. The configurable biases achieve a greater range in representing exponents to improve precision while still maintaining an 8 - bit data size. In some embodiments, the floating - point formats 400 and 410 support non - regular numbers to increase the number of values that can be represented.

[0063] In the example shown, the 8-bit floating-point format 400 includes a single bit for the sign bit 401, four bits for the exponent 403, and three bits for the mantissa 405. The sign bit 401, exponent 403, and mantissa 405 together occupy 8 bits and can be used to represent a floating-point number. Similarly, the 8-bit floating-point format 410 includes a single bit for the sign bit 411, five bits for the exponent 413, and two bits for the mantissa 415. The sign bit 411, exponent 413, and mantissa 415 together occupy 8 bits and can be used to represent a floating-point number. In some embodiments, a configurable bias is used to bias the exponent. For example, the four-bit exponent 403 of format 400 allows the exponent 403 to have 16 different values (i.e., values 0 to 15 inclusive). Using four bits without a bias (or equivalently with the configurable bias set to zero), the exponent 403 can represent exponents with values 2 0 to 2 15 In some embodiments, a configurable bias is used to bias the exponent. For example, the four-bit exponent 403 of format 400 allows the exponent 403 to have 16 different values (i.e., values 0 to 15 inclusive). Using four bits without a bias (or equivalently with the configurable bias set to zero), the exponent 403 can represent exponents with values 2 -5 to 2 10 By using a configurable bias, the range of the exponent can be shifted. For example, using a configurable bias set to the value 5, the exponent 403 can represent exponents with values 2

[0064] In various embodiments, the configurable bias value is limited by the number of bits used to represent the configurable bias. For example, a three-bit configurable bias can have eight different values. In some embodiments, the values represented by the configurable bias are not consecutive. For example, the eight values represented by a three-bit configurable bias are not limited to values 0 to 7. Instead, the bias can be selected from 8 different values. For example, the configurable bias can be selected from eight predetermined values: 1, 3, 5, 7, 9, 11, 15, and 17. In some embodiments, the predetermined values are determined based on the most useful biases. In some embodiments, the predetermined values are at least partially selected to maximize the range of the exponent and minimize the overlap between the ranges of different biases. In some embodiments, the configurable bias is specified by a matrix processor instruction and / or stored in a register (not shown).

[0064] In various embodiments, the matrix processor supports multiple different 8-bit floating-point formats, such as formats 400 and 410. By supporting multiple formats, precision can be exploited in the exponent or mantissa. For example, certain operations such as gradient descent may require additional precision and thus the mantissa requires a larger number of bits. As another example, more bits can be used for the mantissa for operations that cluster values together and do not require additional exponent range. In contrast, for some operations, the range of values may be larger and a larger exponent range is required. Using format 410, fewer bits are dedicated to the mantissa and more bits are dedicated to the exponent. In some embodiments, the format is specified by a matrix processor instruction and can be stored in a register (not shown). In various embodiments, additional floating-point formats not depicted can be supported. For example, a 4-bit mantissa format and a 3-bit exponent format (not shown) can be supported.

[0065] Figure 5 is a block diagram illustrating an embodiment of a 21-bit floating-point format. In the example shown, the floating-point format 500 is a 21-bit floating-point format for representing floating-point numbers using a sign, a mantissa, and an exponent. In some embodiments, node engines such as node engine 300 and matrix processors such as Figure 3 matrix processor 313 utilize a 21-bit floating-point format such as format 500 for certain matrix operations, such as for storing the results (and intermediate results) of matrix multiplication and / or matrix addition. In some embodiments, format 500 is used by accumulators of matrix processors such as Figure 3 output accumulators 329 and 331. For example, if the result is limited to the same 8-bit format, the multiplication result of two 8-bit multiplication operands may cause an overflow error or an underflow error. Using a format greater than 8 bits for the result prevents overflow errors and underflow errors. Similarly, when calculating matrix multiplication using 8-bit matrix elements, using a 21-bit floating-point format to store intermediate and final results prevents overflow errors or underflow errors. Using a result with a bit depth less than 32 bits improves memory usage efficiency. In various embodiments, format 500 with a bit depth of 21 bits is used to optimize both memory usage and accuracy. In some embodiments, format 500 supports a configurable bias. The configurable bias allows for a greater range to improve precision while still maintaining a 21-bit data size. In some embodiments, the configurable bias is specified by matrix processor instructions and / or stored in a register (not shown).

[0066] In the example shown, the 21-bit floating-point format 500 includes a single bit for the sign bit 501, 7 bits for the exponent 503, and 13 bits for the mantissa 505. The sign bit 501, exponent 503, and mantissa 505 together occupy 21 bits and can be used to represent a floating-point number. In some embodiments, a configurable bias is used to bias the exponent. For example, the 7-bit exponent 503 of format 500 allows the exponent 503 to have 128 different values (i.e., values 0 to 127 (inclusive)). Using 7 bits without a bias (or equivalently with the configurable bias set to zero), the exponent 503 can represent exponents with values of 2 0 to 2 127 which correspond to exponent fields with values 0 and 127 respectively.

[0067] In various embodiments, format 500 is used by matrix processors of node engines such as Figure 3 node engine 300 and matrix processor 313 such as Figure 3Used by one or more accumulators such as output accumulators 329 and 331. In some embodiments, registers (not shown) are used to store settings for configurable biases for storing floating-point numbers in a particular accumulator. In some embodiments, multiple 21-bit formats can be used (e.g., with different bit allocations for the exponent field and the mantissa field), and a particular format is specified by matrix processor instructions. The value of the configurable bias can be specified using matrix processor instructions and / or stored in a register.

[0068] Although Figure 5 depicts a 21-bit floating-point format that can be used by the accumulators of a matrix processor such as Figure 3 output accumulators 329 and 331, formats with alternative bit depths can be used. For example, depending on the operation requirements, such as requirements for preventing loss of precision, when supporting operations on certain 16-bit floating-point operations, a 27-bit floating-point format can be used to prevent loss of precision in the quantized results. As an example, a 27-bit floating-point format can include a sign bit for the unit, 9 bits for the exponent, and 17 bits for the mantissa. The 27-bit floating-point format can be used to accumulate multiplication operations on 16-bit floating-point operands. In some embodiments, 16-bit floating-point operands are represented using a unit for the sign bit, 8 bits for the exponent, and 7 bits for the mantissa.

[0069] Figure 6 is a flowchart illustrating an embodiment of a process for performing matrix calculations. Figure 6 The process of Figure 2 is used by a training platform such as Figure 2 training platform 223 to perform matrix calculations through one or more node engines such as Figure 3 node engine 225 or Figure 6 node engine 300. In some embodiments, the training platform receives one or more matrix calculation operations and parallelizes these operations across different node engines. Then, each node engine can also parallelize its operations across different matrix processors. Optionally, the results can be combined at one or more node engines to determine the results, such as the weight matrix of a machine learning model. In some embodiments, Figure 1 the process of

[0070] is performed as part of step 105 of Figure 2A training platform such as training platform 223 receives it. The training platform processes the computing instructions and performs necessary work partitioning and allocation for different node engines. For example, at the server of the training platform that starts a machine learning training process, a computing instruction to convolve an image with a filter is received. In some embodiments, the instruction may include parameters required to execute the computing instruction, including the operations and operands involved. For example, the instruction may include the size of the input operands (e.g., the size of each input matrix), the starting address of each input matrix, the stride parameter, the padding parameter, and / or matrix, vector, and / or post-processing commands. For example, the computing instruction may describe the image data size (e.g., 96×96, 1920×1080, etc.) and bit depth (e.g., 8-bit, 16-bit, etc.) as well as the filter size and bit depth, etc. In many scenarios, the matrix for matrix calculation may be larger than the matrix that can be accommodated inside the matrix processor, so that additional processing can be performed to subdivide the calculation so that it can be executed by different node engines or matrix processors.

[0071] At 603, the matrix operations and operands are determined. In the case where one or more matrices of the computing instruction received at 601 are larger than the input matrix of the matrix processor, the computing instruction at 601 is split into small component operations. At 603, the matrix operations and the operands corresponding to the smaller component operations are determined, and may include slicing, splitting, or dividing the original matrix operands into smaller matrices and performing matrix operations on the smaller matrices. The results of the matrix operations on the smaller matrices can be combined to complete the computing instruction received at 601. Different node engines and matrix processors can be assigned to execute different components of the computing instruction. In some embodiments, the elements of the matrix operand can be converted to an 8-bit floating-point format or targeted to be converted to an 8-bit floating-point format. The node engine uses an 8-bit floating-point format such as Figure 4 Format 400 or format 410 to increase the processing and performance bandwidth and the power efficiency of the matrix processor. In some embodiments, a configurable bias for the corresponding floating-point format is selected or will be selected. For example, a format with a high-precision mantissa is selected for gradient descent operations.

[0072] In various embodiments, the larger matrix is sliced into smaller two-dimensional matrices, where the size is limited to the appropriate dimensions of the matrix processor. In some embodiments, the sliced matrix is a smaller matrix with addresses pointing to the elements of the original matrix. The sliced matrix can be serialized into a vector for processing. In some embodiments, different slices of the matrix can overlap with previous slices. In various embodiments, the matrix can be sliced only at boundaries corresponding to multiples of the read buffer size. For example, in the case where the size of each read buffer is 8 bytes, each row of the sliced matrix must start at an address that is a multiple of 8. If the matrix fits within the computing array, no slicing is required (i.e., the matrix slice used is just the original matrix).

[0073] At 605, matrix operations are allocated and executed. For example, matrix operations corresponding to the matrix operations and operands determined at 603 are allocated to one or more node engines and one or more matrix processors of the node engines. In various embodiments, the matrix operations are executed by one or more matrix processors using matrices of 8-bit elements. The values of the elements of the matrix result are accumulated in a 21-bit floating-point format, a 27-bit floating-point format, or other suitable floating-point format. In various embodiments, the matrix result can be shifted out of the matrix processor in one of several formats including an 8-bit floating-point format, a 16-bit floating-point format, and a 32-bit floating-point format. In various embodiments, each node engine can execute multiple matrix operations in parallel by utilizing multiple matrix processors.

[0074] In some embodiments, references to matrix operands are allocated along with operations on the node engines. In this way, the node engines can perform data reads to load the corresponding elements of the sliced matrix. In some embodiments, the node engines linearize the sliced matrix for loading into memory and / or registers, where the input matrix can then be sent to the matrix processors. In some embodiments, the control unit of the node engine coordinates the scheduling, issuing, and synchronization of operations, which include: loading sliced matrix operands (including specifying strides, padding, and other parameters for addressing matrix operands); and operating the matrix processors. Once a matrix operation is issued to the matrix processor, the matrix processor takes a certain number of clock cycles to complete the matrix operation. In some embodiments, the matrix processor uses Figure 7 and / or Figure 8 processes to perform the matrix operations.

[0075] At 607, post-processing is performed. In some embodiments, the post-processing can be performed by the node engine and can include additional vector operations performed after the matrix operations are completed. The post-processing operations can be performed by a post-processing unit of the node engine such as a vector processor or a vector computing unit. In some embodiments, the vector post-processing includes: performing complex operations such as arithmetic operations, expansion, normalization, and / or applying an activation function such as a rectified linear unit (ReLU) function to each element of the vector. In some embodiments, the elements of the vector can be converted / formatted into 8-bit elements, 16-bit elements, or 32-bit elements depending on the required precision. In various embodiments, the results of the distributed matrix operations of each node engine can be sent back to the training platform server or redirected by the training platform server and used for other processing. For example, the results of the matrix operations allocated and executed at 605 can be combined and used as operands for additional vector or matrix operations. After post-processing is initiated at 607, the processing loop returns to 601 to receive additional computation instructions. In some embodiments, the post-processing does not need to be completed before the processing loop returns to 601 to obtain additional computation instructions.

[0076] Figure 7 is a flowchart showing an embodiment of a process for performing matrix calculations. Figure 7 The process of Figure 3 is performed by matrix processors such as matrix processors 313 and 351 to 357 of a node engine 300 to perform matrix calculations. In some embodiments, each matrix processor of the node engine can perform Figure 7 the process in parallel. For example, matrix processors 313 and 351 to 357 each perform Figure 7 the process on different matrix parameters in parallel, although each matrix processor may be at a different processing step to stagger the completion of their respective operations. In some embodiments, the process is used to perform convolution using a data matrix and a weight matrix. In some scenarios, the input matrix is a slice of a larger matrix. In various embodiments, Figure 7 the process can be initiated by a matrix calculation instruction via a control unit. The instruction can specify two matrix operands (e.g., memory or register locations of the data and weight matrices), a configurable bias, a floating-point format, and a specified accumulator to store the matrix calculation result. In some embodiments, the specified accumulator is zeroed before the matrix calculation begins. In some embodiments, the specified accumulator is Figure 3 output accumulator 329 or 331 of Figure 6 At 605 of Figure 7 the process is performed.

[0077] At 701, a data input matrix is received. For example, the elements of the data input matrix corresponding to training sensor data are linearized and stored in the data input array of the matrix processor. In some embodiments, the data input matrix is stored in a data input array such as Figure 3 data input array 321 of matrix processor 313 of Figure 4format 400 or 410 and includes a configurable bias. The configurable bias can be specified by matrix instructions and / or registers. The received data input matrix can be received from a register or from a memory such as SRAM. In some embodiments, one or more reads are issued to load the entire data input matrix into the matrix processor, but the entire matrix is not immediately available. For example, for a sliced matrix, some rows (or columns) of data may require additional latency before the data is available. Thus, the data of the data input array may arrive one by one. In some embodiments, a single read is sufficient to load the entire data input matrix. In some embodiments, the data input matrix is a gradient input matrix.

[0078] At 703, a weight input matrix is received. For example, the elements of the weight input matrix corresponding to the machine learning weights of a filter are linearized and stored in the weight input array of the matrix processor. In some embodiments, the weight input matrix is stored in a weight input array such as Figure 3 the weight input array 323 of the matrix processor 313 as described. Each weight input array is capable of storing the entire linearized matrix for processing by the corresponding matrix processor by the matrix calculation unit. Thus, a matrix processor capable of multiplying two 8×8 matrices uses a weight input array capable of storing all 64 elements of the input 8×8 weight matrix. For example, in some embodiments, each weight input array is 64 bytes and stores each element as an 8-bit floating point number. The format of the floating point number can use Figure 4 format 400 or 410 and includes a configurable bias. The configurable bias can be specified by matrix instructions and / or registers. The received weight input matrix can be received from a register or from a memory such as SRAM. In some embodiments, one or more reads are issued to load the entire weight input matrix into the matrix processor, but the entire matrix is not immediately available. For example, for a sliced matrix, some rows (or columns) of weight data may require additional latency before the weight data is available. Thus, the weight data of the weight input array may arrive one by one. In some embodiments, a single read is sufficient to load the entire weight input matrix. In some embodiments, the weight input matrix is a gradient input matrix.

[0079] At 705, a pair of vector parameters are loaded into the matrix calculation unit. From each input matrix, the vector corresponding to the rows and the vector corresponding to the columns are loaded as input parameters into such as Figure 3a matrix calculation unit such as matrix calculation unit 325. As part of the loading process, the column vectors are copied across the entire matrix calculation unit, and the row vectors are copied down the entire matrix calculation unit. For example, an entire vector corresponding to a column of the weight input matrix is loaded into the calculation unit. Each element of the column vector is copied across the entire row. Thus, each column of the 8×8 matrix calculation unit receives the same 8-element column vector, and the values of each cell in a row of the matrix calculation unit are the same. Similarly, an entire vector corresponding to a row of the data input matrix is loaded into the calculation unit, and each element of the row vector is copied down the entire column. Thus, each row of the 8×8 matrix calculation unit receives the same 8-element column vector, and the values of each cell in a column of the matrix calculation unit are the same. For an 8×8 matrix calculation unit, one-eighth of the input matrix elements are loaded. At 705, the offload vector pairs from each input matrix are loaded into the matrix calculation unit. The next available columns and rows are loaded from the input weight and data matrices through each subsequent loop of step 705. Thus, an 8×8 matrix requires at least 8 cycles to complete loading, and a 4×4 matrix requires at least 4 cycles to complete loading.

[0080] At 707, the values of the loaded vectors are multiplied. For each calculation unit of a matrix calculation unit such as Figure 3 calculation unit 327, matrix multiplication is performed using the elements loaded at the corresponding calculation unit. In various embodiments, multiplication is performed on two 8-bit floating-point values and stored as a high-precision floating-point value to prevent overflow and maintain precision. In some embodiments, the high-precision floating-point format is Figure 5 the 21-bit floating-point format. In some embodiments, the high-precision floating-point format is a 27-bit floating-point format to further reduce the loss of accuracy of the quantization result. For an 8×8 matrix calculation unit, each of the 64 calculation units performs matrix multiplication.

[0081] At 709, the multiplication results are accumulated into a specified accumulator. For example, the multiplication results at 707 for each calculation unit are each accumulated into one of the accumulators in the accumulator of the matrix processor. In some embodiments, the matrix processor includes such as Figure 3more than one accumulator, such as the two output accumulators 329 and 331. This is beneficial to enable the matrix processor to interleave the operations of different matrix operations. In some embodiments, each computing unit includes an accumulator that adds the current value of the element corresponding to the computing unit in the accumulator to the result of the matrix multiplication of the unit. In various embodiments, the size of the accumulator is designed to store the accumulated result of each element of the matrix. Thus, each accumulator of an 8×8 matrix computing unit has at least 64 elements. In some embodiments, similar to the multiplication result at 707, the elements of the accumulator use a floating-point value with a higher bit than the input of the matrix processor to prevent overflow and maintain precision. In some embodiments, the high-precision floating-point format is Figure 5 a 21-bit floating-point format or another high-precision floating-point format. In some embodiments, the accumulator for an 8×8 matrix computing unit is 168 bytes to allow 64 elements, each storing a 21-bit floating-point number.

[0082] At 711, it is determined whether there are additional vectors remaining for the matrix operation. For example, to multiply two matrices, at most one column from the weight input matrix and one row from the data input matrix are loaded per clock cycle. To complete the entire matrix multiplication, every column and every row must be loaded. An 8×8 matrix requires at least 8 cycles to fully load the two input matrices into the matrix computing unit. Similarly, a 4×4 matrix requires at least 4 cycles to fully load the two input matrices into the matrix computing unit. In the case where there are additional vectors to be loaded remaining, the process returns to 705. In the case where there are no additional vectors to be loaded remaining (both entire input matrices have been loaded), the matrix multiplication is completed and the process continues to 713.

[0083] At 713, the matrix result is loaded from the specified accumulator into the output array. Since the matrix calculation is completed, the matrix result is stored in the specified accumulator. In some embodiments, the elements of the matrix are stored as 21-bit floating-point values in the specified accumulator. Thus, for an 8×8 matrix, the accumulator stores 64 values and has a size of 168 bytes. In some embodiments, multiple move operations are required to move the result from the accumulator to the output array, such as Figure 3 the output array 315. In some embodiments, the width of the output array and the bus to the output array is 64 bytes. The accumulator result is converted from a 21-bit floating-point value to a 16-bit floating-point value, which can be stored in two 64-byte components. Using the 8×8 result matrix as an example, two move operations are required to move the result from the accumulator of the matrix processor. For example, an upward move operation is used to move the high bits (corresponding to 32 elements of the matrix) of the accumulator into the 64-bit output array as 16-bit floating-point values. Once moved in the output array, 32 elements can be stored in, such as Figure 3in a register such as one of the registers in the post - processing unit register file 307 or moved to memory. Subsequently, the down - shift operation is used to shift the lower bits (corresponding to the remaining 32 elements of the matrix) of the accumulator as a 16 - bit floating - point value into the 64 - bit output array. Once in the output array, the remaining 32 elements can be stored in registers. In various embodiments, two or more operations are required to move the matrix result out of the matrix processor. By converting the 21 - bit floating - point value to a 16 - bit floating - point value, only two move operations are required. In some embodiments, these values can be moved out as 8 - bit floating - point values, 16 - bit floating - point values, or 32 - bit floating - point values. In the example described, the values are moved out as 16 - bit values for subsequent processing by a post - processing unit such as Figure 3 the post - processing unit 317. In some embodiments, the post - processing unit is a vector calculation engine. In various embodiments, the output array is connected to the accumulator of each matrix processor of the node engine and acts as a multiplexer to receive the moved results (e.g., up - shift and down - shift instructions) from different matrix processors.

[0084] Figure 8 is a flowchart illustrating an embodiment of a process for performing multiple interleaved matrix calculations. Figure 8 The process is used by matrix processors such as Figure 3 the matrix processors 313 and 351 to 357 of the node engine 300 to interleave multiple matrix calculations, such as two matrix multiplication operations. Each interleaved matrix calculation in the interleaved matrix calculations can be implemented using multiple intermediate matrix multiplications, where the results of the intermediate multiplications are used to compute a larger matrix calculation. To increase processing bandwidth and efficiency, the results of each intermediate matrix multiplication are stored in the matrix processor and are not cleared when interleaving alternative matrix operations. Different matrix operations can be different, and each matrix operation has non - overlapping matrix operands.

[0085] In some embodiments, each matrix processor of the node engine can process more than one matrix operation at a time, and one matrix operation corresponds to each output accumulator of the matrix processor. In some embodiments, the ability to interleave multiple matrix operations allows matrix multiplication operations to be performed on very large matrices. The larger matrix is sliced into smaller matrices that fit the input array of the matrix processor, and the results of the matrix multiplications of the smaller matrices are combined. In various embodiments, for example, while waiting for a memory read to complete, the ability to interleave multiple matrix operations increases the bandwidth and performance of the matrix processor by utilizing the matrix calculation unit. Thus, when the input operands of the outstanding matrix operations of the first set of related matrix operations are not available (e.g., due to the latency of a memory read) but the input operands of the outstanding matrix operations of the second set of related matrix operations are available, the second set of related matrix operations can utilize the matrix calculation unit. By utilizing multiple accumulators, the matrix calculation unit can switch between multiple matrix calculations by storing intermediate results in accumulators dedicated to specific sets of related matrix operations. In some embodiments, the data input array is Figure 3 the data input array 321, the weight input array is Figure 3 the weight input array 323, and the multiple accumulators are Figure 3 the output accumulators 329 and 331. Although two accumulators are shown with respect to Figure 3 the matrix processor 313, additional accumulators can be included to allow interleaving of additional matrix operations.

[0086] Figure 8 The process of Figure 7 is a special variant of the process of Figure 7 which utilizes multiple weight input array operands, multiple data input array operands, and multiple output accumulators to support interleaving of two matrix multiplication operations. As described with respect to Figure 7 the process of Figure 8 implements the Figure 7 steps in a similar manner, which steps include: loading column vectors across the matrix calculation unit, loading row vectors down the matrix calculation unit, multiplying the operands by the calculation unit, and accumulating the multiplication results in the specified accumulators; taking care not to mix or erase the intermediate results of the two interleaved matrix operations. In some embodiments, the process of Figure 6 is performed at 605 of Figure 8 .

[0087] At 801, it is determined whether the matrix processor can receive an additional matrix operation instruction. At Figure 8In the example, the matrix processor is capable of interleaving two matrix operations. Determine whether there are currently two matrix operations being executed. If the matrix processor can receive additional matrix operation instructions, the process continues to 803. For example, the matrix processor can receive additional matrix operation instructions because it is only in the middle of processing a single matrix operation or is idle and not processing any matrix operations. If the matrix processor cannot receive additional matrix operation instructions, the process loops back to 801 until the matrix processor is available to receive new matrix operation instructions. For example, the matrix processor is currently in the middle of processing two matrix operations and cannot receive another operation until at least one of the current operations is completed. In some embodiments, a ready signal is issued to the control unit to signal that the matrix processor is ready to receive additional instructions.

[0088] At 803, the matrix processor receives a matrix instruction and issues a read request for the associated matrix operation. For example, the matrix processor receives a matrix multiplication instruction with two operands corresponding to two input matrices. A read is issued for the values of the matrix operands. These values can be read from registers and / or memory. For example, the matrix parameters can specify addresses in registers and / or memory. In some embodiments, since a memory read may take multiple clock cycles to make the data available, the memory read may stall the matrix calculation. In some embodiments, since the matrices are not stored sequentially in memory, multiple memory reads may be issued. This may be the result of slicing a larger matrix into smaller matrix operands.

[0089] In some embodiments, the received instruction specifies a particular accumulator to store the matrix result. To interleave multiple matrix operations, each operation utilizes its own accumulator. The accumulator is specified to store intermediate and final matrix results. In some embodiments, the specified accumulator uses a higher-precision floating-point format than the format used for the input operands to store the intermediate results. When the results are quantized, the higher-precision format minimizes the loss of accuracy.

[0090] In various embodiments, when the data corresponding to the matrix operands is available, the values are received and prepared for the matrix processor. In some embodiments, the matrix operands are too large for the matrix processor and multiple intermediate matrix operations are performed to complete the matrix instruction. In the case where the data is not available, the matrix calculation unit may stop and idle. Instead of remaining idle, as long as the data for the second operation is available, the second matrix operation can be executed.

[0091] At 803, processing continues to both 801 and 805. Processing loops back to 801 to fetch new instructions while also continuing to 805 to execute the instructions received at 803. In various embodiments, fetching new instructions occurs in parallel with processing the current matrix operation. In some embodiments, the two processing branches to 801 and 805 are implemented using a pipelined approach.

[0092] At 805, it is determined whether data is ready for the current matrix operation. For example, the elements to be loaded from the matrix operands of the current matrix operation must be available for loading into the computing units of the matrix computing unit. In some embodiments, the data loaded into the matrix computing unit is a slice of the matrix operand sized for the input array of the matrix computing unit. For the weight input array, the outstanding columns of elements must be ready. For the data input array, the outstanding rows of elements must be ready. In the case where the elements of the weight columns and data rows of the current matrix operation are available, processing continues to 807. In the case where the outstanding elements of the current matrix operation are not available, processing continues to 813. For example, the outstanding elements may not be available due to the latency of memory reads and / or cache misses. Instead of stopping while waiting for the data to become available, the matrix computing unit can potentially be used for alternative matrix operations.

[0093] At 807, the values from the weight columns and data rows of the current matrix operation are loaded into the respective computing units, a computational operation is performed on the values, and the result of the computation is accumulated into a specified accumulator. In some embodiments, the computational operation is a multiplication operation that multiplies elements from two different matrices. In some embodiments, the process at 807 is described with respect to Figure 7 steps 701, 703, 705, 707, 709, and / or 711. For example, the values are loaded as 8-bit floating-point values with a configurable bias. The result of a computation such as multiplication and accumulation is stored in the first accumulator in a 21-bit floating-point format. In some scenarios, additional configuration related to the matrix operation is performed at 807, such as clearing the accumulator, determining the floating-point format, and / or determining the configurable bias of the floating-point format, etc.

[0094] At 809, it is determined whether the matrix instruction of the current matrix operation is complete. In the case where the matrix instruction is complete, processing continues to 811. In the case where the matrix instruction is not complete, processing continues to 805, where it is determined whether additional data for the current matrix operation is ready to be loaded and processed by the matrix computing unit. In some embodiments, the process at 809 is described with respect to Figure 7 step 711.

[0095] In some alternative embodiments (not shown), in the case where a matrix instruction is not completed, the process continues to 813, where it is determined whether an alternative matrix operation is pending and whether the data for the pending alternative matrix operation is ready to be loaded and processed by the matrix calculation unit. In this alternative embodiment, instead of completing the current matrix operation, as long as the data is available, the matrix calculation unit continues to alternate back and forth between two different matrix operations as long as there are two concurrent matrix operations.

[0096] At 811, the matrix result stored in the designated accumulator is loaded into the output array. Since some embodiments store the resulting matrix using a higher-bit-depth floating-point format such as a 21-bit floating-point format or a 27-bit floating-point format, multiple move instructions may be required to shift the result out of the matrix processor. In some embodiments, the matrix result is shifted via the output array into two 64-byte registers by first converting the matrix elements to 16-bit floating-point values. In some embodiments, the process at 811 is described with respect to Figure 7 step 713. The process loops back to step 805, where if pending, the matrix processor is ready to start a matrix operation or make progress on an alternative matrix operation.

[0097] In some alternative embodiments (shown as dashed lines), the process continues to 813, where it is determined whether an alternative matrix operation is pending and whether the data for the pending alternative matrix operation is ready to be loaded and processed by the matrix calculation unit. In this alternative embodiment, once the current matrix instruction is completed, in the case where there is a pending alternative matrix operation to be completed, the matrix calculation unit switches to the alternative matrix operation.

[0098] At 813, it is determined whether an alternative matrix operation is pending and whether data for the pending alternative matrix operation is ready to be loaded and processed by the matrix calculation unit. For example, in the case where a second matrix operation is received at 803 while processing a first matrix operation, the pending second matrix operation that is completed issues a read for its corresponding matrix arguments. It is determined whether there is a pending second alternative matrix operation and whether its data is ready to be loaded into the matrix calculation unit. When the operand data for the alternative matrix operation is available, the process proceeds to 815. In some embodiments, the operand data is a slice of a larger operand matrix whose size is designed for the input array of the matrix calculation unit. For the weight input array, the pending element columns must be ready. For the data input array, the pending element rows must be ready. In the case where there is no pending alternative matrix operation or the pending elements of the alternative matrix operation are not available, the process proceeds to 805. For example, the pending elements may not be available due to latency from memory reads and / or cache misses. Instead of stopping while waiting for the data to become available, the availability of the data corresponding to the current matrix operation is checked again. The first matrix operation with available data loads its data into the matrix calculation unit for processing.

[0099] At 815, the matrix processor including the matrix calculation unit is switched to perform processing on the pending completed alternative matrix operation. The alternative matrix operation is now designated as the current matrix operation, and the previous current matrix operation is designated as the alternative matrix operation. Since the first matrix operation may have stopped (or in some embodiments, completed), the matrix calculation unit is now dedicated to the pending completed second matrix operation. In various embodiments, as appropriate, the corresponding output accumulator is designated as the source of the previous intermediate result and the destination for accumulating intermediate and final results. The process proceeds to 807, where the calculation progress for the newly designated current matrix operation is performed.

[0100] Although, for the purposes of clear understanding, some of the foregoing embodiments have been described in detail, the present invention is not limited to the details provided. There are many alternative ways to implement the present invention. The disclosed embodiments are illustrative and not restrictive.

Claims

1. A method, comprising: receiving matrix processor instructions from a control unit, wherein the matrix processor instructions specify a first floating-point operand and a second floating-point operand; loading a first half of the first floating-point operand into a first processing element and a second processing element; loading a second half of the first floating-point operand into a third processing element and a fourth processing element; loading a first half of the second floating-point operand into the first processing element and the third processing element; loading a second half of the second floating-point operand into the second processing element and the fourth processing element; determining a first floating-point multiplication result, a second floating-point multiplication result, a third floating-point multiplication result, and a fourth floating-point multiplication result corresponding to each of the first processing unit, the second processing unit, the third processing unit, and the fourth processing element; and storing the first floating-point multiplication result, the second floating-point multiplication result, the third floating-point multiplication result, and the fourth floating-point multiplication result in an output accumulator.

2. The method according to claim 1, further comprising: using a vector calculation unit to add together the first floating-point multiplication result, the second floating-point multiplication result, the third floating-point multiplication result, and the fourth floating-point multiplication result.

3. A microprocessor system, comprising: a matrix calculation unit including a plurality of processing elements; and a control unit configured to provide matrix processor instructions to the matrix calculation unit; wherein the matrix processor instructions specify floating-point operands formatted using a first floating-point representation format, the matrix calculation unit accumulates intermediate result values calculated using the floating-point operands, and the intermediate result values are in a second floating-point representation format.

4. The system according to claim 3, wherein the first floating-point representation format is an 8-bit floating-point format.

5. The system according to claim 3, wherein the second floating-point representation format is a 21-bit floating-point format.

6. The system according to claim 5, wherein the second floating-point representation format allocates 1 bit for a sign bit, 7 bits for an exponent field, and 13 bits for a mantissa field.

7. The system according to claim 3, wherein compared with the first floating-point representation format, the second floating-point representation format uses a greater number of bits for storing floating-point numbers.

8. The system according to claim 7, wherein the greater number of bits prevents an overflow error from occurring and prevents an underflow error from occurring.

9. The system according to claims 3 to 8, wherein the matrix calculation unit outputs the accumulated intermediate result values as an output formatted in a third floating-point representation format.

10. The system according to claim 9, wherein the third floating-point representation format is a 16-bit floating-point format.

11. The system according to claims 3 - 10, wherein the matrix calculation unit is configured to receive two matrix operands, wherein the floating-point operand represents one of the two matrix operands.

12. The system according to claim 11, wherein at least one of the two matrix operands specifies a register value or a memory address location.

13. The system according to claim 11, wherein the two matrix operands are formatted as linearized matrices.

14. The system according to claim 11, wherein the data values of the two matrix operands are stored in a weight input array and a data input array of the matrix processor using the first floating-point representation format.

15. The system according to claims 3 to 14, wherein each of the plurality of processing elements includes a plurality of floating-point accumulators.

16. The system according to claims 3 to 15, wherein the matrix processor instruction specifies a designated accumulator for storing intermediate results of the matrix calculation unit.

17. The system according to claims 3 to 16, wherein a first instruction is used to retrieve a first portion of the matrix result of the matrix processor instruction, and a second instruction is used to retrieve a second portion of the matrix result of the matrix processor instruction, and wherein the matrix result uses the second floating-point representation format.

18. The system according to claim 17, wherein the retrieved first portion of the matrix result and the retrieved second portion of the matrix result use a third floating-point representation format.

19. The system according to claims 3 to 18, wherein each of the plurality of processing elements includes a floating-point multiplier and an accumulator, and is configured to perform floating-point multiplication operations in parallel with other processing elements.

20. A method comprising: receiving a matrix processor instruction from a control unit, wherein the matrix processor instruction specifies a first floating-point matrix operand and a second floating-point matrix operand, and wherein the first floating-point matrix operand and the second floating-point matrix operand are formatted using a first floating-point representation format; receiving data values of the first floating-point matrix operand and the second floating-point matrix operand; storing the data values of the first floating-point matrix operand in a data input array; storing the data values of the second floating-point matrix operand in a weight input array; selecting a single row from the first floating-point matrix operand and a single column from the second floating-point matrix operand; copying the selected single row at each row of a matrix calculation unit of the matrix processor; and copying the selected single column at each column of the matrix calculation unit.

21. The method according to claim 20, further comprising: accumulating intermediate result values at each of a plurality of processing elements of the matrix calculation unit, wherein the intermediate result values are calculated using elements of the single row and elements of the single column.

22. A microprocessor system comprising: a matrix processor, wherein the matrix processor is configured to receive matrix processor instructions that specify floating-point operands formatted using a first floating-point representation format, and the matrix processor is configured to accumulate a matrix result using a second floating-point representation format; An output array, configured to store the matrix result using a third floating-point representation format; A post-processing unit, configured to receive a second floating-point operand using the third floating-point representation format; A control unit, configured to provide post-processing instructions to the post-processing unit and matrix processor instructions to the matrix processor; And A post-processing register file, wherein the post-processing instructions specify post-processing unit operands stored in the post-processing register file.