Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

43 results about "Mixed precision" patented technology

Mixed-precision matrix multiplication

Systems and techniques for providing mixed-precision matrix multiplication in multi-chiplet processors recognize different precision formats of matrices to be multiplied based on, e.g., parameters provided with instructions or start and end memory locations of the matrices. A plurality of different multiplication chains are provided for different formats such that mixed-precision matrix multiplication can be performed using multiplication chains configured to handle multiplication of different precision formats. The multiplication chains are automatically selected based on the precision formats of the matrices to be multiplied, enabling programmers to utilize the chains without having to directly access the individual multiplication chains.
Owner:ADVANCED MICRO DEVICES INC

High-speed rail platform safety determination method and system based on mixed precision inference

PendingCN122333241ARelation graphAlgorithm
The application relates to the technical field of intelligent reasoning and safety judgment, and discloses a high-speed rail platform safety judgment method and system based on mixed-precision reasoning, which comprises the following steps: acquiring multi-source semantic observation records and generating a safety observation element set; constructing a platform safety relation graph; determining relation conflict density, closed residual error, cross-source divergence degree and reasoning difficulty level; generating a mixed-precision bit width scheduling table; performing graph relation reasoning to obtain an initial safety judgment vector and an initial judgment boundary quantity; in step 6, a final judgment vector is determined; and in step 7, a locked safety level is determined and a safety judgment package is output. The application realizes mixed-precision safety judgment and locked output driven by multi-source semantic relation of a high-speed rail platform.
Owner:XIAMEN SILICON TECHNOLOGY CO LTD

Novel translation model reasoning method and system based on rwkv

PendingCN122287660AImplement reasoning methodsscale upComputation complexityTheoretical computer science
The RWKV-based novel translation model inference method and system includes the following steps: 1) Collecting novel translations and extracting parallel corpora, and using dynamic MicroBatch concatenation technology for sequence compression; 2) Introducing a lightweight group query attention mechanism on the basis of the RWKV architecture, directly obtaining KV information from the Embedding layer to build the model; 3) Employing a sublinear complexity hybrid parallel training mode, combined with a global scalar scaling FP16 mixed precision strategy for training; 4) Applying a hierarchical distributed heterogeneous architecture, offloading the optimizer to low-performance devices and performing gradient compression transmission; 5) Outputting the translation using a joint decoder and dynamic batch inference technology. This invention, through the above method and system, effectively reduces the computational complexity and memory usage in the long text translation process, improves model training efficiency and inference throughput, and significantly improves the translation efficiency and contextual coherence of ultra-long texts.
Owner:LIAONING UNIVERSITY

Transform hardware acceleration method and accelerator based on hybrid precision quantization and huffman coding

This invention discloses a hardware acceleration method and accelerator for Transformer based on mixed-precision quantization and Huffman coding. The acceleration method includes: using a genetic algorithm to obtain several configuration schemes for mixed-precision quantization of Transformer network layers; performing mixed-precision quantization on each Transformer network layer based on each quantization configuration scheme to obtain a corresponding KL divergence; training a multilayer perceptron to obtain a quantization configuration prediction network using the quantization configuration scheme and the corresponding KL divergence as the output label and input feature, respectively; receiving a user-set target KL divergence value, using the quantization configuration prediction network to obtain the corresponding quantization configuration scheme, and performing mixed-precision quantization on each network layer based on the quantization scheme; and using Huffman coding to encode and compress all quantization weights before on-chip storage. This invention can reduce storage and computational overhead while maintaining model accuracy.
Owner:HUNAN NORMAL UNIVERSITY

Mixed-precision matrix multiplication

PCT designated stageWO2026143019A1Computational scienceMemory address
Systems and techniques for providing mixed-precision matrix multiplication in multi-chiplet processors recognize different precision formats of matrices to be multiplied based on, e.g., parameters provided with instructions or start and end memory locations of the matrices. A plurality of different multiplication chains (208, 210, 212) are provided for different formats such that mixed-precision matrix multiplication can be performed using multiplication chains configured to handle multiplication of different precision formats. The multiplication chains are automatically selected based on the precision formats of the matrices to be multiplied, enabling programmers to utilize the chains without having to directly access the individual multiplication chains.
Owner:ADVANCED MICRO DEVICES INC

Mixed-precision computations with state compression

A technique for matching the throughput between writing into and reading from a memory can include receiving, in parallel, computational results in a high precision format for storing into the memory at a first frequency, and storing the computational results in the memory. The technique may further include rounding the computational results using round-to-the-nearest-even or stochastic rounding to down-convert the computational results from the high precision format to a low precision format in parallel, and outputting the computational results in the low precision format in parallel from the memory at a second frequency.
Owner:AMAZON TECH INC

Multimedia data compression method and device for narrowband transmission

The application relates to the technical field of data compression, and discloses a multimedia data compression method and device for narrowband transmission. The method comprises the following steps: performing double-threshold contour extraction on an input image to obtain main contour data and secondary contour data; performing mixed-precision chain code coding on the main contour data and the secondary contour data to obtain compressed contour chain codes; performing parameterization decomposition and superframe coding on an input speech signal to obtain compressed speech parameters; allocating a transmission bandwidth according to the compressed speech parameters, and packing the compressed contour chain codes and the compressed speech parameters based on the transmission bandwidth to obtain a target transmission packet sequence. The application ensures the time synchronization and transmission reliability of image and speech data, solves the problems of multi-modal data asynchronous transmission and high bit error rate in a narrowband environment, and guarantees correct reconstruction and playing at a receiving end.
Owner:SHENZHEN YUNTIAN INTELLIGENT COMM CO LTD

Temporally amortized supersampling using a kernel splatting network

One embodiment provides a graphics processor comprising processing resources configured to perform a supersampling anti-aliasing operation via a mixed precision convolutional neural network. The processing resources include circuitry configured to receive, at an input block of a neural network model, a data including previous frame data, current frame data, jitter offset data, and velocity data, pre-process the data to generate pre-processed data, provide pre-processed data to a feature extraction network of the neural network model and an output block of the neural network model, process the first pre-processed data at the feature extraction network via one or more encoder stages and one or more decoder stages, output tensor data from the feature extraction network to the output block, and generate an anti-aliased output frame via the output block based on the current frame data and the tensor data output from the feature extraction network.
Owner:INTEL CORP

In-memory computing circuits, data processing methods, and chips based on mixed precision.

This invention relates to the field of integrated circuits and discloses a mixed-precision in-memory computing circuit, data processing method, and chip. The invention includes: a storage array, an index memory, and a multiply-accumulate calculation circuit. The storage array includes a first data element stored in a first precision format and / or a second data element stored in a second precision format; wherein the bit width of the first precision format is greater than that of the second precision format; the storage locations of the storage array are divided according to the bit width of the second precision format; the first data element occupies multiple consecutive storage locations; the read bit width of the storage array is an integer multiple of the bit width of the first precision format; the index memory stores the position index information of the first data element in the storage array; the multiply-accumulate calculation circuit is used to determine the precision format corresponding to the data read from the storage array based on the position index information, and performs a multiply-accumulate operation corresponding to the precision format, outputting the calculation result. This solves the compatibility problem between mixed-precision data storage and computing circuits.
Owner:SIMINWAY (SHANGHAI) INTEGRATED CIRCUIT CO LTD

Pedestrian fall detection method based on mixed precision quantization and storage medium

ActiveCN116071826Bprocessing speedImprove fall detection speedAlgorithmSimulation
The application discloses a pedestrian fall detection method based on mixed precision quantization and a storage medium, relates to the technical field of computer vision, and solves the technical problem that the mixed quantization precision model used in the existing pedestrian fall detection is difficult to achieve the best balance in model volume, precision and processing speed. The method comprises the following steps: S1, encoding a pedestrian neural network and quantizing the pedestrian neural network into a basic neural network; S2, initializing the basic neural network to obtain K different individuals, and performing selection and crossover operations; S3, performing game mutation operations; and S4, repeating steps S2-S3 until the iteration number reaches a preset maximum iteration number or the iteration termination condition is met, and obtaining an optimal mixed precision quantized pedestrian neural network model. The genetic algorithm improved by the game theory algorithm can make the model after mixed precision quantization achieve the best balance in model volume, precision and processing speed.
Owner:SHENZHEN ICOMM SEMICON CO LTD

A heterogeneous computing power cooperative scheduling system and method for mixed precision training

ActiveCN121579206BExecution planData stream
The application discloses a heterogeneous computing power cooperative scheduling system and method for mixed precision training, and belongs to the technical field of artificial intelligence calculation. The system comprises: a calculation graph analysis and operator image module, which is used for analyzing and dividing a model calculation graph and extracting operator features; a heterogeneous hardware capability sensing and matching module, which is used for managing the performance profile and real-time state of heterogeneous hardware in a cluster, and matching optimal execution hardware for each calculation partition; a data flow coordination and pipeline parallel controller, which is used for generating a global execution plan, managing cross-device data dependency and communication, and optimizing execution efficiency through communication and calculation overlap. The application solves the problem of inefficient scheduling of mixed precision training in a heterogeneous environment, realizes automatic and accurate mapping of calculation tasks to heterogeneous hardware, significantly improves training speed and reduces training cost, and improves the overall resource utilization of the cluster.
Owner:HANHOU (BEIJING) TECH CO LTD

Quantization and inference method and device of mixed precision model, storage medium and equipment

This application discloses a quantization and inference method, apparatus, storage medium, and device for mixed-precision models, belonging to the field of deep learning technology. In the offline stage, a mixed-precision configuration table is generated based on the floating-point model, including configuration information for the optimal static quantization bit width and hardware platform type selected for each layer. Weights are quantized according to the static quantization bit width to obtain the mixed-precision model. During the inference stage of deploying the mixed-precision model, the computational tasks of each layer are scheduled to a target hardware platform matching the hardware platform type within the target terminal. The distribution characteristics of activation values ​​output from historical layers are statistically analyzed in real time, and the quantization parameters for the current layer are dynamically calculated based on the statistical results. The quantization parameters are used to quantize the activation values ​​input to the current layer, and the quantized weights are used for inference on the target hardware platform. This application achieves globally optimal cross-platform inference latency, dynamically adjusts the quantization range, and maintains stable output precision.
Owner:GUANGDONG UCAP INTERNET INFORMATION TECH

Rewriting Model Construction Method Based on Mixed Precision Strategy

This invention relates to the field of artificial intelligence technology and discloses a method for constructing a rewriting model based on a mixed-precision strategy. The method includes instantiating computation nodes in a computation flow graph into multiple intelligent computing units. Each intelligent computing unit maintains a local state vector containing its current computational precision and encapsulates a rule engine. The rule engine performs local inference and generates impact messages containing adjustment proposals. When impact messages are transmitted between intelligent computing units, distributed negotiation based on these messages is triggered, and the computational precision of the corresponding intelligent computing units is collaboratively adjusted according to the negotiation results. A central policy controller collects global state information from the computation flow graph and dynamically optimizes the inference rules used by the rule engine based on this global state information. Through a dual-layer driving approach of local inference and global layer optimization, the mixed-precision strategy and the computation flow graph are dynamically matched and optimized, balancing the allocation of computational precision and hardware resources.
Owner:NINGBO JINWANG INFORMATION IND CO LTD

Mixed precision quantization method, apparatus, device, medium, and program product

This disclosure provides a mixed-precision quantization method, apparatus, device, medium, and program product, relating to the field of deep learning technology. The method is applied to a large language model, where the feedforward network of the large language model includes a multiplicative structure configured with a first weight matrix and a second weight matrix. The method includes: determining the statistical characteristics of the multiplicative structure on each output channel of the feedforward network based on the first and second weight matrices; determining the numerical amplification risk index of the multiplicative structure on each output channel based on the statistical characteristics; and configuring a corresponding computational precision for each output channel according to the numerical amplification risk index; wherein output channels with different numerical amplification risk indices are configured with different computational precisions. This method eliminates the need for real-time scanning of input data, significantly reducing runtime overhead.
Owner:MOORE THREADS TECH CO LTD

A method and system for post-training channel-mixed precision quantization of neural networks

PendingCN122088573AReduce quantization errorAvoid large subsequent quantization errorsHardware monitoringBiological modelsEngineeringNetwork model
This invention relates to the field of deep learning technology, and more particularly to a method and system for post-training channel-mixed precision quantization of neural networks. For each weight channel in each layer of the pre-trained model to be quantized, the invention calculates the initial scaling factor of the weight channel at different bit widths based on the weight range of the weight tensor of that weight channel; optimizes the initial scaling factor of each weight channel in each layer of the pre-trained model to be quantized at different bit widths to obtain the target scaling factor of the weight channel at different bit widths; constructs an optimal bit width allocation integer linear programming problem based on the obtained target scaling factor; obtains the optimal bit width allocated to each weight channel by solving the optimal bit width allocation integer linear programming problem; and quantizes the pre-trained model to be quantized based on the obtained optimal bit width of the channels. This invention effectively improves the processing accuracy of the quantized neural network model for data such as text, images, and audio.
Owner:SUZHOU UNIV

An adaptive test optimization method for AI card mixed precision mode

PendingCN122285409ASelf adaptiveTest scene
This invention discloses an adaptive testing optimization method for AI card hybrid precision mode, belonging to the field of artificial intelligence and AI accelerator card testing technology. The invention first classifies AI tasks and collects load characteristics, constructing an AI business load library containing the operational characteristics of multiple AI tasks. Based on the operator precision sensitivity classification, corresponding precision types are matched. Precision combination test cases are adaptively generated in conjunction with the business load library. Then, the resource constraints of AI card computing power partitioning instances are used to complete the matching, screening, and parameter adjustment of test cases. Finally, testing is completed through dynamic load injection, and an adapted configuration scheme is output after multi-dimensional quantitative evaluation. This invention realizes intelligent testing of AI card hybrid precision mode, improves test scenario coverage and execution efficiency, effectively adapts to single-card multi-instance deployment scenarios, and provides reliable parameter support for the business implementation of hybrid precision mode.
Owner:四川华鲲振宇智能科技有限责任公司

Low-power and real-time mixed-precision 3d-human mesh generation accelerator

A low-power and real-time mixed-precision accelerator structure for 3D-human mesh generation is provided. A mixed-precision accelerator according to one embodiment may include an arithmetic unit comprising a plurality of mixed-precision cores, and a global memory that stores the result of a forward calculation of the arithmetic unit and reuses the stored result for a reverse calculation of the arithmetic unit. Here, each of the plurality of mixed-precision cores may include a low-precision inner product engine that quantizes a first-bit floating-point weight used in a Skinned Multi-Person Linear Model (SMPL) algorithm for 3D-human mesh generation into a second-bit low-precision floating-point weight and performs a matrix multiplication operation using the quantized weight and an input value, and a high-precision SIMD (Single Instruction Multiple Data) engine that processes a high-precision operation using a third-bit high-precision floating-point. Here, the second bit may be smaller than the first bit, and the third bit may be smaller than the first bit and larger than the second bit.
Owner:KOREA ADVANCED INST OF SCI & TECH

Ultra-high-frequency mixed-precision processing elements for zetta-scale all-silicon computing

PCT designated stageWO2026139935A1Computational scienceData stream
A Zetta-scale computing architecture is disclosed, comprising a massive array of ultra-high- frequency, mixed-precision processing elements integrated within a continuous All-Silicon Domain. To achieve ZettaFLOPS performance in a single rack, the processing elements are radically simplified and pipelined to operate at frequencies exceeding 10 GHz (e.g., 15 GHz). The architecture utilizes a mixed-precision W4A8 dataflow, where 4-bit weights are stored locally and 8-bit activations are broadcast through the all-silicon fabric. By eliminating the bandwidth bottlenecks and power penalties of conventional packaging, and by maximizing the transistor utility for arithmetic rather than control, the system enables an unprecedented density of billions of processing elements, offering a 1,000x improvement in performance and efficiency over state-of-the-art packaged GPU clusters.
Owner:SILVEBROOK KIA

A mixed-precision automated optimization method for efficient deployment of neural networks on FPGA

PendingCN122261673ABiological modelsProgram loading/initiatingNeighborhood searchNetwork model
This invention provides an automated optimization method for efficient deployment of neural networks on FPGAs, relating to the fields of deep learning and hardware acceleration technology. The method includes: receiving the original neural network model and hardware constraint information; parsing to obtain quantization operators; generating a candidate precision configuration combination set through two-level filtering; constructing a graph-structured perceptron based on the target FPGA architecture to predict hardware resource consumption under each configuration; constructing a differential objective function that integrates task loss and weighted quantization error; solving for the initial mixed precision configuration using an adaptive penalized augmented Lagrangian method under hardware constraints; progressively expanding the neighborhood and replacing configurations through hierarchical neighborhood search to obtain the final configuration; generating hardware description code for both fixed-point and floating-point formats and calling the backend toolchain to output bitstream files. This invention achieves automated precision configuration while satisfying FPGA resource and timing constraints, significantly improving inference energy efficiency and deployment efficiency.
Owner:NANJING UNIV OF SCI & TECH

A pipeline-based universal multiple-precision matrix multiplication-addition operation circuit

This invention discloses a pipelined, general-purpose multi-precision matrix multiplication and addition circuit, belonging to the field of AI chip design. This invention is designed for a highly parallel multi-precision tensor computation unit based on a many-core architecture. It implements matrix multiplication and addition operations with different data precisions based on a pipelined structure, including four data precisions: int4, int8, fp16, and fp32, as well as four mixed precision operation modes: int4 mixed with int32, int8 mixed with int32, fp16 mixed with fp32, and int8 mixed with fp32. The proposed circuit and method design independent pipelines for different operation modes, allowing different types of data to reuse a single computation module, achieving a hardware resource reuse rate of 88.9%. This invention is widely applicable to tensor computation. By designing independent pipelines for different operation modes and reusing hardware circuits, it eliminates the need for separate circuits for integer and floating-point data operations, thus improving hardware reuse.
Owner:58TH RES INST OF CETC

A mixed-precision dequantization system and method for in-memory computing

ActiveCN122154796ABiological modelsComputing operations for multiplication/divisionFloating pointPipeline (computing)
The application discloses a mixed-precision dequantization system and method for in-memory computing, and belongs to the field of integrated circuits and artificial intelligence hardware. The system comprises a digital domain pre-alignment module, an analog in-memory computing core and a general dequantization DQ-SIMD unit. The pre-alignment module completes floating-point activation processing and generates an activation value compensation term. The analog in-memory computing core completes multiplication and accumulation and eliminates physical offset errors based on a 1T1R single-terminal array. The dequantization unit realizes fixed-point to floating-point conversion and asymmetric quantization inverse transformation through a three-stage pipeline. The application proposes a single-terminal storage scheme of non-negative shift mapping in view of the core contradiction between mixed-precision computing requirements and the non-negativity of device conductance when deploying a large language model inference on a 1T1R structure RRAM array. The single-terminal storage scheme is matched with a double-error elimination mechanism of software and hardware cooperation to realize single-terminal storage of signed weights by superimposing a fixed offset in the analog domain and to eliminate physical errors introduced by the offset in real time through a pipeline circuit in the digital domain.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Heterogeneous mixed precision data processing method, system, device and storage medium

The application provides a heterogeneous mixed-precision data processing method, system, device and storage medium. The method comprises the following steps: generating heterogeneous task description information; configuring precision conversion parameters of a direct data channel based on the heterogeneous task description information, and establishing address alias mapping between an acceleration processing unit and a vector execution unit; transmitting, through the address alias mapping, to-be-processed data in the vector execution unit to the acceleration processing unit for calculation and processing via the direct data channel; and performing, in the process of transmitting the to-be-processed data via the direct data channel, stream online precision conversion on the to-be-processed data according to the precision conversion parameters. The technical scheme can improve the data collaborative processing efficiency of different processors.
Owner:BEIJING VCORE TECH CO LTD

An edge device-oriented large model mixed precision online incremental learning method and system

The application discloses a large model mixed precision online incremental learning method and system for edge devices. Firstly, the application analyzes the distribution offset degree of the newly arrived incremental data, and judges the task priority by calculating the KL divergence. Then, according to the task priority, the sensitivity of each layer of the model, etc., an FP16 and INT8 mixed precision quantization configuration table is generated, and the key layer is given higher precision, and the secondary layer adopts low precision. In the online training process, the quantization error gradient generated by the INT8 update is accumulated, and when the accumulated error exceeds the preset threshold or reaches a certain period, the compensation update is carried out with the FP32 precision, so as to eliminate the error accumulation caused by low precision training. At the same time, by counting the attention weight or the activation contribution degree, the model parameters with continuous low contribution are frozen or eliminated, so as to reduce the memory burden. The application can effectively reduce the memory and computing power consumption of the edge device, while maintaining high model precision and old task memory, and is suitable for various edge scenes.
Owner:SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT) +1

Reconfigurable mixed-precision processing apparatus and operating mehtod thereof

PendingUS20260186741A1AlgorithmOperand
An operating method of a reconfigurable mixed-precision processing apparatus, includes performing a multiplication on a plurality of operands to calculate a first result value in a floating-point format, performing preprocessing on the first result value to calculate a second result value, performing an addition on a mantissa value of the second result value to calculate a third result value, and converting a mantissa of the third result value into a mantissa in a floating-point format to calculate a fourth result value, in which the operand includes any one of a floating-point format, an integer format, and 1 bit.
Owner:AIPS LTD

A multi-precision / hybrid-precision fused multiply-add computing unit circuit

PendingCN122450412APerformance computingBinary tree
The present application relates to the technical field of digital integrated circuit and processor design, in particular to a multi-precision / mixed-precision fusion multiply-add calculation unit circuit.The present application starts from four aspects: a Booth multiplier array with bit width adaptive mapping, a reconfigurable compressor array, a configurable barrel shifter with high hardware multiplexing, and an asymmetric reconfigurable binary tree composed of a leading zero detector;and the four modules are used to realize multi-precision data type support from low precision (INT4, INT8, FP8, FP16, BF16) to high precision (FP32), and the calculation precision is maximally preserved during mixed-precision calculation;compared with the traditional FMA unit designed separately for different precisions, the present application greatly optimizes the hardware multiplexing rate and energy efficiency ratio on the basis of supporting multi-precision / mixed-precision calculation mode, improves the energy efficiency ratio and calculation throughput, and is suitable for artificial intelligence accelerators, general-purpose processors and high-performance computing chips.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Deep learning acceleration with mixed precision

A device for deep learning acceleration with mixed precision may include matrix-vector (MV) components that each include vector-vector (VV) components that are each configured to generate a respective VV output based on an input precision mode, an output precision mode, and an accumulation of products. The accumulation of products may be calculated by adding products based on the input precision mode. Each product may be calculated by multiplying, based on the input precision mode, a map data segment and a kernel data segment. Each MV component may include one or more components configured to concatenate VV outputs to generate a concatenated VV output. The device may include activation function components that are each configured to receive a corresponding concatenated VV output, generate an activation function output based on the corresponding concatenated VV output and the output precision mode, and output the activation function output.
Owner:MICRON TECHNOLOGY INC

A system used to perform mixed-precision inference, a program used to perform mixed-precision inference, and a method for performing mixed-precision inference.

This system provides a mechanism to suppress the decrease in accuracy and the increase in computation time in inference using mixed precision. [Solution] The system 100 used to perform inference using mixed precision includes an information acquisition unit 121 that acquires environmental information, which is information about the surrounding environment of the object of inference, and a data type setting unit 122 that sets the data type used for each layer in inference according to the acquired environmental information.
Owner:DENSO CORP +2

A low-power single-precision floating-point multiplier based on a logarithmic approximation algorithm

The application discloses a low-power single-precision floating-point multiplier based on a logarithmic approximation algorithm and belongs to the technical field of integrated circuits.The application approximates and optimizes an accurate floating-point multiplier in multiple aspects, mainly including a preprocessing module and a mantissa multiplier.The application adopts a hybrid precision architecture combining a logarithmic approximation algorithm and probability statistical truncation compensation, and divides mantissa multiplication into high-weight segmented array multiplication and low-weight segmented logarithmic multiplication.The application retains accurate calculation of high-weight segments and realizes low-weight segment calculation at a low cost, and errors of various approximate circuits can be offset to each other.The approximate multiplier provided by the application can significantly improve accuracy and can achieve a better balance between hardware overhead and accuracy requirements.The floating-point multiplier has reconfigurability and can be reconfigured and designed according to accuracy requirements to realize fine-grained adjustment of accuracy.The application is suitable for fault-tolerant applications with requirements for circuit delay and power consumption.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Configurable dataflow backend for subword parallel SIMT processor

A Single Instruction Multiple Threads (SIMT) processor core comprising: (a) a configurable dataflow backend operative to support subword parallelism for execution of mixed-precision arithmetic operations, wherein the backend includes multiple dataflow network stages, each dataflow network stage including one or more functional units configurable to process multiple subwords within a register in parallel; (b) a configuration memory storing configuration data defining operational configurations of the dataflow network stages in the backend, wherein the configuration memory is accessible via a configuration index mechanism operative to select configuration data from the configuration memory to control the dataflow network stages in the backend, and wherein the configuration index mechanism is configured for enabling the dataflow backend to support multiple variations of mixed-precision arithmetic operations without requiring dedicated instructions for each variation.
Owner:INTERUNIVERSITAIR MICRO ELECTRONICS CENT (IMEC VZW)