Fast matrix multiplication methods and systems

By optimizing matrix data storage and transformation within dedicated caches and local memories, the method addresses data-handling bottlenecks in multithreaded systems, enhancing the efficiency of matrix multiplication through reduced external memory access and improved parallelization.

GB2642197BActive Publication Date: 2026-07-02IMAGINATION TECH LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
IMAGINATION TECH LTD
Filing Date
2024-06-25
Publication Date
2026-07-02

AI Technical Summary

Technical Problem

Matrix multiplication in multithreaded processing systems is bottlenecked by data-handling limitations, particularly in systems with limited cache storage and finite bandwidth, which reduces the parallelization efficiency of large matrices.

Method used

The method involves storing portions of input matrices in dedicated caches and local memories, using workgroups to perform concurrent multiplications while reusing cached matrix subunits, and transforming matrix data to fit register capacities, optimizing data allocation across memory hierarchies to minimize external memory access.

Benefits of technology

This approach enhances the parallelization efficiency of matrix multiplication by reducing the frequency of external memory access and latency, thereby improving processing speed and throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000001_0001
    Figure 00000001_0001
  • Figure 00000002_0000
    Figure 00000002_0000
Patent Text Reader

Abstract

A multithreaded processing system comprises one or more processing units (202 in Figure 2), each processing unit coupled with a dedicated cache (204) and a dedicated local memory (210), and configured
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Machine learning sparse computation mechanism for arbitrary neural networks, arithmetic compute microarchitecture, and sparsity for training mechanism

    US20190205746A1

  • Sparse matrix calculations untilizing ightly coupled memory and gather / scatter engine

    US20220019430A1

  • Assigning processing threads for matrix-matrix multiplication

    US20220138281A1

  • Sparse matrix by vector multiplication

    WO2009037684A2

  • Method executed by accelerator, and electronic device

    WO2023173639A1