Storage structures, access methods, computing devices and media for quantized tensors

By adopting a hierarchical architecture design of block-head-storage area-terminal information, the storage and access methods of large-scale pre-trained models are optimized, solving the problems of insufficient storage capacity and low data access efficiency, and achieving more efficient data processing and resource utilization.

CN121918773BActive Publication Date: 2026-05-26MOFFETT AI TECHNOLOGY SHENZHEN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MOFFETT AI TECHNOLOGY SHENZHEN CO LTD
Filing Date
2026-03-27
Publication Date
2026-05-26

Smart Images

  • Figure CN121918773B_ABST
    Figure CN121918773B_ABST
Patent Text Reader

Abstract

This application provides a storage structure, access method, computing device, and medium for quantization tensors. The storage structure for quantization tensors is applied to devices running large models. The storage structure includes: multiple blocks, the number of which is determined based on the total number of historical lexical units and the storage capacity of the device's memory; the size of each block is determined based on the number of lexical units corresponding to the device's efficient single read granularity; each block includes multiple heads, the number of which is based on the total number of heads in the large model; each head includes multiple storage areas, the number of which is determined based on the quantization method of the quantization tensor; each storage area includes multiple sequentially numbered lexical units, the number of which is based on the block size; and lexical units with the same number in all storage areas of the multiple heads together constitute a single complete lexical unit.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Task execution method and device based on hybrid expert model, equipment, storage medium and program product

    CN121187787A

  • Large model dynamic batch processing method based on sequence splicing

    CN121351991A