Storage structures, access methods, computing devices and media for quantized tensors
By adopting a hierarchical architecture design of block-head-storage area-terminal information, the storage and access methods of large-scale pre-trained models are optimized, solving the problems of insufficient storage capacity and low data access efficiency, and achieving more efficient data processing and resource utilization.
CN121918773BActive Publication Date: 2026-05-26MOFFETT AI TECHNOLOGY SHENZHEN CO LTD
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MOFFETT AI TECHNOLOGY SHENZHEN CO LTD
- Filing Date
- 2026-03-27
- Publication Date
- 2026-05-26
Smart Images

Figure CN121918773B_ABST
Abstract
This application provides a storage structure, access method, computing device, and medium for quantization tensors. The storage structure for quantization tensors is applied to devices running large models. The storage structure includes: multiple blocks, the number of which is determined based on the total number of historical lexical units and the storage capacity of the device's memory; the size of each block is determined based on the number of lexical units corresponding to the device's efficient single read granularity; each block includes multiple heads, the number of which is based on the total number of heads in the large model; each head includes multiple storage areas, the number of which is determined based on the quantization method of the quantization tensor; each storage area includes multiple sequentially numbered lexical units, the number of which is based on the block size; and lexical units with the same number in all storage areas of the multiple heads together constitute a single complete lexical unit.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Task execution method and device based on hybrid expert model, equipment, storage medium and program product
CN121187787A
Large model dynamic batch processing method based on sequence splicing
CN121351991A