Artificial intelligence model inference method and apparatus, readable medium, and terminal device

By modifying the attention mechanism structure in the artificial intelligence model to share the key-value vector cache with the previous layer, the problem of high memory usage was solved, resulting in cost reduction and improved adaptability.

WO2026025704A1PCT designated stage Publication Date: 2026-02-05SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/129949
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2024-11-05
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing artificial intelligence models consume a large amount of GPU memory during inference, resulting in high usage costs.

Method used

By modifying the attention mechanism structure in the artificial intelligence model to share the key-value vector cache with the previous layer of the attention mechanism structure, the requirement for key-value vector cache for each layer of the attention mechanism structure is reduced.

Benefits of technology

It reduces the amount of video memory used, lowers the cost of using the AI ​​model, and improves the model's adaptability to different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129949_05022026_PF_FP_ABST
    Figure CN2024129949_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and in particular to an artificial intelligence model inference method and apparatus, a computer-readable storage medium, and a terminal device. The method comprises: acquiring data to be subjected to inference; and performing inference on said data on the basis of a target artificial intelligence model, to obtain an inference result of said data, wherein the target artificial intelligence model is an artificial intelligence model obtained by pre-training and based on each layer of transformer encoder, each layer of transformer encoder comprises a preset number of layers of target attention mechanism structures, and the target attention mechanism structures share a key-value vector cache with a layer of attention mechanism structure above the target attention mechanism structures.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Inference method and device and electronic equipment

    CN118036741A

  • Model reasoning method and device, electronic equipment, storage medium and program product

    CN118070905A

  • Inference model, data processing method and device and medium

    CN118095430A

  • Artificial intelligence inference apparatus and method

    US20220374740A1