Artificial intelligence model inference method and apparatus, readable medium, and terminal device
By modifying the attention mechanism structure in the artificial intelligence model to share the key-value vector cache with the previous layer, the problem of high memory usage was solved, resulting in cost reduction and improved adaptability.
WO2026025704A1PCT designated stage Publication Date: 2026-02-05SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
Patent Information
- Application Number
- PCT/CN2024/129949
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2024-11-05
- Publication Date
- 2026-02-05
AI Technical Summary
Technical Problem
Existing artificial intelligence models consume a large amount of GPU memory during inference, resulting in high usage costs.
Method used
By modifying the attention mechanism structure in the artificial intelligence model to share the key-value vector cache with the previous layer of the attention mechanism structure, the requirement for key-value vector cache for each layer of the attention mechanism structure is reduced.
Benefits of technology
It reduces the amount of video memory used, lowers the cost of using the AI model, and improves the model's adaptability to different scenarios.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN2024129949_05022026_PF_FP_ABST
Abstract
The present application relates to the technical field of computers, and in particular to an artificial intelligence model inference method and apparatus, a computer-readable storage medium, and a terminal device. The method comprises: acquiring data to be subjected to inference; and performing inference on said data on the basis of a target artificial intelligence model, to obtain an inference result of said data, wherein the target artificial intelligence model is an artificial intelligence model obtained by pre-training and based on each layer of transformer encoder, each layer of transformer encoder comprises a preset number of layers of target attention mechanism structures, and the target attention mechanism structures share a key-value vector cache with a layer of attention mechanism structure above the target attention mechanism structures.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Inference method and device and electronic equipment
CN118036741A
Model reasoning method and device, electronic equipment, storage medium and program product
CN118070905A
Inference model, data processing method and device and medium
CN118095430A
Artificial intelligence inference apparatus and method
US20220374740A1