Large language model calculation device, large language model acceleration device, and large language model calculation method

By designing a large language model computing device with multiple computing units, using broadcasting and merge processing technology, the problems of low computing efficiency and high cost in the decoding process of large language model are solved, and efficient and low-cost computing effects are achieved.

CN119990215AInactive Publication Date: 2025-05-13BEIJING PINGXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164374.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

During the decoding process, large language models require a large number of large-scale parameters to perform operations, resulting in low operation efficiency, high latency, and relying on expensive graphics processors, increasing computing pressure and cost.

Method used

A large language model computing device is designed, including a plurality of first computing units and at least one second computing unit. Each second computing unit is in communication and connected with a plurality of first computing units. The operation of the target network layer of the large language model is realized through broadcasting and merging processing, reducing dependence on the graphics processor.

Benefits of technology

By independently synchronously performing partial operations of the target network layer of the large language model, decoding costs and computing pressure are reduced, operating efficiency is improved, delay is reduced, and flexible communication interconnection between computing units is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990215A_ABST
    Figure CN119990215A_ABST
Patent Text Reader

Abstract

The invention provides a large language model operation device, a large language model acceleration device and a large language model operation method, and the large language model operation device comprises a plurality of first calculation units and at least one second calculation unit. Each first calculation unit can independently and synchronously execute part of operation of the target network layer of the large language model, and large-scale operation required by the large language model does not need to be completed by adopting a graphics processor, so that decoding calculation of the large language model is completed by using each calculation unit with relatively low cost instead of an expensive graphics processor; the decoding cost and the calculation pressure of the large language model are effectively reduced, and the data volumes of the matrix data stored in the first calculation units are the same, so that the calculation speeds of the first calculation units are similar when the first calculation units perform calculation based on the stored matrix data; the second calculation unit does not need to consume long time to specially wait for feedback of a certain first calculation unit, so that delay is reduced, and the overall operation efficiency of the large language model is improved.
Need to check novelty before this filing date? Find Prior Art