Large language model calculation device, large language model acceleration device, and large language model calculation method
By designing a large language model computing device with multiple computing units, using broadcasting and merge processing technology, the problems of low computing efficiency and high cost in the decoding process of large language model are solved, and efficient and low-cost computing effects are achieved.
Patent Information
- Application Number
- CN202510164374.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the decoding process, large language models require a large number of large-scale parameters to perform operations, resulting in low operation efficiency, high latency, and relying on expensive graphics processors, increasing computing pressure and cost.
A large language model computing device is designed, including a plurality of first computing units and at least one second computing unit. Each second computing unit is in communication and connected with a plurality of first computing units. The operation of the target network layer of the large language model is realized through broadcasting and merging processing, reducing dependence on the graphics processor.
By independently synchronously performing partial operations of the target network layer of the large language model, decoding costs and computing pressure are reduced, operating efficiency is improved, delay is reduced, and flexible communication interconnection between computing units is realized.
Smart Images

Figure CN119990215A_ABST