Electronic device, operation method of electronic device and optimization method of large language model

By coupling external storage circuitry to the edge device and optimizing the quantization process of large language models, the startup latency and inference accuracy issues of the edge device are resolved, thus improving the user experience.

CN121882244APending Publication Date: 2026-04-17SIGMASTAR TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SIGMASTAR TECH LTD
Filing Date
2025-12-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The limited hardware resources of edge devices lead to severe delays in the inference operation startup of large language models, and the extreme values ​​of tokens cause abnormal quantization ranges, reducing inference accuracy.

Method used

An external storage circuit is coupled into the electronic device to store the initial prompt statement database and the quantized startup template. Through standardization processing and matching of prompt words, the token calculation of the initial prompt words is skipped. The core calculation structure is quantized using calibrated quantization parameters, the token sequence characteristics are pre-calculated, and the quantized startup template is generated.

Benefits of technology

It reduces the startup latency of large language models, improves the user experience, avoids quantization range anomalies, and improves inference accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882244A_ABST
    Figure CN121882244A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic device. The electronic device is coupled to an external storage circuit and comprises an intelligent processing unit and a computing circuit. The external storage circuit stores an initial prompt statement database and a quantized startup template. The calculation circuit is used for executing the following steps: carrying out standardization processing on an input text to generate a preprocessed text containing M tokens; searching prompt words and sentences matched with the preprocessed text in the initial prompt statement database to obtain a matched prompt word and sentence containing N tokens; obtaining a target token sequence feature corresponding to the matched prompt word and sentence from the quantized startup template; controlling the intelligent processing unit to mark a calculation state of the N tokens of the matched prompt word and sentence as a calculated state; generating data to be processed according to the N value and the preprocessed text; and controlling the intelligent processing unit to perform a reasoning operation based on the to-be-processed data and the target token sequence feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to electronic devices, and more particularly to electronic devices that execute large language models (LLMs). Background Technology

[0002] The need to deploy large language models on edge devices (e.g., personal computers, mobile phones, etc.) is growing. However, due to the limited hardware resources of edge devices, large language models are a significant burden, especially during the inference startup phase. Existing edge devices perform token-by-token computation on the input text regardless of its type, resulting in severe startup delays and a poor user experience.

[0003] Furthermore, extreme values ​​of tokens (corresponding to some common phrases, such as "(please) control", "(please) turn on", "(please) turn off", etc.) can cause abnormal quantization ranges when performing quantization operations on large language models, which can reduce the inference accuracy of large language models. Summary of the Invention

[0004] In view of the shortcomings of the prior art, one of the objectives of the present invention is to provide an electronic device, an operation method of the electronic device, and an optimization method for a large language model, so as to improve the shortcomings of the prior art.

[0005] One embodiment of the present invention provides an electronic device. The electronic device is coupled to an external storage circuit. The external storage circuit stores an initial prompt statement database and a quantized startup template. The electronic device includes an intelligent processing unit and a computing circuit. The intelligent processing unit is coupled to the external storage circuit and includes a storage circuit for executing a large language model. The computing circuit is coupled to the external storage circuit and is used to perform the following steps: (A) performing a normalization process on an input text to generate a preprocessed text, wherein the preprocessed text includes M tokens, where M is a positive integer; (B) searching the initial prompt statement database for a prompt phrase that matches the preprocessed text to obtain a matching prompt phrase, wherein the matching prompt phrase includes N tokens, where N is a positive integer; (C) obtaining a target token sequence feature corresponding to the matching prompt phrase from the quantized start template and storing the target token sequence feature in the storage circuit; (D) controlling the intelligent processing unit to mark the computation status of the N tokens of the matching prompt phrase as computed; (E) determining a range to be processed for the preprocessed text based on the value of N to generate data to be processed; and (F) controlling the intelligent processing unit to perform an inference operation based on the data to be processed and the target token sequence feature.

[0006] Another embodiment of the present invention provides a method for operating an electronic device. The electronic device includes an intelligent processing unit coupled to an external storage circuit. The intelligent processing unit includes a storage circuit and is used to execute a large language model. The external storage circuit stores an initial prompt statement database and a quantized startup template. The operation method includes: (A) performing a standardization process on an input text to generate a preprocessed text, wherein the preprocessed text includes M tokens, where M is a positive integer; (B) searching the initial prompt statement database for prompt phrases that match the preprocessed text to obtain a matching prompt phrase, wherein the matching prompt phrase includes N tokens, where N is a positive integer; (C) obtaining a target token sequence feature corresponding to the matching prompt phrase from the quantized start template, and storing the target token sequence feature in the storage circuit; (D) controlling the intelligent processing unit to mark a calculation state of the N tokens of the matching prompt phrase as calculated; (E) determining a range to be processed for the preprocessed text based on the value of N to generate data to be processed; and (F) controlling the intelligent processing unit to perform an inference operation based on the data to be processed and the target token sequence feature.

[0007] Another embodiment of the present invention provides an optimization method for a large language model, comprising: (A) inputting a plurality of words and phrases from a quantization calibration dataset into the large language model, wherein the quantization calibration dataset includes an input text and a corresponding response; (B) controlling the large language model to skip at least one token corresponding to an initial prompt phrase in an inference operation, wherein the inference operation generates a numerical distribution; (C) calculating and determining a quantization scaling factor and a zero point based on the numerical distribution to obtain a calibrated quantization parameter; (D) quantizing a core computational structure in the large language model; (E) pre-calculating a token sequence feature corresponding to the initial prompt phrase; and (F) quantizing the token sequence feature using the calibrated quantization parameter to generate a quantized start template.

[0008] The technical means embodied in the embodiments of the present invention can improve at least one of the shortcomings of the prior art, and therefore the present invention can improve the user experience compared with the prior art.

[0009] The features, implementation, and effects of this invention are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0010] Figure 1 This is a functional block diagram of one embodiment of the electronic device of the present invention; Figure 2 This is a flowchart of one embodiment of the data collection phase of the present invention; Figure 3 This is a flowchart of one embodiment of the large language model optimization method of the present invention; and Figures 4A-4B This is a flowchart of one embodiment of the operation method of the electronic device 100 of the present invention.

[0011] Explanation of reference numerals in the attached figures: 100: Electronic devices; 105: External storage circuit; 106: Initial prompt statement database; 107: Quantized startup template; 110: Computational circuits; 112: Preprocessing and Analysis Matching Module; 114: Data Processing and Model Control Module; 120: IPU (Intelligent Processing Unit); 122: Storage circuit; 124: Large Language Model; KVC: Target Token Sequence Characteristics; N: Number of tokens; TK_tmp: Intermediate results; Ts[1:M]: Preprocessed text; Ts[N+1:M]: Data to be processed; Txt_in: Input text; Txt_out: Output text; S210, S220, S230, S240, S310, S312, S314, S316, S320, S330, S340, S410, S420, S430, S440, S450, S460, S470, S480, S490, S495: Steps. Detailed Implementation

[0012] The technical terms used in the following description are based on the customary terms in this technical field. If this specification provides explanations or definitions for certain terms, the explanations or definitions in this specification shall prevail.

[0013] This invention discloses electronic devices, methods for operating the electronic devices, and methods for optimizing large language models. Since some components of the electronic device of this invention may be known individually, details of known components will be omitted in the following description without affecting the full disclosure and implementability of the invention. Furthermore, some or all of the procedures of the operation method of the electronic device of this invention may be in the form of software and / or firmware, and can be executed by the electronic device of this invention or its equivalents. Without affecting the full disclosure and implementability of the method invention, the following description of the method invention will focus on the steps rather than the hardware.

[0014] This invention addresses the problems encountered in existing technologies through data collection, model optimization, and hardware implementation.

[0015] Figure 1 This is a functional block diagram of one embodiment of the electronic device of the present invention. The electronic device 100 may be a system on a chip (SoC), including a computing circuit 110 and an intelligence processing unit (IPU) 120 coupled to each other. The computing circuit 110 may be a circuit or electronic component with program execution capability, such as a central processing unit, microprocessor, microprocessor unit, digital signal processor, application-specific integrated circuit (ASIC), or equivalent circuits thereof. In some embodiments, the IPU 120 may be a neural processing unit (NPU).

[0016] Electronic device 100 is coupled to external storage circuit 105. In some embodiments, external storage circuit 105 may be dynamic random access memory (DRAM).

[0017] The computing circuit 110 is coupled to the external storage circuit 105 and includes a preprocessing and analysis matching module 112 and a data processing and model control module 114. In some embodiments, the computing circuit 110 implements the functions of the preprocessing and analysis matching module 112 and the data processing and model control module 114 by executing program code and / or program instructions, and the program code and / or program instructions can be stored in the external storage circuit 105.

[0018] The IPU 120 is coupled to an external storage circuit 105 and includes a storage circuit 122 and a large language model 124. In some embodiments, the storage circuit 122 may be a static random access memory (SRAM).

[0019] Figure 2 This is a flowchart of one embodiment of the data collection phase of the present invention. Figure 2 The process can be executed by a general computer and includes the following steps.

[0020] Step S210: Analyze typical use cases of the large language model in edge devices (including but not limited to home control, information query, chat dialogue, etc.), and establish multiple preset initial prompts for each use case. The initial prompts can correspond to the 1st to the Kth characters of the input text Txt_in. For example, the initial prompt for the input text Txt_in "Please turn on the living room air conditioner" is "Please turn it on" (i.e., K = 3). For the home control scenario, the initial prompts include, but are not limited to, "Please control," "Please turn on," and "Please turn off." For the information query scenario, the initial prompts include, but are not limited to, "Please query," "Please tell me," and "Why."

[0021] Step S220: Generate token sequence features (i.e., key vectors and value vectors) corresponding to these initial prompts. This step can be achieved by inputting the initial prompts into a large language model to obtain the corresponding token sequence features.

[0022] Step S230: A structured initial prompt statement database 106 is established based on usage scenarios, initial prompt phrases, and token sequence features. The initial prompt statement database 106 includes multiple usage scenarios, multiple initial prompt phrases corresponding to these usage scenarios, and token sequence features corresponding to each initial prompt phrase. Therefore, the corresponding usage scenario and token sequence features can be queried in the initial prompt statement database 106 based on the initial prompt phrase. In some embodiments, the initial prompt statement database 106 can be stored in an external storage circuit 105.

[0023] Step S240: Collect a labeled dataset for quantization. This step selects multiple target input texts (Txt_in) and corresponding target standard answers from input texts (Txt_in) and their corresponding standard answers in multiple typical use cases of edge devices as a quantization labeled dataset. The quantization labeled dataset provides data support based on real-world use cases for subsequent calculations of calibrated quantization parameters.

[0024] Figure 3 This is a flowchart of one embodiment of the large language model optimization method of the present invention. Figure 3 It is part of the model optimization phase, can be performed by a general computer, and includes the following steps.

[0025] Step S310: Calibrate the quantization parameters to generate calibrated quantization parameters for the quantization operation. Quantization refers to converting a number from floating-point format to fixed-point format. Step S310 includes sub-steps S312, S314, and S316.

[0026] Step S312: Input the words and phrases in the quantization calibration dataset into the large language model in a loop.

[0027] Step S314: Control the large language model to skip tokens corresponding to the initial prompt during inference. That is, this step only calculates based on tokens not corresponding to the initial prompt to statistically analyze the distribution of model weights (including the QKV matrix, feed-forward network (FFN) linear layer weights) and activation values. The "QKV matrix" represents the "query-key-value matrix," where "Q," "K," and "V" represent the query, key, and value, respectively.

[0028] Step S316: Calculate and determine the quantization scaling factor and zero point based on the numerical distribution to obtain the final calibrated quantization parameters. Since step S314 skips the tokens corresponding to the initial prompt phrase, extreme token values ​​(caused by the tokens corresponding to the initial prompt phrase) will not occur in subsequent quantization operations. Therefore, the quantization operation will not encounter quantization range anomalies, thus solving the problem of decreased inference accuracy in existing large language models.

[0029] Step S320: Quantize the core computational structures in the large language model 124. These computational structures include, but are not limited to, the QKV matrix and the FFN linear layer weights. More specifically, this step converts the core computational structures in the large language model 124 from floating-point (e.g., single-precision floating-point (FP32)) operations to fixed-point (e.g., INT8) operations based on the instruction set format of the IPU 120 (e.g., 8-bit integer (INT8)) and the calibrated quantization parameters. This step ensures that the computational logic of the large language model matches or is identical to the instruction set format of the IPU 120 (e.g., INT8).

[0030] Step S330: Pre-calculate the token sequence features corresponding to the initial prompt phrase.

[0031] Step S340: Quantize the token sequence features using the calibrated quantization parameters to generate a quantized startup template 107. That is, the data format of the quantized startup template 107 matches or is identical to the hardware instruction format of the IPU 120, so that the IPU 120 can use the quantized startup template 107. In some embodiments, the quantized startup template 107 can be stored in an external storage circuit 105.

[0032] Figures 4A-4B This is a flowchart of one embodiment of the operation method of the electronic device 100 of the present invention. Figures 4A-4B It is part of the hardware implementation phase and includes the following steps.

[0033] Step S410: The computation circuit 110 (more specifically, the preprocessing and analysis matching module 112) standardizes the input text Txt_in input by the user to generate preprocessed text Ts. Standardization includes adding system information (e.g., role information, usage scenario information, location information, etc.), word segmentation (segmenting text according to the vocabulary rules of the large language model), tokenization (e.g., converting to integer identification recognizable by the large language model), and filtering invalid characters (e.g., removing punctuation marks, special symbols, and other non-critical information). The operational details of standardization are well known to those skilled in the art and will not be elaborated further. The preprocessed text Ts is presented in the form of tokens. It is assumed that the number of tokens in the preprocessed text Ts is M (i.e., the preprocessed text Ts has a total of M tokens), and the entire preprocessed text Ts can be represented as "Ts[1:M]" ("[1:M]" represents the 1st to the Mth tokens).

[0034] Step S420: The calculation circuit 110 (more specifically, the preprocessing and analysis matching module 112) searches the initial prompt statement database 106 for prompt phrases that match the preprocessed text Ts to obtain the matching prompt phrases, and records the number N of tokens for the matching prompt phrases. In some embodiments, the preprocessing and analysis matching module 112 compares tokens one by one, starting from the first token of the preprocessed text Ts, until the tokens are different. That is, if a matching token is found, the preprocessing and analysis matching module 112 compares the next token; if no matching token is found, the preprocessing and analysis matching module 112 ends the comparison. For example, if the initial prompt statement database 106 includes tokens corresponding to the prompt phrase "Please turn off the air conditioner", and the preprocessed text Ts includes tokens corresponding to the prompt phrase "Please close the curtains", then the matching prompt phrase is the token corresponding to "Please close". For example, if the matching prompt phrase "Please close" corresponds to 5 tokens, then N = 5.

[0035] Step S430: The computing circuit 110 obtains the corresponding target token sequence feature KVC (e.g., INT8 format) from the quantized start template 107 based on the matched prompt words.

[0036] Step S440: Store the target token sequence feature KVC in the storage circuit 122. Because the target token sequence feature KVC is part of the quantized start template 107, it was originally stored in the external storage circuit 105. In some embodiments, this step can be performed by moving or copying the target token sequence feature KVC from the external storage circuit 105 to the storage circuit 122 via a Direct Memory Access (DMA) circuit (not shown).

[0037] Step S450: The computing circuit 110 (more specifically, the data processing and model control module 114) sends a memory mapping instruction to the IPU 120 to map the target token sequence feature KVC from the storage circuit 122 to the token sequence feature buffer (not shown) inside the IPU 120. As described in step S340, because the data format of the quantized start template 107 matches or is the same as the hardware instruction format of the IPU 120, the target token sequence feature KVC matches or is the same as the hardware instruction format of the IPU 120. In other words, the target token sequence feature KVC can be directly used by the IPU 120 without further format conversion.

[0038] Step S460: The computation circuit 110 (more specifically, the data processing and model control module 114) controls the IPU 120 to mark the computation status of the N tokens corresponding to the matching prompt phrase obtained in step S430 (i.e., the first N tokens out of the M tokens of the preprocessed text Ts, N ≦ M) as "computed," indicating that these N tokens do not need to be computed again. In this way, the IPU 120 (more specifically, the large language model 124) will skip the generation process of the token sequence features of these N tokens in the next inference operation.

[0039] Step S470: The calculation circuit 110 (more specifically, the data processing and model control module 114) determines the range to be processed of the preprocessed text Ts based on the number of tokens N, to generate the data to be processed Ts[N+1:M]. Since the calculation status of the 1st to Nth tokens of the preprocessed text Ts has been marked as "calculated", the data to be processed Ts[N+1:M] is the N+1th to Mth tokens in the preprocessed text Ts.

[0040] Step S480: The computing circuit 110 controls the IPU 120 to perform inference operations based on the data to be processed Ts[N+1:M] and the target token sequence feature KVC to obtain the intermediate result TK_tmp.

[0041] Step S490: The computing circuit 110 (more specifically, the data processing and model control module 114) analyzes the intermediate result TK_tmp to determine whether to control the IPU 120 to continue the inference operation. In some embodiments, the computing circuit 110 controls the IPU 120 to continue the inference operation based on the intermediate result TK_tmp until one of the following three conditions occurs: (1) when the length of the intermediate result TK_tmp reaches a preset upper limit (e.g., 512 or 2048 tokens); (2) when the intermediate result TK_tmp includes a termination token (e.g., "..."). <end> ”、" <eos>(2) when; or (3) when the user triggers an interrupt.

[0042] Step S495: The computation circuit 110 (more specifically, the data processing and model control module 114) converts the final intermediate result TK_tmp into natural language output text Txt_out. The details of step S495 are well known to those skilled in the art and will not be repeated here.

[0043] In summary, because this invention skips tokens corresponding to the initial prompt words during the model optimization stage, subsequent quantization operations will not encounter quantization range anomalies, thus solving the problem of decreased inference accuracy in existing large language models. Furthermore, because the electronic device 100 of this invention skips the first few tokens of the preprocessed text Ts when running the large language model 124, this invention reduces the startup latency of the large language model and improves the user experience.

[0044] While the embodiments of the present invention have been described above, these embodiments are not intended to limit the present invention. Those skilled in the art can make changes to the technical features of the present invention based on the explicit or implicit content of the present invention. All such changes may fall within the scope of patent protection sought by the present invention. In other words, the scope of patent protection of the present invention shall be determined by the scope of the patent application in this specification.< / eos> < / end>

Claims

1. An electronic device coupled to an external storage circuit, characterized in that, The external storage circuit stores an initial prompt statement database and a quantized startup template. The electronic device includes: An intelligent processing unit, coupled to the external storage circuitry and including a storage circuitry, is used to execute a large language model; and A computing circuit, coupled to the external storage circuit, is used to perform the following steps: (A) A normalization process is performed on an input text to produce a preprocessed text, wherein the preprocessed text comprises M tokens, where M is a positive integer; (B) Search the initial prompt statement database for prompt words that match the preprocessed text to obtain a matching prompt word, wherein the matching prompt word includes N tokens, where N is a positive integer; (C) Obtain a target token sequence feature corresponding to the matched prompt phrase from the quantized startup template, and store the target token sequence feature in the storage circuit; (D) Control the intelligent processing unit to mark a computed state of the N tokens of the matched prompt phrase as computed; (E) Determine a range of the preprocessed text to be processed based on the N value to generate data to be processed; and (F) Control the intelligent processing unit to perform a reasoning operation based on the data to be processed and the characteristics of the target token sequence.

2. The electronic device as claimed in claim 1, characterized in that, The N tokens are the first to the Nth tokens among the M tokens.

3. The electronic device as claimed in claim 1, characterized in that, Step (B) involves comparing the first of the M tokens with the initial prompt statement database.

4. The electronic device as claimed in claim 1, characterized in that, The data format of the target token sequence features matches or is the same as the hardware instruction format of the intelligent processing unit.

5. The electronic device as claimed in claim 1, characterized in that, The large language model skips the generation process of the token sequence features of the N tokens in the inference operation, and the token sequence features include a key vector and a value vector.

6. A method of operating an electronic device, the electronic device comprising an intelligent processing unit coupled to an external storage circuit, the intelligent processing unit including a storage circuit and used to execute a large language model, characterized in that, The external storage circuit stores an initial prompt statement database and a quantized startup template. The operation method includes: (A) A normalization process is performed on an input text to produce a preprocessed text, wherein the preprocessed text comprises M tokens, where M is a positive integer; (B) Search the initial prompt statement database for prompt words that match the preprocessed text to obtain a matching prompt word, wherein the matching prompt word includes N tokens, where N is a positive integer; (C) Obtain a target token sequence feature corresponding to the matched prompt phrase from the quantized startup template, and store the target token sequence feature in the storage circuit; (D) Control the intelligent processing unit to mark a computed state of the N tokens of the matched prompt phrase as computed; (E) Determine a range of the preprocessed text to be processed based on the N value to generate data to be processed; and (F) Control the intelligent processing unit to perform a reasoning operation based on the data to be processed and the characteristics of the target token sequence.

7. The method as described in claim 6, characterized in that, The N tokens are the first to the Nth tokens among the M tokens.

8. The method as described in claim 6, characterized in that, Step (B) involves comparing the first of the M tokens with the initial prompt statement database.

9. The method as described in claim 6, characterized in that, The data format of the target token sequence features matches or is the same as the hardware instruction format of the intelligent processing unit.

10. The method as described in claim 6, characterized in that, The large language model skips the generation process of the token sequence features of the N tokens in the inference operation, and the token sequence features include a key vector and a value vector.

11. An optimization method for a large language model, characterized in that, include: (A) Input a plurality of words and phrases from a quantized calibration dataset into the large language model, wherein the quantized calibration dataset includes an input text and a corresponding response; (B) Controlling the large language model to skip at least one token corresponding to an initial cue word in an inference operation, wherein the inference operation produces a numerical distribution; (C) Calculate and determine a quantization scaling factor and a zero point based on the numerical distribution to obtain a calibrated quantization parameter; (D) Quantify a core computational structure in the large language model; (E) Predicting a token sequence feature corresponding to the initial prompt phrase; and (F) Quantize the token sequence features using the calibrated quantization parameters to generate a quantized start template.

12. The method as described in claim 11, characterized in that, The large language model is executed by an intelligent processing unit. Step (D) is to quantize the core computational structure based on an instruction set format of the intelligent processing unit and the calibrated quantization parameters.

13. The method as described in claim 12, characterized in that, The core computational structure includes a query-key-value matrix and the weights of a linear layer of a feedforward neural network.

14. The method as described in claim 12, characterized in that, The token sequence features include a key vector and a value vector, and the intelligent processing unit does not generate the token sequence features of the initial prompt phrase when performing the reasoning operation.

15. The method as described in claim 11, characterized in that, The numerical distribution is the numerical distribution of a model weight and an activation value of the large language model.