Heterogeneous assisted text generation with cpu and ai accelerator
Patent Information
- Application Number
- CN202480088421.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2024-10-31
- Publication Date
- 2026-09-25
Smart Images

Figure CN122826577A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This application claims priority to International Application No. PCT / CN2024 / 083243, filed on March 22, 2024. Background Technology
[0002] Large Language Models (LLMs) are gaining popularity in a variety of text generation use cases, such as text-to-text, image-to-text, and speech-to-text applications (e.g., automatic speech recognition). While the generation quality of LLMs may be sufficient for some end users, economics related to the profitability of LLMs are becoming a consideration. For example, from a monetary perspective, the cost of LLM inference operations can be high because large-scale LLMs (e.g., tens or hundreds of billions of parameters) typically involve the use of multiple graphics processing units (GPUs) and longer GPU execution times, which increases the cost of lexical generation (e.g., predicting the next word in a sentence). Attached Figure Description
[0003] Various advantages of the embodiments will become apparent to those skilled in the art upon reading the following description and appended claims and referring to the following drawings, in which:
[0004] Figure 1 This is a block diagram of an example of an auxiliary text generation solution according to an embodiment;
[0005] Figure 2 This is a block diagram illustrating an example of an accelerated assisted text generation solution with extended look-ahead lexical width in assisted decoding, according to an embodiment.
[0006] Figure 3 This is a block diagram illustrating an example of an auxiliary text generation solution with extended lookahead lexical length in auxiliary decoding;
[0007] Figure 4 Example results according to embodiments of the present disclosure are shown;
[0008] Figure 5 This is a flowchart illustrating an example of a method for an operational performance-enhanced computing system according to an embodiment;
[0009] Figure 6 This is a block diagram of an example of a performance-enhanced computing system according to an embodiment;
[0010] Figure 7 This is an illustration of an example of a semiconductor packaging apparatus according to an embodiment;
[0011] Figure 8 This is a block diagram of an example processor according to an embodiment; and
[0012] Figure 9 This is a block diagram of an example of a multiprocessor-based computing system according to an embodiment. Detailed Implementation
[0013] To reduce lexical generation costs, many schemes utilize model compression techniques such as quantization (4-bit and 8-bit), which can reduce lexical latency. However, quantization can negatively impact generation quality. For example, when the BLOOM-176B Large Language Model (LLM) is used for the cue " Once upon a time, there was a little girl who loved adventure. She wanted to go to different places. A place to meet new friends and have fun. When generating long sequences:
[0014] - For the BF16 (16-bit brain floating point) model, the generated message is: "She wants to see the world, and she wants to be a part of it. She wants to make a difference, and she wants to make the world..."
[0015] - For the INT8 (8-bit integer) model, generate "She wants to learn new things. She wants to do things she has never done before. She wants to do things she has never done before. She wants to."
[0016] Therefore, in the INT8 model, the solution begins to repeat previous outputs after approximately 16 words, which presents an additional challenge for supporting long output scenarios. Thus, the technique described in this paper addresses the problem of accelerating inference without negatively impacting generation quality.
[0017] The assisted text generation described in this paper addresses this technical problem. The core understanding of assisted decoding is that the autoregressive nature of decoder-only LLMs is a performance bottleneck, and this nature imposes a hard constraint on model FLOPS (floating-point operations per second) utilization (MFU), resulting in underutilization of AI accelerators in terms of FLOPS. For example, the time required to decode 128 terms in an autoregressive manner using a graphics processing unit (GPU) can be up to approximately 100 times longer than performing a sequence-level forward pass on the same number of terms, highlighting the significant inefficiency of LLM text generation. Therefore, in this case, the FLOPS efficiency of the LLM forward pass operation is 100 times that of LLM decoding.
[0018] The auxiliary text generation described in this paper introduces a small language model (e.g., a "draft model") that takes over the decoding process at a lower cost (e.g., because the draft model has fewer parameters) by first drafting K lexical units and then feeding the drafted sequence to the original large language model (e.g., a "validation model") to run regular model forward pass operations to validate the output and obtain the final output.
[0019] For example, Figure 1 The diagram illustrates a draft model 10 (e.g., a small language model / LM) that drafts three lexical outputs for the cue "that orange cat," namely "ate that fish," and sends the entire input "that orange cat ate that fish" to a validation model 12 (e.g., a large LM) for a sequential forward pass (e.g., a single forward pass operation). The result is that the validation model 12's forward pass accepts "ate," but rejects "that," replacing "that" with "my." Therefore, in this case, two valid lexicals can be generated in a single operation comprising three small LM decoding operations and one large LM forward pass operation. In this case, if it can be guaranteed that the three small LM decoding operations take less time than one large LM decoding operation, the auxiliary lexical delay can be accelerated (e.g., achieving a 1 / 5 or 1 / 3 speedup).
[0020] Therefore, the assisted text generation described in this paper is a promising solution for accelerating LLM inference while maintaining the same generation quality. Current assisted text generation practices may involve pure AI accelerator solutions (e.g., both the draft and validation models run on AI accelerators), which are not cost-efficient.
[0021] 1. From a solution perspective, the small draft model 10 and the large validation model 12 running on the same device will contend for limited accelerator resources (e.g., compute and memory resources) and add additional context switching overhead, which will result in suboptimal latency and additional orchestration engineering costs.
[0022] 2. From a system perspective, each AI instance contains some central processing unit (CPU), but in the current solution, the CPU remains idle except for some lightweight lexicalization tasks. This solution results in wasted resources, especially for CPUs with AI capabilities (e.g., Advanced Matrix Extensions (AMX)). Therefore, fully utilizing the CPU (e.g., CPUs with AMX) can reduce system-level costs.
[0023] Therefore, previous solutions may have placed the entire ancillary text generation solution on an AI accelerator. This approach is not optimal in terms of latency and cost, and it does not fully utilize the entire system, especially when the CPU has powerful AI acceleration capabilities.
[0024] The technique described herein employs heterogeneous assisted text generation, where the decoding of draft model 10 runs on a CPU, while the forward pass of validation model 12 runs on an AI accelerator. The technique also provides a novel drafting scheme that achieves a similar number of effective tokens per operation using only one operation of draft model 10, thereby further reducing the latency of draft model 10 and end-to-end latency. This method offers unique value by combining CPU and AI accelerator to achieve system-level cost reduction and provide a performance-optimized assisted text generation solution. Furthermore, by utilizing the CPU in the system optimally, customers can experience lower latency and lower system costs without compromising text generation quality.
[0025] As already noted, one example of the technology described herein runs a draft model 10 on a CPU and a validation model 12 on one or more AI accelerators. Furthermore, embodiments run the draft model 10 in only a single operation / step and generate multiple candidates for validation by the validation model 12.
[0026] Figure 2 A heterogeneous intra-node assisted text generation system 20 is illustrated. In this system 20, a draft model runs on a CPU 22, which generates multiple candidate lexical units 24 (“candidates”) and sends them to a large validation model on AI-accelerated hardware 26 (e.g., a PCIe (Fast Peripheral Component Interconnect, e.g., Fast PCI Foundation Specification 6.0, Version 1.0, January 11, 2022, PCI Special Interest Group) interface or a CXL (Fast Compute Link, e.g., CXL Specification, Revision 2.0, October 2020) interface) via an interface 28. The validation model performs a model forward pass to select matching lexical units and predict an additional lexical unit immediately following the last matching lexical unit.
[0027] The latency speedup ratio of assisted generation relative to the "conventional" generation scheme can be defined as:
[0028] here, It is the lexical generation delay of the validation model (large LM). It is the lexical generation delay of the draft model (small LM). This refers to the number of decoding operations performed in the draft model for each validation model forward pass operation (“forward”), and This is the percentage of tokens generated by the draft model that are accepted by the validation model (e.g., acceptance rate). Furthermore, Defined as compression ratio, it refers to the number of tokens generated in each forward pass operation of the validation model.
[0029] According to the above formula, the acceleration ratio increases with the increase of the compression ratio, and... It decreases as it increases. Current practice may primarily focus on increasing. To improve the compression ratio (e.g., to allow the draft model to decode more tokens in one round and then send those tokens together to the validation model in a single forward pass).
[0030] For example, Figure 3 An auxiliary text generation solution with extended lookahead lexical length in auxiliary decoding is presented. However, this method has two problems:
[0031] 1. Increase It doesn't necessarily increase the compression ratio. Figure 3 Even in As the number increases from three to five, since the validation model 32 does not recognize the second output word "the", all subsequent words generated by the draft model 30 will be discarded by the validation model 32, regardless of their length.
[0032] 2. Increase This will increase the time spent on the draft model by 30 (e.g., increase the time required). The denominator of the acc_ratio is reduced, thus decreasing the acc_ratio. This problem is particularly pronounced if draft model 30 runs on a CPU; under TFLOPS (trillions of floating-point operations per second) and memory bandwidth constraints, the CPU... higher and It performs better at lower levels.
[0033] Therefore, the technology described in this paper proposes a CPU-friendly auxiliary decoding method, which only... It can be improved under certain circumstances The intuitive understanding behind this method is that, while top-1 alignment between the draft and validation models cannot be guaranteed, if the draft model outputs top-k words, the probability that the top-1 words of the validation model are among the top-k words of the draft model is much higher. Therefore, a single draft model decoding operation can increase the probability of top-1 alignment. This is very suitable for CPUs.
[0034] return Figure 2This illustrates an enhanced accelerated assisted text generation solution according to an embodiment, featuring an expanded width of the "look-ahead" candidate lexical unit 24 in the assisted decoding. Expanding the width, rather than the length, of the candidate lexical unit 24 in the assisted decoding increases the acceptance ratio with only one assisted model decoding operation, resulting in a better acc_ratio in a "CPU + accelerator" heterogeneous assisted text generation solution. The process is as follows: prompt = "xxx" / / Prompt for any input for i in range(new_tokens_to_generate): prompt = prompt.to("cpu") outputs = draft_model.forward(prompt) candidate_tokens = get-top-k(outputs) inputs = cat(input.repeat(top-k), candidate_tokens) inputs = inputs.to("accelerater_x") outputs = verify_model.forward(inputs) prompt = get_max_len_sequence(outputs) return prompt
[0035] result
[0036] Now go to Figure 4 Table 40 shows the results, where, in one example, the BLOOM-560M model was used as a draft model and the BLOOM-176B model was used as a validation model, and experiments were conducted on the top-200 WIKIPEDIA dataset. The BLOOM-560M model ran on a single slot, while the BLOOM-176B model ran using tensor parallelism on eight A100-PCIE-80GB cards.
[0037] Compared to the baseline, heterogeneous assisted text generation achieved a 30% latency improvement. However, increasing K worsened the latency instead of improving it, as sequential verification and increased CPU operations negatively impacted overall performance. In the proposed solution, however, the acceptance rate was improved without additional latency to the draft model. Furthermore, the high throughput of the GPU allowed for scaling up the batch size without negatively affecting latency. Therefore, end-to-end (“e2e”) steady-state latency improvement was achieved by increasing the candidate lexical width from four to sixteen lexicals. The techniques described in this paper leverage the characteristics and capabilities of both CPUs and AI accelerators to create a better overall solution and maximize the potential of the entire system.
[0038] Figure 5 A method 50 for operating a performance-enhanced computing system is illustrated. Method 50 can be implemented in one or more modules as multiple logic instructions stored in a machine-readable or computer-readable storage medium (e.g., random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc.), hardware, or any combination thereof. For example, hardware implementations may include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic (e.g., configurable hardware) include appropriately configured programmable logic arrays (PLAs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and general-purpose microprocessors. Examples of fixed-function logic (e.g., fixed-function hardware) include appropriately configured application-specific integrated circuits (ASICs), combinational logic circuits, and sequential logic circuits. Configurable logic or fixed-function logic can be implemented using complementary metal-oxide-semiconductor (CMOS) logic circuits, transistor-transistor (TTL) logic circuits, or other circuits.
[0039] For example, the computer program code used to perform the operations shown in method 50 can be written using any combination of one or more programming languages, including object-oriented programming languages such as JAVA, SMALLTALK, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. Additionally, the logic instructions can include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, status setting data, configuration data for integrated circuit systems, and status information that personalizes the electronic circuit system and / or other structural components native to the hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.).
[0040] Box 52 is used by a first processor to generate multiple candidate lexical units based on input prompts to a draft language model, wherein the multiple candidate lexical units are generated in a single decoding operation. In one example, the first processor includes a CPU without AI capabilities (e.g., AMX). Additionally, the width of the multiple candidate lexical units can be greater than one (e.g., four to sixteen). Box 54 is used by one or more second processors to generate an output prompt based on the multiple candidate lexical units, input prompts, and a validation language model, wherein the output prompt is generated in a single forward pass operation. In one example, the second processor(s) includes an AI accelerator. Additionally, the number of parameters in the validation language model can be greater than the number of parameters in the draft language model. In embodiments, the time spent performing a single decoding operation is less than the time spent performing a single forward pass operation. Furthermore, during a single decoding operation, the computing system can deprive the first processor of the memory and computing resources of the second processor(s). Box 56 is used to send the output prompt to the draft language model. Therefore, Method 50 enhances performance at least in the following sense: performing a single decoding operation on the CPU and a single forward pass operation on one or more AI accelerators provides lower latency and system cost without compromising text generation quality.
[0041] Now go to Figure 6 The diagram illustrates a performance-enhanced computing system 280. System 280 can typically be part of an electronic device / platform having computing capabilities (e.g., personal digital assistant / PDA, laptop, tablet, convertible tablet, server), communication capabilities (e.g., smartphone), imaging capabilities (e.g., camera, camcorder), media playback capabilities (e.g., smart TV), wearable capabilities (e.g., watch, glasses, headwear, footwear, jewelry), vehicle capabilities (e.g., car, truck, motorcycle), robotic capabilities (e.g., autonomous robot), Internet of Things (IoT) capabilities, drone capabilities, and any combination thereof.
[0042] In the illustrated example, system 280 includes a host processor 282 (e.g., a CPU) having an integrated memory controller (IMC) 284 coupled to system memory 286 (e.g., a dual in-line memory module / DIMM). In this embodiment, an I / O module 288 is coupled to the host processor 282. The illustrated I / O module 288 communicates, for example, with a display 290 (e.g., a touchscreen, liquid crystal display / LCD, or light-emitting diode / LED display) and a network controller 292 (e.g., for wired and / or wireless communication). The host processor 282 may be combined with the I / O module 288, the graphics processor 294, and one or more AI accelerators 296 into a system-on-a-chip (SoC) 298.
[0043] In an embodiment, one or more AI accelerators 296, host processors 282, and / or SoCs 298 execute multiple executable program instructions 300 retrieved from mass storage device 302 and / or system memory 286 to perform the aforementioned method 50. Figure 5 One or more aspects of the language model. Therefore, execution instruction 300 causes host processor 282 to generate multiple candidate lexical units based on input prompts to a draft language model, wherein the multiple candidate lexical units are generated in a single decoding operation. Execution instruction 300 also causes AI accelerator(s)(s)296 to generate output prompts based on the multiple candidate lexical units, input prompts, and a validation language model, and to send the output prompts to the draft language model, wherein the output prompts are generated in a single forward pass operation. Therefore, computing system 280 is considered performance-enhanced at least in the sense that performing a single decoding operation on host processor 282 and a single forward pass operation on one or more AI accelerators 296 provides lower latency and system cost without compromising text generation quality.
[0044] Figure 7 A semiconductor device 350 (e.g., a chip, die, package) is shown. The device 350 includes one or more substrates 352 (e.g., silicon, sapphire, gallium arsenide) and logic 354 (e.g., a transistor array and other integrated circuit / IC components) coupled to the substrate(s) 352. In an embodiment, the logic 354 implements the aforementioned method 50. Figure 5 One or more aspects of ).
[0045] Logic 354 can be implemented at least partially in configurable hardware or fixed-function hardware. In one example, logic 354 includes a transistor channel region located (e.g., embedded) within (one or more) substrates 352. Therefore, the interface between logic 354 and (one or more) substrates 352 may not be an abrupt junction. Logic 354 can also be considered to include an epitaxial layer grown on an initial wafer of (one or more) substrates 352.
[0046] Figure 8 A processor core 400 according to one embodiment is shown. The processor core 400 can be the core of any type of processor (e.g., a microprocessor, embedded processor, digital signal processor (DSP), network processor, or other device for executing code). Although Figure 8 Only one processor core 400 is shown, but the processing element may alternatively include more than one such core. Figure 8The processor core 400 is shown. The processor core 400 may be a single-threaded core, or, in at least one embodiment, the processor core 400 may be a multi-threaded core, since each core may include more than one hardware thread context (or "logical processor").
[0047] Figure 8 Also shown is memory 470 coupled to processor core 400. Memory 470 may be any of a variety of memories (including layers of memory hierarchy) known to those skilled in the art or otherwise available. Memory 470 may include one or more instructions of code 413 to be executed by processor core 400, wherein code 413 may implement the aforementioned method 50. Figure 5 The processor core 400 follows the instruction sequence indicated by code 413. Each instruction can enter the front-end section 410 and be processed by one or more decoders 420. The decoder 420 can generate micro-operations as its output, such as fixed-width micro-operations in a predefined format, or it can generate other instructions, micro-instructions, or control signals that reflect the original code instructions. The front-end section 410 also includes register renaming logic 425 and scheduling logic 430, which typically allocate resources and queue operations corresponding to the translation instructions for execution.
[0048] The processor core 400 shown includes execution logic 450 having a set of execution units 455-1 to 455-N. Some embodiments may include multiple execution units dedicated to a specific function or set of functions. Other embodiments may include only one execution unit, or include a single execution unit capable of performing a specific function. The execution logic 450 shown performs operations specified by code instructions.
[0049] After completing the execution of the operation specified by the code instructions, the backend logic 460 deprecates the instructions of code 413. In one embodiment, processor core 400 allows out-of-order execution but requires instructions to be deprecated in order. Deprecation logic 465 can take various forms known to those skilled in the art (e.g., reorder buffers, etc.). In this way, during the execution of code 413, processor core 400 is transformed at least with respect to the output generated by the decoder, the hardware registers and tables used by register renaming logic 425, and any registers (not shown) modified by execution logic 450.
[0050] although Figure 8Not shown, but the processing element may include other components on the same chip as the processor core 400. For example, the processing element may include memory control logic along with the processor core 400. The processing element may include I / O control logic and / or may include I / O control logic integrated with the memory control logic. The processing element may also include one or more caches.
[0051] Now for reference Figure 9 The diagram shows a block diagram of an embodiment of a computing system 1000 according to an embodiment. Figure 9 A multiprocessor system 1000 including a first processing element 1070 and a second processing element 1080 is shown. Although two processing elements 1070 and 1080 are shown, it should be understood that embodiments of the system 1000 may also include only one such processing element.
[0052] System 1000 is shown as a point-to-point interconnect system, wherein a first processing element 1070 and a second processing element 1080 are coupled via a point-to-point interconnect 1050. It should be understood that... Figure 9 Any or all of the interconnects shown can be implemented as multipoint buses rather than point-to-point interconnects.
[0053] like Figure 9 As shown, each of processing elements 1070 and 1080 can be a multi-core processor, including a first processor core and a second processor core (i.e., processor cores 1074a and 1074b and processor cores 1084a and 1084b). These cores 1074a, 1074b, 1084a, and 1084b can be configured to combine with the above. Figure 8 The instruction code is executed in a similar manner to the method described above.
[0054] Each processing element 1070, 1080 may include at least one shared cache 1896a, 1896b. Shared caches 1896a and 1896b may store data (e.g., instructions) used by one or more components of the processor (e.g., cores 1074a, 1074b and cores 1084a, 1084b). For example, shared caches 1896a, 1896b may locally cache data stored in memories 1032, 1034 for faster access by processor components. In one or more embodiments, shared caches 1896a, 1896b may include one or more intermediate level caches, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, Last Level Cache (LLC), and / or combinations thereof.
[0055] Although only two processing elements 1070 and 1080 are shown, it should be understood that the scope of the embodiments is not limited thereto. In other embodiments, one or more additional processing elements may be present in a given processor. Alternatively, one or more of the processing elements 1070 and 1080 may be elements other than the processor, such as accelerators or field-programmable gate arrays. For example, the additional processing elements(s) may include one or more additional processors identical to the first processor 1070, one or more additional processors heterogeneous or asymmetric to the first processor 1070, accelerators (e.g., graphics accelerators or digital signal processing (DSP) units), field-programmable gate arrays, or any other processing element. There may be various differences between the processing elements 1070 and 1080 in a range of evaluation metrics including architecture, microarchitecture, thermal characteristics, power consumption characteristics, etc. These differences can manifest in the asymmetry and heterogeneity between the processing elements 1070 and 1080. In at least one embodiment, the respective processing elements 1070 and 1080 may be located in the same die package.
[0056] The first processing element 1070 may further include a memory controller logic (MC) 1072 and point-to-point (PP) interfaces 1076 and 1078. Similarly, the second processing element 1080 may include an MC 1082 and PP interfaces 1086 and 1088. Figure 9 As shown, MC 1072 and MC 1082 couple the processor to the corresponding memories, namely memories 1032 and 1034, which may be portions of the main memory locally attached to the respective processor. Although MC 1072 and 1082 are shown as integrated into processing elements 1070 and 1080, in alternative embodiments, the MC logic may be discrete logic external to processing elements 1070 and 1080, rather than integrated therein.
[0057] The first processing element 1070 and the second processing element 1080 can be coupled to the I / O subsystem 1090 via PP interconnects 1076 and 1086, respectively. For example... Figure 9 As shown, the I / O subsystem 1090 includes PP interfaces 1094 and 1098. Furthermore, the I / O subsystem 1090 includes an interface 1092 for coupling the I / O subsystem 1090 to a high-performance graphics engine 1038. In one embodiment, the graphics engine 1038 can be coupled to the I / O subsystem 1090 using a bus 1049. Alternatively, point-to-point interconnects can be used to couple these components.
[0058] Furthermore, the I / O subsystem 1090 can be coupled to the first bus 1016 via interface 1096. In one embodiment, the first bus 1016 may be a peripheral component interconnect (PCI) bus, or a bus such as a fast PCI bus or another third-generation I / O interconnect bus, but the scope of the embodiment is not limited thereto.
[0059] like Figure 9 As shown, various I / O devices 1014 (e.g., biometric scanners, speakers, cameras, sensors) can be coupled to a first bus 1016 along with a bus bridge 1018, which in turn couples the first bus 1016 to a second bus 1020. In one embodiment, the second bus 1020 may be a low pin count (LPC) bus. Various devices can be coupled to the second bus 1020, including, for example, a keyboard / mouse 1012, one or more communication devices 1026, and a data storage unit 1019 such as a disk drive or other mass storage device, which in one embodiment may include code 1030. The code 1030 shown can implement the aforementioned method 50 ( Figure 5 In addition, the audio I / O 1024 can be coupled to the second bus 1020, and the battery 1010 can power the computing system 1000.
[0060] It should be noted that other embodiments are also contemplated. For example, as Figure 9 As an alternative to the point-to-point architecture, the system can implement a multi-point bus or another such communication topology. Alternatively, it can use... Figure 9 The diagram shows more or fewer integrated chips to divide the data. Figure 9 The components in.
[0061] Additional notes and examples:
[0062] Example 1 includes a performance-enhanced computing system comprising: a network controller; a first processor coupled to the network controller; one or more second processors coupled to the network controller; and a memory coupled to the first processor and the one or more second processors, the memory including a plurality of executable program instructions that, when executed by the computing system, cause the computing system to: generate a plurality of candidate lexical units by the first processor based on an input cue to a draft language model, wherein the plurality of candidate lexical units are generated in a single decoding operation; generate an output cue by the one or more second processors based on the plurality of candidate lexical units, the input cue, and a validation language model, wherein the output cue is generated in a single forward pass operation; and send the output cue to the draft language model.
[0063] Example 2 includes the computational system of Example 1, wherein the number of parameters in the validation language model is greater than the number of parameters in the draft language model.
[0064] Example 3 includes the computing system of Example 1, wherein the time spent performing the single decoding operation is less than the time spent performing the single forward pass operation.
[0065] Example 4 includes a computing system of any one of Examples 1 to 3, wherein, when the plurality of executable program instructions are executed, the computing system further causes the memory and computing resources of the one or more second processors to be unavailable to the first processor during the single decoding operation.
[0066] Example 5 includes a computing system of any one of Examples 1 to 3, wherein the width of the plurality of candidate lexical units is greater than one, wherein the first processor includes a central processing unit without artificial intelligence capabilities, and wherein the one or more second processors include an artificial intelligence accelerator.
[0067] Example 6 includes at least one computer-readable storage medium comprising a plurality of executable program instructions that, when executed by a computing system, cause the computing system to: generate a plurality of candidate lexical units by a first processor based on an input cue to a draft language model, wherein the plurality of candidate lexical units are generated in a single decoding operation; generate an output cue by one or more second processors based on the plurality of candidate lexical units, the input cue, and a validation language model, wherein the output cue is generated in a single forward pass operation; and send the output cue to the draft language model.
[0068] Example 7 includes at least one computer-readable storage medium of Example 6, wherein the number of parameters in the validation language model is greater than the number of parameters in the draft language model.
[0069] Example 8 includes at least one computer-readable storage medium of Example 6, wherein the time spent performing the single decoding operation is less than the time spent performing the single forward pass operation.
[0070] Example 9 includes at least one computer-readable storage medium of Example 6, wherein, when the plurality of executable program instructions are executed, the computing system further causes the memory and computing resources of the one or more second processors to be unavailable to the first processor during the single decoding operation.
[0071] Example 10 includes at least one computer-readable storage medium of Example 6, wherein the first processor does not have artificial intelligence capabilities.
[0072] Example 11 includes at least one computer-readable storage medium of any one of Examples 6 to 10, wherein the width of the plurality of candidate morphemes is greater than one.
[0073] Example 12 includes at least one computer-readable storage medium of any one of Examples 6 to 10, wherein the first processor includes a central processing unit, and the one or more second processors include an artificial intelligence accelerator.
[0074] Example 13 includes a semiconductor device comprising: one or more substrates; and logic coupled to the one or more substrates, wherein the logic is implemented at least in part with one or more of configurable hardware and fixed-function hardware, the logic comprising: a first processor configured to generate a plurality of candidate lexical units based on an input cue to a draft language model, wherein the plurality of candidate lexical units are generated in a single decoding operation; and one or more second processors configured to generate an output cue based on the plurality of candidate lexical units, the input cue, and a validation language model, and to send the output cue to the draft language model, wherein the output cue is generated in a single forward pass operation.
[0075] Example 14 includes the semiconductor device of Example 13, wherein the number of parameters in the verification language model is greater than the number of parameters in the draft language model.
[0076] Example 15 includes the semiconductor device of Example 13, wherein the time required to perform the single decoding operation is less than the time required to perform the single forward pass operation.
[0077] Example 16 includes the semiconductor device of Example 13, wherein the logic is configured to deprive the memory and computing resources of the one or more second processors of use by the first processor during the single decoding operation.
[0078] Example 17 includes the semiconductor device of Example 13, wherein the first processor does not have artificial intelligence capabilities.
[0079] Example 18 includes a semiconductor device of any of Examples 13 to 17, wherein the width of the plurality of candidate tokens is greater than one.
[0080] Example 19 includes a semiconductor device of any one of Examples 13 to 17, wherein the first processor includes a central processing unit, and the one or more second processors include an artificial intelligence accelerator.
[0081] Example 20 includes a semiconductor device of any one of Examples 13 to 17, wherein the logic coupled to the one or more substrates includes a transistor region located within the one or more substrates.
[0082] Example 21 includes a method for operating a performance-enhanced computing system, the method comprising: generating a plurality of candidate lexical units by a first processor based on an input cue to a draft language model, wherein the plurality of candidate lexical units are generated in a single decoding operation; generating an output cue by one or more second processors based on the plurality of candidate lexical units, the input cue, and a validation language model, and sending the output cue to the draft language model, wherein the output cue is generated in a single forward pass operation.
[0083] Example 22 includes an apparatus comprising a unit for performing the method of Example 21.
[0084] The embodiments are applicable to all types of semiconductor integrated circuit (“IC”) chips. Examples of these IC chips include, but are not limited to, processors, controllers, chipset assemblies, programmable logic arrays (PLAs), memory chips, network chips, system-on-a-chip (SoCs), SSD / NAND controllers (ASICs), etc. Furthermore, in some of the figures, signal conductors are represented by lines. Some lines may be different to indicate more constituting signal paths; some lines may have numerical labels to indicate the number of constituting signal paths; and / or some lines may have arrows at one or more ends to indicate the primary direction of information flow. However, this should not be interpreted in a limiting manner. Rather, such additional details may be used in conjunction with one or more exemplary embodiments to facilitate a more readily understood understanding of the circuit. Any signal line represented, whether or not it has additional information, may in practice include one or more signals that can travel in multiple directions and can be implemented using any suitable type of signaling scheme, such as digital or analog lines implemented using differential pairs, fiber optic lines, and / or single-ended lines.
[0085] Example dimensions / models / values / ranges may be given, but the embodiments are not limited thereto. As manufacturing technologies (e.g., photolithography) mature, it is expected that devices with smaller dimensions will be manufactured. Furthermore, to simplify illustrations and discussion and to avoid obscuring certain aspects of the embodiments, well-known power / ground connections for IC chips and other components may or may not be shown in the drawings. Additionally, arrangements may be shown in block diagram form to avoid obscuring the embodiments, also taking into account the fact that details regarding the implementation of such block diagram arrangements are highly dependent on the computing system in which the embodiments will be implemented; that is, these details should be entirely within the capabilities of those skilled in the art. Where specific details (e.g., circuitry) are set forth for the purpose of describing exemplary embodiments, it will be apparent to those skilled in the art that the embodiments can be practiced without these specific details or with modifications to them. Therefore, this specification should be considered illustrative rather than restrictive.
[0086] The term “coupling” may be used in this document to refer to any type of direct or indirect relationship between the components under discussion, and may be applied to electrical, mechanical, fluid, optical, electromagnetic, electromechanical, or other connections. Furthermore, the terms “first,” “second,” etc., may be used herein for ease of discussion only, and unless otherwise stated, these terms do not have any specific temporal or sequential meaning.
[0087] As used in this application and claims, a list of items connected by the term "one or more of..." can represent any combination of the listed terms. For example, the phrase "one or more of A, B, or C" can mean A; B; C; A and B; A and C; B and C; or A, B, and C.
[0088] Those skilled in the art will understand from the foregoing description that the overall technology of the embodiments can be implemented in various forms. Therefore, although embodiments have been described in conjunction with specific examples, the actual scope of the embodiments should not be so limited, as other modifications will become apparent to those skilled in the art upon studying the drawings, specification, and appended claims.
Claims
1. A computing system, comprising: Network controller; A first processor coupled to the network controller; One or more second processors coupled to the network controller; as well as A memory coupled to the first processor and the one or more second processors, the memory including a plurality of executable program instructions, which, when executed by the computing system, cause the computing system to: The first processor generates multiple candidate lexical units based on input prompts to the draft language model, wherein the multiple candidate lexical units are generated in a single decoding operation; The output prompt is generated by the one or more second processors based on the plurality of candidate lexical units, the input prompt, and the validation language model, wherein the output prompt is generated in a single forward pass operation; and The output prompt is sent to the draft language model.
2. The computing system according to claim 1, wherein, The number of parameters in the validation language model is greater than the number of parameters in the draft language model.
3. The computing system according to claim 1, wherein, The time spent performing the single decoding operation is less than the time spent performing the single forward pass operation.
4. The computing system according to any one of claims 1 to 3, wherein, When the plurality of executable program instructions are executed, the computing system further causes the memory and computing resources of the one or more second processors to be unavailable to the first processor during the single decoding operation.
5. The computing system according to any one of claims 1 to 3, wherein, The width of the plurality of candidate lexical units is greater than one, wherein the first processor includes a central processing unit without artificial intelligence capabilities, and wherein the one or more second processors include an artificial intelligence accelerator.
6. At least one computer-readable storage medium comprising a plurality of executable program instructions, which, when executed by a computing system, cause the computing system to: The first processor generates multiple candidate lexical units based on input prompts to the draft language model, among which, The multiple candidate lexical units are generated in a single decoding operation; An output prompt is generated by one or more second processors based on the plurality of candidate lexical units, the input prompt, and the validation language model, wherein the output prompt is generated in a single forward pass operation; as well as The output prompt is sent to the draft language model.
7. The at least one computer-readable storage medium according to claim 6, wherein, The number of parameters in the validation language model is greater than the number of parameters in the draft language model.
8. The at least one computer-readable storage medium according to claim 6, wherein, The time spent performing the single decoding operation is less than the time spent performing the single forward pass operation.
9. The at least one computer-readable storage medium according to claim 6, wherein, When the plurality of executable program instructions are executed, the computing system further causes the memory and computing resources of the one or more second processors to be unavailable to the first processor during the single decoding operation.
10. The at least one computer-readable storage medium according to claim 6, wherein, The first processor does not have artificial intelligence capabilities.
11. At least one computer-readable storage medium according to any one of claims 6 to 10, wherein, The width of the plurality of candidate lexical units is greater than one.
12. At least one computer-readable storage medium according to any one of claims 6 to 10, wherein, The first processor includes a central processing unit, and the one or more second processors include an artificial intelligence accelerator.
13. A semiconductor device, comprising: One or more substrates; as well as Logic coupled to the one or more substrates, wherein the logic is implemented at least in part in one or more of configurable hardware and fixed-function hardware, the logic comprising: A first processor is configured to generate a plurality of candidate lexical units based on input prompts to a draft language model, wherein the plurality of candidate lexical units are generated in a single decoding operation; and One or more second processors are configured to generate output prompts based on the plurality of candidate lexical units, the input prompts, and the validation language model, and to send the output prompts to the draft language model, wherein the output prompts are generated in a single forward pass operation.
14. The semiconductor device according to claim 13, wherein, The number of parameters in the validation language model is greater than the number of parameters in the draft language model.
15. The semiconductor device according to claim 13, wherein, The time spent performing the single decoding operation is less than the time spent performing the single forward pass operation.
16. The semiconductor device according to claim 13, wherein, The logic is used to prevent the memory and computing resources of the one or more second processors from being used by the first processor during the single decoding operation.
17. The semiconductor device according to claim 13, wherein, The first processor does not have artificial intelligence capabilities.
18. The semiconductor device according to any one of claims 13 to 17, wherein, The width of the plurality of candidate lexical units is greater than one.
19. The semiconductor device according to any one of claims 13 to 17, wherein, The first processor includes a central processing unit, and the one or more second processors include an artificial intelligence accelerator.
20. The semiconductor device according to any one of claims 13 to 17, wherein, The logic coupled to the one or more substrates includes transistor regions located within the one or more substrates.