Pipelined execution of generative artificial intelligence models

By employing a pipelined execution method for generative artificial intelligence models, where model portions are loaded and deserialized in parallel, the inefficiencies and high latency in generating responses are addressed, resulting in improved resource utilization and response speed.

WO2025111787A1PCT designated stage expired Publication Date: 2025-06-05QUALCOMM INC +9
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/134656
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Generative artificial intelligence models, such as large language models, are computationally expensive and inefficient in generating responses to queries, especially on devices with limited resources, leading to high latency and significant resource utilization.

Method used

The method involves pipelined execution of inference operations using generative artificial intelligence models, where portions of the model are loaded and deserialized in parallel, allowing for simultaneous generation of inferences and reduction of latency.

Benefits of technology

This approach significantly reduces the latency in generating responses to queries, optimizes computational resource utilization, and allows for more efficient deployment of generative AI models on devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023134656_05062025_PF_FP_ABST
    Figure CN2023134656_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Techniques and apparatus for efficiently generating a response to an input query using a generative artificial intelligence model in a pipelined execution environment. An example method generally includes loading a first portion of a machine learning model, wherein the first portion of the machine learning model is associated with a first inference; loading a second portion of the machine learning model, wherein the second portion of the machine learning model is associated with a second inference; and while loading the second portion of the machine learning model, generating the first inference based on an input data set and the first portion of the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

PIPELINED EXECUTION OF GENERATIVE ARTIFICIAL INTELLIGENCE MODELS

[0001] INTRODUCTION

[0002] Aspects of the present disclosure relate to generative artificial intelligence models, and more specifically to pipelined execution of inferencing operations using generative artificial intelligence models.

[0003] Generative artificial intelligence models can be used in various environments in order to generate a response to an input query. For example, generative artificial intelligence models can be used in chatbot applications in which large language models (LLMs) are used to generate an answer, or at least a response, to an input query. Other examples in which generative artificial intelligence models can be used include stable diffusion, in which a model generates an image from an input text description of the content of the desired image, and decision transformers, in which future actions are predicted based on sequences of prior actions within a given environment.

[0004] Generally, generating a response to a query using generative artificial intelligence models may be computationally expensive. For example, in a chatbot deployment in which a large language model is used to generate a response to a query formatted as a text query, a response to the query may be generated using a pass through the large language model for each token (e.g., word or part of word) generated as part of the response. The output of each pass may be a probability distribution on a set of tokens (e.g., words or parts of words) from which the next token (e.g., word or part of word) may be selected, either by sampling or based on maximum likelihood, for example. Because a pass through a large language model is used to generate each word (or token (s) ) in a response to a query, the computational expense may be modeled as the product of the number of words included in the response and the computational resource expense (e.g., in terms of processing power, memory bandwidth, and / or other compute resources used) of performing a pass through the large language model, which generally increases as the number of parameters within the large language model increases.

[0005] BRIEF SUMMARY

[0006] Certain aspects of the present disclosure provide a method for efficiently generating a response to an input query using a generative artificial intelligence model in  a pipelined execution environment. The method generally includes loading a first portion of a machine learning model, wherein the first portion of the machine learning model is associated with a first inference; loading a second portion of the machine learning model, wherein the second portion of the machine learning model is associated with a second inference; and while loading the second portion of the machine learning model, generating the first inference based on an input data set and the first portion of the machine learning model.

[0007] Other aspects provide processing systems configured to perform the aforementioned methods as well as those described herein; non-transitory, computer-readable media comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the aforementioned methods as well as those described herein; a computer program product embodied on a computer-readable storage medium comprising code for performing the aforementioned methods as well as those further described herein; and a processing system comprising means for performing the aforementioned methods as well as those further described herein.

[0008] The following description and the related drawings set forth in detail certain illustrative features of one or more aspects.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The appended figures depict only certain aspects of this disclosure and are therefore not to be considered limiting of the scope of this disclosure.

[0010] FIG. 1 illustrates a timeline for generating a response to an input query using a generative artificial intelligence model.

[0011] FIG. 2 illustrates a timeline for generating a response to an input query using a pipelined generative artificial intelligence model, according to aspects of the present disclosure.

[0012] FIG. 3 illustrates a timeline for generating a response to an input query using a pipelined generative artificial intelligence model and selective loading of the generative artificial intelligence model, according to aspects of the present disclosure.

[0013] FIG. 4 illustrates example operations for generating a response to an input query using a pipelined generative artificial intelligence model, according to aspects of the present disclosure.

[0014] FIG. 5 illustrates an example processing system for generating a response to an input query using a pipelined generative artificial intelligence model, according to aspects of the present disclosure.

[0015] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the drawings. It is contemplated that elements and features of one aspect may be beneficially incorporated in other aspects without further recitation.DETAILED DESCRIPTION

[0016] Aspects of the present disclosure provide apparatus, methods, processing systems, and computer-readable mediums for efficiently generating responses to input queries using generative artificial intelligence models.

[0017] Generally, generative artificial intelligence models generate a response to a query input into the model. For example, a large language model (LLM) deployed within a chatbot can generate a response to a query using multiple passes through the large language model, with each successive pass being based on the query and the tokens (or words or parts thereof) generated using previous passes through the large language model. Generally, these large language models may include millions, or even billions, of weights or parameters within the model. Because of the size of these models and the operations performed on each token to predict what should be the next token generated in response to a query and the previously generated tokens, it may not be practical, or even possible, to deploy large language models on a variety of devices which may have limited memory, storage, and / or processing capabilities relative to cloud compute instances on which large language models typically operate. Further, the computational complexity involved in generating a response to a query provided as input into a model may involve significant energy expenditure, processing time, memory utilization, and / or other resource utilization which may prevent compute resources from being used for other tasks.

[0018] Generally, tokens generated by a generative artificial intelligence model are generated sequentially. That is, a first token is generated, then a second token is generated using the first token as context, a third token is generated using the first and second tokens  as context, and so on. Because tokens are generally generated sequentially, delays in generating the first token generally cause cascading delays in generating subsequent tokens.

[0019] Aspects of the present disclosure provide techniques for efficiently generating responses to a query input into a generative artificial intelligence model. Generally, execution of the generative artificial intelligence model may be pipelined using a multi-stage pipeline including at least a first stage for loading the generative artificial intelligence model (or a relevant portion thereof) into memory and a second stage for generating an inference based on the generative artificial intelligence model and an input query (and, in some aspects, previously generated tokens serving as context for the next token to be generated using the generative artificial intelligence model) . By pipelining execution of inference operations using the generative artificial intelligence model, a latency between receiving a request to generate a response to an input query and outputting a first token in the response may be significantly reduced, which may allow for more rapid generation of responses to queries using generative artificial intelligence models. In turn, this may reduce overall computational resource utilization on a device on which these generative artificial intelligence models operate, which may reduce the amount of power consumed by the device during inference operations and allow for computational resources (e.g., processor time, memory, etc. ) to be allocated for other operations.

[0020] Response Generation in Generative Artificial Intelligence Models

[0021] Generally, autoregressive token generation (e.g., in large language models) may take historical tokens as an input in order to generate an output. That is, autoregressive token generation may be represented by the expression: xt ~ p (x|x0, x1, …, xt-1) → xt+1 ~ p (x|x0, x1, …, xt-1, xt)

[0022] where xt represents a sequence of tokens generated at time t, having a conditional probability p conditioned on the selection of tokens x0 through xt-1, and xt+1 represents a sequence of tokens generated at time t + 1, having a conditional probability p conditioned on on the selection of tokens x0 through xt. Generally, a single token may be generated each time an autoregressive model is executed, which means that N inferences may be performed to generate a sequence of N tokens. As discussed above, speculative decoding techniques can be used to accelerate token generation by using a draft model,  smaller in size than the target model, that speculatively generates tokens, with the target model being used to verify the tokens (speculatively) generated by the draft model.

[0023] In a speculative decoding pipeline, the draft model may speculatively generate n tokens autoregressively, according to the expression:

[0024] where t corresponds to a point in time and corresponds to the conditional probability distribution associated with a selected token x at time t conditioned on the selection of tokens x0 through xt-1.

[0025] The target model takes the generated n tokens and processes the n tokens in parallel to generate probability distributions for each of the n tokens, according to the expression:

[0026] where k corresponds to a token index relative to the generated n tokens.

[0027] The target model can then verify the tokens generated by the draft model by comparing distributions from the draft model and target model to determine whether a token is accepted or rejected. A given token may be accepted when  for some function f and some threshold α (also known as an acceptance rate) . Otherwise, the token may be rejected. The final token may then be generated at the first rejection position or at the last position n based on some function

[0028] Speculative decoding, with an acceptance rate of α, may result in cost reductions relative to using a single autoregressive model to generate tokens iteratively. Inference cost savings, relative to iterative token generation, may be represented by the expression:

[0029] Consider the example, for N = 1000, Ctarget=10, Cdraft=1, n=4, α=3, wherein N corresponds to a number of tokens, CAR corresponds to a computational cost using an acceptance rate of α, Ctarget corresponds to a computational cost of  generating a set of tokens using the target model, CSD corresponds to a computational cost of generating a set of tokens using speculative decoding techniques, Cdraft corresponds to a computational cost of generating a set of tokens using the draft model, and n corresponds to a number of tokens generated speculatively generated tokens generated through a single pass through an autoregressive model. In such an example, speculative decoding may result in a 35%reduction in computational expense relative to autoregressive iterative token generation alone.

[0030] However, speculative decoding on a per-token basis, as discussed, may impose limits on the rate at which tokens are generated, as a first token may be sampled individually by a draft model and then verified by a target model before the next token is sampled by the draft model and verified by the target model. That is, generating response to an input query using per-token speculative decoding techniques may involve executing the draft model and target model for each token generated as part of a response to the input query, which may use significant amounts of computational resources (e.g., processor time, memory, memory bandwidth, etc. ) in order to generate the response.

[0031] Example Execution of Inference Operations Using Generative Artificial Intelligence Models

[0032] FIG. 1 illustrates an example timeline 100 for execution of inference operations using a generative artificial intelligence model.

[0033] Generally, as illustrated in the timeline 100, execution of inference operations using a generative artificial intelligence model may be divided into a model initialization phase 110 and an inference phase 120. The model initialization phase 110 is generally executed prior to execution of the inference phase. Within the model initialization phase 110, a binary file (e.g., executable application) may be loaded into memory at block 112. The binary may include the generative artificial intelligence model component and other components that may be used to generate a response to an input query. For example, in a scenario in which a large language model is executed on a device, the generative artificial intelligence model may be a transformer neural network, such as a Bidirectional Encoder Representations from Transformers (BERT) model, and additional components loaded into memory from the binary may include various components used to project data used within the machine learning model into keys and values. After loading the binary into memory at block 112, the binary may be deserialized to create an executable context for the machine learning model at block 114.

[0034] The deserialized generative artificial intelligence model may subsequently be used in the inference phase 120 to generate a response to an input query, with each inference performed by the generative artificial intelligence model corresponding to a token (e.g., word or part of word) to be included in the response to the input query. As illustrated, a series of inferences may be generated sequentially using the generative artificial intelligence model at inference blocks 122, 124, 126, and 128. For example, as illustrated, a first inference may be generated at index 0 (corresponding to a first inference block 122) , a second inference may be generated at index 1 (corresponding to a second inference block 124) based on the first inference, a third inference may be generated at index 2 (corresponding to a third inference block 126) based on the first and second inferences, and a third inference may be generated at index 3 (corresponding to a fourth inference block 128) based on the first, second, and third inferences.

[0035] However, a token may not be output from the generative artificial intelligence model until the tokens are passed sequentially through key-value projection blocks 130, 132, 134, and 136 (corresponding to the inference blocks 122, 124, 126, and 128, respective) , as each inference generated by the generative artificial intelligence model may correspond to a key from which a corresponding value is projected. Thus, the first token outputted by the generative artificial intelligence model in the timeline 100 may be output after execution of the key-value projection block 130 at index 0, which executes after the inference block executes at index 3 (discussed above) . Thus, the total latency involved in generating the initial token output as a response to the input query may include the time spent in the model initialization phase and the time spent generating inferences using the generative artificial intelligence model. The second token may subsequently be output by the generative artificial intelligence model in the timeline 100 after execution of the key-value projection block 132; the third token may subsequently be output by the generative artificial intelligence model in the timeline 100 after execution of the key-value projection block 134; and the fourth token may subsequently be output by the generative artificial intelligence model in the timeline 100 after execution of the key-value projection block 136.

[0036] For example, assume that a generative artificial intelligence model split into four stages, indexed 0 through 3, has the execution timing properties illustrated in Table 1 below:

[0037] Table 1: Execution Timing Properties of a Generative Artificial Intelligence Model

[0038] The total latency involved in generating the first token included in a response to an input query using sequential execution techniques, as illustrated in the timeline 100, may be calculated as the sum of the time spent in reading the binary, the time spent deserializing the machine learning model from the binary, and the time spent in performing inferencing operations using the machine learning model. In this example, the total binary read time, calculated as the sum of loading each portion of the machine learning model from storage, may be 2, 700 milliseconds. The total deserialization time may be 2, 700 milliseconds. Finally, the total inferencing time may be 2,000 milliseconds. Thus, the total latency prior to generation and output of an initial token in a response to an input query may be 7, 400 milliseconds (or 7.4 seconds) .

[0039] However, it should be noted that many of the operations performed in sequential execution of a generative artificial intelligence model do not depend on each other. For example, once a first portion of the generative artificial intelligence model (e.g., the portion corresponding to Split #0 in Table 1) is loaded into memory, the first portion may be deserialized without depending upon loading a second portion of the generative artificial intelligence model (e.g., the portion corresponding to Split #1 in Table 1) into memory. Similarly, after the first portion of the generative artificial intelligence model is deserialized, a first inference can be generated without depending on deserialization of the second portion of the generative artificial intelligence model. Thus, it can be seen that sequential loading and execution of inference operations using a generative artificial intelligence model may result in an inefficient use of resources and introduce unnecessary latency into the generation of responses to input queries using generative artificial intelligence models.

[0040] FIG. 2 illustrates an example timeline 200 for pipelined execution of inferencing operations using a generative artificial intelligence model, according to aspects of the present disclosure.

[0041] As illustrated, a pipeline for efficient execution of inferencing operations using a generative artificial intelligence model may be divided into three pipeline stages: a first stage corresponding to reading a binary including the generative artificial intelligence model, a second stage corresponding to deserializing the binary, and a third stage corresponding to inferencing using the deserialized binary.

[0042] To generate the first token included in a response to an input query, during a first time block 210, a first portion of the generative artificial intelligence model may be read from a binary at a model read block 212. During a second time block 220, the first portion of the generative artificial intelligence model may be deserialized, at a deserialization block 224, into executable instructions that, when executed, cause a processor to generate an inference based on an input into the generative artificial intelligence model. While the first portion of the generative artificial intelligence model is deserialized at the deserialization block 224, a second portion of the generative artificial intelligence model may be read from the binary at a model read block 222. Finally, during a third time block 230, a first inference may be generated based on an input into the generative artificial intelligence model at an inference block 236. Simultaneously, or at least substantially simultaneously, the second portion of the generative artificial intelligence model may be deserialized at a deserialization block 234, and a third portion of the generative artificial intelligence model may be loaded from the binary at a model read block 232.

[0043] Subsequent operations may similarly be performed for the remaining portions of the generative artificial intelligence model until a defined set of inferences have been performed. For example, as illustrated, the pipelined execution of the machine learning model, including the loading of portions of a generative artificial intelligence model, deserialization of a previously loaded portion of the generative artificial intelligence model, and inferencing based on deserialized portions of the generative artificial intelligence model may continue until an inference is generated for each of the portions of the generative artificial intelligence model corresponding to splits 0 through 3 illustrated in FIG. 2 and described in Table 1 above. That is, during a fourth time block 240, a fourth portion of the generative artificial intelligence model may be loaded from  the binary in a model read block 242, the third portion of the generative artificial intelligence model may be deserialized in a deserialization block 244, and a second inference may be generated in an inference block 246 based on an input into the generative artificial intelligence model and the deserialized second portion of the machine learning model. During a fifth time block 250, because there may be no further portions of the generative artificial intelligence model to be loaded from the binary, the operations performed in the fifth time block may be reduced to deserialization of the fourth portion of the machine learning model in a deserialization block 252 and generation of a third inference in an inference block 254 based on an input into the generative artificial intelligence model and the deserialized third portion of the machine learning model. Finally, during a sixth time block 260, because there are no further portions of the generative artificial intelligence model to be deserialized, the operations performed in the sixth time block may be reduced to generation of a fourth inference in an inference block 262 based on an input into the generative artificial intelligence model and the deserialized fourth portion of the machine learning model. Subsequently, the inferences may be processed (e.g., sequentially) through one or more key-value projection blocks 270, 272, 274, and 276, and tokens may be output for each of the first through fourth inferences.

[0044] The latency involved in generating an initial token included in a response to an input query using the pipelining techniques illustrated in FIG. 2 may be significantly reduced relative to the latency involved in generating an initial token included in a response to an input query using the sequential execution techniques illustrated in FIG. 1. In this example, the total latency may be calculated as the sum of the execution times at each of the pipeline stages prior to projection of inferences into output tokens using the key-value projection blocks illustrated in the timeline 200. Thus, the total latency involved in generating the first token may correspond to the sum of (1) the time to read the first portion of the generative artificial intelligence model from a binary, (2) the larger of the time to read the second portion of the generative artificial intelligence model from the binary and the time to deserialize the first portion of the generative artificial intelligence model, (3) the largest of the time to read the third portion of the generative artificial intelligence model from the binary, the time to deserialize the second portion of the generative artificial intelligence model, and the time to generate the first inference, (4) the largest of the time to read the fourth portion of the generative artificial intelligence model from the binary, the time to deserialize the third portion of the generative artificial  intelligence model, and the time to generate the second inference, (5) the larger of the time to deserialize the fourth portion of the generative artificial intelligence model and the time to generate the third inference, and (6) the time to generate the fourth inference. This total latency may also be represented as the sum of the durations of the first time block 210, the second time block 220, the third time block 230, the fourth time block 240, the fifth time block 250, and the sixth time block 260. Based on the timing discussed in Table 1 above, thus, the total latency may be represented as: 800 + max (550, 460) +max(550, 470, 500) + max (800, 820, 500) + max (950, 500) + 500 = 4, 200 milliseconds (or 4.2 seconds) . This represents a total time savings of 3, 200 ms (3.2 seconds) relative to the sequential illustration illustrated in FIG. 1 and discussed above.

[0045] FIG. 3 illustrates a timeline 300 for generating a response to an input query using a pipelined generative artificial intelligence model and selective loading of the generative artificial intelligence model, according to aspects of the present disclosure.

[0046] In the timeline 300, an execution pipeline for generating a response to an input query using a generative artificial intelligence model may be divided into two stages: a first stage for initializing a portion of a generative artificial intelligence model based on selective reading of data from a memory address to which the portion of the machine learning model is mapped and a second stage for generating an inference based on the initialized portion of the machine learning model. As illustrated, during a first time block 310, thus, a first portion of the generative artificial intelligence model may be loaded into memory, and a buffer may be allocated for use across different portions of the generative artificial intelligence model. During a second time block 320, a second portion of the generative artificial intelligence model may be loaded into memory. Simultaneously, or at least substantially simultaneously, a first inference may be generated based on the first portion of the generative artificial intelligence model and an input into the generative artificial intelligence model. After the first inference is generated, the first portion of the generative artificial intelligence model, or at least portions of the model that are not shared with other portions of the generative artificial intelligence model, may be deinitialized.

[0047] This process may continue for each portion of the generative artificial intelligence model, similar to the process discussed above with respect to FIG. 2. That is, during a third time block 330, as illustrated, a third portion of the generative artificial intelligence model may be loaded into memory simultaneously with the generation of a second inference using the second portion of the generative artificial intelligence model.  During a fourth time block 340, a fourth portion of the generative artificial intelligence model may be loaded into memory simultaneously with the generation of a third inference using the third portion of the generative artificial intelligence model. Additionally, during the time block in which the last inference is generated (e.g., as illustrated, at the fifth time block 350) , a first key-value projection block may be loaded into memory. A first token will be generated after the different portions of the generative artificial intelligence model execute. Subsequently, as illustrated in the sixth time block 360, an autoregressive token generation loop may be initiated with the key-value projection block loaded into memory at the fifth time block 350. A first inference (corresponding to the outputs of the generative artificial intelligence model) may be generated by the key-value projection block loaded into memory at the fifth time block 350, and at least a second key-value projection block may simultaneously or at least substantially simultaneously loaded into memory. The at least the second key-value projection block may be subsequently used at key-value projection blocks 370, 380, and 390 to identify and output tokens in each iterative step corresponding to the inferences generated by the various portions of the generative artificial intelligence model and key-value projection block itself in a final time block, discussed above.

[0048] The techniques illustrated in FIG. 3 may reduce the memory footprint involved in generating responses to input queries using generative artificial intelligence models and may further reduce token output latency, as smaller amounts of data may be loaded into memory prior to generation of inferences and projection of these inferences into tokens.

[0049] Example Operations for Pipelined Execution of Inference Operations Using Generative Artificial Intelligence Models

[0050] FIG. 4 illustrates example operations 400 that may be performed by a computing device to generate a response to an input query using pipelined generative artificial intelligence models, according to aspects of the present disclosure. The computing device for performing the operations 400 may be a device on which at least a draft model can be deployed, such as a smartphone, a tablet computer, a laptop computer, a desktop, a server, a cloud compute instance hosted in a distributed computing environment, or the like.

[0051] As illustrated, the operations 400 begin at block 410, with loading a first portion of a machine learning model. Generally, the first portion of the machine learning model may be associated with a first inference.

[0052] At block 420, the operations 400 proceed with loading a second portion of the machine learning model. The second portion generally may be associated with a second inference.

[0053] At block 430, while loading the second portion of the machine learning model, operations 400 proceed with generating the first inference based on an input data set and the first portion of the machine learning model.

[0054] In some aspects, loading the first portion of the machine learning model may include loading a serialized version of the first portion of the machine learning model during a first time period. The serialized version of the first portion of the machine learning model may be deserialized during a second time period subsequent to the first time period.

[0055] In some aspects, loading the second portion of the machine learning model may include loading a serialized version of the second portion of the machine learning model during the second time period. The serialized version of the second portion of the machine learning model may be deserialized during a third time period subsequent to the second time period, and the first inference may be generated during the third time period.

[0056] In some aspects, the operations 400 may further include loading a serialized version of a third portion of the machine learning model during the third time period. Generally, the third portion of the machine learning model may be associated with a third inference. The serialized version of the third portion of the machine learning model may be deserialized during a fourth time period subsequent to the third time period, and the third inference may be generated during a fifth time period subsequent to the fourth time period.

[0057] In some aspects, the operations 400 may further include generating a second inference based on the input data set, the loaded second portion of the machine learning model, and the generated first inference without concurrently loading a third portion of the machine learning model.

[0058] In some aspects, loading the first portion of the machine learning model may include initializing the first portion of the machine learning model based on a memory  address associated with the first portion of the machine learning model. An input / output buffer shared by at least the first portion of the machine learning model and the second portion of the machine learning model may be allocated. In some aspects, loading the second portion of the machine learning model may include initializing the second portion of the machine learning model based on a memory address associated with the second portion of the machine learning model. Generally, the second portion of the machine learning model uses the allocated input / output buffer. Components associated with the first portion of the machine learning model that are not shared by the second portion of the machine learning model may be deallocated.

[0059] In some aspects, the first portion of the machine learning model comprises a transformer neural network block and a key-value projection block. The first inference may be generated using the transformer neural network block, and an output of the machine learning model may be generated based on the first inference and the key-value projection block.

[0060] Example Processing Systems for Pipelined Execution of Inference Operations Using Generative Artificial Intelligence Models

[0061] FIG. 5 depicts an example processing system 500 for generating a response to an input query using a pipelined generative artificial intelligence model, such as described herein for example with respect to FIG. 4.

[0062] The processing system 500 includes a central processing unit (CPU) 502, which in some examples may be a multi-core CPU. Instructions executed at the CPU 502 may be loaded, for example, from a program memory associated with the CPU 502 or may be loaded from a memory partition (e.g., of memory 524) .

[0063] The processing system 500 also includes additional processing components tailored to specific functions, such as a graphics processing unit (GPU) 504, a digital signal processor (DSP) 506, a neural processing unit (NPU) 508, and a connectivity component 512.

[0064] An NPU, such as the NPU 508, is generally a specialized circuit configured for implementing control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs) , deep neural networks (DNNs) , random forests (RFs) , and the like. An NPU may sometimes alternatively be referred to as a neural signal processor (NSP) , tensor processing unit  (TPU) , neural network processor (NNP) , intelligence processing unit (IPU) , vision processing unit (VPU) , or graph processing unit.

[0065] NPUs, such as the NPU 508, are configured to accelerate the performance of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, a plurality of NPUs may be instantiated on a single chip, such as a system on a chip (SoC) , while in other examples such NPUs may be part of a dedicated neural-network accelerator.

[0066] NPUs may be optimized for training or inference, or in some cases configured to balance performance between both. For NPUs that are capable of performing both training and inference, the two tasks may still generally be performed independently.

[0067] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly compute-intensive operation that involves inputting an existing dataset (often labeled or tagged) , iterating over the dataset, and then adjusting model parameters, such as weights and biases, in order to improve model performance. Generally, optimizing based on a wrong prediction involves propagating back through the layers of the model and determining gradients to reduce the prediction error.

[0068] NPUs designed to accelerate inference are generally configured to operate on complete models. Such NPUs may thus be configured to input a new piece of data and rapidly process this new piece through an already trained model to generate a model output (e.g., an inference) .

[0069] In some implementations, the NPU 508 is a part of one or more of the CPU 502, the GPU 504, and / or the DSP 506. These may be located on a user equipment (UE) in a wireless communication system or another computing device.

[0070] In some examples, a connectivity component 512 may include subcomponents, for example, for third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., Long-Term Evolution (LTE) ) , fifth generation (5G) connectivity (e.g., New Radio (NR) ) , Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. The connectivity component 512 may be further coupled to one or more antennas 514.

[0071] The processing system 500 may also include one or more sensor processing units 516 associated with any manner of sensor, one or more image signal processors  (ISPs) 518 associated with any manner of image sensor, and / or a navigation processor 520, which may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.

[0072] The processing system 500 may also include one or more input and / or output devices 522, such as screens, touch-sensitive surfaces (including touch-sensitive displays) , physical buttons, speakers, microphones, and the like.

[0073] In some examples, one or more of the processors of the processing system 500 may be based on an ARM or RISC-V instruction set.

[0074] The processing system 500 also includes a memory 524, which is representative of one or more static and / or dynamic memories, such as a dynamic random access memory, a flash-based static memory, and the like. In this example, the memory 524 includes computer-executable components, which may be executed by one or more of the aforementioned processors of the processing system 500.

[0075] In particular, in this example, the memory 524 includes a model loading component 524A, an inference generating component 524B, and a machine learning model component 524C. The depicted components, and others not depicted, may be configured to perform various aspects of the methods described herein.

[0076] Generally, the processing system 500 and / or components thereof may be configured to perform the methods described herein.

[0077] Example Clauses

[0078] Implementation details of various aspects of the present disclosure are set forth in the following numbered clauses:

[0079] Clause 1: A processor-implemented method, comprising: loading a first portion of a machine learning model, wherein the first portion of the machine learning model is associated with a first inference; loading a second portion of the machine learning model, wherein the second portion of the machine learning model is associated with a second inference; and while loading the second portion of the machine learning model, generating the first inference based on an input data set and the loaded first portion of the machine learning model.

[0080] Clause 2: The method of Clause 1, wherein loading the first portion of the machine learning model comprises: loading a serialized version of the first portion of the  machine learning model during a first time period; and deserializing the serialized version of the first portion of the machine learning model during a second time period subsequent to the first time period.

[0081] Clause 3: The method of Clause 2, wherein loading the second portion of the machine learning model comprises: loading a serialized version of the second portion of the machine learning model during the second time period; and deserializing the serialized version of the second portion of the machine learning model during a third time period subsequent to the second time period, wherein the first inference is generated during the third time period.

[0082] Clause 4: The method of Clause 3, further comprising: loading a serialized version of a third portion of the machine learning model during the third time period, the third portion of the machine learning model being associated with a third inference; deserializing the serialized version of the third portion of the machine learning model during a fourth time period subsequent to the third time period; and generating the third inference during a fifth time period subsequent to the fourth time period.

[0083] Clause 5: The method of any of Clauses 1 through 4, further comprising generating a second inference based on the input data set, the loaded second portion of the machine learning model, and the generated first inference without concurrently loading a third portion of the machine learning model.

[0084] Clause 6: The method of any of Clauses 1 through 5, wherein: the first portion of the machine learning model comprises a transformer neural network block and a key-value projection block, the first inference is generated using the transformer neural network block, and an output of the machine learning model is generated based on the first inference and the key-value projection block.

[0085] Clause 7: The method of any of Clauses 1 through 6, wherein loading the first portion of the machine learning model comprises: initializing the first portion of the machine learning model based on a memory address associated with the first portion of the machine learning model; and allocating an input / output buffer shared by at least the first portion of the machine learning model and the second portion of the machine learning model.

[0086] Clause 8: The method of Clause 7, wherein loading the second portion of the machine learning model comprises: initializing the second portion of the machine  learning model based on a memory address associated with the second portion of the machine learning model, wherein the second portion of the machine learning model uses the allocated input / output buffer; and deallocating components associated with the first portion of the machine learning model that are not shared by the second portion of the machine learning model.

[0087] Clause 9: A processing system, comprising: at least one memory having executable instructions stored thereon; and one or more processors configured to execute the executable instructions in order to cause the processing system to perform the operations of any of Clauses 1 through 8.

[0088] Clause 10: A processing system comprising means for performing the operations of any of Clauses 1 through 8.

[0089] Clause 11: A non-transitory computer-readable medium having executable instructions stored thereon which, when executed by one or more processors, perform the operations of any of Clauses 1 through 8.

[0090] Additional Considerations

[0091] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not limiting of the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For instance, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method that is practiced using other structure, functionality, or structure and functionality in addition to, or other than, the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0092] As used herein, the word “exemplary” means “serving as an example, instance, or illustration. ” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0093] As used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c) .

[0094] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure) , ascertaining, and the like. Also, “determining” may include receiving (e.g., receiving information) , accessing (e.g., accessing data in a memory) , and the like. Also, “determining” may include resolving, selecting, choosing, establishing, and the like.

[0095] The methods disclosed herein comprise one or more steps or actions for achieving the methods. The method steps and / or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Further, the various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component (s) and / or module (s) , including, but not limited to a circuit, an application specific integrated circuit (ASIC) , or processor. Generally, where there are operations illustrated in figures, those operations may have corresponding counterpart means-plus-function components with similar numbering.

[0096] The following claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims. Within a claim, reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more. ” Unless specifically stated otherwise, the term “some” refers to one or more. No claim element is to be construed under the provisions of 35 U.S.C. § 112 (f) unless the element is expressly recited using the phrase “means for” or, in the case of a method claim, the element is recited using the  phrase “step for. ” All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

Claims

1.A processing system, comprising:at least one memory having executable instructions stored thereon; andone or more processors configured to execute the executable instructions to cause the processing system to:load a first portion of a machine learning model, wherein the first portion of the machine learning model is associated with a first inference;load a second portion of the machine learning model, wherein the second portion of the machine learning model is associated with a second inference; andwhile loading the second portion of the machine learning model, generate the first inference based on an input data set and the loaded first portion of the machine learning model.2.The processing system of Claim 1, wherein to load the first portion of the machine learning model, the one or more processors are configured to cause the processing system to:load a serialized version of the first portion of the machine learning model during a first time period; anddeserialize the serialized version of the first portion of the machine learning model during a second time period subsequent to the first time period.3.The processing system of Claim 2, wherein to load the second portion of the machine learning model, the one or more processors are configured to cause the processing system to:load a serialized version of the second portion of the machine learning model during the second time period; anddeserialize the serialized version of the second portion of the machine learning model during a third time period subsequent to the second time period, wherein the first inference is generated during the third time period.4.The processing system of Claim 3, wherein the one or more processors are further configured to cause the processing system to:load a serialized version of a third portion of the machine learning model during the third time period, the third portion of the machine learning model being associated with a third inference;deserialize the serialized version of the third portion of the machine learning model during a fourth time period subsequent to the third time period; andgenerate the third inference during a fifth time period subsequent to the fourth time period.5.The processing system of Claim 1, wherein the one or more processors are further configured to cause the processing system to generate a second inference based on the input data set, the deserialized second portion of the machine learning model, and the generated first inference without concurrently loading a third portion of the machine learning model.6.The processing system of Claim 1, wherein:the first portion of the machine learning model comprises a transformer neural network block and a key-value projection block,the first inference is generated using the transformer neural network block, andan output of the machine learning model is generated based on the first inference and the key-value projection block.7.The processing system of Claim 1, wherein to load the first portion of the machine learning model, the one or more processors are configured to cause the processing system to:initialize the first portion of the machine learning model based on a memory address associated with the first portion of the machine learning model; andallocate an input / output buffer shared by at least the first portion of the machine learning model and the second portion of the machine learning model.8.The processing system of Claim 7, wherein to load the second portion of the machine learning model, the one or more processors are configured to cause the processing system to:initialize the second portion of the machine learning model based on a memory address associated with the second portion of the machine learning model, wherein the second portion of the machine learning model uses the allocated input / output buffer; anddeallocate components associated with the first portion of the machine learning model that are not shared by the second portion of the machine learning model.9.A processor-implemented method, comprising:loading a first portion of a machine learning model, wherein the first portion of the machine learning model is associated with a first inference;loading a second portion of the machine learning model, wherein the second portion of the machine learning model is associated with a second inference; andwhile loading the second portion of the machine learning model, generating the first inference based on an input data set and the loaded first portion of the machine learning model.10.The method of Claim 9, wherein loading the first portion of the machine learning model comprises:loading a serialized version of the first portion of the machine learning model during a first time period; anddeserializing the serialized version of the first portion of the machine learning model during a second time period subsequent to the first time period.11.The method of Claim 10, wherein loading the second portion of the machine learning model comprises:loading a serialized version of the second portion of the machine learning model during the second time period; anddeserializing the serialized version of the second portion of the machine learning model during a third time period subsequent to the second time period, wherein the first inference is generated during the third time period.12.The method of Claim 11, further comprising:loading a serialized version of a third portion of the machine learning model during the third time period, the third portion of the machine learning model being associated with a third inference;deserializing the serialized version of the third portion of the machine learning model during a fourth time period subsequent to the third time period; andgenerating the third inference during a fifth time period subsequent to the fourth time period.13.The method of Claim 9, further comprising generating a second inference based on the input data set, the loaded second portion of the machine learning model, and the generated first inference without concurrently loading a third portion of the machine learning model.14.The method of Claim 9, wherein:the first portion of the machine learning model comprises a transformer neural network block and a key-value projection block,the first inference is generated using the transformer neural network block, andan output of the machine learning model is generated based on the first inference and the key-value projection block.15.The method of Claim 9, wherein loading the first portion of the machine learning model comprises:initializing the first portion of the machine learning model based on a memory address associated with the first portion of the machine learning model; andallocating an input / output buffer shared by at least the first portion of the machine learning model and the second portion of the machine learning model.16.The method of Claim 15, wherein loading the second portion of the machine learning model comprises:initializing the second portion of the machine learning model based on a memory address associated with the second portion of the machine learning model, wherein the second portion of the machine learning model uses the allocated input / output buffer; anddeallocating components associated with the first portion of the machine learning model that are not shared by the second portion of the machine learning model.17.A processing system, comprising:means for loading a first portion of a machine learning model, wherein the first portion of the machine learning model is associated with a first inference;means for loading a second portion of the machine learning model, wherein the second portion of the machine learning model is associated with a second inference; andmeans for generating, while loading the second portion of the machine learning model, the first inference based on an input data set and the loaded first portion of the machine learning model.18.The processing system of Claim 17, wherein the means for loading the first portion of the machine learning model comprise:means for loading a serialized version of the first portion of the machine learning model during a first time period; andmeans for deserializing the serialized version of the first portion of the machine learning model during a second time period subsequent to the first time period.19.The processing system of Claim 18, wherein the means for loading the second portion of the machine learning model comprise:means for loading a serialized version of the second portion of the machine learning model during the second time period; andmeans for deserializing the serialized version of the second portion of the machine learning model during a third time period subsequent to the second time period, wherein the first inference is generated during the third time period.20.The processing system of Claim 19, further comprising:means for loading a serialized version of a third portion of the machine learning model during the third time period, the third portion of the machine learning model being associated with a third inference;means for deserializing the serialized version of the third portion of the machine learning model during a fourth time period subsequent to the third time period; andmeans for generating the third inference during a fifth time period subsequent to the fourth time period.21.The processing system of Claim 17, further comprising means for generating a second inference based on the input data set, the loaded second portion of the machine learning model, and the generated first inference without concurrently loading a third portion of the machine learning model.22.The processing system of Claim 17, wherein:the first portion of the machine learning model comprises a transformer neural network block and a key-value projection block,the first inference is generated using the transformer neural network block, andan output of the machine learning model is generated based on the first inference and the key-value projection block.23.The processing system of Claim 17, wherein the means for loading the first portion of the machine learning model comprise:means for initializing the first portion of the machine learning model based on a memory address associated with the first portion of the machine learning model; andmeans for allocating an input / output buffer shared by at least the first portion of the machine learning model and the second portion of the machine learning model.24.The processing system of Claim 23, wherein the means for loading the second portion of the machine learning model comprise:means for initializing the second portion of the machine learning model based on a memory address associated with the second portion of the machine learning model, wherein the second portion of the machine learning model uses the allocated input / output buffer; andmeans for deallocating components associated with the first portion of the machine learning model that are not shared by the second portion of the machine learning model.25.A computer-readable medium having executable instructions stored thereon which, when executed by one or more processors, perform an operation comprising:loading a first portion of a machine learning model, wherein the first portion of the machine learning model is associated with a first inference;loading a second portion of the machine learning model, wherein the second portion of the machine learning model is associated with a second inference; andwhile loading the second portion of the machine learning model, generating the first inference based on an input data set and the loaded first portion of the machine learning model.

Citation Information

Patent Citations

  • Optimizing machine learning model performance

    CN114008636A

  • Debugging and profiling of machine learning model training

    CN114503132A

  • Deep neural networks (DNN) inference using practical early exit networks

    US20230342278A1