Systems and methods for controlling transducers in artificial intelligence

By introducing cyclic exit technology into the transformer of AI learning model, the problems of high computing costs and delays in the prior art are solved, and more efficient computing resource usage and lower power consumption are achieved.

CN120068942APending Publication Date: 2025-05-30SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411738253.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-15
Filing Date
2024-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing AI learning model technologies have problems with high computing costs, increased latency and processing time, and increased power consumption of computing components used in the architecture.

Method used

By introducing a loop exit technology into the transformer, the termination condition is determined using the interaction between the system scheduler and the processing unit scheduler, the threshold is developed, and the iterative loop exit is realized when the loop exit condition is less than the threshold.

Benefits of technology

Faster response, reduced latency and processing time, lower computing resource usage, and lower computing component power consumption are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068942A_ABST
    Figure CN120068942A_ABST
Patent Text Reader

Abstract

A system and method for loop exit is disclosed. The cell scheduler includes a command queue, a report queue, an interface, and a cell controller. The command queue is configured to store commands for execution in an iterative process for an application using a transformer model with a multi-head attention (MHA) mechanism and a decoder. The report queue is configured to store status reports regarding execution of commands. The interface is configured to communicate with a host processor to receive a command and send a status report. The cell controller is configured to determine a change in the iterative process based on a loop exit condition being satisfied. The cell controller reports a loop exit condition in a report queue.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of priority of U.S. Provisional Patent Application Serial No. 63 / 604,877, filed on November 30, 2023, the disclosure of which is incorporated herein by reference in its entirety as if fully set forth herein. Technical Field

[0003] This disclosure generally relates to artificial intelligence (AI). More specifically, the subject matter disclosed herein relates to controlling a transformer. Background Art

[0004] Artificial intelligence (AI) is becoming increasingly popular in many applications including natural language processing, vision, content creation, art, pattern recognition, robotics, etc. Generative AI is a technique that uses deep learning models to generate new content (e.g., text, images, audio, video) based on input data. In a typical query - response system, these deep learning models generally require large databases that store information from various sources. The deep learning models help the system learn queries and respond to queries based on the knowledge collected from these sources.

[0005] There are several drawbacks in the prior art of AI learning models. These include high computational costs, increased latency and processing time, and increased power consumption of the computing elements used in the architecture. Summary of the Invention

[0006] To overcome these problems, systems and methods for loop exit techniques are described herein. The technique is aimed at terminating iterations in a processing layer in a transformer. The technique is based on the interaction between a system scheduler and a processing unit (PU) scheduler to determine termination conditions. Thresholds are developed according to the application and the operating environment. When the loop exit condition is less than the threshold, a loop exit of the iteration is achieved. The above - mentioned methods improve upon prior methods as they provide faster responses, reduced latency and processing time, lower computational resource usage, and lower power consumption of computing elements.

[0007] In an embodiment, the unit scheduler includes a command queue, a report queue, an interface, and a unit controller. The command queue is configured to store commands for execution during an iteration process for an application using a transformer model with a multi - head attention (MHA) mechanism and a decoder. The report queue is configured to store status reports regarding the execution of the commands. The interface is configured to communicate with a host processor to receive commands and send status reports. The unit controller is configured to determine a change in the iteration process based on the loop exit condition being met. The unit controller reports the loop exit in the report queue.

[0008] In another embodiment, the host scheduler includes a command queue, a report queue, an interface, and a host controller. The command queue is configured to store commands for execution during an iteration process for an application that uses a transformer model with a multi-head attention (MHA) mechanism and a decoder. The report queue is configured to store status reports regarding the execution of the commands. The interface is configured to communicate with at least one processing unit when processing the commands and status reports. The host controller is configured to issue commands from the command queue to at least one processing unit via the interface and read status reports from the report queue. The host controller determines a change in the iteration process based on a loop exit condition being satisfied as reported in the status report. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the following sections, aspects of the subject matter disclosed herein will be described with reference to exemplary embodiments shown in the accompanying drawings, in which:

[0010] Figure 1 is a block diagram showing a system according to an embodiment.

[0011] Figure 2 is a diagram showing a loop exit (LE) platform with multiple processing units according to an embodiment.

[0012] Figure 3 is a view showing an LE platform with multiple processing unit clusters according to an embodiment.

[0013] Figure 4 is a diagram showing a host processor according to an embodiment.

[0014] Figure 5 is a diagram showing a processing unit (PU) according to an embodiment.

[0015] Figure 6 is a diagram showing a computing element in a PU according to an embodiment.

[0016] Figure 7 is a flowchart showing an LE process at a PU according to an embodiment.

[0017] Figure 8 is a diagram showing an LE process at a host processor according to an embodiment. DETAILED DESCRIPTION

[0018] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, those skilled in the art will understand that the aspects disclosed herein may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the subject matter disclosed herein.

[0019] References to "one embodiment" or "an embodiment" in the present specification mean that the particular features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment disclosed herein. Thus, the phrases "in one embodiment", "in an embodiment", "according to one embodiment" (or other phrases with similar meanings) that appear throughout the present specification may not necessarily all refer to the same embodiment. In addition, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, as used herein, the word "exemplary" means "serving as an example, instance, or illustration". Any embodiment described herein as "exemplary" should not be construed as necessarily being preferred or advantageous over other embodiments. Additionally, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Further, depending on the context discussed herein, singular terms may include their corresponding plural forms, and plural terms may include their corresponding singular forms. Similarly, hyphenated terms (e.g., "two-dimensional", "pre-determined", "pixel-specific", etc.) may occasionally be used interchangeably with their corresponding non-hyphenated versions (e.g., "two dimensional", "predetermined", "pixelspecific", etc.), and capitalized entries (e.g., "Counter Clock", "Row Select", "PIXOUT", etc.) may be used interchangeably with their corresponding non-capitalized versions (e.g., "counterclock", "row select", "pixout", etc.). Such occasional interchangeable use should not be regarded as inconsistent with each other.

[0020] Further, depending on the context discussed herein, singular terms may include their corresponding plural forms, and plural terms may include their corresponding singular forms. It should also be noted that the various figures (including component diagrams) shown and discussed herein are for illustrative purposes only and are not drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Additionally, if deemed appropriate, reference numerals are repeated in the figures to indicate corresponding and / or similar elements.

[0021] The terms used herein are for the purpose of describing particular example embodiments only and are not intended to limit the claimed subject matter. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0022] It should be understood that when an element or layer is referred to as being "on", "connected to" or "coupled to" another element or layer, it can be directly on, connected or coupled to the other element or layer, or intervening elements or layers may be present. In contrast, when an element is referred to as being "directly on", "directly connected to" or "directly coupled to" another element or layer, there are no intervening elements or layers. The same reference numerals always refer to the same elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0023] As used herein, the terms "first", "second", etc. are used as labels for the nouns that follow them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.), unless explicitly defined as such. Additionally, the same reference numerals may be used across two or more figures to refer to parts, components, blocks, circuits, units, or modules having the same or similar functionality. However, this usage is merely for the purpose of simplifying the description and facilitating discussion; it does not mean that the construction or architectural details of such components or units are the same in all embodiments, or that such commonly referenced parts / modules are the only way to implement some of the example embodiments disclosed herein.

[0024] All terms used herein, unless otherwise defined, have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense, unless expressly so defined herein.

[0025] As used herein, the term "module" refers to any combination of software, firmware, and / or hardware that is configured to provide the functionality described herein in connection with the module. For example, software can be embodied as a software package, code, and / or instruction set or instructions, and the term "hardware" as used in any of the embodiments described herein can include (e.g., individually or in any combination) components, hardwired circuitry, programmable circuitry, state machine circuitry, and / or firmware that stores instructions executed by the programmable circuitry. A module can be embodied jointly or individually as circuitry that forms part of a larger system, such as, for example, but not limited to, an integrated circuit (IC), a system-on-chip (SoC), components, and the like.

[0026] As used herein, the term "iteration" refers to one computational pass as part of a loop. It can correspond to a "layer" in a multi-layer architecture. As used herein, the term "loop exit" refers to the act of leaving the iterative process of a loop at the end of the loop, in the middle of the loop, or at the end of an iteration.

[0027] One type of learning model is a large language model (LLM), which involves huge data sets. Examples of these LLMs include Generative Pretrained Transformer (GPT) and Large Language Model Meta AI (LLAMA). The basic component of these models is the transformer. For language applications, the transformer takes a sequence of text as input and produces another sequence of text as output. In many query-response systems, a typical transformer includes several operations, such as tokenization, positional encoding, multi-head attention (MHA), and a feed-forward network. These operations are repeated across multiple layers to refine the quality of the output based on the model's understanding of the context embedded in the data. The number of layers is fixed (e.g., 96).

[0028] In an embodiment, the unit scheduler includes a command queue, a report queue, an interface, and a unit controller. The command queue is configured to store commands for execution during an iterative process for an application that uses a transformer model with a multi-head attention (MHA) mechanism and a decoder. The report queue is configured to store status reports regarding the execution of the commands. The interface is configured to communicate with a host processor to receive commands and send status reports. The unit controller is configured to determine a change in the iterative process based on a loop exit condition being met. The unit controller reports the loop exit in the report queue.

[0029] Figure 1 is a block diagram showing a system 100 according to an embodiment. The system 100 includes a development environment 110 and an operating environment 120. The development environment 110 is a place for developing, testing, and evaluating algorithms. The result of the development includes instructions or commands that will be executed in the operating environment 120. The operating environment 120 can include a platform on which an application will be executed to perform a specified task.

[0030] In one embodiment, the development environment 110 may include a software framework 112, a compiler 114, and an instruction formatter 118. The development environment 110 may include more or fewer elements than those described above. The software framework 112 provides development tools such as libraries and pre-built functions to facilitate algorithm development in a machine learning or deep learning environment. Examples of the software framework 112 may include Pytorch and TensorFlow. The software framework 112 generates code for an application. The compiler 114 compiles the code from the software framework 112 and generates an instruction set to be executed by a processor in the operating environment 120. The instruction formatter 118 formats the instruction set into commands in a suitable format to be executed on the processor in the operating environment 120.

[0031] The operating environment 120 may include an application 123, a large language model (LLM) transformer 125, and a loop exit (LE) platform 127. The application 123 may include generative AI-based applications such as query and response systems, chatbots, search engines, story generation, audio and video content creation, or any application that creates new content based on large amounts of domain-specific data. The LLM transformer 125 transforms input data (e.g., text data) into a different form for the purpose of serving the application 123. The LLM transformer 125 generally includes an encoder section and a decoder section. The encoder section encodes input data (such as text) called tokens into a digital representation so that they can be manipulated by a processor. The decoder section generates an output based on the digital representation and a learning process. The format of the output is the same as the format of the input. For example, if the input data is text, the output is also text. The LLM transformer 125 generally employs an attention mechanism such as a multi-head attention (MHA) mechanism to generate relationships between tokens. The process is an iterative process involving many computations such as matrix multiplication, softmax, and dot products. The number of iterations is usually fixed. For a typical chatbot application, the number of iterations is 96. In addition, the process involves a very large amount of data. Due to the large amount of data and the number of iterations, there is a motivation to reduce the computational workload.

[0032] Figure 2 is a diagram showing a loop exit (LE) platform 127 having multiple processing units according to an embodiment. In this configuration, the processing units (PUs) are separate units that run asynchronously. The LE platform 127 includes a storage element 210 that stores control and status words, a host processor 220, a communication interface 230, a network interface 240, and N PUs 250 1 to 250 N , where N is a positive integer.

[0033] The storage element 210 can be a register or a memory location. A control and status (CS) word can be issued or read by the application 123. The CS word can include an LE enable flag 212, a threshold 214, and an LE status 216. The LE enable flag 212 indicates whether the LE is active. It can be a single bit, where 0 indicates the LE is off, and 1 indicates the LE is on. When the LE is on, the host processor 220 and the N PUs 250 1 to 250 N are all ready to perform an LE operation, which can include the calculation of a similarity metric that shows the similarity between consecutive results in an iterative process. The threshold 214 is a value used to compare with the similarity metric indicating the progress of the LE. When the iterative process approaches the end of the loop, its result may only improve slightly. This improvement may not be commensurate with the computational cost. Therefore, loop exit or early termination can be employed to obtain reasonable results while saving computational resources. To determine this condition, consecutive results calculated during the iterative process are compared and a similarity metric is determined. If each iteration produces a result in the form of a vector comprising several components, any suitable similarity metric (such as cosine similarity) can be used to calculate the similarity metric. The similarity can be normalized such that its range is from 0 to 1, where 0 indicates identical similarity, and 1 indicates complete dissimilarity. The calculated similarity can be adjusted to fall within this range. For example, the resulting cosine similarity metric has values ranging from -1.0 to +1.0, where -1.0 indicates complete dissimilarity, and +1.0 indicates identical similarity. The cosine similarity α can be adjusted to the [0,1] range as follows:

[0034] S = (1 – α) / 2 (1)

[0035] The similarity metric S is compared with the threshold T. If it is less than T, it indicates that the result of the current iteration is not very different from the previous iteration, and the process may be approaching the end of the loop. Therefore, it is time to terminate the loop early to save computational resources.

[0036] The threshold can be fixed or variable depending on the application (e.g., text, image, audio). For example, for visual applications, the threshold can be higher (to allow "early" loop exits) because images tend to have a high tolerance for noise. For text applications, the threshold can be lower (to allow "late" loop exits) because text tends to be more precise. It can also be variable or adaptive based on the progress of learning or the nature of the output of the layer. If the nature is a broad concept, the threshold is higher (to allow "early" loop exits) because broad concepts have a higher tolerance level (e.g., "sad", "heartbroken", and "unhappy" are similar). If the nature is narrow, the threshold is lower (to allow "late" loop exits) (e.g., "Los Angeles", "Seoul", "Washington D.C.", "Paris" are very specific and require more stringent tolerance). For example, assume the application is text and the process is predicting the next word. Assume the next word to be predicted is related to "emotion" as a broad concept, then the threshold is made higher, etc.

[0037] The LE state 216 indicates the condition of the LE. In other words, it indicates the result of the comparison of the similarity metric and the threshold. It can be a single bit indicating whether the LE has been reached. For example, if it is 0, the LE has not been reached; if it is 1, the LE has been reached.

[0038] The host processor 220 can be any programmable device capable of executing a program to work with N PUs 250 1 to 250 N together. It can be a high-performance microprocessor, a graphics processing unit (GPU), a programmable application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA). The communication interface 230 allows communication between the host processor 220 and any one of the PUs 250 1 to 250 N It can include direct memory access (DMA), an interrupt mechanism, double-buffered memory, a bidirectional bus transceiver, or any element or function that allows asynchronous communication between two processors. The network interface 240 provides another interface to allow the PUs 250 1 to 250 N to exchange information with each other or with the host processor. In one embodiment, the network interface 240 can be the Internet or a local area network. Any one of the PUs 250 1 to 250 N can be referred to as a network processing unit (NPU).

[0039] PU 250 k(k = 1, …, N) can be a high-performance microprocessor, a graphics processing unit (GPU), a programmable application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a microcontroller, a signal processor, an image processor, or any programmable device or circuit that can perform calculations quickly. The N PUs 250 1 through 250 N can work independently of each other or cooperate to perform some common tasks. They can form a multiprocessor system and operate synchronously or asynchronously. In one embodiment, the N PUs 250 1 through 250 N can represent threads in a multithreaded environment. They can also be software functions. Additionally, the N PUs 250 1 through 250 N can be a hybrid of hardware and software elements. They interact with the LLM transformer 125.

[0040] The communication between the host processor and the N PUs 250 1 through 250 N can be parallel and asynchronous, where the processors can issue commands, push reports, update status, or send notifications independently of other processors. Similarly, they can also retrieve or pull reports, read status, and receive notifications independently of other devices. The N PUs 250 1 through 250 N can share the workload among themselves under the control of the host processor 220. The workload and commands can be configured and prepared by the development environment 110. By providing a parallel and asynchronous operating environment, the total throughput can be increased. In particular, when the N PUs 250 1 through 250 N participate in the calculations in the processing chain of the LLM transformer 125, tasks that contribute to the overall system performance can be assigned to each PU. For example, a matrix multiplication problem can be decomposed into multiple row and column multiplications, and one of the multiple row and column multiplications can be assigned to the PU 250 k . When there are no dependencies between these calculations, they can be executed in parallel.

[0041] To increase the communication throughput, the host processor 220 and each of the N PUs 250 1 through 250 N have their own command queues and report queues. These queues are dynamically changed and provide buffered data so that they can be updated asynchronously. In a typical environment, the host processor 220 maintains its command queue to buffer the commands provided by the application 123. It will dispatch the commands to the N PUs 250 1 through 250 NScheduling can be conditional on the status included in reports stored in the report queue. It will also retrieve reports from the report queue. These reports are sent by N PUs 250 1 to 250 N and contain the status of operations performed by the PUs.

[0042] Figure 3 is a diagram showing the LE platform 127 with multiple clusters of processing units according to an embodiment. Figure 3 The configuration of the LE platform 127 in Figure 2 is almost the same as that in k The main difference is that the individual PUs 250 k are replaced by clusters k 260 k (k = 1, …, N). A cluster is a group of PUs from the same or different dies. In each cluster 260 k1 there can be P PUs 265 kP to 265

[0043] By having synchronized PUs in each cluster and asynchronous clusters, the LE platform 127 can provide optimized performance with high flexibility. In addition, the PUs in each cluster can send a common report to the host processor 220, and the host processor 220 can send commands to each cluster. This will significantly reduce the traffic at the network interface 240 and the communication interface 230.

[0044] Figure 4 is a diagram showing the host processor 220 according to an embodiment. The host processor 220 includes host logic and processing elements (LPEs) 401 and a host scheduler 405. The host processor 220 can include more or fewer elements than those described above.

[0045] The host LPE 401 includes all elements not involved in LE processing. It can include programmable executable elements, memory, and an input / output (IO) interface to IO devices. The memory can include instructions or programs that, when executed by the programmable elements, cause the programmable elements to perform the operations described below.

[0046] The host scheduler 405 schedules activities related to LE. It can include a host controller 410, a host command queue 420, a host report queue 430, and a host interface 440. The host interface 440 communicates directly with the communication interface 220. The host scheduler 405 can include more or fewer components than those described above.

[0047] The host controller 410 controls all activities within the host scheduler 405. It can be a programmable processor that can execute the instructions or programs described below. These instructions or programs can be fetched from the memory in the host LPE 401. The host controller 410 can read the CS word 210 to obtain the values of the LE enable flag 212 and the threshold 214. It can also update the LE status 216 based on reports from the host report queue 430. The host command queue 420 stores commands sent from the application 123 or the development environment 110. These commands can include instructions to perform tasks, execute calculations, or check status. The host controller 410 can also fetch commands from the host command queue 420 to send to the PU 250 k or the cluster 260 k . The host report queue 430 stores reports sent by the PU250 k or the cluster 260 k . Reports can contain the status of operations during an iterative process. Reports can contain the identifier (ID) of the PU such that the host processor will know where it came from and will act accordingly. In particular, when the report contains the status of the LE indicating that the LE has been satisfied, the host processor 220 will identify the sending PU and will provide advice or send commands to the sending PU. The host interface 440 provides an interface to the communication interface 220. It can include various components for communication, such as buffers, bidirectional drivers, and local storage devices. If these components are available in the communication interface 220, the host interface 440 may not be required.

[0048] Figure 5 is a diagram showing the processing units (PUs) 250 k and 265 k according to an embodiment. For clarity, subscripts may be removed. The PU 250 / 265 can include unit logic and processing elements (LPEs) 501 and a unit scheduler 505. The PU 250 / 265 can include more or fewer elements than those described above.

[0049] The unit LPE 501 includes all elements not involved in LE processing. It can include programmable executable elements, memory, and an input / output (IO) interface to IO devices. The memory can include instructions or programs that, when executed by the programmable elements, cause the programmable elements to perform the operations described below.

[0050] The unit scheduler 505 schedules activities involving the LE. It can include a storage element 510, a computing element 520, a unit controller 530, a unit command queue 540, a unit report queue 550, and a unit interface 560. The unit interface 440 communicates directly with the network interface 240. The unit scheduler 505 can include more or fewer components than those described above.

[0051] The storage element 510 can be a register or a memory that stores a control / status (CS) word. The CS word 510 includes a LE enable flag 512, a threshold 514, and a LE status 516 from the application 123. These elements have the same functions as the LE enable flag 212, the threshold 214, and the LE status 216, but they are kept private in the PU 250 / 265 to allow for fast access.

[0052] The computing element 520 performs computations during the iterative process and LE processing. This will be further described in Figure 6 The unit controller 530 controls all activities within the unit scheduler 505. It can be a programmable processor that can execute the instructions or programs described below. These instructions or programs can be fetched from the memory in the unit LPE 501. The unit controller 530 can read the CS word 510 to obtain the values of the LE enable flag 512 and the threshold 514. It can also update the LE status 516 based on the result of the computing element 520. The unit command queue 540 stores commands sent from the host processor 220. These commands can include instructions to perform tasks, execute computations, or check status. The unit report queue 550 stores reports completed by the unit controller 530, which reflect the status of the operations performed by the computing element 520. The report can contain the status of the operations during the iterative process. The report can contain the identifier (ID) of the PU 250 / 260 such that when it is sent to the host processor 220, the host processor 220 will know where it came from and will act accordingly. In particular, when the report contains a LE status indicating that the LE has been satisfied, the host processor 220 will identify the sending PU 250 / 260 and will provide advice or send commands to the sending PU 250 / 260. The unit interface 560 provides an interface to the network interface 240. It can include various components for communication, such as buffers, bidirectional drivers, and local storage devices. If these components are available in the network interface 240, the unit interface 560 may not be needed.

[0053] Figure 6 FIG. is a diagram showing the computing element 520 in the PU 250 / 260 according to an embodiment. The computing element 520 can include a memory 610, an arithmetic logic unit (ALU) 620, a storage element 630, and M storage elements 635 1 to 635 M and a comparator circuit 640. The computing element 520 can include more or fewer elements than those described above.

[0054] The memory 610 stores data used in the calculations during the iterative process. This data can include parameters obtained from the LLM transformer 125, such as values of vectors, values of matrices, etc. in the encoder or decoder. It can also store temporary data generated during the calculations. It can also include instructions to the ALU 620.

[0055] The ALU 620 receives the LE enable flag 512 so that it can gate or direct the calculation results to the registers 630, 635 1 to 635 M . If LE is not enabled, then the registers 630, 635 1 to 635 M and the comparator circuit 640 are not used. If LE is enabled, then the registers 630, 635 1 to 635 M and the comparator circuit 640 are used to determine whether LE has been reached. The ALU 620 is designed to perform calculations during the iterative process. These calculations can include tokenization 621, embedding 622, positional encoding 624, softmax 625, matrix multiplication 627, and dot product 628. This list is merely an example. The ALU 620 can perform additional calculations or logical operations.

[0056] The register 630 stores the current result vector V 0 . The register 635 1 to 635 M stores the past result vectors V 1 to V M . The registers 630 and 635 1 to 635 M are arranged as shift registers such that after each iteration, the values in each register are shifted to the register to its left: V M ←V M -1,…,V 1 ←V 0 , and V 0 will be loaded with the new (current) result so that the registers always store the most recent M past results. Generally, in order to determine whether LE has been reached, the comparator circuit 640 may need to compare only the current result V 0 and the immediately preceding result V 1 . However, sometimes it is better to check the past M results to determine whether the trend is indeed the correct trend. In other words, by checking the M past results instead of only the single previous result, the comparator circuit 640 can avoid errors due to noisy results. Alternatively, the comparator circuit 640 can determine the average of the past M results and use that average to compare with the current value V 0 to determine the similarity metric.

[0057] Depending on the type of similarity metric, the similarity metric S can be adjusted or scaled to the range for threshold processing. For example, as mentioned above in equation (1), if the similarity is cosine similarity with a range of [-1, +1], then equation (1) can be used to scale the value S to the range [0, 1], where 0 corresponds to the same similarity and 1 corresponds to complete dissimilarity. If the scaled similarity metric is less than the threshold T, then LE has been reached and the LE state 516 is asserted.

[0058] Numerical example: The following is a numerical example for illustrating the calculation of the similarity metric. For this example, assume that the output result of each iteration is a vector with three components (x1, x2, x3). The cosine similarity α between vectors Vx and Vy is equal to the ratio of the dot product of Vx and Vy to the product of their magnitudes. The LE similarity is given in equation (1): S = (1 - α) / 2. Table 1 shows the values of six nearest vectors. Assume the threshold is T = 0.001.

[0059]

[0060] Table 1 : Values of V0 to V5 and the resulting similarity

[0061] In Table 1, V0 is the current vector at time t = 0, V1 is the vector at time t = -1, V2 is the vector at time t = -2, etc. The dot product of V0 and V1 is 2.5105, the dot product of V1 and V2 is -1.3303, etc. The cosine similarities between V0 and V1, V1 and V2, V2 and V3, V3 and V4, V4 and V5 are 0.9991, -0.48, 0.396, -0.4702, -0.7576 respectively. Looking at the values of V1 to V5, it seems that the iterative process is unstable because the values change significantly, as reflected by the S similarity. However, assume that due to some noise conditions, V0 and V1 are very similar, with a resulting S value of 0.0004. This value exceeds the threshold T = 0.001. Therefore, the comparator 640 can declare that LE has been reached. In fact, this is due to the noise conditions.

[0062] When using the average value, the result is significantly different. Assume that the average value is calculated over V1, V2, V3, V4, and V5. Then this value is compared with V0. Table 2 shows the results.

[0063]

[0064] Table 2 : Using the average value to determine similarity.

[0065] As can be seen from Table 2, the similarity measure S has now become 0.567, which is significantly different from 0.0004. This example shows that it is beneficial to filter the result values before using them to calculate the similarity.

[0066] Figure 7 FIG. is a flowchart showing the LE process 700 at the PU according to an embodiment. At the beginning, the process 700 communicates with the host processor to receive commands for execution during an iterative process for an application using a transformer model with a multi-head attention (MHA) mechanism and a decoder (block 710). The host processor may issue one or more commands to the PU. Then, the process 700 stores the commands in a command queue (block 720). Next, the process 700 receives a control word from the host processor (block 730). The control word includes at least a threshold and a loop exit (LE) enable flag. This flag indicates that LE is enabled during the iterative process when asserted. LE results from a comparison between the similarity measure and the threshold.

[0067] Next, the process 700 performs operations during the iterative process (block 740). These operations may include calculations in the encoder-decoder operations of the LLM transformer 125, such as tokenization, matrix multiplication, softmax, etc. The process 700 determines whether LE is enabled or whether the LE flag bit is asserted (block 750). If not, the process 700 determines whether the current iteration is the end of the iterative process (block 760). If not (the "no" branch at block 760), the process 700 returns to block 740 to continue the process. If it is the end of the loop (the "yes" branch on block 760), the process 700 terminates.

[0068] If LE is enabled (the "yes" branch on block 750), the process 700 determines whether the similarity measure is less than the threshold (block 770). The similarity measure is obtained from the consecutive results of the decoder during the iterative process. The threshold depends on the application and can be set adaptively. If the similarity measure is not less than the threshold, the process 700 returns to block 740 to continue the next operation or the next iteration. The process 700 may also go to block 760 to determine whether the end of the iteration has been reached. However, since LE is most likely to be reached before the end of the loop, block 760 may not be necessary. If the similarity measure is less than the threshold, the process 700 determines a change or stop of the iterative process based on the loop exit of the operation and the LE enable flag (block 780). Next, the process 700 reports the status of LE to the host processor (block 790), and then terminates.

[0069] Figure 8FIG. 0 is a diagram illustrating an LE process 800 at a host processor according to an embodiment. The LE process 800 generally runs in parallel with the LE process 700 in an asynchronous manner. At the beginning, the process 800 communicates with at least one processing unit via a communication interface (block 810). The communication interface can be a direct memory access (DMA), a two-way buffer, an interrupt mechanism, or any other structure that allows the host processor to communicate with the PU. Next, the process 800 issues commands for execution in an iterative process to at least one processing unit from a command queue via the communication interface (block 820). The commands can be pre-loaded in the host command queue. The iterative process can be used for applications using a transformer model with a multi-head attention (MHA) mechanism and a decoder.

[0070] Then, the process 800 reads a status report from a report queue (block 830). The status report is sent by at least one processing unit regarding the execution status. Next, the process 800 determines whether an LE enable word or bit is asserted, i.e., whether LE is turned on (block 840). If not (the "no" branch at block 840), the process 800 processes the report in a normal manner (block 850) and then terminates. Otherwise (the "yes" branch at block 840), the process 800 determines whether the status report indicates that LE has been confirmed (block 860). LE results from a comparison between a similarity metric and a threshold. The similarity metric is obtained from consecutive results of the decoder in the iterative process. The threshold depends on the application. The threshold can be fixed or adaptive. If the status report indicates that LE has not been confirmed (the "no" branch at block 860), the process 800 terminates. Otherwise (the "yes" branch at block 860), the process 800 determines a change or stop of the iterative process (block 870).

[0071] Then, the process 800 executes a system LE process (block 880). This can include any operations that the application wants to perform, such as continuing the calculation on the next level of the transformer, generating a response to a query, etc. Then the process 800 terminates.

[0072] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium can be, or include, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of them. Moreover, although a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer storage medium can also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or be included in them. Additionally, the operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0073] Although this specification may include many specific implementation details, the implementation details should not be construed as limitations on the scope of any claimed subject matter, but rather as descriptions of features specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excluded from the combination, and the claimed combination can be directed to a sub-combination or variation of a sub-combination.

[0074] Similarly, although the operations are depicted in the drawings in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the above embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0075] Accordingly, specific embodiments of the subject matter have been described herein. Other embodiments are within the scope of the following claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures need not be in the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.

[0076] As those skilled in the art will recognize, the innovative concepts described herein can be modified and varied over a wide range of applications. Accordingly, the scope of the claimed subject matter should not be limited to any of the specific exemplary teachings discussed above, but instead is defined by the appended claims.

Claims

1. A device for controlling a transformer in artificial intelligence, comprising: a device command queue configured to store commands for execution in an iterative process for an application using a transformer model with a multi-head attention (MHA) mechanism and a decoder; a device report queue configured to store a status report regarding the execution of the command; an interface configured to communicate with a host processor to receive the commands and to send the status reports; a device controller configured to determine a change in the iterative process based on a loop exit condition being satisfied, Wherein, the device controller reports the loop exit condition in the device report queue.

2. The device according to claim 1, wherein: The loop exit condition comes from the comparison between the similarity measure and a threshold value.

3. The device according to claim 2, wherein: The similarity measure is determined by cosine similarity.

4. The device according to claim 2, wherein: The similarity measure is obtained from at least one result of the decoder during the iteration.

5. The device according to claim 2, wherein: The threshold is based at least in part on the application.

6. The device according to claim 2, wherein: The threshold is set adaptively.

7. The apparatus according to claim 1, further comprising: A computation element is configured to perform computations during the iteration process and the loop exit condition.

8. The device according to claim 1, wherein: The host processor includes a host scheduler having a host command queue for storing the commands and a host report queue configured to store the status report regarding the execution of the commands.

9. The device according to claim 8, wherein: The host scheduler issues the commands from the host command queue.

10. The device according to claim 8, wherein: The host scheduler reads the status report and determines the changes to the iterative process.

11. A method for controlling a transformer in artificial intelligence, comprising: Communicating with a host processor to receive commands for execution in an iterative process for an application using a transformer model with a multi-head attention (MHA) mechanism and a decoder; storing the command in a device command queue; receiving a control word from the host processor, the control word having a loop exit LE enable flag; Execute the operations in the iterative process; determining a change in the iterative process based on an LE condition for the operation being satisfied and the LE enable flag; and The status of the LE condition is reported to the host processor. The method according to claim 11 , wherein the loop exit condition comes from a comparison between a similarity measure and a threshold value. The method of claim 12 , wherein the similarity measure is determined by cosine similarity. The method of claim 12 , wherein the similarity measure is obtained from at least one result of the decoder during the iteration.

15. The method according to claim 12, wherein: The threshold is based at least in part on the application. The method according to claim 12 , wherein the threshold is set adaptively.

17. The method according to claim 11, further comprising: Calculations are performed during the iteration process and the loop exit condition.

18. The method according to claim 11, wherein: The host processor includes a host scheduler having a host command queue for storing the commands and a host report queue configured to store the status report regarding the execution of the commands.

19. The method according to claim 18, wherein: The host scheduler issues the commands from the host command queue.

20. A system for controlling a transformer in artificial intelligence, comprising: A host processor having a host scheduler, the host scheduler comprising: a host command queue configured to store commands for execution in an iterative process for an application using a transformer model with a multi-head attention (MHA) mechanism and a decoder; a host report queue configured to store status reports regarding the execution of the command; a host controller configured to issue the commands from the command queue to at least one processing unit via an interface and to read the status reports from the report queue, wherein the host controller determines a change to the iterative process based on a loop exit condition reported in the status report; and At least one processing unit has a unit scheduler configured to determine the change of the iterative process based on the loop exit condition being satisfied.