Query processing method and system

By processing data in the PNM storage device and generating activation values, combined with D2D communication and integrated circuit technology, the problem of insufficient memory bandwidth and capacity in the prior art is solved, efficient AI query processing and memory utilization are achieved, and system costs are reduced.

CN120144613APending Publication Date: 2025-06-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411581872.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-12
Filing Date
2024-11-07
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the artificial intelligence query processing of the storage device, it is difficult to effectively utilize near-memory processing (PNM) storage devices, resulting in insufficient memory bandwidth and capacity, and it is impossible to adapt to the increased query length and number of concurrent users.

Method used

By receiving data in the PNM storage device, processing values ​​using transposed query values, determining the probability distribution of the processing results, and generating activation values ​​based on the probability distribution, indicating the correlation between the text unit in the query and the data. In addition, memory stacking is used to use the D2D communication interface and integrated circuit to realize high-bandwidth extended bus communication and reduce dependence on expensive GPU systems.

Benefits of technology

It realizes scalable memory bandwidth and capacity, adapts to the increased query length and number of concurrent users, reduces system costs, and performs softmax and hierarchical normalization operations in PNM SSD, avoiding backup of inactivated KV data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144613A_ABST
    Figure CN120144613A_ABST
Patent Text Reader

Abstract

The invention provides a query processing method and system. In one or more examples, the method includes receiving data at a first near memory processing (PNM) storage device; and processing, at the first PNM storage, a first value from the data using the transposed query value from the data. In one or more examples, the method includes determining, at a first PNM storage, a probability distribution of results of a process; and generating, at the first PNM storage, an activation value based on the probability distribution, the activation value indicating a correlation between text units in a query associated with the data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 608,821, filed on Dec. 11, 2023, and U.S. Patent Application No. 18 / 634,880, filed on Apr. 12, 2024, which are hereby incorporated by reference for all purposes. Technical Field

[0002] The disclosure generally relates to memory systems and, more particularly, to artificial intelligence query processing for storage devices through near-memory processing. Background Art

[0003] This background art section is intended only to provide context and the disclosure of any concepts herein does not constitute an admission that such concepts are prior art.

[0004] Artificial intelligence (AI) is a branch of computing that mimics human intelligence and can learn from and adapt to data inputs. Through AI search, platforms learn from data about users to automatically generate the most accurate and relevant search experiences. AI search can include query processing, retrieval, and ranking. Query processing involves analyzing a user's query to understand its intent, scope, and constraints. AI search uses machine learning and algorithms to search and classify large amounts of data. AI search engines use natural language processing (NLP) and other AI-based algorithms to understand search queries and provide accurate results.

[0005] The above information disclosed in this background art section is only for enhancing the understanding of the background of the disclosure and thus may contain information that does not constitute prior art. Summary of the Invention

[0006] In various embodiments, systems, methods, and devices for artificial intelligence query processing for storage devices through near-memory processing are described herein.

[0007] In some aspects, the techniques described herein relate to a method of query processing, the method including: receiving data at a first processing-in-near-memory (PNM) storage device; processing a first value from the data using a transposed query value from the data at the first PNM storage device; determining a probability distribution of a result of the processing at the first PNM storage device; and generating an activation value at the first PNM storage device based on the probability distribution, the activation value indicating a correlation between text units in a query associated with the data.

[0008] In some aspects, the techniques described herein relate to a method, wherein the first PNM storage device includes a memory storing key-value data from the data. In some cases, the data includes attention data.

[0009] In some aspects, the techniques described herein relate to a method in which a memory and a processor of a first PNM storage device include integrated circuits, the integrated circuits are stacked on top of the processor based on the memory, and the memory is communicatively connected to the processor.

[0010] In some aspects, the techniques described herein relate to a method further comprising: receiving at least a second activation value from a second PNM storage device via a die-to-die (D2D) communication interface; and forming a unified activation value based at least in part on a combination of an activation value of the first PNM storage device and the second activation value of the second PNM storage device, wherein the D2D communication interface enables communication between the first PNM storage device and the second PNM storage device.

[0011] In some aspects, the techniques described herein relate to a method further comprising: receiving a trigger from a processing device before receiving the data, wherein the trigger includes at least one of a user identifier associated with the data and layer count information.

[0012] In some aspects, the techniques described herein relate to a method in which the data is part of multi-head attention data distributed among a plurality of PNM storage devices.

[0013] In some aspects, the techniques described herein relate to a method in which the data includes a portion of query attention data from a first attention layer, a portion of key attention data from a second attention layer, and a portion of value attention data from a third attention layer.

[0014] In some aspects, the techniques described herein relate to a method in which the data is based on an iteration of activation values, and the activation values are generated based on partial outputs generated by a plurality of PNM storage devices that are reduced to a unified output.

[0015] In some aspects, the techniques described herein relate to a method in which the plurality of PNM storage devices includes an array of solid state drives.

[0016] In some aspects, the techniques described herein relate to a method in which the first PNM storage device is a system-on-chip die that includes a solid state drive and at least one processor.

[0017] In some aspects, the techniques described herein relate to a method in which the first PNM storage device is communicatively connected to a processing device via a high bandwidth expansion bus.

[0018] In some aspects, the techniques described herein relate to a method in which the processing device includes at least one graphics processing unit (GPU) communicatively connected to high bandwidth memory.

[0019] In some aspects, the techniques described herein relate to a method in which the data includes attention data, the first value includes a key value, the second value includes a query value, and the key value data includes a key value matrix.

[0020] In some aspects, the techniques described herein relate to a query processing system that includes: a processing device communicatively connected to a first near-memory processing (PNM) storage device, the processing device for sending data to the first PNM storage device; and a plurality of near-memory processing (PNM) storage devices, the first PNM storage device for: processing a first value from the data using a transposed query value from the data; determining a probability distribution of the result of the processing; and generating an activation value based on the probability distribution, the activation value indicating a correlation between text units in a query associated with the data.

[0021] In some aspects, the techniques described herein relate to a query processing system in which the first PNM storage device includes a memory that stores key value data from the data.

[0022] In some aspects, the techniques described herein relate to a query processing system in which the memory and processor of the first PNM storage device include an integrated circuit, the integrated circuit based on a memory stack on top of the processor and the memory communicatively connected to the processor.

[0023] In some aspects, the techniques described herein relate to a query processing system in which the first PNM storage device is configured to: receive at least a second activation value from a second PNM storage device via a die-to-die (D2D) communication interface; and form a unified activation value based at least in part on a combination of the activation value of the first PNM storage device and the second activation value of the second PNM storage device, where the D2D communication interface enables the first PNM storage device and the second PNM storage device to communicate.

[0024] In some aspects, the techniques described herein relate to a query processing system in which the first PNM storage device is configured to: receive a trigger from the processing device before receiving data, where the trigger includes at least one of a user identifier associated with the data and layer number information.

[0025] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium storing code that includes instructions executable by a processor of a first proximal memory processing (PNM) storage device to: receive data; process a first value from the data using a transposed query value from the data at the first PNM storage device; determine a probability distribution of the result of the processing at the first PNM storage device; and generate an activation value at the first PNM storage device based on the probability distribution, the activation value indicating a correlation between text units in a query associated with the data.

[0026] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium, wherein the first PNM storage device includes a memory storing key-value data from the data, and the memory and the processor of the first PNM storage device include an integrated circuit, the integrated circuit being stacked on top of the processor based on the memory and the memory being communicatively coupled to the processor.

[0027] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium, wherein the code further includes instructions executable by the processor to cause the first PNM storage device to: receive at least a second activation value from a second PNM storage device via a die-to-die (D2D) communication interface; and form a unified activation value based at least in part on a combination of the activation value of the first PNM storage device and the second activation value of the second PNM storage device, wherein the D2D communication interface enables communication between the first PNM storage device and the second PNM storage device.

[0028] A computer-readable medium is disclosed. The computer-readable medium may store instructions that, when executed by a computer, cause the computer to perform operations substantially the same as or similar to those described herein. Similarly, a non-transitory computer-readable medium, apparatus, and system for performing operations substantially the same as or similar to those described herein are also disclosed.

[0029] The techniques described herein include several advantages and benefits. For example, the AI inference delegation technique provides scalable memory bandwidth and scalable memory capacity to accommodate increasing query lengths and increasing numbers of concurrent users. The AI inference delegation technique reduces system costs by reducing the use of expensive Graphics Processing Unit (GPU) systems. Additionally, the AI inference delegation technique supports relatively long query sizes with sufficient storage space for advanced generative pre-trained transformers. The AI inference delegation technique enables query key-value (QKV) vector multiplication through near-memory processing (PNM) Solid State Drives (SSDs). Further, the AI inference delegation technique enables softmax and layer normalization operations to be performed in the PNM SSD. The AI inference delegation technique avoids backing up non-activated KV data to external storage devices. Additionally, the AI inference delegation technique minimizes the communication overhead between the main computing node (e.g., a GPU system) and the memory bandwidth-intensive system (e.g., a PNM SSD). Thus, the AI inference delegation technique reduces system costs and provides scalable memory bandwidth and capacity for advanced generative pre-trained transformers by enabling the ability to simply add more PNM SSDs to a given system. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The above-mentioned aspects and other aspects of the technology will be better understood when reading this application in view of the following drawings, in which the same reference numerals indicate similar or equivalent elements. Additionally, the drawings provided herein are for the purpose of illustrating specific embodiments only; other embodiments that may not be explicitly illustrated are not excluded from the scope of the present disclosure.

[0031] These and other features and advantages of the present disclosure will be appreciated and understood with reference to the specification, claims, and the appended drawings.

[0032] Figure 1 Illustrates an example system according to one or more embodiments described herein.

[0033] Figure 2 Illustrates details of a system according to one or more embodiments described herein Figure 1 of the system.

[0034] Figure 3 Illustrates an example system according to one or more embodiments described herein.

[0035] Figure 4 Illustrates an example system according to one or more embodiments described herein.

[0036] Figure 5 Illustrates an example system according to one or more embodiments described herein.

[0037] Figure 6A flowchart depicting an example method associated with the disclosed system in accordance with example embodiments described herein.

[0038] Figure 7 A flowchart depicting an example method associated with the disclosed system in accordance with example embodiments described herein.

[0039] While the present technology is susceptible to various modifications and alternative forms, specific embodiments of the technology are shown by way of example in the drawings and will be described herein. The drawings may not be to scale. However, it should be understood that the drawings and the detailed description thereof are not intended to limit the technology to the particular form disclosed, but on the contrary, are intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the technology as defined by the appended claims. Detailed Description

[0040] Details of one or more embodiments of the subject matter described herein are set forth in the drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

[0041] Various embodiments of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all embodiments are shown. In fact, the disclosure may be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Unless otherwise indicated, the term "or" is used herein in both an alternative and a conjunctive sense. The terms "exemplary" and "example" are used for illustration and do not denote a quality level. Like reference numerals always denote like elements. The arrows in each drawing depict two-way data flow and / or two-way data flow capabilities. The terms "path", "passageway", and "route" may be used interchangeably herein.

[0042] Embodiments of the present disclosure may be implemented in various ways, including as a computer program product including a product. A computer program product may include a non-transitory computer-readable storage medium storing applications, programs, program components, scripts, source code, program code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, etc. (also referred to herein as executable instructions, instructions for execution, computer program products, program code, and / or similar terms that may be used interchangeably herein). Such non-transitory computer-readable storage media include all computer-readable media (including volatile and non-volatile media).

[0043] In one embodiment, the non-volatile computer-readable storage medium may include a floppy disk, a flexible disk, a hard disk, a solid-state storage device (SSD) (e.g., a solid-state drive (SSD), a solid-state card (SSC), a solid-state module (SSM)), an enterprise flash drive, magnetic tape, or any other non-transitory magnetic medium, etc. The non-volatile computer-readable storage medium may include punched cards, paper tapes, optical mark sheets (or any other physical medium having a pattern of holes or other optically recognizable marks), compact disc read-only memory (CD-ROM), rewritable compact disc (CD-RW), digital versatile disc (DVD), Blu-ray disc (BD), any other non-transitory optical medium, etc. Such non-volatile computer-readable storage medium may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory (e.g., serial, NAND (negative AND), NOR (negative OR), etc.), multimedia memory card (MMC), secure digital (SD) memory card, smart media card, compact flash (CF) card, memory stick, etc. In addition, the non-volatile computer-readable storage medium may include conductive-bridging random access memory (CBRAM), phase-change random access memory (PRAM), ferroelectric random access memory (FeRAM), non-volatile random access memory (NVRAM), magnetoresistive random access memory (MRAM), resistive random access memory (RRAM), silicon-oxide-nitride-oxide-silicon memory (SONOS), floating junction gate random access memory (FJGRAM), Millipede memory, racetrack memory, etc.

[0044] In one embodiment, a volatile computer-readable storage medium may include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), fast page mode dynamic random access memory (FPM DRAM), extended data output dynamic random access memory (EDO DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), second generation double data rate synchronous dynamic random access memory (DDR2 SDRAM), third generation double data rate synchronous dynamic random access memory (DDR3 SDRAM), Rambus dynamic random access memory (RDRAM), two-transistor RAM (TTRAM), thyristor RAM (T-RAM), zero-capacitor (Z-RAM), Rambus in-line memory module (RIMM), dual in-line memory module (DIMM), single in-line memory module (SIMM), video random access memory (VRAM), caches (including various levels), flash memory, register memory, etc. It will be understood that in cases where an embodiment is described as using a computer-readable storage medium, other types of computer-readable storage media may be substituted or used in addition to the above computer-readable storage media.

[0045] It should be understood that the various embodiments of the present disclosure may also be implemented as a method, device, system, computing device, computing entity, etc. Thus, the embodiments of the present disclosure may take the form of a device, system, computing device, computing entity, etc. that executes instructions stored on a computer-readable storage medium to perform specific steps or operations. Accordingly, the embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete computer program product embodiment, and / or an embodiment including a combination of a computer program product and hardware that executes specific steps or operations.

[0046] Embodiments of the present disclosure will be described below with reference to block diagrams and flowcharts. Accordingly, it should be understood that each block of the block diagrams and flowcharts can be implemented in the form of a computer program product, a fully hardware embodiment, a combination of hardware and a computer program product, and / or a device, system, computing device, computing entity, etc. that executes instructions, operations, steps, and similar words that can be used interchangeably (e.g., executable instructions, instructions for execution, program code, etc.) on a computer-readable storage medium for execution. For example, the obtaining, loading, and execution of the code can be performed sequentially such that one instruction is obtained, loaded, and executed at a time. In some example embodiments, the obtaining, loading, and / or execution can be performed in parallel such that multiple instructions are obtained, loaded, and / or executed together. Thus, such embodiments can produce a specially configured machine that executes the steps or operations specified in the block diagrams and flowcharts. Accordingly, the block diagrams and flowcharts support various combinations of embodiments for executing the specified instructions, operations, or steps.

[0047] The following description is presented to enable a person having ordinary skill in the art to make and use the subject matter disclosed herein and to incorporate it into the context of a particular application. While the following is directed to specific examples, other and further examples can be designed without departing from the basic scope of the disclosure.

[0048] Various modifications and various uses in different applications will be apparent to those skilled in the art, and the general principles defined herein can be applied to a wide range of embodiments. Accordingly, the subject matter disclosed herein is not intended to be limited to the presented embodiments, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0049] In the provided description, numerous specific details are set forth in order to provide a more thorough understanding of the subject matter disclosed herein. However, it will be apparent to those skilled in the art that the subject matter disclosed herein can be practiced without necessarily being limited to these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring the subject matter disclosed herein.

[0050] Unless otherwise clearly stated, all features disclosed in this specification (e.g., any of the appended claims, abstract, and drawings) can be replaced by alternative features for the same, equivalent, or similar purpose. Accordingly, unless otherwise clearly stated, each feature disclosed is only one example of a general series of equivalent or similar features.

[0051] Various features are described herein with reference to the accompanying drawings. It should be noted that the drawings are only intended to facilitate the description of the features. The various features described are not intended to be an exhaustive description of the subject matter disclosed herein or a limitation on the scope of the subject matter disclosed herein. Additionally, the examples shown need not have all aspects or advantages shown. Aspects or advantages described in connection with a particular example are not necessarily limited to that example and may be practiced in any other example, even if not so shown or if not so expressly described.

[0052] Moreover, any element in a claim that does not expressly state “means for” performing a specific function or “step for” performing a specific function should not be construed as a “means” or “step” clause as specified in paragraph 6 of 35 U.S.C. § 112. Specifically, the use of “step of...” or “act of...” in the claims herein is not intended to invoke the provisions of paragraph 6 of 35 U.S.C. § 112.

[0053] It should be noted that the use of the labels left, right, front, back, top, bottom, forward, reverse, clockwise, and counterclockwise is for convenience only and is not intended to imply any particular fixed direction. Instead, the labels are used to reflect the relative position and / or orientation between the various parts of an object.

[0054] Any data processing may include data buffering, aligning incoming data from multiple communication paths, forward error correction (“FEC”), and / or others. For example, data may first be received by an analog front end (AFE) that prepares the incoming for digital processing. The digital portion of the transceiver (e.g., a DSP) may provide skew management, equalization, reflection cancellation, and / or other functions. It will be understood that the processing described herein may provide many benefits including both power and cost savings.

[0055] Moreover, the terms “system,” “component,” “module,” “interface,” “model,” etc. generally intend to denote a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of example, both an application running on a controller and the controller can be a component. One or more components may reside within a process and / or a thread of execution, and a component may be located on one computer and / or distributed between two or more computers.

[0056] Unless otherwise expressly stated, each numerical value and range can be interpreted as approximate, as if the word “about” or “approximately” preceded the value or values of the range. A signal and the corresponding node or port may be represented by the same name and are interchangeable herein for purposes.

[0057] Although embodiments may have been described with respect to circuit functionality, embodiments of the subject matter disclosed herein are not so limited. Possible implementations may be embodied in a single integrated circuit, a multi-chip module, a single card, a system-on-a-chip, or a multi-card circuit pack. It will be apparent to those skilled in the art that various embodiments may also be implemented as part of a larger system. Such embodiments may be employed in conjunction with, for example, a digital signal processor, a microcontroller, a field programmable gate array, an application specific integrated circuit, or a general purpose computer.

[0058] It will be apparent to those skilled in the art that the various functions of circuit elements may also be implemented as processing blocks in a software program. Such software may be employed in, for example, a digital signal processor, a microcontroller, or a general purpose computer. Such software may be embodied in the form of program code embodied in a tangible medium such as, for example, a magnetic recording medium, an optical recording medium, a solid state memory, a floppy disk, a CD-ROM, a hard disk drive, or any other non-transitory machine-readable storage medium. When the program code is loaded into and executed by a machine such as, for example, a computer, the machine becomes an apparatus for practicing the subject matter disclosed herein. When implemented on a general purpose processor, the program code segments combine with the processor to provide a unique apparatus that operates similarly to a specific logic circuit. The described embodiments may also be presented in other sequences of bitstreams or signal values transmitted electrically or optically through a medium, magnetic field changes stored in a magnetic recording medium, etc., generated using the methods and / or apparatuses described herein.

[0059] Artificial intelligence (AI) includes the concept of creating intelligent machines that can sense, reason, act, and adapt. Machine learning (ML) can be a subset of AI that helps build AI-driven applications. Deep learning can be a subset of machine learning that uses artificial neural networks to mimic the learning processes of the human brain. Deep learning algorithms can use large amounts of data and complex algorithms to train models. Neural networks can be the basis of deep learning algorithms. In machine learning, AI inference can include the process of making predictions using a trained model. In some cases, AI training is generally the first step in a two-part process of machine learning. Since inference does not include the model adjusting its parameters based on new data, inference can be faster than training. Inference also uses less processing power than a training cluster. AI can include AI inference delegation. AI inference delegation techniques provide scalable memory bandwidth and scalable memory capacity to accommodate increasing query lengths and increasing numbers of concurrent users.

[0060] AI search may include query processing, retrieval, and ranking. In some cases, the systems and methods described herein may include AI query processing of a storage device via near-memory processing. AI search can process large amounts of data and queries in real time, anticipate user needs based on previous search patterns, deliver accurate and relevant results quickly, automatically improve itself over time, learn from data about the user to automatically generate the most accurate and relevant search experience, etc. An AI search system can process various types of inputs including natural language queries, voice commands, images, context information, and the like.

[0061] The systems and methods described herein may include AI processing based on a neural processing unit (NPU). The NPU can be a dedicated processor that executes machine learning algorithms. The NPU may also be referred to as an AI accelerator or an intelligent processing unit (IPU). The NPU improves the inference performance of neural networks. The NPU can be configured to work similarly to the human brain. The NPU can be composed of nerve cells and synapses that send and receive signals to each other. The NPU can use a data-driven parallel computing architecture to process large amounts of multimedia data (e.g., images and videos). The NPU can be used to offload specific workloads, allowing the dedicated hardware to focus on more specialized tasks.

[0062] In some examples, the systems and methods may include an attention network. The attention network may include machine learning techniques that identify the strongest correlations between words in a sentence. The attention network can do this by learning patterns from a training corpus. The attention model can evaluate the input to identify the most important components and assign weights to each component. For example, when translating a sentence, the attention model can select the most important words and assign them higher weights. The attention mechanism can be additive or dot product. Additive attention can use a feed-forward neural network to compute the compatibility between a query and a key vector. Dot product attention can use the dot product to measure their similarity. The attention mechanism can also be self-attention. Self-attention can include mechanisms used in machine learning, particularly in natural language processing (NLP) and computer vision tasks. Self-attention can allow the model to identify and weigh the importance of different parts of an input sequence and how different parts are related to each other (e.g., the correlations between different parts of an input sequence or tokens). In some examples, the systems and methods of the present application may incorporate an attention network to perform the AI inference delegation techniques described herein.

[0063] In some cases, systems and methods may include machine learning (ML)-based attention. ML-based attention can be a mechanism that mimics cognitive attention. An attention network can compute "soft" weights for each word (more precisely, for its embedding) within a context window. The attention network can compute either in parallel (such as in a Transformer) or sequentially (such as in a recurrent neural network). The soft weights can change during each runtime, as compared to hard weights that are (pre)trained and fine-tuned and then kept frozen. Assuming that the attention network has learned those patterns through training, the attention network can be designed to identify the highest correlations between the words of a sentence. Such correlations can be captured in the neuron weights through backpropagation from self-supervised pre-training or supervised fine-tuning.

[0064] In machine learning, queries, keys, and values (QKV) in an attention network can be used to model the relationships between words. In some examples, the systems and methods described herein enable QKV vector multiplication through a processing-in-memory (PNM) solid-state drive (SSD). The attention network can assign weights to the words in a sentence, giving greater weights to the relevant words. The assigned weights can help preserve the context of the sentence and improve the accuracy of predictions. QKV enables the attention network to focus on what is considered the most important part of the input and generate a relevant and coherent output. The attention operation can be considered a retrieval process. For example, when searching for a video online, the query is the text in the search bar, the keys are the video title and description, and the values are the videos that match the query. In an AI Transformer, the query is the information being searched for, the keys are the context or references, and the values are the content being searched for. The query and the key can be multiplied together to produce an attention score, which is then used to calculate the weighted sum of the values.

[0065] In some examples, the systems and methods described herein may incorporate multi-head attention (MHA). MHA can be an attention mechanism that processes information from an input sequence using multiple attention layers. MHA can allow a neural network to control how information is mixed between segments of the input sequence. This can result in richer representations and improved performance on machine learning tasks. In some cases, MHA includes one or more modules configured for the attention mechanism, and one or more modules configured for the attention mechanism run through the attention mechanism in parallel several times. In some cases, the independent attention outputs are concatenated and linearly transformed to the desired dimension. MHA allows different parts of the sequence to be focused on (e.g., long-term dependencies versus short-term dependencies). Multi-head attention may include multiple attention layers (heads) in parallel. Each head may have a different linear transformation for queries, keys, values, and outputs. For example, one head may focus on the relationships between people, while another head may focus on the context of a sentence, etc. Multi-head attention combines the knowledge of the same attention pooling via different representation subspaces for queries, keys, and values. To compute the multiple heads of multi-head attention in parallel, appropriate tensor manipulation can be used.

[0066] In some examples, the systems and methods described herein enable softmax (flexible maximization) and layer normalization operations to be performed in the PNM SSD. The softmax function can be a function that transforms a vector of K real values into a vector of K real values that sum to 1 (e.g., transforms into a probability distribution of K possible outcomes). The input values can be positive, negative, zero, or greater than one, and softmax transforms them into values between 0 and 1 so that they can be interpreted as probabilities. The softmax function can provide probability values ranging from 1 to 0 (e.g., 0, 0.0925, 0.1, 0.95, 1, etc.), while the Max function can only give a binary output that is 1 for the maximum and 0 otherwise, with no possible values between 0 and 1. The softmax activation function can be a mathematical function that transforms a vector of real numbers into a probability distribution. The softmax function can be used as the activation function in the output layer of a neural network model that predicts a multinomial probability distribution. The softmax activation function can be used for multi-class classification problems where class membership is used on more than two class labels. The softmax function exponentiates each element to make them positive and then normalizes them by dividing by the sum of all exponentiated values. The output of Softmax can be a vector with the probabilities of each possible outcome. The sum of the probabilities in the vector v is 1 for all possible outcomes or classes. The softmax function can be an extension of the Sigmoid function. Sigmoid can be used for binary classification methods where there are two classes.

[0067] The systems and methods described herein may include activation functions. An activation function can be a function in a neural network for determining the output of a neuron. In some cases, the activation value may be based on a probability distribution. The activation value may indicate the correlation between text units in a query associated with attention data. An AI system may determine whether a neuron is activated or not based on a weighted sum of inputs. In some cases, the systems and methods may include layer normalization. Layer normalization can be a technique for normalizing the activations of a neural network layer. Layer normalization works by normalizing the activations of each individual sample in a batch by subtracting the mean and dividing by the standard deviation. Examples of activation functions may include the tanh function (hyperbolic tangent function), sigmoid function, exponential linear unit (ELU), linear activation function, maxout, and binary step activation function. The tanh function can be a non-linear activation function that can be used between layers of a neural network. The tanh function has a similar shape to the sigmoid function, but its range is from -1 to 1. The sigmoid activation function can be used in a neural network. The sigmoid function can be applied to the output of each neuron, allowing the network to introduce non-linearity into the model. The ELU activation function can be used for non-linear estimation. The output layer of the ELU activation function may include a single node and produce an estimated level from a given input. A linear activation function is a simple straight-line activation function where the function is proportional to the weighted sum of a neuron or input. With the maxout function, non-linearity can be applied as the dot product between the weights of a neural network and the data. When creating a binary classifier, the binary step function can be used as an activation function.

[0068] In some examples, the systems and methods may include a large language model (LLM), which may include an AI algorithm (e.g., a deep learning algorithm) that can understand, summarize, generate, and predict new content. The LLM may use statistical models to analyze large amounts of data, learn patterns, and connections between words and phrases. The LLM may be built on machine learning, specifically a type of neural network called a transformer model. In some cases, the systems and methods implement the LLM using a transformer model.

[0069] In some cases, the LLM may include a feed-forward layer (FFN) consisting of multiple fully-connected layers that transform the input embeddings. In doing so, these layers enable the model to gather higher levels of abstraction (e.g., to understand the user's intent for the text input). The systems and methods described herein may implement the Gaussian Error Linear Unit (GELU) as the activation function. The GELU activation function can be xΦ(x), where Φ(x) is the standard Gaussian cumulative distribution function. Dropout regularization randomly multiplies the inputs of neurons by 0, randomly deactivating them. The ReLU activation deterministically multiplies the input by 0 or 1 depending on the value of the input. GELU combines dropout regularization and the rectified linear unit (ReLU) by multiplying the input by a value from 0 to 1. However, although the zero-to-one mask values are determined randomly, they can also depend on the value of the input.

[0070] In some cases, the systems and methods may implement an LLM that includes a reduction layer. For example, a map-reduce file chain first applies the LLM chain separately to each document (the map step), treating the chain output as a new document. After that, the LLM passes all the new documents to a separate combined document chain to obtain a single output (the reduce step).

[0071] In some cases, the systems and methods described herein may be based on a Recurrent Neural Network (RNN). An RNN can be an artificial neural network designed to process sequential data. The RNN can identify sequential characteristics in the data and use patterns to predict the next possible scenario. The RNN may have feedback connections that allow them to retain information from previous time steps. This enables the RNN to capture temporal dependencies. The RNN can consist of a series of repeating neural network units connected in a chain-like structure. The output of one unit is passed as input to the next unit.

[0072] In some cases, the AI inference delegation systems and methods described herein are based on tokens. A token can be the basic unit for an AI model (e.g., an LLM) to process and generate language in text or code. Tokenization is the splitting of input / output text into smaller units for LLM AI processing. The vocabulary size is the number of tokens used by each model, and the vocabulary size varies between different models. Examples of tokens include characters, words, subwords, other segments of text or code, punctuation marks, parts of words, parts of sentences, phrases, etc. The tokenization method or scheme used determines the type of token. In some examples, the phrase [I love you.] may have five tokens: [I], [love], [you], [ ], and [.]. Tokens can be transformed into embeddings, which the LLM model then processes to understand the text. Each LLM has a maximum limit on the number of tokens it can process.

[0073] In some examples, the systems and methods may be based on a key-value cache (KV cache). The KV cache may include ways to temporarily store data to improve application response time. The KV cache involves caching frequently accessed data in a key-value store. The KV cache reduces database queries, complex calculations, and saves computing resources by reusing previously computed attention key-value pairs instead of recomputing attention key-value pairs for each generated token. The key and value states are used to compute scaled dot-product attention. The decoding phase may generate a single token at each time step, but each token is based on the key and value tensors of all previous tokens (including the input token KV tensors computed during pre-fill, and any new KV tensors computed up to the current time step). At each token generation step, the query vector of a single current token may be multiplied by the key vectors of all previous tokens in the sequence to create attention scores, and the scores are also multiplied by the value vectors of all previous tokens. Thus, instead of recomputing the key and value vectors of all previous tokens at each token generation step, the KV cache may perform incremental calculations based only on the current token and reuse the previously computed key / value vectors from the KV cache. The KV vectors of the current token may also be added to the KV cache for the next token generation step. The AI inference delegation systems and methods described herein may support relatively long query sizes with sufficient storage for advanced generative pre-trained transformers.

[0074] In some cases, the systems and methods implement near-memory computing (NMC). NMC can be a system architecture that moves computing capabilities to or near memory (e.g., random access memory (RAM)). This allows memory-centric computing and addresses the central processing unit (CPU)-memory bandwidth bottleneck. NMC may include near-memory processing (PNM) and processing-in-memory (PIM). PIM can improve performance and energy efficiency by offloading some of the data computation tasks from the CPU to inside the memory. PIM may allow computations and processing to be performed inside the memory of a computer, server, or similar device. PNM may incorporate memory and logic chips (e.g., processing units) into an integrated circuit package (e.g., system-on-chip (SoC)) that reduces data movement between the CPU and memory by leveraging the memory for data computation, resulting in improved system performance and increased energy efficiency. PNM may enable computational functions (e.g., AI processing) to execute closer to the memory in order to reduce the bottleneck that occurs between CPU and memory data transfers. PNM may be applied to caches, multi-threading, embedded random access memory, and AI processing to alleviate CPU memory bottleneck issues.

[0075] A solid-state drive (SSD) may include a non-volatile storage medium that stores persistent data on solid-state flash memory. A PNM SSD is an SSD incorporating PNM. Thus, a PNM SSD can incorporate memory and logic chips into an integrated circuit package (e.g., a SoC, a SoC SSD with memory, a logic chip, a processing unit, a microcontroller, solid-state flash memory, etc.). A system-on-chip (SoC) can be an integrated circuit (IC) that includes components of a computer system. The SoC can include at least one of a central processing unit (CPU), a field-programmable gate array (FPGA), a graphics processing unit (GPU), random access memory (RAM), a storage device (e.g., an SSD), a network interface, an input / output (I / O) port, a peripheral interface, an auxiliary storage device, and / or an I / O driver, etc. A system-in-package (SiP) can be a method for bundling multiple ICs (e.g., multiple SoCs) and passive components into a single package. The SiP can perform the functions of an entire system. The AI inference delegation systems and methods described herein reduce system costs by using a PNM SSD and reducing the use of expensive graphics processing unit (GPU) systems.

[0076] In some cases, the systems and methods described herein can implement RAM channels. A RAM channel can represent an aspect of a memory architecture referred to as a memory channel. A memory channel can be the number of channels of communication available between a RAM module and a memory controller. A memory controller can be a digital circuitry that manages the data flow to and from a RAM module. The rank of a RAM can be related to the number of memory chip sets included in a module. The memory channels of a RAM enable a memory controller to access the respective ranks of the RAM to provide a data pipeline between the RAM and the CPU. A multi-channel memory architecture can increase the number of channels available to a memory controller, resulting in an increase in the data transfer rate. A dynamic RAM (DRAM) channel can include a controller interface that can communicate with one or more ranks. A DRAM channel can be a common group of address / data lines that act together. The systems and methods can implement through-silicon vias (TSVs, also referred to as through-silicon vias). A TSV can include a chip packaging technology that vertically connects integrated chip dies (e.g., a processor die, a DRAM die, etc.). A TSV can be an alternative to wire bonding and flip chip for creating 3D packages and 3D integrated circuits. Hybrid copper bonding is a process that uses copper-to-copper connections to connect dies in a package. Hybrid copper bonding can be used for packages with a relatively small pitch (e.g., a 10 μm pitch and below).

[0077] Some AI inference systems (e.g., generative AI, generative pre-trained transformers, LLMs, etc.) may encounter performance obstacles associated with memory bandwidth and memory capacity. As the length of queries grows and / or as the number of concurrent users increases, the AI inference system can reach the limits of memory bandwidth and / or memory capacity. In some cases, QKV data processing can be a limiting factor for concurrent user and / or query length scalability.

[0078] Using a transformer attention system, user query processing (e.g., QKV processing) can include relatively high memory bandwidth and / or memory capacity constraints based on the likelihood of encountering relatively long query lengths. However, in some cases, QKV processing can be associated with relatively low computational constraints (e.g., low computational load).

[0079] Some methods use a main processing unit (e.g., a GPU, CPU, etc. of a high-performance computing system) to compute weight operations and QKV operations (e.g., all processing of a transformer attention system). As generative pre-trained transformers have evolved, query lengths have steadily increased (e.g., progressing from 2 kilobyte (K) query lengths to 120K query lengths). Storage space constraints are increasing to handle longer query lengths and the increasing number of concurrent users. High-bandwidth memory (HBM) can be used, but HBM capacity is limited, and host cloud storage can be too slow to handle the increasing memory bandwidth and memory capacity constraints.

[0080] Since weight operations are more computationally intensive compared to memory-intensive (because the computational intensity of weight operations is higher than the memory intensity of weight calculations) (e.g., weight operations involve a relatively high level of processing), weight operations are better executed by the main processing unit (e.g., CPU, GPU) of a compute node. Since QKV operations are more memory-intensive compared to computationally intensive (e.g., QKV operations involve relatively high memory bandwidth and / or memory capacity), QKV operations are better executed by a near-memory processing (PNM) storage device.

[0081] Figure 1 An example system 100 is shown in accordance with one or more embodiments described herein. In Figure 1 is shown a machine 105, which may be referred to as a host, system, or server. Although Figure 1 machine 105 is depicted as a tower computer, the disclosed embodiments can scale to any form factor or type of machine. For example, machine 105 can be a rack server, blade server, desktop computer, tower computer, mini-tower computer, desktop server, laptop computer, notebook computer, tablet computer, etc.

[0082] Machine 105 may include a processor 110, a memory 115, and a storage device 120. The processor 110 can be any kind of processor. Note that, for ease of illustration, the processor 110 along with other components discussed below are shown outside the machine: the disclosed embodiments may include these components within the machine. Although Figure 1 a single processor 110 is shown, machine 105 may include any number of processors, each of which can be a single-core or multi-core processor, each of which can implement a reduced instruction set computer (RISC) architecture or a complex instruction set computer (CISC) architecture (and other possibilities), and can be mixed in any desired combination.

[0083] The processor 110 may be coupled to the memory 115. The memory 115 can be any kind of memory (such as, flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), persistent random access memory, ferroelectric random access memory (FRAM), or non-volatile random access memory (NVRAM) (such as, magnetoresistive random access memory (MRAM), phase change memory (PCM), or resistive random access memory (ReRAM))). The memory 115 may include volatile memory and / or non-volatile memory. The memory 115 may use any desired form factor: for example, single in-line memory module (SIMM), dual in-line memory module (DIMM), non-volatile DIMM (NVDIMM), etc. The memory 115 can be any desired combination of different memory types and can be managed by a memory controller 125. The memory 115 can be used to store data that may be referred to as "short-term": that is, data not expected to be stored for long periods of time. Examples of short-term data may include temporary files, data used locally by an application (which may have been copied from other storage locations), etc.

[0084] The processor 110 and the memory 115 may support an operating system under which various applications can run. These applications may issue requests (which may be referred to as commands) to read data from or write data to the memory 115 or the storage device 120. When the storage device 120 is used to support applications that read or write data via some file system, a device driver 130 may be used to access the storage device 120. Although Figure 1A storage device 120 is shown, but any number (one or more) of storage devices may be present in machine 105. The storage device 120 may support any desired one or more protocols, and the one or more protocols include, for example, the Non-Volatile Memory Express (NVMe) protocol, the Serial Attached SCSI (SAS) protocol, or the Serial ATA (SATA) protocol. The storage device 120 may include any desired interface, and the interface includes, for example, a Peripheral Component Interconnect Express (PCIe®) interface or a Compute Express Link (CXL) interface. The storage device 120 may adopt any desired form factor, and the form factor includes, for example, the U.2 form factor, the U.3 form factor, the M.2 form factor, the Enterprise and Data Center Standard Form Factor (EDSFF) (including all its varieties (such as, E1 short, E1 long, and E3 varieties)), or an Add-in Card (AIC).

[0085] Although Figure 1 the term "storage device" is used, the disclosed embodiments may include any storage device format that may benefit from the use of a compute storage unit, and examples of the compute storage unit may include a hard disk drive, a solid state drive (SSD), or a persistent memory device (such as, PCM, ReRAM, or MRAM). Any reference to "storage device" or "SSD" below should be understood to include such other disclosed embodiments and other types of storage devices. In some cases, the term "storage unit" may encompass the storage device 120 and the memory 115. The machine 105 may include a power supply 135. The power supply 135 may supply power to the machine 105 and its components.

[0086] The machine 105 may include a transmitter 145 and a receiver 150. The transmitter 145 or the receiver 150 may be used to transmit or receive data (such as, query processing data), respectively. In some cases, the transmitter 145 and / or the receiver 150 may be used to communicate with the memory 115 and / or the storage device 120. The transmitter 145 may include a write circuit 160, and the write circuit 160 may be used to write data to a storage device (such as, a register in the memory 115 and / or the storage device 120). In a similar manner, the receiver 150 may include a read circuit 165, and the read circuit 165 may be used to read data from a storage device (such as, a register in the memory 115 and / or the storage device 120).

[0087] In one or more examples, machine 105 can be implemented using any type of device. Machine 105 can be configured as one or more of a server (such as, a compute server, a storage server, a storage node, a network server, a supercomputer, a data center system, etc., or any combination thereof) (e.g., a host configured as one or more of a server (such as, a compute server, a storage server, a storage node, a network server, a supercomputer, a data center system, etc., or any combination thereof)). Additionally or optionally, machine 105 can be configured as one or more of a computer (such as, a workstation, a personal computer, a tablet, a smart phone, etc., or any combination thereof) (e.g., a host configured as one or more of a computer (such as, a workstation, a personal computer, a tablet, a smart phone, etc., or any combination thereof)). Machine 105 can be implemented using any type of device that can be configured as an apparatus, which includes, for example, an accelerator device, a storage device, a network device, a memory expansion and / or buffer device, a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), etc., or any combination thereof.

[0088] Any communication between devices including machine 105 (e.g., a host, a compute storage device, and / or any intermediate device) can occur through an interface, which can be implemented using any type of wired and / or wireless communication medium, interface, protocol, etc. Any type of wired and / or wireless communication medium, interface, protocol, etc. includes PCIe, NVMe, Ethernet, NVMe over Fabrics (NVMe-oF), Compute Express Link (CXL), and / or coherence protocols (such as, CXL.mem, CXL.cache, CXL.IO, etc.), Gen-Z, Open CAPI (Open Coherent Accelerator Processor Interface), Cache Coherent Interconnect for Accelerators (CCIX), Advanced eXtensible Interface (AXI), etc., or any combination thereof, Transmission Control Protocol / Internet Protocol (TCP / IP), Fibre Channel, InfiniBand, Serial ATA (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, any generation of wireless network including 2G, 3G, 4G, 5G, etc., any generation of Wi-Fi, Bluetooth, Near Field Communication (NFC), etc., or any combination thereof. In some embodiments, the communication interface can include a communication network fabric, which includes one or more links, buses, switches, hubs, nodes, routers, translators, repeaters, etc. In some embodiments, system 100 can include one or more additional devices having one or more additional communication interfaces.

[0089] Any functionality described herein (including any host functionality, device functionality, query processing controller 140 functionality, etc.) can be implemented using hardware, software, firmware, or any combination thereof, including, for example, a hardware and / or software combination logic, timing logic, timer, counter, register, state machine, volatile memory (such as dynamic random access memory (DRAM) and / or static random access memory (SRAM)), non-volatile memory (including flash memory), persistent memory (such as cross-grid non-volatile memory), memory with body resistance change, phase change memory (PCM), etc. and / or any combination thereof, complex programmable logic device (CPLD), field programmable gate array (FPGA), application specific integrated circuit (ASIC), CPU (including complex instruction set computer (CISC) processors (such as x86 processors) and / or reduced instruction set computer (RISC) processors (such as RISC-V and / or ARM processors)), graphics processing unit (GPU), neural processing unit (NPU), tensor processing unit (TPU), etc. In some embodiments, one or more components of the query processing controller 140 can be implemented as a system on a chip (SoC).

[0090] In some examples, the query processing controller 140 can include any one or combination of logic (e.g., logic circuits), hardware (e.g., processing unit, memory, storage device), software, firmware, etc. In some cases, the query processing controller 140 can execute one or more functions in conjunction with the processor 110. In some cases, at least a portion of the query processing controller 140 can be implemented in or by the processor 110 and / or the memory 115. One or more logic circuits of the query processing controller 140 can include any one or combination of a multiplexer, register, logic gate, arithmetic logic unit (ALU), cache, computer memory, microprocessor, processing unit (CPU, GPU, NPU, and / or TPU), FPGA, ASIC, etc., which enables the query processing controller 140 to provide artificial intelligence query processing through near-memory processing storage devices.

[0091] In one or more examples, the query processing controller 140 may provide artificial intelligence query processing in conjunction with a near-memory processing storage device. For example, the query processing controller 140 may provide the AI inference delegation techniques described herein. In one or more examples, the query processing controller 140 receives attention data from a GPU processing device, processes key values from the attention data using transposed query values from the attention data, determines a probability distribution of the result of the processing, and generates an activation value (e.g., activation value, activation vector) based on the probability distribution, the activation value indicating the correlation between text units in the query associated with the attention data. Thus, the query processing controller 140 provides scalable memory bandwidth and scalable memory capacity to accommodate increasing query lengths and / or increasing numbers of concurrent users, thereby reducing system costs by reducing the use of relatively expensive graphics processing unit (GPU) compute-intensive systems, avoiding backing up non-activated KV data to external storage devices, and minimizing the communication overhead between the main compute node (e.g., GPU system) and a relatively high memory bandwidth system (e.g., PNM SSD array).

[0092] Figure 2 illustrating details of a Figure 1 machine 105 according to the examples described herein. In Figure 2 general, the machine 105 includes one or more processors 110, which may include a memory controller 125 and a clock 205, and the memory controller 125 and the clock 205 may be used to coordinate the operation of the components of the machine. The processor 110 may be coupled to a memory 115, which, by way of example, may include random access memory (RAM), read only memory (ROM), or other state-preserving media. The processor 110 may be coupled to a storage device 120 and to a network connector 210, which may be, for example, an Ethernet connector or a wireless connector. The processor 110 may be connected to a bus 215, to which, among other components, a user interface 220 and an input / output (I / O) interface port may be attached, and the user interface 220 and the input / output (I / O) interface port may be managed using an I / O engine 225. As shown, the processor 110 may be coupled to a query processing controller 230, which may be an Figure 1 example of the query processing controller 140. Additionally or alternatively, the processor 110 may be connected to the bus 215 to which the query processing controller 230 may be attached.

[0093] Figure 3Illustrates an example system 300 according to one or more embodiments described herein. In one or more examples, system 300 illustrates an example system for artificial intelligence query processing via a near-memory processing storage device. In the example shown, system 300 may include a PNM SSD SoC 305. As shown, the PNM SSD SoC 305 may include a key-value (KV) cache 310, at least one NPU 315, a memory channel 320, a communication interface 325 (e.g., a PCIe communication interface), at least one interconnect 330 (e.g., one or more die-to-die (D2D), chip-to-chip (C2C), chiplet-to-chiplet, SoC interconnect, etc.). As shown, the PNM SSD SoC 305 may include at least one NPU (e.g., NPU 315a, NPU 315b, NPU 315c to NPU 315N), where N is a positive integer. The memory channel 320 may include a persistent memory channel (e.g., flash memory, SSD, etc.) and / or a DRAM channel (e.g., RAM, SRAM, etc.), etc. In some examples, one or more components of system 300 (e.g., PNM SSD SoC 305, KV cache 310, at least one NPU, memory channel 320, etc.) may be incorporated into Figure 1 query processing controller 140 of Figure 2 query processing controller 230 of Figure 1 query processing controller 140 of Figure 2 query processing controller 230 of and operate in conjunction with

[0094] Weight operations (e.g., weight matrix operations) are more constrained by computing performance (e.g., measured in teraflops) than by memory bandwidth. However, QKV operations are more constrained by memory bandwidth and / or memory capacity than by computing performance. Thus, the techniques described herein delegate weight processing to one or more processing units (e.g., the CPU, GPU, etc. of a main computing node) and delegate QKV processing to one or more PNM storage devices (e.g., PNM SSD, PNM SSD SoC, PNM SSD SoC array) (such as PNM SSD SoC 305). Thus, the PNM SSD SoC 305 may be configured to perform QKV processing for LLM queries, while the processing unit (e.g., a GPU compute-intensive system) may process the weight operations of the LLM queries.

[0095] In one or more examples, the PNM SSD SoC 305 can be configured to support relatively long query lengths with sufficient storage for advanced AI inference systems, generative AI, generative pre-trained transformers, LLMs, etc. The PNM SSD SoC of system 300 is configured to perform query processing (e.g., QKV vector multiplication) for one or more queries, while a compute-intensive system (e.g., a GPU system with high-bandwidth memory) performs weight processing for one or more queries. The PNM SSD SoC 305 can be configured to perform softmax functions and layer normalization operations (e.g., offload the processing unit of the main processing unit, the compute-intensive hardware system).

[0096] Figure 4 FIG. 400 shows an example system 400 according to one or more embodiments described herein. In the example shown, system 400 may include a GPU system 405 and a QKV system 410. The GPU system 405 can be configured as a compute-intensive GPU-based system for weight matrix multiplication. The QKV system 410 can be configured as a scalable QKV processing system. As shown, the GPU system 405 may include at least one GPU (e.g., GPUs 415a to 415K, where K is a positive integer). In some examples, one or more components of system 400 (e.g., the GPU system 405, the QKV system 410, one or more components of the GPU system 405, and / or one or more components of the QKV system 410) may be incorporated into Figure 1 query processing controller 140 and / or Figure 2 query processing controller 230 and / or in conjunction with Figure 1 query processing controller 140 and / or Figure 2 query processing controller 230 to act together.

[0097] As shown, GPU 415a may include at least one HBM (e.g., HBMs 420a, 420b, 420c, 420d, 420e to 420L, where L is a positive integer). As shown, GPU 415K may include at least one HBM (e.g., HBMs 425a, 425b, 425c, 425d, 425e to 425L). As shown, each GPU may include L HBM units, where L is a positive integer. In some cases, GPU 415K may have fewer or more HBM units than GPU 415a.

[0098] In the example shown, the GPU system 405 is connected to the QKV system 410 via at least one instance of a communication interface (or high - bandwidth expansion bus) 480. Communication via the communication interface 480 can flow from the GPU system 405 to the QKV system 410 and / or from the QKV system 410 to the GPU system 405. In some cases, the communication interface 480 can implement the NVMe protocol (e.g., 16GB / s, up to 300 users per instance of the communication interface 480). In some cases, the GPU system 405 can communicate the values or vectors of Q, K, and / or V (e.g., vector Q[N], vector K[N], vector V[N], where N for the QKV values is a positive integer based on the query length and / or the number of users and can be independent of other instances of N used as an integer here) to the QKV system 410. In some cases, the QKV system 410 can communicate activation values (e.g., vector activation[N], where N for the activation values is a positive integer) to the GPU system 405. In this case, N can be independent of other uses or other instances of the variable N here (e.g., the number N of memory modules is different from N in activation[N]). In terms of Q[N], K[N], V[N], and activation[N], N can be the number of concurrent users and / or queries in a batch process. In some cases, N can change (e.g., increase, decrease) with each iteration of query processing. For each iteration, the QKV values received by the QKV system 410 can be based on attention data from a previous iteration (e.g., updated attention data) (e.g., the QKV values received by the QKV system 410 can be updated based on attention data from a previous iteration (e.g., updated attention data) and adjusted based on attention data from a previous iteration (e.g., updated attention data)).

[0099] As shown, the QKV system 410 includes at least one PNM SSD SoC (e.g., a PNM SSD array including PNM SSD SoC 430a, PNM SSD SoC 430b, etc.). In some cases, the PNM SSD SoC 430a is a system - on - chip die including a solid - state drive and at least one processor. In some cases, the PNM SSD SoC 430b is a system - on - chip die including a solid - state drive and at least one processor. The number of PNM SSD SoCs in the QKV system 410 can be based on the determined demand at any given time (e.g., adding and / or activating PNM SSD SoCs based on demand, pluggable PNM SSD SoCs, hot - pluggable PNM SSD SoCs, etc.). In some cases, the number of PNM SSD SoCs in the QKV system 410 can increase over time according to demand (e.g., the number of queries, the number of users).

[0100] In the illustrated example, the PNM SSD SoC 430a includes a KV cache 435 and at least one NPU (e.g., NPU 440a, NPU 440b, NPU 440c to NPU 440M, where M is a positive integer). As shown, the PNM SSD SoC 430a may include a communication interface 445 (e.g., a PCIe communication interface) and an interconnect 450 (e.g., D2D, C2C, SoC interconnect, etc.). The communication interface 445 may be an example of the communication interface 325. The interconnect 450 may be an example of the interconnect 330. As shown, the communication interface 445 may enable communication between the GPU system 405, the QKV system 410, the PNM SSD SoC 430a, and / or the PNM SSD SoC 430b.

[0101] In the illustrated example, the PNM SSD SoC 430b includes a KV cache 465 and at least one NPU (e.g., NPU 470a, NPU 470b, NPU 470c to NPU 470M, where M is a positive integer). In some cases, the PNM SSD SoC 430b may have more or fewer NPUs than the PNM SSD SoC 430a. In some cases, the PNM SSD SoC 430b may include an interconnect 460a and an interconnect 460b (e.g., D2D, C2C, SoC interconnect). As shown, the interconnect 460a may enable the PNM SSD SoC 430b to communicate with the PNM SSD SoC 430a (e.g., D2D, C2C, SoC interconnect communication, etc.). As shown, the interconnect 460b may enable the PNM SSD SoC 430b to communicate with another PNM SSD SoC added to the QKV system 410 (e.g., via D2D, C2C, SoC interconnect communication, etc.). In some examples, the QKV system 410 includes multiple PNM SSD SoC dies (e.g., PNM SSD SoC 430a, PNM SSD SoC 430b, etc.) clustered in a system-in-package (SiP) via D2D interconnects (e.g., interconnect 450, interconnect 460a, interconnect 460b, etc.).

[0102] In one or more examples, each PNM SSD SoC may include an array of memory modules (e.g., NAND flash memory, DRAM, other types of persistent or non-persistent memory). In the example shown, PNM SSD SoC 430a may be connected to and / or include at least one memory cell (e.g., memory 455a, memory 455b, memory 455d to memory 455N, where N is a positive integer). The memory of PNM SSD SoC 430a (e.g., memory 455a to memory 455N) may be an example of memory channel 320. In the example shown, PNM SSD SoC 430b may be connected to and / or include at least one memory cell (e.g., memory 475a, memory 475b, memory 475d to memory 475N, where N is a positive integer).

[0103] In some examples, system 400 delegates the calculation of weight processing of queries (e.g., a batch of queries) to GPU system 405 and delegates the QKV processing of queries (e.g., a batch of queries) to QKV system 410. In some cases, QKV system 410 stores the key-value matrix (the KV matrix associated with QKV processing) via the PNM SSD array, enabling the processing of relatively large query lengths and / or concurrent users associated with increasingly advanced generative pre-trained transformers. Based on QKV system 410, there is no need to back up the inactive KV data to an external storage device as in some methods. Instead, the KV data is held in the KV cache (e.g., KV cache 435, KV cache 465).

[0104] In one or more examples, the KV cache 435 and / or the KV cache 465 can be on-die caches (e.g., caches, SRAMs, etc.), and the on-die caches (e.g., caches, SRAMs, etc.) are configured to buffer or store key-value matrices (KV matrices and / or key-value tables associated with QKV processing). In some cases, the KV cache 435 and / or the KV cache 465 can be implemented to avoid redundantly loading key-value tables from NAND (e.g., NAND that is slower than a cache such as the KV cache). In some examples, the KV cache 435 and / or the KV cache 465 can include on-die caches built on corresponding PNM SSD SoCs based on a 3D stacking process (e.g., through-silicon vias (TSVs), hybrid copper bonding, etc.). In some cases, the KV cache 435 and the processor PNM SSD SoC 430a (e.g., at least one of the NPUs 440a, 440b, 440c to 440M) can be formed as a stacked integrated circuit, and the stacked integrated circuit is stacked on top of the processor based on the KV cache 435 and the KV cache 435 is communicatively connected to the processor. For example, at least a portion of the KV cache 435 can be stacked on top of the NPU 440a, and the KV cache 435 can be communicatively connected to the NPU 440a through one or more vertical connections running between the KV cache 435 and the NPU 440a.

[0105] The memories of the PNM SSD SoC 430b (e.g., memories 475a, 475b, 475d to 475N) may be examples of memory channels 320. In some cases, the number of memory units (e.g., memories 475a, 475b, 475d to 475N) connected to the PNM SSD SoC 430b may be less than or more than the number of memory units (e.g., memories 455a, 455b, 455d to 455N) connected to the PNM SSD SoC 430a. In some examples, at least one memory unit of the PNM SSD SoC 430a may be at least partially incorporated on the PNM SSD SoC 430a. In some cases, at least one memory unit of the PNM SSD SoC 430b may be at least partially incorporated on the PNM SSD SoC 430b. In some examples, at least one memory unit of the PNM SSD SoC 430a may be NAND flash memory, DRAM memory, or another type of persistent memory and / or non-persistent memory. In some cases, at least one memory unit of the PNM SSD SoC 430b may be NAND flash memory or another type of persistent and / or non-persistent memory.

[0106] In one or more examples, system 400 illustrates an example of a QKV processing system for artificial intelligence query processing via a near-memory processing storage device. As shown, the GPU system 405 may be configured to process weight operations (e.g., weight quantization) associated with a query (e.g., a batch of queries). In some examples, the PNM SSD arrays of the QKV system 410 (e.g., PNM SSD SoC 430a, PNM SSD SoC 430b, etc.) may be configured to process QKV operations (e.g., QKV matrix multiplication) associated with a query (e.g., a batch of queries). Each PNM SSD SoC of the QKV system 410 provides both relatively high memory / storage device bandwidth and memory / storage device capacity. Thus, increased memory / storage device bandwidth and memory / storage device capacity constraints can be met by simply adding one or more additional PNM SSD SoCs to the QKV system 410 (e.g., pluggable PNM SSDs, hot-pluggable PNM SSDs). Accordingly, the QKV system 410 provides a scalable system configured to accommodate increased query lengths and / or the number of concurrent users at a relatively low cost.

[0107] Based on the QKV system 410, multiple PNM SSD dies can be clustered via D2D interconnect. For example, the interconnect 450 enables the PNM SSD SoC 430a to be connected to the PNM SSD SoC 430b via the interconnect 460a. Similarly, the interconnect 460b enables the PNM SSD SoC 430b to be connected to another PNM SSD SoC to provide a cluster of interconnected PNM SSD SoCs for QKV processing. In some cases, the PNM SSD SoCs (e.g., die of the SoC) of the QKV system 410 can be clustered in a system-in-package (SiP), enabling a PNM SSD SoC (e.g., PNM SSD SoC 430a) to send partial computation results to one or more other PNM SSD SoCs (e.g., PNM SSD SoC 430b) and / or receive partial computation results from one or more other PNM SSD SoCs (e.g., PNM SSD SoC 430b) without host communication (e.g., bypassing the GPU system 405 and / or the host).

[0108] Based on the QKV system 410, the KV matrix (e.g., one or more queries, a batch of queries, etc.) can be distributed to one or more of the PNM SSD SoCs. For example, the first part of the KV matrix of one or more queries can be distributed to the PNM SSD SoC 430a, the second part of the KV matrix of one or more queries can be distributed to the PNM SSD SoC 430b, the third part of the KV matrix of one or more queries can be distributed to another PNM SSD SoC (e.g., connected to the PNM SSD SoC 430b), and so on.

[0109] Each PNM SSD SoC of the QKV system 410 can be configured to perform query processing (e.g., perform partial matrix multiplication) on a part of the KV matrix. The PNM SSD SoC can be configured to perform one or more reduction operations in coordination with one or more other PNM SSD SoCs (e.g., based on the partial matrix multiplication performed by two or more PNM SSD SoCs). The reduction operation between two or more PNM SSDs (e.g., between at least the PNM SSD SoC 430a and the PNM SSD SoC 430b, etc.) can be performed based on the D2D interconnect of two or more PNM SSD SoCs.

[0110] To minimize the load latency for a memory (e.g., NAND flash) of a PNM SSD array, the QKV system 410 can pre-load at least the PNM SSD SoC 430a and / or the PNM SSD SoC 430b based on an "early hint" operation (e.g., issued from a compute-intensive die, issued from the GPU system 405, issued from the host system, etc.). For example, the GPU system 405 can provide a hint (e.g., a pre-load trigger) to the QKV system 410 (e.g., before the QKV system 410 receives QKV data, receives the KV matrix, and / or starts QKV matrix multiplication, etc.). In some cases, the hint can include a user ID (e.g., one or more user IDs) associated with one or more queries and / or users and / or layer count information (e.g., user ID and layer count information associated with a batch of queries / users). In one embodiment, the QKV system 410 receives a trigger from a processing device (e.g., the GPU system 405) based on an early hint operation, where the trigger includes at least one of a user identifier and layer count information associated with the attention data.

[0111] The techniques described herein include query processing logic for providing artificial intelligence query processing via a near-memory processing storage device. The query processing logic includes any combination of hardware (e.g., at least one memory, at least one processor), logic circuitry, firmware, and / or software for providing artificial intelligence query processing via a near-memory processing storage device. The query processing can involve artificial intelligence inference (e.g., related to an attention network).

[0112] To accommodate increased query lengths and an increasing number of concurrent users, the techniques described divide query processing into two parts rather than using only a single system (e.g., an expensive GPU-based system). Weight operations are more compute-intensive compared to memory-intensive (e.g., weight operations have relatively high processing constraints). Thus, weight operations are performed by a high-performance computing node (e.g., the GPU system 405). QKV operations are more memory-intensive compared to compute-intensive (e.g., QKV operations have relatively high memory bandwidth and / or memory capacity constraints). Thus, QKV operations can be delegated to a high-memory bandwidth system (e.g., the QKV system 410, the PNM SSD SoC 430a, the PNM SSD SoC 430b, etc.) and performed by a high-memory bandwidth system (e.g., the QKV system 410, the PNM SSD SoC 430a, the PNM SSD SoC 430b, etc.).

[0113] Compared with some methods, the delegation of weight operations and QKV operations provides better performance and scalability for LLM inference. Based on the techniques described herein, the communication overhead between the main computing node (e.g., GPU system) and the memory bandwidth-intensive system (e.g., PNMSSD) is minimized.

[0114] In some examples, the PNM SSD SoC 430a receives attention data from the GPU system 405 (e.g., via PCIe, communication interface 445). In some cases, the PNM SSD SoC 430a processes the key values from the attention data using the transposed query values from the attention data and determines the probability distribution of the processed result. In some examples, the PNM SSD SoC430a generates activation values (e.g., activation values, activation vectors, activation[N]) based on the probability distribution. The activation values can indicate the correlations between text units in the queries associated with the attention data.

[0115] In one or more examples, the PNM SSD SoC 430a receives at least a second activation value (e.g., second activation value, second activation vector) from the PNM SSD SoC 430b via the interconnect 450 and forms a unified activation value (e.g., vector activation2[N]) at least in part based on a combination of the activation value of the PNM SSD SoC430a and the second activation value of the PNM SSD SoC 430b.

[0116] In one or more examples, the attention data received by the PNM SSD SoC 430a is part of the multi-head attention data distributed among multiple PNM storage devices of the QKV system 410. In some cases, the attention data includes a part of the query attention data from the query multi-head attention layer, a part of the key attention data from the key multi-head attention layer, and a part of the value attention data from the value multi-head attention layer. In some examples, the attention data is based on an iteration of activation values that are generated based on previous partial outputs generated by multiple PNM storage devices that are reduced to a unified output.

[0117] Figure 5 Illustrates an example system 500 according to one or more embodiments described herein. The system 500 illustrates an example system for artificial intelligence query processing via a near-memory processing storage device. As shown, Figure 5 Depicts a PNM storage device.

[0118] In the example shown, system 500 includes query MHA (MHA(Q)) 505a, key MHA (MHA(K)) 505b, value MHA (MHA(V)) 505c, QKV system 510, and output MHA (MHA(O)) 530. In some examples, QKV system 510 is Figure 4 an example of QKV system 410. As shown, QKV system 510 includes query processing controller 515. Query processing controller 515 can be Figure 1 an example of query processing controller 140 and / or Figure 2 an example of query processing controller 230. In some cases, query processing controller 515 can be incorporated into one or more PNM SSD SoCs (e.g., PNM SSD SoC array, PNM SSD SoC 305, PNM SSD SoC 430a, PNM SSD SoC 430b) and / or operate in conjunction with one or more PNM SSD SoCs (e.g., PNM SSD SoC array, PNM SSD SoC 305, PNM SSD SoC 430a, PNM SSD SoC 430b).

[0119] In one or more examples, an input query is assigned to a given system (e.g., lexical unit index [N], where N indicates the number of concurrent users and / or queries in a batch process). The input query can be provided to an embedding position encoding process that filters the input query. In some examples, query data from a batch process can be divided into MHA(K) 505a, MHA(V) 505b, and MHA(Q) 505c. In some cases, the input to MHA(Q) 505c includes activation values (e.g., vector activation 1[N] based on lexical unit index [N]) based on query weights (Wq1) applied to activation values (e.g., query weights applied to a set of activation values). In some examples, the input to MHA(K) 505a includes activation values based on key weights (Wk1) applied to activation values (e.g., key weights applied to a set of activation values). In some examples, the input to MHA(V) 505b includes activation values based on value weights (Wv1) applied to activation values (e.g., value weights applied to a set of activation values). In some cases, a computationally intensive system (e.g., GPU system 405) calculates weight values (e.g., Wq1, Wk1, Wv1).

[0120] As shown, the MHA(Q) outputs Q1[N], the MHA(K) outputs K1[N], and the MHA(V) outputs V1[N] into the QKV system 510 (e.g., a near-memory processing (PNM) storage device, an array of PNM SSD SoCs). In some examples, the query processing controller 515 performs QKV processing based on the input Q1[N], K1[N], and V1[N].

[0121] As shown, the QKV system 510 receives query-based inputs (e.g., input vectors {K[N], V[N], Q[N]}, where N indicates the number of concurrent users / queries in a batch process). In some cases, N indicates the user_ID associated with the query (e.g., the value of N is the user_ID for a given query). In the example shown, the queries / responses (e.g., all queries / responses {K[0:QL], V[0:QL]}, where QL is the query length (such as 2048, 4096, or 8192 characters)) are stored and processed in the near-memory processing storage device. For example, the vector K1[N] and the vector V1[N] are stored in the QKV system 510 (e.g., in an array of PNM SSD SoCs).

[0122] In some examples, the QKV processing may include reading the QKV history (e.g., K[0:QL][N], V[0:QL][N]), matrix multiplication, scoring via softmax, matrix multiplication, and outputting the calculation result. The QKV processing may include matrix multiplication operations (e.g., K[0:QL][N] × Q-T[N], K1{1 to 2048} × Q1[N]). The QKV processing may include scoring via the softmax function (e.g., softmax[N], softmax[0:QL][N]). In some cases, the matrix multiplication operation may include softmax[0:QL][N] × V[0:QL][N] and / or softmax[N] × V1{1 to 2048}. In some examples, K1{1 to 2048} and V1{1 to 2048} represent a memory data set (e.g., historical query metrics). When the query length increases, the given memory data set increases (e.g., from 1 to 2048 characters based on a system limit of 2048 characters per query). K1{1 to 4096} and V1{1 to 4096} represent data sets that can be from 1 to 4096 characters, and so on. When new metrics are added, the QKV data is multiplied by the historical query metrics (e.g., the result of a previous matrix multiplication). For example, K1[N] may be multiplied by the historical query metric 520 to obtain K1{1 to 2048} (e.g., K1 has 1 to 2048 characters). The historical query metric 520 may include K 1 1…n(12288 x n). Similarly, V1[N] can be multiplied by the historical query metric 525 to obtain V1{1 to 2048} (e.g., V1 has 1 to 2048 characters). The historical query metric 525 can include V 1 1…n (12288 x n).

[0123] In one or more examples, QKV processing can include outputting a calculation result (e.g., outputting {activation2[N]}). In some cases, the query processing controller 515 can generate an output based on the QKV processing (e.g., MHA outputs MHA(O)). In some cases, the processing can include one or more reduction functions. In some examples, one or more activation functions can be implemented to obtain an activation output (e.g., vector activation1[N], vector activation2[N], vector activation3[N]). In some examples, activation1[N] is the activation output from the embedding positional encoding processing, and activation2[N] is the output of the QKV system 510 for at least one iteration of the QKV processing based on the input query (e.g., lexical unit index[N]).

[0124] In some examples, QKV processing can identify and weigh the importance of different parts of the input sequence, and the output can indicate how different parts of the input sequence are related to each other (e.g., the correlation between different parts of the input sequence or lexical units). In some cases, the calculation result can be based on a previous iteration of matrix multiplication (e.g., the second iteration of matrix multiplication based on the first iteration of matrix multiplication, the final iteration of matrix multiplication based on the penultimate iteration of matrix multiplication, or the vertex of a previous iteration of matrix multiplication). In some cases, the query processing of the QKV system 510 can include layer normalization. In some cases, the query processing can include feed-forward processing. The query processing can include a decoding function. As shown, the query processing can output an answer to the query. In some cases, one or more steps of the query processing can be repeated (e.g., repeating a specific number of iterations of matrix multiplication (such as 96 times), etc.).

[0125] Figure 6 A flowchart depicting an example method 600 associated with the disclosed system according to the example embodiments described herein. In some configurations, method 600 can be performed by Figure 1 query processing controller 140, Figure 2 query processing controller 230, Figure 3 PNM SSD SoC 305, Figure 4 QKV system 410, and / or Figure 5implemented by the query processing controller 515. In some configurations, method 600 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. Method 600 is merely one implementation, and one or more operations of method 600 may be rearranged, reordered, omitted, and / or otherwise modified such that other implementations are possible and contemplated.

[0126] At 605, method 600 may include receiving attention data from a GPU processing device. For example, query processing controller 140 may receive attention data from a GPU processing device. The attention data may include QKV data (e.g., vector Q1[N], vector K1[N], vector V1[N]) of a query that may be based on attention data from a previous iteration.

[0127] At 610, method 600 may include processing key values from the attention data using transposed query values from the attention data. For example, query processing controller 140 may use transposed query values from the attention data to process key values from the attention data.

[0128] At 615, method 600 may include determining a probability distribution of the result of the processing. For example, query processing controller 140 may determine a probability distribution of the result based on the processing.

[0129] At 620, method 600 may include generating activation values (e.g., a set of activation values, an activation vector) based on the probability distribution, the activation values indicating correlations between text units in the query associated with the attention data. For example, query processing controller 140 may generate activation values based on the probability distribution, where the activation values indicate correlations between text units in the query associated with the attention data.

[0130] Figure 7 A flowchart depicting an example method 700 associated with the disclosed system in accordance with example embodiments described herein. In some configurations, method 700 may be implemented by Figure 1 query processing controller 140, Figure 2 query processing controller 230, Figure 3 PNM SSD SoC 305, Figure 4 QKV system 410, and / or Figure 5 query processing controller 515. In some configurations, method 700 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. Method 700 is merely one implementation, and one or more operations of method 700 may be rearranged, reordered, omitted, and / or otherwise modified such that other implementations are possible and contemplated.

[0131] At 705, method 700 may include receiving attention data from a GPU processing device. For example, query processing controller 140 may receive attention data from a GPU processing device.

[0132] At 710, method 700 may include processing key values from the attention data using transposed query values from the attention data. For example, query processing controller 140 may process key values from the attention data using transposed query values from the attention data.

[0133] At 715, method 700 may include determining a probability distribution of the result of the processing. For example, query processing controller 140 may determine a probability distribution of the result based on the processing.

[0134] At 720, method 700 may include generating activation values (e.g., a set of activation values, an activation vector) based on the probability distribution, where the activation values indicate the correlation between text units in the query associated with the attention data. For example, query processing controller 140 may generate activation values based on the probability distribution, where the activation values indicate the correlation between text units in the query associated with the attention data.

[0135] At 725, method 700 may include receiving at least a second activation value (e.g., a second activation vector) via a die-to-die (D2D) communication interface. For example, query processing controller 140 of a first PNM storage device may receive at least a second activation value from a second PNM storage device via the D2D communication interface.

[0136] At 730, method 700 may include forming a unified activation value based at least in part on a combination of multiple activation values. For example, query processing controller 140 may form a unified activation value based at least in part on a combination of a first activation value of a first PNM storage device and a second activation value of a second PNM storage device, where the D2D communication interface enables the first PNM storage device and the second PNM storage device to communicate independently of the GPU processing device or the host.

[0137] In the examples described herein, the configurations and operations are example configurations and example operations, and may involve various additional configurations and additional operations not explicitly shown. In some examples, one or more aspects of the shown configurations and / or operations may be omitted. In some embodiments, one or more operations may be performed by components other than those shown herein. Additionally or alternatively, the sequence and / or temporal order of the operations may be changed.

[0138] Certain embodiments may be implemented in one or a combination of hardware, firmware, and software. Other embodiments may be implemented as instructions stored on a computer-readable storage device that can be read and executed by at least one processor to perform the operations described herein. The computer-readable storage device may include any non-transitory memory mechanism for storing information in a form readable by a machine (e.g., a computer). For example, the computer-readable storage device may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and other storage devices and media.

[0139] As used herein, the word “exemplary” means “serving as an example, instance, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments. As used herein, the terms “computing device,” “user device,” “communication station,” “station,” “handheld device,” “mobile device,” “wireless device,” and “user equipment (UE)” denote a wireless communication device (such as, a cellular phone, smartphone, tablet computer, netbook, wireless terminal, laptop computer, femtocell, high data rate (HDR) subscriber station, access point, printer, point-of-sale device, access terminal, or other personal communication system (PCS) device. The device may be mobile or stationary.

[0140] As used within this document, the term “communicate” is intended to include sending, or receiving, or both sending and receiving. This may be particularly useful in claims when describing the organization of data sent by one device and received by another device but only requiring the functionality of one of those devices to infringe the claim. Similarly, when only the functionality of one of those devices is claimed, a two-way data exchange between the two devices (where both devices send and receive during the exchange) may be described as “communicating.” As used herein with respect to wireless communication signals, the term “communicate” includes sending wireless communication signals and / or receiving wireless communication signals. For example, a wireless communication unit capable of communicating wireless communication signals may include a wireless transmitter for sending wireless communication signals to at least one other wireless communication unit, and / or a wireless communication receiver for receiving wireless communication signals from at least one other wireless communication unit.

[0141] Some embodiments may be used in conjunction with a variety of devices and systems such as, for example: personal computers (PCs), desktop computers, mobile computers, laptop computers, notebook computers, tablet computers, server computers, handheld computers, handheld devices, personal digital assistant (PDA) devices, handheld PDA devices, on-board devices, off-board devices, hybrid devices, vehicle-mounted devices, non-vehicle-mounted devices, mobile or portable devices, consumer devices, non-mobile or non-portable devices, wireless communication stations, wireless communication devices, wireless access points (APs), wired or wireless routers, wired or wireless modems, video devices, audio devices, audio-video (A / V) devices, wired or wireless networks, wireless local area networks, wireless video area networks (WVANs), local area networks (LANs), wireless LANs (WLANs), personal area networks (PANs), wireless PANs (WPANs), etc.

[0142] Some embodiments may be used in conjunction with one-way and / or two-way radio communication systems, cellular wireless telephone communication systems, mobile telephones, cellular telephones, wireless telephones, personal communication system (PCS) devices, PDA devices incorporating wireless communication devices, mobile or portable global positioning system (GPS) devices, devices incorporating GPS receivers or transceivers or chips, devices incorporating RFID elements or chips, multiple-input multiple-output (MIMO) transceivers or devices, single-input multiple-output (SIMO) transceivers or devices, multiple-input single-output (MISO) transceivers or devices, devices having one or more internal antennas and / or external antennas, digital video broadcast (DVB) devices or systems, multi-standard radio devices or systems, wired or wireless handheld devices (e.g., smart phones), wireless application protocol (WAP) devices, etc.

[0143] Some embodiments may be used in conjunction with one or more wireless communication protocols such as, for example, radio frequency (RF), infrared (IR), frequency division multiplexing (FDM), orthogonal FDM (OFDM), time division multiplexing (TDM), time division multiple access (TDMA), extended TDMA (E-TDMA), general packet radio service (GPRS), extended GPRS, code division multiple access (CDMA), wideband CDMA (WCDMA), CDMA 2000, single-carrier CDMA, multi-carrier CDMA, multi-carrier modulation (MDM), discrete multi-tone (DMT), Bluetooth TM , global positioning system (GPS), Wi-Fi, Wi-Max, ZigBee TM), one or more types of wireless communication signals and / or systems such as Ultra-Wideband (UWB), Global System for Mobile Communications (GSM), 2G, 2.5G, 3G, 3.5G, 4G, Fifth Generation (5G) mobile network, 3GPP, Long Term Evolution (LTE), Advanced LTE, Enhanced Data Rates for GSM Evolution (EDGE), etc. Other embodiments may be used in a variety of other devices, systems, and / or networks.

[0144] Although example processing systems have been described above, embodiments of the subject matter and functional operations described herein may be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware including the structures disclosed in this specification and their structural equivalents, or in a combination of one or more of them.

[0145] Embodiments of the subject matter and operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware including the structures disclosed in this specification and their structural equivalents, or in a combination of one or more of them. Embodiments of the subject matter described herein may be implemented as one or more computer programs (i.e., one or more components of computer program instructions), the one or more computer programs being encoded on a computer storage medium for execution by, or to control the operation of, an information / data processing device. Optionally or additionally, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode the information / data for transmission to a suitable receiver device for execution by the information / data processing device. A computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or a computer storage medium may be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Additionally, although a computer storage medium is not a propagated signal, a computer storage medium may be the source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium may also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or be included in one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0146] The operations described herein may be implemented as operations performed by an information / data processing device on information / data stored on one or more computer-readable storage devices or received from other sources.

[0147] The term "data processing equipment" includes all kinds of equipment, devices and machines for processing data. By way of example, all kinds of equipment, devices and machines include programmable processors, computers, systems on a chip, or multiple programmable processors, multiple computers, multiple systems on a chip, or combinations of the foregoing. The equipment may include dedicated logic circuitry (e.g., FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit)). In addition to the hardware, the equipment may also include code that creates an execution environment for the computer program being discussed, e.g., code that constitutes processor firmware, protocol stack, database management system, operating system, cross-platform runtime environment, virtual machine, or a combination of one or more of them. The equipment and the execution environment may implement various different computing model infrastructures (such as, network services, distributed computing, and grid computing infrastructures).

[0148] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language including compiled or interpreted languages, declarative or procedural languages, and a computer program (also referred to as a program, software, software application, script, or code) can be deployed in any form including as a stand-alone program or as a component, module, subroutine, object, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored as part of a file that holds other programs or information / data (e.g., one or more scripts stored in a markup language document), stored in a single file dedicated to the program being discussed, or stored in multiple coordinated files (e.g., files that store one or more components, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site, or distributed across multiple sites and interconnected by a communication network.

[0149] The processes and logical flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input information / data and generating output. As an example, processors suitable for executing computer programs include both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and information / data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or be operatively coupled to receive information / data from one or more mass storage devices or transfer information / data to one or more mass storage devices or both. However, a computer need not have such devices. Devices suitable for storing computer program instructions and information / data include all forms of non-volatile memory, media, and memory devices, as examples, all forms of non-volatile memory, media, and memory devices include semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0150] To provide for interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information / data to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form including, but not limited to, acoustic, speech, or tactile input. Additionally, a computer can interact with the user by sending documents to and receiving documents from the devices used by the user; for example, by sending a web page in response to a request received from a web browser on a client device of the user.

[0151] Embodiments of the subject matter described herein can be implemented in a computing system that includes backend components (e.g., as an information / data server), or includes middleware components (e.g., an application server), or includes frontend components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with embodiments of the subject matter described herein), or any combination of one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital information / data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0152] The computing system can include a client and a server. The client and the server are typically located remotely from each other and typically interact through a communication network. The relationship between the client and the server arises by virtue of computer programs running on the respective computers and having a client-server relationship with each other. In some embodiments, the server sends information / data (e.g., a HyperText Markup Language (“HTML”) page) to the client device (e.g., for the purpose of displaying the information / data to a user interacting with the client device and receiving user input from the user interacting with the client device). Information / data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.

[0153] Although this specification contains many specific implementation details, these implementation details should not be construed as limitations on any embodiment or the scope of any possible claims, but rather as descriptions of features specific to particular embodiments. The specific features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although the features may be described above as acting in a particular combination and even initially claimed as such, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.

[0154] Similarly, although the operations are depicted in the drawings in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Additionally, the separation of the various system components in the embodiments described above should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0155] Accordingly, specific embodiments of the subject matter have been described herein. Other embodiments are within the scope of the appended claims. In some instances, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures need not be in the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing may be advantageous.

[0156] Benefiting from the foregoing description and the teachings presented in the associated drawings, those skilled in the art to which these embodiments pertain will envision many modifications and other examples of the embodiments described herein. Accordingly, it is to be understood that the embodiments are not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. A method for query processing, the method comprising: receiving data at a first near memory processing storage device; processing, at a first near memory processing storage device, a first value from the data using a transposed second value from the data; determining, at a first near memory processing storage device, a probability distribution of results of processing a first value from the data using a transposed second value from the data; as well as An activation value is generated at the first near memory processing storage device based on the probability distribution, the activation value indicating a relevance between text units in a query associated with the data.

2. The method according to claim 1, wherein: The first near memory processing storage device includes a memory for storing key-value data from the data.

3. The method according to claim 2, wherein: The memory and processor of the first near memory processing storage device include an integrated circuit based on the memory stacked on top of the processor and the memory communicatively connected to the processor.

4. The method according to claim 1, further comprising: receiving at least a second activation value from a second near memory processing storage device via a die-to-die communication interface; as well as A third activation value is formed based at least in part on a combination of an activation value of a first near memory processing storage device and a second activation value of a second near memory processing storage device, wherein the die-to-die communication interface enables the first near memory processing storage device and the second near memory processing storage device to communicate.

5. The method according to claim 1, further comprising: A trigger is received from a processing device based on the advance prompt operation, wherein the trigger includes at least one of a user identifier and layer quantity information associated with the data.

6. The method according to claim 1, wherein: The data is part of the multi-head attention data distributed among multiple near-memory processing storage devices.

7. The method according to claim 1, wherein: The data includes a portion of the query attention data from the first attention layer, a portion of the key attention data from the second attention layer, and a portion of the value attention data from the third attention layer.

8. The method according to claim 1, wherein: The data is based on iterations of activation values ​​generated based on partial outputs generated by a plurality of near memory processing storage devices that are combined to form a unified output.

9. The method according to claim 8, wherein: The plurality of near memory processing storage devices comprises an array of solid state drives.

10. The method according to claim 1, wherein: The first near memory processing storage device is a system-on-chip die including a solid state drive and at least one processor that performs query processing.

11. The method according to claim 1, wherein: The first near memory processing storage device is communicatively connected to the processing device via a high bandwidth expansion bus.

12. The method according to claim 5, wherein: The processing device includes at least one graphics processor communicatively connected to the high bandwidth memory.

13. The method according to any one of claims 2 to 12, wherein: The data includes attention data, the first value includes a key value, the second value includes a query value, and the key-value data includes a key-value matrix.

14. A query processing system, the query processing system comprising: a processing device communicatively connected to a first near memory processing storage device among the plurality of near memory processing storage devices, the processing device sending data to the first near memory processing storage device; as well as The plurality of near memory processing storage devices, a first near memory processing storage device is configured to: processing a first value from the data using a transposed second value from the data; determining a probability distribution of outcomes of the processing; and An activation value is generated based on the probability distribution, the activation value indicating a relevance between text units in a query associated with the data.

15. The query processing system according to claim 14, wherein: The first near memory processing storage device includes a memory for storing key-value data from the data.

16. The query processing system according to claim 15, wherein: The memory and processor of the first near memory processing storage device include an integrated circuit based on the memory stacked on top of the processor and the memory communicatively connected to the processor.

17. The query processing system according to claim 14, wherein: The first near memory processing storage device is configured to: receiving at least a second activation value from a second near memory processing storage device via a die-to-die communication interface; and A unified activation value is formed based at least in part on a combination of an activation value of a first near memory processing storage device and a second activation value of a second near memory processing storage device, wherein a die-to-die communication interface enables the first near memory processing storage device and the second near memory processing storage device to communicate.

18. A non-transitory computer readable medium storing code, the code comprising instructions executable by a processor of a first near memory processing storage device to: processing, at a first near memory processing storage device, a first value from the data using a transposed query value from the data; determining, at the first near memory processing storage device, a probability distribution of results of the processing; as well as An activation value is generated at the first near memory processing storage device based on the probability distribution, the activation value indicating a relevance between text units in a query associated with the data.

19. The non-transitory computer readable medium of claim 18, wherein: The first near memory processing storage device includes a memory storing key value data from the data, and The memory and processor of the first near memory processing storage device include an integrated circuit based on the memory stacked on top of the processor and the memory communicatively connected to the processor.