Three-dimensional stacking, storage and calculation integrated mobile artificial intelligence acceleration system and reasoning method
By employing a deep coupling of a three-dimensional stacked storage module and a computing unit array on mobile devices, and pre-storing and utilizing the key-value cache of system prompt words, the memory and computing bottlenecks of mobile devices in the large language model pre-filling stage are solved, achieving low-latency and efficient AI Agent inference.
Patent Information
- Application Number
- CN202511348880.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-11-21
AI Technical Summary
Mobile devices face severe bottlenecks in memory and computing resources when running large language models, especially during the system prompt word pre-filling stage, resulting in high latency and high power consumption, which affects the user experience.
A three-dimensional stacked storage module is deeply coupled with the computing unit array to pre-store the key-value cache of system prompt words and directly access the cache during inference calculations, avoiding repeated pre-filling calculations and combining dynamic user input for calculations.
It significantly reduces the initial response time of the mobile AI Agent from several seconds to sub-seconds, reduces the power consumption of the LLM core computing stage by 65-75%, extends battery life, and improves user experience.
Smart Images

Figure CN120996203A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a three-dimensional stacked in-memory computing integrated mobile artificial intelligence acceleration system and inference method. Background Technology
[0002] With the rapid development of Large Language Model (LLM) technology, deploying AI Agents with multi-step reasoning and tool usage capabilities to mobile devices has become a key direction for achieving inclusive, highly private, and low-latency intelligent experiences. However, mobile devices face a severe "memory wall" technical bottleneck when running LLM, which is composed of their limited memory, power consumption, and heat dissipation capabilities.
[0003] The inference process of LLM is typically divided into two stages: prefill and decoding. In the prefill stage, the model needs to process the complete input sequence, including system prompts, to generate an initial key-value (KV) cache. For mobile AI agent applications, the system prompts are usually predefined templates used to set the agent's role, capabilities, and behavioral norms, and can be thousands of tokens in length.
[0004] This process incurs significant resource overhead on mobile devices. For example, in a 13B parameter model, the key-value cache for each token occupies approximately 800KB of memory. When the system prompt word length reaches 4000 tokens, its key-value cache alone requires about 3.2GB of memory, posing a significant challenge to mobile devices with typical memory capacities of 8-16GB. Even more serious is the computational overhead. The computational load in the pre-filling stage increases quadratically with the input length. A system prompt word of 4000 tokens may require 5-10 seconds of processing time on a mobile device's NPU or GPU, which is unacceptable for interactive agent applications requiring instant response and severely damages the user experience.
[0005] To address the issue of large KV cache size, such as Figure 1 As shown, existing methods typically reduce KV capacity by having multiple Qs share a single KV set. In the diagram, Q1, Q2, ..., QN represent multiple different query heads in the multi-head attention mechanism.
[0006] Three-dimensional (3D) stacked memory technologies (such as Samsung's HBM-PIM) and compute-in-memory (CIM) architectures offer new solutions to data movement bottlenecks. 3D stacking combined with distributed computing is a viable architectural choice for large-model inference applications. 3D stacking effectively increases the bandwidth of both storage and computational cores, while distributed computing and memory access methods match the self-attention and high parallelism characteristics of large models, allowing the full utilization of hardware performance advantages.
[0007] In particular, attention computation is currently performed using a multi-head computation method, where the self-attention mechanisms between heads are independent of each other. Therefore, it can be deployed on multiple computing cores for parallel execution.
[0008] However, large-scale model application solutions based on this architecture are not deeply optimized for the specific workloads of LLM, particularly neglecting the critical characteristic of frequently reused system cue word templates in mobile AI agent scenarios. Existing systems require repetitive and expensive pre-filling calculations on almost identical system cue word templates at the start of each session. For each fixed piece of information, the activation value matrix is very large during large-scale model inference due to the large number of tokens.
[0009] Specifically, the tokens generated from extremely long system prompts need to be copied to each core via the on-chip network for multi-head parallel computation, which significantly increases the data transfer volume of the on-chip network. Meanwhile, the number of computational operations and DRAM data accesses during the Attention phase is proportional to the square of the number of tokens. However, given that the tokens contain a large amount of duplicate information, wasting computational and memory resources on processing duplicate information will severely impact application interactivity.
[0010] Therefore, there is an urgent need in this field for a novel technical solution that can combine advanced hardware architecture with the workload characteristics of AI Agents to fundamentally eliminate the computational bottleneck in the system prompt word pre-filling stage, so as to achieve truly efficient and low-latency AI Agent inference under the strict resource constraints of mobile devices. Summary of the Invention
[0011] To address the above problems, this application provides a three-dimensional stacked in-memory computing integrated mobile artificial intelligence acceleration system, comprising: A three-dimensional stacked memory module comprising multiple vertically stacked DRAM layers communicating via a high-density vertical interconnect structure; An array of computing units, directly communicatively coupled to at least one layer of the three-dimensional stacked storage module via three-dimensional integration technology, and configured to perform at least partial inference computation of a large language model or a multimodal large language model; and A pre-stored key-value cache management module, configured as follows: Within one or more designated physical regions of the three-dimensional stacked storage module, a key-value cache generated by pre-filling calculations of predefined system prompts is pre-stored; wherein, the computing unit array is further configured to: access the pre-stored key-value cache during the execution of the inference calculations and combine it with data generated based on dynamic user input, thereby avoiding repeated pre-filling calculations of the system prompts.
[0012] Optionally, the key-value cache is generated based on a static input prefix, which includes: a system prompt word template, one or more few-shot learning examples, or a retrieved context information.
[0013] The three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system provided in this embodiment deeply couples the key-value cache template storage area with the three-dimensional stacked in-memory computing integrated artificial intelligence acceleration hardware architecture, eliminating the high-latency and high-power prefill computation bottleneck in mobile AI Agent interaction, and reducing the first response time (TTFT) of the mobile AI Agent from several seconds to sub-seconds.
[0014] In addition, the in-memory computing architecture physically integrates computing and storage through 3D integration, avoiding frequent data movement between DRAM and external computing units (GPU / NPU), which can reduce the power consumption of the LLM inference core computing stage by 65-75% and significantly extend the battery life of mobile devices.
[0015] Optionally, the computing unit array is configured to perform the inference computation, the computing unit array realizes data interaction between computing arrays through on-chip network, and realizes distributed memory access through the three-dimensional stacked storage module, the inference computation includes matrix-vector multiplication computation in the attention mechanism of the model and / or matrix multiplication computation in the feedforward network layer.
[0016] By directly executing the most data-intensive operations (attention mechanism and FFN) in the in-memory computing unit, the round-trip data transfer between storage and computing units is minimized, solving the "memory wall" problem and thus significantly reducing the power consumption of the LLM core computing stage.
[0017] Optionally, the pre-stored key-value cache management module further includes: A template index table is configured to record metadata for each pre-stored key-value cache template, the metadata including a template identifier, a physical address in the storage module, and length information; and A dynamic loading controller is configured to look up and load the corresponding key-value cache template from the template index table into the local cache of the computing unit array based on the application's request.
[0018] The AI acceleration system provided in this application offers flexible, efficient, and scalable template management capabilities. Through index tables and dynamic loading controllers, the system can quickly and orderly manage and call multiple different KV cache templates, enabling the acceleration solution to flexibly support various AI Agent applications and ensuring a rapid response to user requests.
[0019] Optionally, it also includes a mobile optimization controller, the mobile optimization controller comprising: A power management unit is configured to switch between multiple operating modes based on the real-time power consumption status of the device to dynamically adjust the operating frequency and the number of active units of the computing unit array; and / or A memory bandwidth allocator is configured to dynamically and preferentially allocate bandwidth of the high-density vertical interconnect structure among access requests to the key-value cache template and other memory operation requests.
[0020] This application optimizes the controller settings on mobile devices, ensuring the practicality and stability of the solution. Power management and bandwidth allocation units enable this high-performance system to intelligently adapt to the strict and dynamically changing power consumption, heat dissipation, and memory bandwidth limitations of mobile devices, guaranteeing a smooth user experience and extended battery life.
[0021] Optionally, a template learning module is also included, which is configured to automatically identify frequently used system prompt words or their variations by analyzing the user's historical usage patterns, and trigger the pre-stored key-value cache management module to generate new key-value cache templates for pre-storage.
[0022] This design enables personalized, adaptive acceleration. By learning users' unique usage habits and automatically generating new templates, the system can evolve from a general-purpose accelerator into a personalized AI engine that continuously improves efficiency, tailored to specific users, further enhancing user engagement and experience.
[0023] To achieve the above-mentioned objectives, this application provides a reasoning method using an artificial intelligence acceleration system, comprising the following steps: Initialization steps: When the system starts up or the application is installed, one or more predefined system prompt word templates are loaded into a three-dimensional stacked storage module, and offline calculations are performed on them to generate corresponding key-value caches. The key-value caches are then stored as templates in a designated area of the three-dimensional stacked storage module. Inference preparation steps: When a user's inference request is received, the corresponding pre-stored key-value cache template is loaded from the specified area of the three-dimensional stacked storage module according to the system prompt word type specified in the request; Core computational steps: Within a digital in-memory computing unit coupled to the DRAM via three-dimensional integration technology, attention mechanism computation or other LLM / MLLM level computations are directly performed using the loaded key-value cache template and data generated solely from the dynamic user input portion of the request; and Incremental update steps: As the session progresses, only the key-value cache generated by the dynamic content newly entered by the user is calculated and appended, while the key-value cache template of the system prompt words remains unchanged.
[0024] The inference method provided in this application defines a complete, end-to-end low-latency interaction process. This method covers the entire process from offline initialization to online incremental updates, ensuring that costly computations are executed only once, while frequent user interactions can continuously enjoy the near-latency KV cache reuse advantage.
[0025] Optionally, in the core computation step: The query vector generated by the dynamic user input portion and the key vector from the pre-stored key-value cache template are subjected to matrix-vector multiplication within the digital storage and computing unit. By reusing the pre-stored key-value cache, the latency and power consumption of pre-filling calculations for the system prompt words are eliminated, thereby avoiding the on-chip network data communication, DRAM memory access data volume and computational load of a large amount of repetitive information during the self-attention calculation process.
[0026] Optionally, it also includes an adaptive optimization step, comprising: Monitor the usage frequency of different key-value cache templates to identify personalized user usage patterns; and During system idle periods, based on the personalized usage pattern, new key-value cache templates are dynamically generated and stored, or the priority of existing templates in the index table is adjusted.
[0027] To achieve the above-mentioned objectives, this application provides an electronic device, including the three-dimensional stacked in-memory computing integrated mobile artificial intelligence acceleration system described above.
[0028] To achieve the above-mentioned objectives, this application provides a non-transitory computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the method described in any of the preceding descriptions. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of a shared key-value cache scheme in the prior art; Figure 2 This is a schematic diagram of the structure of an artificial intelligence acceleration system provided in one embodiment; Figure 3 This is a schematic diagram illustrating the steps of an inference method for an artificial intelligence acceleration system provided in one embodiment. Figure 4 The figure shows an experimental comparison of the first word response time performance of the inference method provided in one embodiment and the inference method of the prior art. Figure 5 This is a schematic diagram illustrating the steps of an inference method for an artificial intelligence acceleration system provided in one embodiment. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0031] like Figure 2 As shown, this embodiment provides a three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system 100, including: a main processing unit 101, which can be any combination of CPU, GPU and NPU, connected to and controlling the operation of a three-dimensional stacked storage module 110; a three-dimensional stacked storage module 110, which includes multiple vertically stacked DRAM layers communicating through a high-density vertical interconnect structure and a pre-stored key-value cache management module 111, wherein the pre-stored key-value cache management module 111 is configured to: pre-store key-value caches generated by pre-filling calculations of predefined system prompt words in one or more designated physical regions of the three-dimensional stacked storage module 110; a computing unit array 120, which is directly communicatively coupled to at least one layer of the three-dimensional stacked storage module through three-dimensional integration technology, and is configured to perform at least part of the inference calculation of a large language model; wherein the computing unit array 120 is further configured to: access the pre-stored key-value cache management module 111 when performing inference calculations, and combine the key-value cache with data generated according to dynamic user input, thereby avoiding repeated pre-filling calculations of system prompt word templates.
[0032] Optionally, the key-value cache is generated based on a static input prefix, which may include: a system prompt word template, one or more few-shot learning examples, or a retrieved contextual information.
[0033] The three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system provided in this embodiment deeply couples the key-value cache template storage area 111 with the three-dimensional stacked in-memory computing integrated artificial intelligence acceleration hardware architecture, eliminating the high-latency and high-power prefill computation bottleneck in mobile AIAgent interaction, and reducing the response time from several seconds to sub-seconds.
[0034] The three-dimensional stacked storage module 110 can adopt a 16-layer stacked structure with a total capacity of 32GB, of which 12GB is allocated as a key-value cache template storage area 111. The computing unit array 120 contains 512 parallel computing units 121, each unit integrating an INT8 MAC unit, a 512KB local SRAM cache 122, and a dedicated Softmax accelerator 123. The hybrid bonding interface uses Cu-Cu direct bonding technology, providing an internal bandwidth of over 4TB / s.
[0035] This embodiment does not limit the implementation form of the computing unit array 120, which can be any in-memory computing array capable of matrix calculation, such as a digital computing unit array, a pulse array, or a PE array.
[0036] Optionally, continue to refer to Figure 2 The computational unit array 120 is configured to perform inference computations, including matrix-vector multiplication computations in the model’s attention mechanism and / or matrix multiplication computations in feedforward network layers.
[0037] In this embodiment, by directly executing the most data-intensive operations (attention mechanism and FFN) in the in-memory computing unit, the round-trip data transfer between the storage and computing units is minimized, solving the "memory wall" problem and thus significantly reducing the power consumption of the LLM core computing stage.
[0038] Optionally, the key-value cache management module 111 includes a key-value cache template storage area 112 for storing system prompt word templates for multiple tokens; a template index table 113 configured to record metadata for each pre-stored key-value cache template, including template identifier, physical address in the storage module, and length information; and a dynamic loading controller 114 configured to look up and load the corresponding key-value cache template from the template index table 113 into the local cache of the computing unit array 120 according to the application's request.
[0039] The AI acceleration system provided in this embodiment offers flexible, efficient, and scalable template management capabilities. Through an index table and a dynamic loading controller, the system can quickly and systematically manage and call multiple different KV cache templates, enabling the acceleration solution to flexibly support various AI Agent applications and ensuring a rapid response to user requests. In addition, the dynamic loading controller enables a predictive loading algorithm based on the context of the user application, improving the cache hit rate.
[0040] Optionally, such as Figure 2 As shown, it also includes a mobile optimization controller 130, which includes: A power management unit 131 is configured to switch between multiple operating modes based on the real-time power consumption status of the device to dynamically adjust the operating frequency and the number of active units of the computing unit array 120; and / or A memory bandwidth allocator 132 is configured to dynamically and preferentially allocate the bandwidth of the three-dimensional stacked storage module 110 among access requests to the key-value cache template storage area 112 and other memory operation requests.
[0041] In this embodiment, the practicality and stability of the solution on mobile devices are ensured by optimizing the settings of the mobile terminal controller 130. The power management and bandwidth allocation unit enables this high-performance system to intelligently adapt to the strict and dynamically changing power consumption, heat dissipation, and memory bandwidth limitations of mobile devices, guaranteeing a smooth user experience and device battery life.
[0042] Optionally, the AI acceleration system also includes a template learning module 140, which is configured to automatically identify frequently used system prompt words or their variations by analyzing the user's historical usage patterns, and trigger the pre-stored key-value cache management module to generate new key-value cache templates for pre-storage.
[0043] This design enables personalized, adaptive acceleration. By learning users' unique usage habits and automatically generating new templates, the system can evolve from a general-purpose accelerator into a personalized AI engine that continuously improves efficiency, tailored to specific users, further enhancing user engagement and experience.
[0044] Optionally, the artificial intelligence acceleration system 100 provided in this embodiment also includes a mobile AI Agent application interface 150.
[0045] This embodiment also provides a reasoning method using an artificial intelligence-accelerated system, such as... Figure 3 As shown, it includes the following steps: The initialization step involves loading one or more predefined system prompt word templates into a three-dimensional stacked storage module during system startup or application installation, performing offline calculations on them to generate corresponding key-value (KV) caches, and storing the KV caches as templates in a designated area of the storage module. The reasoning preparation step involves loading the corresponding pre-stored key-value cache template from the specified area of the storage module based on the system prompt word type specified in the request when a user's reasoning request is received. The core computational steps, within a digital in-memory computing unit directly coupled to the storage module via 3D integration technology, utilize loaded key-value cache templates and data generated from dynamic user input in the request to perform attention mechanism computations or other LLM / MLLM-level computations; and The incremental update step only calculates and appends the key-value cache generated by the dynamic content newly entered by the user as the session progresses, while keeping the key-value cache template of the system prompt words unchanged.
[0046] The inference method provided in this implementation defines a complete, end-to-end low-latency interaction process. This method covers the entire process from offline initialization to online incremental updates, ensuring that costly computations are executed only once, while frequent user interactions can continuously enjoy the near-latency advantage of KV cache reuse.
[0047] Optionally, the method for using an AI-accelerated system for inference described above also includes, in the core computing phase: The query (Q) vector generated from the dynamic user input portion and the key (K) vector from the pre-stored key-value cache template are subjected to matrix-vector multiplication within the digital in-memory unit; By reusing the pre-stored KV cache, the latency and power consumption of pre-filling calculations for system prompt words are eliminated.
[0048] This method eliminates the largest amount of matrix multiplication computation in the attention mechanism by reusing pre-stored key (K) vectors, providing direct support for achieving the key advantage of "eliminating more than 70% of pre-filled computation".
[0049] Optionally, the method for inference using an AI-accelerated system described above also includes an adaptive optimization step, including: monitoring the usage frequency of different KV cache templates to identify personalized user usage patterns; and During system idle periods, new KV cache templates are dynamically generated and stored based on personalized usage patterns, or the priority of existing templates in the index table is adjusted.
[0050] Specifically, dynamically generating KV cache templates can be based on template generation based on usage frequency and context, or it can be based on pre-generated templates based on user profiles.
[0051] This method makes the system no longer static, but capable of dynamically adjusting its internal caching strategy through continuous monitoring and learning, so as to always maintain the best acceleration state for the most frequently used user scenarios.
[0052] The following section provides a detailed explanation of the method implementation process in the application of the code assistant AI Agent: During the initialization phase, when the application is installed, a system prompt word containing general programming language specifications and code style is used. Its key-value cache is calculated offline, quantized with INT4, and then stored as a template in the storage key-value cache management module of the three-dimensional stacked storage module.
[0053] During the inference phase, the user initiates a code completion request. The dynamic loading controller recognizes the "code assistant" scenario and loads the corresponding key-value cache template into the computation unit array. The main processing unit 101 processes only the code snippets of a few dozen tokens input by the user to generate the Q matrix. Attention calculation is completed within the computation unit array; the Q matrix is queried and a pre-stored K matrix is used for dot product operation, and the result is weighted and summed with the V matrix. The entire process avoids repeated pre-filling of system prompts.
[0054] For distributed computing and memory access architectures based on 3D stacking, the pre-filling stage can significantly reduce on-chip network communication during the process of copying input to each computing engine in self-attention computation. Simultaneously, it can reduce DRAM data access and computational load during self-attention computation by more than an order of magnitude. Taking a 13-B parameter LLM model, a mobile SoC, 16GB of LPDDR5 memory, and 100 test runs as an example, the experimental data is as follows: Figure 4 As shown, under different system prompt word lengths, the reasoning method of the embodiment of this application has a TTFT of 0.6-0.8 seconds, while the conventional solution requires 3-12.5 seconds. In addition, from the perspective of power consumption, the power consumption of the conventional solution is about 50W, while the power consumption of this embodiment is only 2W, saving at least 70% energy.
[0055] Optionally, this embodiment provides a reasoning method, such as... Figure 5 As shown, the steps include: Offline pre-computation steps: Perform a one-time pre-filling calculation on one or more system prompts to generate their corresponding key-value (KV) caches, and store the KV caches as templates in a designated area of a three-dimensional stacked storage module; and Online inference steps: Upon receiving a request containing dynamic user input, the KV cache template is loaded from the designated area, and computation is performed only on the dynamic user input to generate its own query, key, and value vector. Subsequently, in a storage-computing integrated unit directly coupled to the storage module, the core inference computation is completed by combining the vector generated by the KV cache template and the dynamic user input.
[0056] The process is detailed below. Phase 1: Offline Pre-computation Phase. The goal of this phase is to execute the static, high-cost computation task only once and permanently solidify its results into a reusable asset. Input and Processing: The input is a pre-defined system prompt template for a specific AI Agent. This template typically contains thousands of tokens used to define the Agent's roles and capabilities. This template is fed into the Transformer model for complete prefill computation. Generation and Storage: The prefill computation generates corresponding key-value vectors for each token in the system prompt. All these KV vectors together form a large "KV cache".
[0057] Quantization and Permalinking: To significantly reduce storage size, the generated KV cache undergoes low-precision quantization techniques such as INT4. Subsequently, this quantized, complete KV cache template is permanently stored in a designated physical region of the 3D stacked DRAM.
[0058] The core advantage lies in its one-time offline computation process, which eliminates the expensive step—repeated in every dialogue in traditional solutions, lasting up to 8.3 seconds, consuming up to 12W of power, and requiring 414 GFLOPs of computation—by preemptively handling and eliminating this step. The result can then be reused an unlimited number of times, achieving "compute once, benefit forever."
[0059] Phase Two: Online Inference Phase (per request) This phase is executed every time a user makes a request. Its core idea is to maximize the reuse of offline computing results and minimize the amount of online computing.
[0060] Receive user input: The system receives the dynamic query from the user's current interaction, such as the 50 tokens in the flowchart.
[0061] Parallel task initiation: At this point, the system initiates two critical tasks in parallel: Task A - Loading Pre-stored KV Cache (Zero Computational Cost): The system rapidly loads pre-stored system cue word KV cache templates (K_sys, V_sys) from 3D DRAM via a high-bandwidth internal interconnect (TSV provides 4TB / s). This process takes only about 0.2 milliseconds, is almost instantaneous, and involves no computational overhead.
[0062] Task B - Processing User Input (Minimizing Computation): Meanwhile, only the user's 50 token inputs are processed to generate their corresponding query vector Q_user, key vector K_user, and value vector V_user. This is the only pre-filling computation required in the online phase, and its computational complexity depends only on the length of the user input (50 tokens), not on the sum of the system prompt length and the user input length.
[0063] Data concatenation and fusion: The newly generated K_user and V_user are concatenated with K_sys and V_sys loaded from DRAM, respectively, to form a complete key vector K_all = [K_sys; K_user] and value vector V_all = [V_sys; V_user]. This is an extremely low-cost memory operation.
[0064] Core attention calculation: Finally, within the in-memory computing unit, the attention calculation Attention(Q_user, K_all, V_all) is performed using the query vector Q_user from the user input and the concatenated complete key / value vectors K_all and V_all.
[0065] Entering the decoding stage: After the calculation is completed, a context vector is generated. The subsequent decoding stage is exactly the same as the traditional method.
[0066] Key advantages: This process reduces the original 8.3-second prefill process to approximately 0.7 seconds, achieving an 11.8x speedup in initial response. The user experience is transformed from a lengthy "wait-for-response" to a smooth "instant conversation." The computational load, memory usage, and peak power consumption of the entire online inference process are all reduced by more than 75%.
[0067] Optionally, this embodiment provides an electronic device that applies the three-dimensional stacked in-memory computing integrated mobile artificial intelligence acceleration system provided in this embodiment.
[0068] Specifically, the system can be applied to consumer electronics products such as smartphones and tablets, as well as to electronic devices with corresponding functional requirements in the industrial field.
[0069] Alternatively, a non-transitory computer-readable storage medium stores a computer program that, when executed by one or more processors, implements the method of reasoning using an artificial intelligence-accelerated system as described above.
[0070] The technical solution of the present invention has now been described in conjunction with the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to the specific embodiments described above. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions resulting from such changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system, characterized in that, include: A three-dimensional stacked memory module comprising multiple vertically stacked DRAM layers communicating via a high-density vertical interconnect structure; The computing unit array is directly communicatively coupled to at least one layer of the three-dimensional stacked storage module via three-dimensional integration technology and is configured to perform at least part of the inference computation of a large language model or a multimodal large language model. as well as The pre-stored key-value cache management module is configured as follows: Within one or more designated physical regions of the three-dimensional stacked storage module, a key-value cache generated by pre-filling calculations of predefined system prompts is pre-stored; wherein, the computing unit array is further configured to: access the pre-stored key-value cache during the execution of the inference calculations and combine it with data generated based on dynamic user input, thereby avoiding repeated pre-filling calculations of the system prompts.
2. The three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system according to claim 1, characterized in that, The key-value cache is generated based on a static input prefix, which includes: a system prompt word template, one or more few-shot learning examples, or a retrieved context information.
3. The three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system according to claim 1, characterized in that, The computing unit array enables data interaction between computing arrays through an on-chip network and distributed memory access through the three-dimensional stacked storage module. The computing unit array is configured to perform the inference computation, including matrix-vector multiplication computation in the attention mechanism of the model and / or matrix multiplication computation in the feedforward network layer.
4. The three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system according to claim 1, characterized in that, The pre-stored key-value cache management module further includes: A template index table is configured to record metadata for each pre-stored key-value cache template, the metadata including a template identifier, a physical address in the storage module, and length information; and A dynamic loading controller is configured to look up and load the corresponding key-value cache template from the template index table into the local cache of the computing unit array based on the application's request.
5. The three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system according to claim 1, characterized in that, It also includes a template learning module, which is configured to automatically identify frequently used system prompt words or their variations by analyzing the user's historical usage patterns, and trigger the pre-stored key-value cache management module to generate new key-value cache templates for pre-storage.
6. A reasoning method using an artificial intelligence acceleration system, characterized in that, Includes the following steps: Initialization steps: When the system starts up or the application is installed, one or more predefined system prompt word templates are loaded into a three-dimensional stacked storage module, and offline calculations are performed on them to generate corresponding key-value caches. The key-value caches are then stored as templates in a designated area of the three-dimensional stacked storage module. Inference preparation steps: When a user's inference request is received, the corresponding pre-stored key-value cache template is loaded from the specified area of the three-dimensional stacked storage module according to the system prompt word type specified in the request; Core computation steps: Within a digital in-memory computing unit coupled to the DRAM via three-dimensional integration technology, attention mechanism computation or other LLM / MLLM level computation is performed using the loaded key-value cache template and data generated from the dynamic user input portion of the request; as well as Incremental update steps: As the session progresses, only the key-value cache generated by the dynamic content newly entered by the user is calculated and appended, while the key-value cache template of the system prompt words remains unchanged.
7. The reasoning method for applying an artificial intelligence acceleration system according to claim 6, characterized in that, In the core calculation steps: The query vector generated by the dynamic user input portion and the key vector from the pre-stored key-value cache template are subjected to matrix-vector multiplication within the digital storage and computing unit. By reusing the pre-stored key-value cache, the latency and power consumption of pre-filling calculations for the system prompt words are eliminated, thereby avoiding the on-chip network data communication, DRAM memory access data volume and computational load of a large amount of repetitive information during the self-attention calculation process.
8. The reasoning method for applying an artificial intelligence acceleration system according to claim 6, characterized in that, It also includes an adaptive optimization phase, which includes: Monitor the usage frequency of different key-value cache templates to identify personalized user usage patterns; and During system idle periods, based on the personalized usage pattern, new key-value cache templates are dynamically generated and stored, or the priority of existing templates in the index table is adjusted.
9. A reasoning method using an artificial intelligence acceleration system, characterized in that, Includes the following steps: Offline pre-computation steps: Perform a one-time pre-filling calculation on one or more system prompt words to generate their corresponding key-value cache, and store the key-value cache as a template in a designated area of a three-dimensional stacked storage module; as well as Online inference steps: Upon receiving a request containing dynamic user input, the key-value cache template is loaded from the designated area, and computation is performed only on the dynamic user input to generate its own query, key, and value vector. Subsequently, in a storage-computing integrated unit directly coupled to the three-dimensional stacked storage module, the core inference computation is completed by combining the vector generated by the key-value cache template and the dynamic user input.
10. An electronic device, characterized in that, The system includes the three-dimensional stacked in-memory computing integrated artificial intelligence acceleration system according to any one of claims 1 to 5.
11. A non-transitory computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by one or more processors, implements the method of any one of claims 6 to 10.
Citation Information
Cited By
Large model reasoning acceleration method and system
CN121920550A
Model reasoning task processing system and model reasoning task processing method
CN122114197A