Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

519 results about "Cache management" patented technology

Cache Management. Magento’s cache management system is an easy way to improve the performance of your site. Whenever a cache needs to be refreshed, a notice appears at the top of the workspace to guide you through the process. Follow the link to Cache Management, and refresh the invalid caches.

AI model automatic deployment platform based on containerization technology

The invention discloses an AI model automatic deployment platform based on a containerization technology, which relates to the technical field of automatic deployment of an artificial intelligence model, and comprises a container construction and intelligent configuration module, a transmission and cache management module, a security and multi-version warehouse module, a resource scheduling and optimization module and a deployment and interface management module, the container construction and intelligent configuration module adopts a three-layer mirror image construction strategy of a base layer, a framework layer and a model layer. According to the invention, all the modules cooperate to form a closed loop, after the container construction module generates an incremental packet and the incremental packet is subjected to security signature verification, the transmission module distributes the incremental packet according to network quality, the resource scheduling module dynamically adjusts bandwidth and quota, and the deployment module starts the container and monitors the container in real time. The problems of efficiency, safety and resource optimization of model deployment in an edge environment are solved, and cross-platform compatibility and service continuity are improved.
Owner:SEEYA TECH CORP

System and Method for Endpoint-Aware Adaptive Protocol Caching with Semantic Deduplication in Heterogeneous Networks

A system for adaptively caching network communication protocols enhances efficiency across heterogeneous device environments through a multi-level cache architecture with device-capability-based tiers. The system collects endpoint telemetry data including device capabilities and operational constraints to classify endpoints and generate context-aware protocol variants optimized for specific device types. Protocol optimization opportunities are determined through structural analysis of message patterns and state transitions. The system performs protocol deduplication by identifying functionally equivalent variants and maintaining canonical representations to reduce cache redundancy. Cache synchronization across distributed nodes uses enhanced Merkle tree structures with protocol normalization processing. The system predicts communication needs based on historical patterns, network context, and endpoint constraints, enabling proactive cache management tailored to device capabilities. Integration with event-driven data communication systems enables seamless protocol selection and translation while maintaining compatibility between diverse endpoint types, from high-performance servers to resource-constrained IoT devices.
Owner:ATOMBEAM TECH INC

Cache techniques for large language model processing

Techniques for cache management for reducing latency in LLM inferencing are described. In some embodiments, a system caches encoded data of portions of a prompt so that the encoded data is available for use by the LLM across dialog turns of a dialog session. Within a dialog session, a portion of the LLM prompt may be the same across dialog turns, and instead of recomputing the attention / encodings for such portions, the cached encodings can be used by the LLM during processing. In some embodiments, user inputs for the dialog session may be routed to the same LLM container and encoded data for the dialog session may be stored at the same cache associated with the LLM. In some embodiments, the system enables asynchronous prompt encoding while performing ASR processing.
Owner:AMAZON TECH INC

Page progressive rendering method and system based on streaming data

The invention relates to the technical field of page rendering, and discloses a progressive page rendering method and system based on streaming data, and the method comprises the steps: carrying out the data segmentation through obtaining a data stream, user interaction data and equipment performance parameters, and obtaining data blocks; then, performing cache management in combination with the data, and constructing a multi-level cache pool; thirdly, performing priority grading and sorting on the data blocks to form a rendering task queue; and according to the equipment performance parameters, optimizing a task sequence and obtaining an optimized task sequence. Next, combining the optimized task sequence and user interaction data, predicting data about to enter a viewport, and generating a viewport pre-rendering task; and finally, according to the viewport pre-rendering task, the multi-level cache pool and the equipment performance parameters, performing rendering strategy optimization to obtain a dynamically adjusted rendering task flow. The method can realize dynamic resource scheduling.
Owner:DEEP BLUE INTERNET (BEIJING) TECHNOLOGY CO LTD

Cache management method and device, storage medium and electronic equipment

The invention provides a cache management method, a cache management device, a computer storage medium and electronic equipment, and relates to the technical field of computers. The method comprises the steps of receiving a reasoning task request and distributing the reasoning task request to a target storage page; key value cache information of the first round of reasoning task is stored in a hard disk cache, and when the second round of reasoning task is executed, key value cache information generated before the second round of reasoning task is preloaded layer by layer from the hard disk cache; when the last round of reasoning task is received, storing first target key value cache information correspondingly generated by the last round of reasoning task into the matched target physical block; and performing hybrid grouping compression on key cache information and value cache information in the first target key value cache information to obtain second target key value cache information after quantization compression. According to the invention, triple balance of video memory-calculation performance-precision can be realized.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Method for improving long text processing efficiency and accuracy

The invention discloses a method for improving long text processing efficiency and accuracy, and relates to the technical field of natural language processing and large language models.According to the method, text word segmentation embedding, sliding block preprocessing, YaRN position code injection, dynamic sparse attention calculation, multi-level attention fusion, graded KV cache management and output generation are sequentially executed; position drift is inhibited through logarithmic scaling, and key contexts are adaptively screened according to the attention activeness, so that the attention calculation complexity is close to linearity; in million-level Token reasoning, the video memory occupation of the method is reduced, the remote dependency recall rate is improved, and the method is suitable for scenes such as document analysis, code auditing and multi-mode streaming understanding.
Owner:BEI JING JING YUE KE JI YOU XIAN GONG SI

Cache techniques for large language model processing

Techniques for cache management for LLM processing are described. Example embodiments include a signal hashing model that generates a key for particular context data. An LLM output corresponding to the context data is stored in a cache along with the key. For a user input received by the system, a cache lookup is performed using a key for context data corresponding to the received user input. For a cache hit, the stored output is used to respond to the user input. For a cache miss, a LLM processes the context data and the user input to generate an output within a first timeout. If the LLM is unable to generate an output within the first timeout, then in some cases, the LLM is allowed to continue processing until a second timeout, and a final or partial output from the LLM is stored in the cache.
Owner:AMAZON TECH INC

Dynamic script compiling and hot updating system based on Grouping and implementation method of dynamic script compiling and hot updating system

PendingCN121029183AVersion controlCode compilationDynamic compilationDowntime
The invention discloses a dynamic script compiling and hot updating system based on Grouping and an implementation method of the dynamic script compiling and hot updating system. The system comprises a script source code management module, a dynamic compiling engine, an intelligent cache manager, an exception handling and recovery system and a hot update controller. And the script source code management module receives and preprocesses the Grouping script to generate a unique identifier for subsequent management. And the dynamic compilation engine executes compilation operation according to the identifier to ensure that a compilation result accords with a specification. And the intelligent cache manager efficiently stores and retrieves the compiling result, so that the repeated compiling cost is reduced. And the exception handling and recovery system monitors the execution state in real time, and automatically starts a recovery process in case of exception to guarantee stable operation of the system. The hot update controller realizes seamless update and version switching of scripts, and ensures service continuity. Through cooperative work of the modules, high-reliability, high-performance and zero-shutdown deployment is realized.
Owner:SICHUAN CHANGHONG JIAHUA INFORMATION PROD CO LTD

Metadata access method and apparatus, device, storage medium, and program product

A metadata access method, apparatus, and computer-readable storage medium for efficient metadata retrieval through cache management. The method receives metadata query requests including target index information from processes and performs matching operations on a global cache file containing records with index information and slot identifiers. Each cache description array corresponds to memory blocks caching metadata. Upon successful matching, the target cache description array is accessed using the target slot identifier. The data state of target metadata is determined from the cache description array, and target address information indicating the location of the target memory block in shared memory is obtained and returned to the requesting process, enabling efficient shared memory-based metadata access.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Key-value cache management, model reasoning, and data processing methods and apparatuses for large language models

Implementations of this specification provide key-value cache management, model reasoning, and data processing methods and apparatuses for large language models. In an implementation, a method comprises allocating a virtual memory block in a virtual address slot to newly-added token key-value data of a model reasoning request, in response to determining that a scheduling result of the model reasoning request indicates the model reasoning request is scheduled for execution, maintaining a mapping relationship between an occupied virtual address slot and a physical graphics memory block allocated to the model reasoning request, and copying the newly-added token key-value data to the physical graphics memory block.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Cache management system, server, method, electronic equipment and storage medium

The invention discloses a cache management system and method, a server, electronic equipment and a storage medium, and relates to the technical field of storage, and the cache management system comprises an acquisition engine used for acquiring a data request stream of an access target; the processing assembly is composed of a load balancer, a plurality of nodes and a processing sub-assembly of each node, the load balancer distributes a data request stream to the plurality of nodes, and the processing sub-assembly is used for providing multi-dimensional features, predicting probability distribution of access targets, determining a pre-fetching target set and determining pre-fetching window parameters at the next moment; and the collaboration component is used for generating a global thermodynamic diagram according to the pre-fetching window parameters of each node at the next moment, finally generating a pre-fetching instruction, pre-fetching the cache from the storage system by using the pre-fetching instruction, and performing multi-dimensional feature fusion through a dynamic window mechanism. The technical problems that a static strategy cannot adapt to a dynamically changing IO mode, resource utilization is low in efficiency and the like in related technologies are solved, and the technical effects of adapting to the dynamic IO mode, improving the resource utilization rate and the like are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Model reasoning cache management method and device based on software and hardware collaboration, equipment and medium

The embodiment of the invention provides a model reasoning cache management method based on software and hardware collaboration, which comprises the following steps: a user side sends a model access request, executes reasoning calculation, and stores a newly generated key value cache in a video memory cache layer; in the video memory cache layer, updating the access popularity corresponding to each unit-level key value cache, and sorting according to the access popularity from high to low to form a first popularity sequence; and updating the hierarchical access popularity corresponding to each hierarchical key value cache, and sorting according to the hierarchical access popularity from high to low to form a second popularity sequence. And when the space utilization rate of the video memory cache layer reaches a set threshold value, determining and evicting a unit-level key value cache and / or a hierarchical key value cache to be evicted based on the first heat sequence and / or the second heat sequence. The cross-layer data migration efficiency is optimized by utilizing the characteristics of a multi-level memory and through mechanisms such as cache elimination prediction and the like, and efficient utilization of a video memory, reduction of reasoning delay and improvement of the throughput capacity of a system are realized.
Owner:RED BRICK INTELLIGENT MODEL (SHANGHAI) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

External plug-in system based on retrieval enhancement

The invention provides an external plug-in system based on retrieval enhancement. The external plug-in system comprises an indexing module, a retrieval module and a generation module, the index module is used for performing cleaning, segmentation, organization, semantic coding and storage tasks of text data; the index module comprises a dynamic data management sub-module, a strategy decision case library management sub-module, a Chinese character recursive segmentation sub-module, a tree index structure construction sub-module and a vectorization storage sub-module; the retrieval module is used for completing semantic matching and sorting on the basis of the structured index; the retrieval module comprises a double-tower rough arrangement sub-module, a cross fine arrangement sub-module, a real-time update sensing sub-module and a cache management sub-module; the generation module is used for carrying out integration and semantic modeling on the high-quality intelligence information provided by the retrieval module; the generation module comprises an information integration and reasoning sub-module and a credibility evaluation and feedback sub-module; the method supports fact alignment, real-time updating, logic interpretability and output evaluability.
Owner:杭州智元研究院有限公司

Method and device for AI automatic testing

The invention relates to the technical field of artificial intelligence and automatic testing, and particularly discloses an AI-based automatic testing method and device. The method comprises the following steps: capturing and recording complete execution historical information in a process of executing a test task by an AI agent, and serializing and storing the complete execution historical information as reusable cache data; before the test is executed, the validity of the cache is judged through a multi-dimensional verification mechanism including data integrity, timeliness, functionality and the like; when the cache verification is passed, a result is quickly obtained based on the cache data playback test process; when cache verification is not passed or execution fails, triggering an AI agent to perform new reasoning and steps to generate and update cache data; and meanwhile, a caching strategy is dynamically adjusted according to test case changes and resource conditions. The device comprises a test executor, an AI agent manager, an intelligent cache manager and a dynamic cache management module, and all the modules work cooperatively to improve the test efficiency, reduce the calling cost of a large language model and enhance the stability of a test result. The method is suitable for Web / APP automatic testing, cross-platform testing, CI / CD process testing and other scenes, the execution efficiency can be remarkably improved, the cost is reduced, and resource utilization is optimized.
Owner:EZUZHIHUI (BEIJING) TECH CO LTD

Data cache management method and system based on big data

The invention discloses a data cache management method and system based on big data, and relates to the technical field of data cache management, and the method comprises the steps: collecting access mode metadata, constructing a data cache state space and an action space, and screening an effective action set of data cache; on the basis of the state space, calculating a context feature weight, and constructing a CO-IAM model to predict a data caching action; performing data caching action implementation based on a CO-IAM model prediction result; by constructing a state space and an action space of data caching, introducing a CO-IAM model, performing precise modeling and dynamic adjustment based on context feature weight, generating an invalid action mask, filtering useless caching actions, optimizing a caching hit rate and access delay, and improving the performance of the whole caching system.
Owner:YICHANG YOUZHI TECH CO LTD

Layered storage read cache management method and system, storage medium and product

The invention provides a hierarchical storage read cache management method and system, a storage medium and a product, and relates to the field of data storage, the method comprises the following steps: dividing all data blocks into three data hierarchies of hot data, warm data and cold data for respective storage according to access popularity, the data storage positions, the storage data proportions and the information recording precision of the three data levels are different; when the cold data is accessed, the accessed target cold data is upgraded to the temperature data, and the temperature data is updated; when the hot data is accessed, the complete metadata information of the accessed target hot data is cached in the memory, and the hot data is sorted and updated again; and when the temperature data is accessed, determining whether to replace the target hot data with the target temperature data or not according to the size relationship between the first historical access time of the accessed target temperature data and the second historical access time of the target hot data with the lowest access heat. By implementing the method, the occupation of memory resources can be reduced when the popularity of mass data is tracked.
Owner:北京志凌海纳科技股份有限公司

Structured data storage method and system based on natural language transformation

The invention discloses a structured data storage method based on natural language transformation, which comprises the following steps of: a system initialization configuration stage: deploying a protocol adapter in a local memory of a PC (Personal Computer) client, and loading natural language processing pipeline configuration parameters; a heterogeneous data acquisition stage: capturing a multi-source text data stream through the protocol adapter, uniformly converting the multi-source text data stream into a standardized data packet, and sending the standardized data packet to a message queue theme; a text cleaning stage: a named entity recognition stage: inputting the pure text data into an NER module deployed with a language model loader; in the conditional feature extraction stage, feature vectors are generated for texts meeting preset conditions on the basis of entity type tags in the entity recognition result; and a consistent storage stage: inserting the entity identification results into a relational database in batches, and updating the entity mapping relationship in the cache. According to the method, the intelligent level of cache management is remarkably improved, and the access fluency of the key data of the user is guaranteed.
Owner:TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD

Marketing system based on intelligent analysis of multi-modal literature data

The invention provides a marketing system based on multi-modal literature data intelligent analysis, and the system comprises the steps: automatically recognizing and analyzing various formats of scientific research literatures through a literature analysis and preprocessing module, and extracting structured information; the AI enhancement analysis module deeply mines literature semantics by using a large model; the data processing and integration module cleans and integrates the information, constructs a researcher portrait including five dimensions of identity, interest, technology, equipment and demand prediction, and further generates a personalized marketing strategy; the cache management module optimizes result storage and retrieval efficiency, and the batch processing module coordinates a multi-file parallel processing flow; according to the system, the automatic conversion from the original literature to the precision marketing strategy is realized. According to the method, deep mining and precise marketing of the multi-source scientific research literature are realized, and the information extraction accuracy and the marketing conversion efficiency are remarkably improved.
Owner:HANGZHOU INST FOR ADVANCED STUDY UCAS

Data migration method, data migration processing device, chip and electronic equipment

The invention provides a data migration method, which is applied to a data migration processing device and comprises the following steps: polling each cache region of a doorbell cache, determining a target accelerator corresponding to an accelerator identifier under the condition that the accelerator identifier is acquired from a target cache region, and acquiring a first target address of first target data from the target cache region; the first target data is to-be-moved data, and the accelerator identifier and the first target address are sent by the host; and sending a first data migration instruction carrying the first target address and the address of the target accelerator to the DMA, so that the DMA migrates the first target data from the host to the target accelerator. The independent data migration processing device is used for controlling DMA to achieve cache space management, participation of a CPU can be reduced, CPU resource consumption caused by cache management is reduced, the cache management process and the CPU are further decoupled, and the flexibility of accelerator cache space management is improved. The invention further provides a data migration processing device, a chip and electronic equipment.
Owner:SANECHIPS TECH CO LTD

Key value cache scheduling method and device, data processing method and device, server and medium

The invention provides a key value cache scheduling method and device, a data processing method and device, a server and a medium. Relates to the technical field of data processing. The key value cache scheduling method is applied to a cache manager in a server for deploying a large model, and is used for uniformly scheduling cache resources when a heterogeneous accelerator executes large model reasoning. The method comprises the steps of obtaining a first request sent by a heterogeneous accelerator based on to-be-reasoned data, and determining the number of required cache blocks; selecting a plurality of target cache blocks from a preset memory cache region according to the cache block occupation state; and aggregating the logic addresses to generate a key value cache address and sending the key value cache address to the heterogeneous accelerator for accessing the key value data. According to the method, through a multi-level cache division and state awareness scheduling mechanism, the cache resource utilization rate and the cross-platform compatibility are effectively improved, the problems that an existing method is low in resource allocation efficiency and poor in compatibility are solved, and the large model reasoning efficiency and the system stability are remarkably improved.
Owner:CHINA UNIONPAY

Key value cache management system and method in model, equipment and medium

The invention discloses a key value cache management system and method in a model, equipment and a medium, and relates to the technical field of computers, which comprises the following steps: identifying a shared prefix part and a non-shared part of an input text, calculating key value cache data of the non-shared part, and storing the key value cache data in the model; inputting the size of a storage space required by the key value cache data of the non-shared part into the affinity estimation model to obtain an affinity estimation value for storing the key value cache data of the non-shared part into each memory node, and determining a memory allocation strategy according to the affinity estimation value, the key value cache data of the non-shared part is stored according to the memory allocation strategy, the problem that the key value cache data occupies a large amount of memory and affects the service performance of model reasoning is solved, and the technical effects of increasing the memory capacity, meeting the memory requirement of model reasoning, improving the service performance of model reasoning and improving load balancing are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Three-dimensional stacking, storage and calculation integrated mobile artificial intelligence acceleration system and reasoning method

The invention relates to a three-dimensional stacking storage and calculation integrated artificial intelligence acceleration system and inference method, and the system comprises a three-dimensional stacking storage module which comprises a plurality of vertically stacked DRAM layers which carry out the communication through a high-density vertical interconnection structure; the calculation unit array is in direct communication coupling with at least one layer of the three-dimensional stacked storage module through a three-dimensional integration technology and is configured to execute at least partial reasoning calculation of a large language model or a multi-mode large language model; and the pre-stored key value cache management module is configured to pre-store key value caches generated by pre-filling and calculating predefined system cue words in one or more specified physical areas of the three-dimensional stack storage module, and the pre-stored key value caches are stored in one or more specified physical areas of the three-dimensional stack storage module. The computing unit array is further configured to access a pre-stored key value cache and combine it with data generated from dynamic user input when the inference computations are performed, thereby avoiding repeated pre-fill computations on system cues.
Owner:HANG ZHOU NANO CORE CHIP ELECTRONIC TECH CO LTD

Cache maintenance system of heterogeneous computing system and electronic equipment

The invention discloses a cache maintenance system of a heterogeneous computing system and electronic equipment, and relates to the technical field of data processing, a first directory controller manages the consistency state of a host end, and a second directory controller manages the consistency state of an equipment end. The consistency maintenance operation is firstly efficiently processed by the corresponding directory controller in the storage domain, cross-device frequent coordination communication is reduced, and when cross-domain data access occurs, the two directory controllers cooperate through an established communication coupling mechanism and execute global consistency maintenance as required. The method not only can adapt to a high-burst and high-parallel access mode of a device end, but also can avoid direct conflicts with a cache management mechanism of a host end, and realizes global consistency through distributed collaboration. The technical problem of performance bottleneck of cache consistency in the heterogeneous system is solved, and the technical effect that the overall data access efficiency of the heterogeneous system is improved while the correctness is maintained is achieved.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Efficiency self-adaptive optimization method and system for data retrieval

The invention discloses an efficiency self-adaptive optimization method and system for data retrieval, relates to the technical field of data management, and discloses an efficiency self-adaptive optimization system for data retrieval. Comprising a mode acquisition module, a task evaluation module, a cost prediction module, a strategy selection module, a consensus retrieval module, an execution decision module, a cache management module and a model updating module. Risk perception support is provided for strategy selection and execution modes; according to the method, resource consumption, execution time delay and path complexity under a plurality of standard retrieval strategies are predicted by using a query cost prediction model, and a confidence score is output, so that prospective evaluation of a strategy execution effect is realized.
Owner:HUNAN INT ECONOMICS UNIV

KV Cache compression and hierarchical management method and device for RAG acceleration and medium

The invention discloses a KV Cache compression and hierarchical management method and device for RAG acceleration and a medium, and belongs to the technical field of large model reasoning acceleration. In order to solve the problems that in an existing RAG KV Cache management technology, storage occupation is too large, and long-term importance and instant accessibility of data cannot be considered at the same time, the invention provides a hierarchical management strategy, and according to the method, a multi-dimensional popularity vector containing global popularity, session popularity and context popularity is defined for a KV Cache block. The method is characterized in that the compression precision of a cache block is independently determined by using global popularity so as to balance fidelity and performance; and meanwhile, the session heat and the context heat are used for independently determining the placement positions in heterogeneous hierarchies such as a GPU (Graphics Processing Unit), a CPU (Central Processing Unit) fixed memory and a CPU paging memory. According to the method, placement is guided through the real-time dimension, compression is guided through the importance dimension, context intelligent prefetching is achieved, and the reasoning performance of an RAG system, especially in a session and in a context pursuit scene is remarkably improved.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD QUZHOU POWER SUPPLY CO

Front-end cache management method, system and equipment for conversation state of lightweight large model and medium

The invention discloses a front-end cache management method, system and device for a lightweight large-model dialogue state and a medium, belongs to the technical field of front-end cache management of a large-model dialogue system, and aims at solving the technical problem of how to overcome the defects that in a traditional scheme, long context cache is low in efficiency, storage redundancy and insufficient in dynamic semantic adaptation capacity, and the large-model dialogue state cannot be managed easily. In order to realize dialogue context volume compression, improve semantic similar request hit rate and reduce cross-end synchronization delay, the adopted technical scheme is as follows: data acquisition and preprocessing: capturing user interaction behaviors in real time through front-end burying points, and performing preprocessing operation on the acquired user behavior data; semantic normalization processing: performing embedded vector conversion and semantic clustering on the text input by the user to generate a unique semantic identifier and a context vector; querying and updating the multi-level cache; and dynamic collaborative updating: dynamically adjusting the cache based on the cache hit rate, the response delay and the user feedback, and optimizing the cache effect in real time.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

System and method for adaptive protocol caching in event-driven data communication networks

A system for adaptively caching network communication protocols provides improved efficiency and performance through a multi-level cache architecture. The system monitors performance metrics and synchronizes cache contents across distributed nodes using a hierarchical structure. Protocol cache requirements are predicted based on usage patterns, enabling proactive cache management. The system compresses cached protocols using existing codebooks to optimize storage and transmission efficiency. Integration with event-driven data communication systems enables seamless protocol selection and translation. The caching system works in conjunction with transaction managers and protocol predictors to enhance network communication performance. The system maintains cache coherency across local, regional, and global cache levels while minimizing network overhead and optimizing protocol availability. By predicting and pre-caching likely needed protocols, the system reduces protocol negotiation latency and improves overall network communication efficiency.
Owner:ATOMBEAM TECH INC

Compilation optimization device and method supporting Page Attention and medium

The invention discloses a compiling optimization device and method supporting Page Attention and a medium, and belongs to the field of compiling optimization devices.The device comprises a multi-core AI accelerator, and each calculation core comprises a tensor component, a vector component, a DMA component and an IMM component; the compiling optimization device is optimized through the following modules: (a) a KV Cache management module, which is used for storing page table items corresponding to different KV Cache blocks according to a batch sequence and generating standardized page table items; (b) a DMA control module which supports a linear reading and remapping mode, carries discrete KV Cache blocks from an external memory to an on-chip weight buffer area in a segmented manner according to an index table, and hides data loading delay by adopting a double-buffer technology; and (c) a multi-batch reasoning scheduling module which dynamically allocates computing resources according to the data scale in a code stage. According to the invention, through deep cooperation of a hardware architecture and functional logic, the problems of low KV Cache management efficiency, complex multi-batch dynamic scheduling and insufficient computing resource utilization rate on an ASIC platform are solved.
Owner:BEIJING YIXIN YIYU MICROELECTRONICS TECH CO LTD