Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

75 results about "Cache optimization" patented technology

Large language model long text reasoning acceleration method and device based on speculative key value cache sparse technology, medium, terminal and program product

The invention provides a large language model long text reasoning acceleration method and device based on a speculative key value cache sparse technology, a medium, a terminal and a program product, the method is applied to electronic equipment comprising a GPU and a CPU, and the method comprises the steps that an index list of most important historical tokens is generated based on a distillation language model according to an obtained context sequence; comparing the index list generated at the current moment with the index list at the previous moment, and calculating to obtain a difference set part; asynchronously prefetching the key value cache of the difference set part from the CPU to the GPU; according to the asynchronously prefetched key value cache, performing parallel execution based on a large language model to generate a new token; obtaining a new sequence length according to the generated new token, and judging whether the new sequence length exceeds a preset threshold value or not; and if the threshold value is exceeded, executing unloading operation. According to the method, the performance and the stability of processing long text reasoning by the large language model can be improved, and key value cache optimization in the long context reasoning process is ensured to be always effective.
Owner:SHANGHAI JIAOTONG UNIV

KV cache optimization method and device, computer equipment, readable storage medium and program product

The invention relates to a KV cache optimization method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: calculating a key vector and a value vector corresponding to each element in a text input sequence input into a large language model; through a multi-head potential attention mechanism, performing low-rank joint compression on the key vector and the value vector to obtain a potential vector, and storing the potential vector in a KV cache space; based on a scaling law, determining an optimal compression dimension, regenerating an adaptive potential vector and updating a KV cache space; for the same text input sequence, generating corresponding query vectors, and grouping the query vectors according to a preset grouping rule; calculating a semantic association weight between each group and the correspondingly called potential vector, and taking the semantic association weight as a group attention calculation result; in the reasoning process, potential vectors and grouping attention calculation results are calculated to calculate attention weights. By adopting the method, the storage requirement of the KV cache can be further reduced.
Owner:CHINA TELECOM CLOUD TECH CO LTD

Efficient query acceleration and cache optimization device for data knitting

The invention provides an efficient query acceleration and cache optimization device for data knitting, which comprises a semantic partitioning component used for performing dynamic partitioning on a multi-source heterogeneous data set based on multi-source heterogeneous semantic metadata and embedding cross-domain association identifiers of data knitting to generate a semantic partitioning index table; the query intention mapping component is used for executing intention extraction on a query request initiated by a user to generate a query intention representation, and establishing a mapping relationship of the query intention representation to-be-queried target data partitions according to the semantic partition index table so as to generate a query-partition mapping table; the cache dynamic sorting component is used for generating a cache priority sorting table according to the query-partition mapping table and historical query frequency statistical data; and the scheduling and query execution component is used for optimizing the cache mechanism of the to-be-queried target data partition based on the cache priority sorting table and a preset proportion threshold value. According to the method, the time consumption of cache calling and source data repeated loading in the later query period is reduced.
Owner:BEIJING ZHONGSHURUIZHI TECH CO LTD

Distributed storage cache hotspot prediction method and system based on access mode

The invention relates to a distributed storage cache hotspot prediction method and system based on an access mode. The method comprises the following steps: acquiring an access log to obtain original data; the gradient change of the access frequency between the adjacent data blocks is analyzed in real time to divide the hot spot range and identify continuous hot spot data blocks; constructing a time sequence based on the historical access record, and pre-judging an access hotspot in a future time period by applying a time sequence trend prediction algorithm; constructing a multi-dimensional feature vector and inputting the multi-dimensional feature vector into a machine learning prediction model to output a hotspot probability of the corresponding data object in a future time period; according to the hotspot probability and a popularity threshold value, scheduling the data predicted as the hotspot into a cache, and allocating corresponding cache levels or storage paths for the data objects with different popularity at the same time; and periodically updating the prediction model and the popularity statistics to obtain dynamic prediction and cache optimization of the hotspot data. According to the invention, the prediction of the hotspot data is realized, and the hit rate and the system performance of the distributed cache are improved.
Owner:BANGYAN TECH

Federal edge communication and calculation optimization method, system and device based on task prediction and medium

The invention discloses a federal edge communication and calculation optimization method, system and device based on task prediction and a medium, and belongs to the technical field of edge intelligent collaborative optimization, and the method comprises the steps: collecting historical task data, carrying out the modeling of a communication and calculation process, building a multi-dimensional task feature modeling mechanism, extracting the heterogeneous features of the communication and calculation process, and carrying out the calculation of the communication and calculation process. Performing task demand prediction and load perception to obtain a prediction result; and establishing an integer programming model, performing approximate solution through a heuristic algorithm to complete service cache optimization, performing calculation unloading and resource joint allocation, calculating key performance indexes, performing periodic acquisition, and performing dynamic adjustment on a prediction result. According to the method, an efficient cache strategy is generated by adopting a heuristic algorithm, the defects of single resource allocation, decision lag and lack of global coordination in a traditional method are overcome through joint optimization of resource allocation and a real-time performance monitoring feedback mechanism, and the system response efficiency and the resource utilization rate are improved while the task success rate and reliability are guaranteed.
Owner:GUIZHOU POWER GRID CO LTD

Control system optimization problem compiling method, device and system and computer readable storage medium

The invention relates to the technical field of automatic control, in particular to a control system optimization problem compiling method, device and system and a computer readable storage medium, and the method comprises the steps: obtaining a to-be-compiled optimization problem, and recognizing the structure information and feature information of the to-be-compiled optimization problem; obtaining parameter information of the embedded hardware, and matching and optimizing a corresponding numerical solution algorithm in combination with the structure and the feature information to obtain an optimized numerical solution algorithm; converting the optimized algorithm, structure information and parameter information into a first embedded execution code without dynamic memory allocation; performing memory and cache optimization on the memory layout of the first embedded execution code to generate a second embedded execution code; compiling the second embedded execution code into a binary file adaptive to embedded hardware, and deploying the binary file; the environment parameters are obtained and transmitted into embedded hardware, the target control solution is obtained by operating the binary file and solving based on the environment parameters, and the calculation efficiency and hardware compatibility in an embedded scene are improved.
Owner:YOUDI ROBOT (WUXI) CO LTD

Data compression and cache optimization method and system for edge nodes of electric power internet of things

The invention discloses a data compression and cache optimization method and system for an electric power internet of things edge node. The method comprises the steps of performing normalization adaptation on electric power internet of things edge node data to obtain unified structure data; performing silent report suppression and window deduplication on the unified structure data to obtain redundancy suppression data; inputting the redundancy suppression data into a compression controller for compression to obtain compressed data; and performing layered caching on the compressed data. According to the method, systematic advantages are formed in the aspects of bandwidth, time delay, reliability and operability and maintainability through self-adaptive compressor selection according to point locations, hierarchical caching, congestion self-adaptive transmission, consistency and idempotence and complete decoupling, and the method is particularly adaptive to the characteristics of multi-source, strong real-time, weak network and high-consistency constraint of the electric power Internet of Things.
Owner:BEIJING GUODIANTONG NETWORK TECH CO LTD +1

A blockchain data traceability query optimization method

The application discloses a kind of blockchain data traceability query optimization methods, by introducing the method of cache optimization, utilize cache to reduce the number of disk IO in traceability, improve the efficiency of traceability search, while multi-level cache structure is designed, the problem that cache hit rate is not high under the condition that the memory resource of full node is limited is solved, that is, while giving consideration to the consumption of memory resource in the improvement of cache hit rate, reduce the burden of full node.A kind of blockchain data traceability query optimization method first by the node in network to full node of blockchain initiates traceability query request, then in full node query cache, finally carries out consistency check to full node, realizes blockchain data traceability query.The application improves the traceability query efficiency of blockchain system, query credibility, with certain practicality.
Owner:DALIAN UNIV OF TECH

A read cache optimization method and device in an AI model training scenario

The application provides a read cache optimization method and device in an AI model training scene, comprising: during AI model training, setting each Epoch reading training set data process as two rounds of iterations; in the first round of iteration, each Mini-Batch queries S*N items of data to the cache in random order; S items of data are read from the data in the cache after the query for the current Mini-Batch, and the remaining data is marked; after all the training set data is traversed once, the first round of iteration ends, the second round of iteration starts, and the marked data in the first round of iteration is traversed; each Mini-Batch queries S items of data to the cache, the data is directly read in the cache, or the data is read from a remote storage node. The application improves the read cache hit rate when loading the training set, thereby reducing the time delay of reading data, improving the utilization rate of GPU and other scarce computing power devices, saving computing power, and improving the model training efficiency.
Owner:KYLIN CORP

Spring Session performance optimization method based on partition cache

The invention provides a Spring Session performance optimization method based on partition cache, and relates to the technical field of session management and cache optimization in a distributed system, and the method comprises the steps: storing a mapping relation between a session ID and a final access timestamp by designing concurrent hash mapping caches C1 and C2 which are alternately used, and identifying an active cache by using a current cache label; configuring a timed task to periodically scan the cache, judging a session state according to the critical zone time, and dynamically updating or cleaning cache data; a session object obtained by a thread local variable cache Redis is introduced in Http request processing, so that repeated access is avoided; and a Spring session filter and a warehouse are rewritten to reconstruct session access logic, so that the dependence on Redis is effectively reduced, and the concurrency performance and response efficiency of the system are improved.
Owner:INSPUR GENERSOFT CO LTD

Geometric Baking Cache System In Universal Scene Description Framework

The present application relates to the field of computer graphics animation, and more particular to a system and method for geometric cache optimization. One aspect of the present invention relates to a method for optimizing a geometric data cache. The method comprises importing geometric data into a Universal Scene Description framework, applying an advanced compression algorithm to the geometric data cache, indexing the compressed geometric data, selecting a portion of the compressed geometric data for decompression, and decompressing the compressed geometric data.
Owner:GOOD CREATIVE LLC

Cache optimization method and device based on distributed storage

The invention discloses a cache optimization method and device based on distributed storage, and relates to the technical field of distributed storage systems. The method comprises the following steps: dynamically calculating a weight factor of a cache file, attenuating the weight factor based on the access frequency and timeliness of the file, and compensating and adjusting according to the size of the file; the maximum file threshold value set by the system is dynamically adjusted according to the size distribution of the files in the cache, and the maximum file threshold value is determined based on a preset quantile range so as to eliminate interference of extremely large files on weight calculation; when the weight factor of the cache file is lower than a preset threshold value, the file is removed from the cache, and the cache space is released; when the cache capacity is insufficient, the file with the minimum weight is selected for replacement based on the weight factor of the current cache file, and the high-value file is preferentially reserved in the cache. According to the method, the cache hit rate of small files in the distributed storage system is effectively increased, cache space waste and metadata operation overhead are reduced, and system performance and resource utilization rate are remarkably improved.
Owner:JINAN INSPUR DATA TECH CO LTD

DNS cache optimization system based on multi-node collaboration

The invention discloses a DNS (Domain Name Server) cache optimization system based on multi-node collaboration, which is suitable for public DNS, operator recursion, CDN (Content Delivery Network) scheduling and cloud native scenes. The system is composed of an edge recursion node, a shared second-level cache layer, a cooperative forwarding and near-source rollback module, a distributed active refreshing and lease cooperation module, a cache key generation and routing module and a toughness mechanism module. The domain name and the client network segment are taken as main keys, and the ECS and the two-stage consistent Hash are combined to realize request viscosity and fragmentation stability; when a local miss occurs, preferentially forwarding to a same network segment / same city-same operator / same region collaborative node, and then returning to share cache and authority; a coordination node is elected and refreshed according to domain name fragments, active refreshing of hotspots and immediate entries is driven by lease + priority, and group frightening is avoided through push / pull diffusion; toughness capabilities such as request merging, fusing degradation, health detection and the like are provided, and automatic near-source fallback and removal of abnormities are achieved. According to the method, the hit rate and the first packet delay can be remarkably improved.
Owner:JINAN DIXUN INFORMATION TECH CO LTD

A parallel optimization system for mass video processing

PendingCN122340293AImprove parallel efficiencyImproved parallel throughputRate limitingComputer architecture
This invention discloses a parallel optimization system for massive video processing. The system adopts a four-layer integrated parallel and collaborative processing architecture, comprising, from top to bottom: a parallel task layer for video stream access, grouping, splitting, encapsulation, and queue management; a resource abstraction layer for unified hardware modeling, status acquisition, topology construction, and capability assessment; a parallel scheduling layer for task-hardware matching, load balancing, priority scheduling, and dynamic adjustment; and an execution optimization layer for data stream localization, parallel read / write, cache optimization, and zero-copy processing. The parallel task layer, resource abstraction layer, parallel scheduling layer, and execution optimization layer form a complete parallel processing link: task input, resource awareness, accurate scheduling, and optimized execution. This invention implements concurrent rate limiting and smooth access for the input video stream to prevent traffic surges; and sets synchronization points and timing control for parallel tasks to ensure orderly output.
Owner:北京中科通量科技有限公司

A method, medium, and system for fast loading and caching of online map tiles

This invention provides a method, medium, and system for fast loading and caching optimization of online map tiles, belonging to the field of map processing technology. The invention establishes a tile index library and exchanges a priority matrix, calculates the viewport movement trend vector based on user swiping operations to predict future viewport range, and filters the list of pre-loaded tiles. It checks the cache validity of tiles and establishes a small change detection vector to determine whether updates are needed. A tile hit rate statistical matrix is ​​established, and a heat score is calculated by fusing time decay factors and spatial correlation factors. A thermodynamic entropy reduction algorithm is used to cluster high-heat tiles in fast-access areas to achieve ordered caching. Furthermore, it predicts the user's next tile demand. This invention solves the technical problem of disordered distribution of cached resources during online map tile loading, which leads to excessively long access paths for high-value tiles and reduces the overall loading response speed.
Owner:山东省地图院 +1

Large model reasoning cache optimization method and system based on power transaction system knowledge base

According to the big model reasoning cache optimization method and system based on the power transaction system knowledge base, the big model reasoning efficiency is improved through cooperative work of the initialization stage and the background continuous optimization stage. In the initialization stage, according to document types, a knowledge base document and a time sequence data stream are subjected to fragment division, pre-filling calculation and key value cache storage, and an initial cache pool is constructed. The background continuous optimization stage comprises time sequence data optimization and self-adaptive cache elimination: performing incremental pre-filling on newly-added time sequence data, splicing the newly-added time sequence data with an existing cache, and keeping the cache effective through a sliding window and an overdue elimination strategy; meanwhile, a hotspot function value is calculated according to the access frequency and the recent access time, low-priority fragments are eliminated when the cache exceeds the limit, and cache metadata are periodically refreshed. According to the method, repeated calculation is effectively reduced, the first Token delay is reduced, the response speed is improved in a high-concurrency scene, and the method is suitable for large-scale knowledge base application of electric power, medical treatment and the like.
Owner:CPI INFORMATION TECH CO LTD

Data cache optimization method and device, electronic equipment and storage medium

The invention discloses a data caching optimization method and device, electronic equipment and a storage medium, and belongs to the technical field of data processing. According to the data caching optimization method and device, a model can be constructed based on multi-dimensional user behavior data to predict a target resource subsequently accessed by a user, and a prefetching list matched with an actual demand is generated and cached in combination with a resource dependency relationship and a priority; when the loading condition is met, the cache resource is preferentially called, and when the cache resource is not hit, the prefetching quantity and the lazy loading advance can be dynamically adjusted according to the loading progress, the network state and the like, so that blind prefetching is avoided; the technical problems that due to the fact that an existing prefetching technology lacks user behavior modeling, prefetching resources are not matched with actual requirements, bandwidth is wasted, key resource channels are occupied, and Web application loading efficiency and interaction fluency are affected can be solved. The technical effects of improving the resource loading efficiency, reducing the bandwidth waste, guaranteeing the key resource loading channel, enhancing the Web application interaction fluency and optimizing the user experience are achieved.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Cache communication and deployment scheme generation method and device for air-ground integrated wireless network, computer device and storage medium

Embodiments of the present application relate to the field of unmanned aerial vehicles, and provide a cache communication and deployment scheme generation method and device for air-ground integrated wireless networks, computer equipment and a storage medium. The method comprises: obtaining initial model parameters; constructing an unmanned aerial vehicle communication and deployment optimization model based on the initial model parameters; solving the unmanned aerial vehicle communication and deployment optimization model by using a combination algorithm of a particle swarm optimization and a Hungarian algorithm; constructing an unmanned aerial vehicle cache optimization model based on a function relationship between optimal values of the unmanned aerial vehicle communication and deployment optimization model and decision variables; solving the unmanned aerial vehicle cache optimization model by using a Monte Carlo simulation and an improved particle swarm optimization algorithm; and generating a cache communication and deployment scheme for air-ground integrated wireless networks according to a communication and deployment sub-scheme and a cache sub-scheme. The method can reduce the implementation cost of unmanned aerial vehicle-assisted wireless networks and improve the implementation effect of air-ground integrated wireless networks.
Owner:GUANGDONG YIJING INFORMATION TECH CO LTD

SVC hierarchical edge cache optimization method and device based on unmanned aerial vehicle cooperation

The invention discloses an SVC hierarchical edge cache optimization method and device based on unmanned aerial vehicle cooperation, and belongs to the technical field of wireless communication. A system applied to the method comprises a remote server, a ground base station, unmanned aerial vehicles and ground users. The unmanned aerial vehicle service stage keeps hovering, a ground user initiates a video image quality request to the associated unmanned aerial vehicle, and the associated unmanned aerial vehicle selects a video transmission link mode according to a cache hit condition; the optimization method comprises the following steps: establishing an optimization problem of minimizing the total video transmission delay of a system under the constraint conditions of the cache capacity and the coverage range of the unmanned aerial vehicle; based on the cache capacity and coverage constraint of the unmanned aerial vehicle, the optimization problem is decomposed into sub-problems of user unmanned aerial vehicle association and unmanned aerial vehicle cooperative association, unmanned aerial vehicle cache strategy and unmanned aerial vehicle service hovering position optimization; the optimal solution of each sub-problem is solved and obtained, and the global optimal solution is found through alternate iteration, so that the total video transmission delay of the system is minimized.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-level KV cache optimization method based on compensation token

The invention relates to the technical field of computers, in particular to a multi-level KV cache optimization method based on a compensation token. The method comprises the following specific steps that in the pre-filling stage, a simplified KV cache is constructed, and the simplified KV cache comprises an initial token, a first compensation token set and a nearest token; in the decoding stage, according to whether the simplified KV cache reaches the upper limit value or not and when the nearest window starts to slide out the oldest token, updating operation is activated, a second compensation token set is generated according to the abandoned token, the second compensation token set is inserted into the position between the first compensation token set of the simplified KV cache and the nearest token, and the KV cache is dynamically updated. By adopting the optimization method, the precision loss caused by discarding the long-distance token is effectively recovered by introducing the compensation token mechanism, the memory occupation of the KV cache is reduced, and the storage cost is reduced.
Owner:HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Method, system and application of input data sharing and cache optimization in matrix multiplication calculation

The application discloses a kind of input data sharing and cache optimization method in matrix multiplication calculation, the method includes: step one, input data matrix and weight matrix are divided and distributed according to the number N of computing core on chip;Step two, a plurality of the computing core is sequentially connected to form data transmission ring structure;Step three, each computing core is calculated to the matrix multiplication distributed to it, and the input submatrix of current computing core is passed to the next computing core;Step four, the input submatrix of transmission is operated with the weight submatrix in the next computing core;Step five, the transmission and calculation of the above input submatrix are iterated, and the matrix multiplication operation of input submatrix and weight submatrix is carried out for N rounds, to complete the whole operation process.The application also discloses a system for implementing the above method, which has wide application value.
Owner:SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD

Audio and video stream real-time processing method integrating key frame control, timestamp synchronization and cache optimization

The invention discloses an audio and video stream real-time processing method integrating key frame control, timestamp synchronization and cache optimization. According to the method, multiple types of audio and video data streams are received, and frame types, key frame identifiers, timestamps and media attribute parameters are extracted, so that synchronous processing of the audio and video streams is ensured. Through a key frame guiding mechanism, it is ensured that the first frame is a key frame capable of being independently decoded, and decoding failure is avoided. And a dual-threshold timestamp jump detection and compensation technology is adopted, so that accurate synchronization of audio and video streams is ensured, and the problem of asynchronization caused by network fluctuation is avoided. The audio processing solves the problem of audio frame breakage through intelligent splicing and standardized packaging, and improves the continuity of audio streams. In combination with a memory management strategy of a single-flow dynamic buffer area and a global memory pool, memory utilization is optimized, and fragmentization is reduced. Through priority scheduling and intelligent congestion control, low-delay transmission of real-time audio and video streams is ensured, and the stability of the system in a high-concurrency and network fluctuation environment is improved.
Owner:HANGZHOU ARTECH

Method and apparatus for optimizing cache during large language model inference

PCT designated stageWO2026138159A1Cache optimizationAlgorithm
Provided in the embodiments of the present description is a method for optimizing a cache during large language model inference. The method comprises: in a pre-filling stage, performing cache operations layer by layer for a plurality of attention layers in a large language model. The cache operation for any ith layer comprises: acquiring a target attention matrix of the ith layer; on the basis of the distributions of row data and column data of the target attention matrix, determining a first indicator value and a second indicator value, respectively; on the basis of the first indicator value and the second indicator value, determining an ith preference score corresponding to the ith layer; on the basis of the ith preference score, determining from a total cache area a target cache area allocated to the ith layer, and storing in the target cache area attention data of target characters in input text; and on the basis of the ith preference score, updating prior cache areas of layers preceding the ith layer, and updating attention data of characters stored therein.
Owner:ALIPAY (HANGZHOU) DIGITAL SERVICE TECHNOLOGY CO LTD

A method and system for starting a driving recorder

The present application relates to the technical field of recording device control, and particularly relates to a driving recorder starting control method and system. The method comprises the following steps: based on the driving recorder being powered on, a boot loader starts a real-time operation unit of the driving recorder; the real-time operation unit of the driving recorder loads a main operation unit kernel from a storage medium to a memory unit point by point, the main operation unit kernel responds and starts to run a main control application program; when the main control application program runs, a camera and an encoder are immediately initialized, encoded video data is written into a memory cache area through non-waiting cache response, and a mounting operation of an external storage device is intermittently and asynchronously executed. Through asynchronous mounting storage and memory cache optimization technology, the application realizes starting recording in a short time and data lossless, thereby improving the starting speed and reliability of the driving recorder.
Owner:SHENZHEN MEITONG VIDEO TECH CO LTD

Separated memory architecture-oriented swap space cache optimization system and method

The invention relates to a separate memory architecture-oriented swap space cache optimization system and method, and the method comprises the steps: receiving a page swap request sent by a host memory space, and converting the page swap request into a swap space memory access request; a cardinal number tree cache index is generated according to the logic block number in the exchange space memory access request, searching is conducted according to the index, and if searching succeeds, a data block is obtained from a cache space and sent to a host memory space; if the search is not successful, adopting an asynchronous active cache replacement strategy to select a data block from the original cache region, writing the data block back to the remote memory pool, and meanwhile, adopting a dynamic pre-fetching strategy based on historical access to read pre-fetching data from the remote memory, and storing the pre-fetching data into a pre-fetching cache region. On the premise that an existing switching mechanism of an operating system is not changed, the access efficiency of the switching space of the separated memory is improved, the access frequency of the remote memory is reduced, and therefore the memory access time delay of the whole system is reduced, and the I / O performance of the system is improved.
Owner:XIAN INSTITUE OF SPACE RADIO TECH

A database-based event trigger implementation method, system and electronic device

This application discloses a database-based event trigger implementation method and system. The method includes: embedding a trigger call entry point into the database kernel error handling process; automatically triggering the SERVERERROR handling process when a runtime error event is detected; filtering and judging the error event by type; if it does not belong to the excluded type, retrieving the trigger list from the cache; executing the trigger function in an autonomous transaction; committing or rolling back the transaction based on the execution result; and sending an error report to the client. This invention achieves unified real-time capture and automated response to kernel-level error events, ensures the isolation of trigger execution from the main transaction through an autonomous transaction mechanism, and employs recursive protection and cache optimization design to ensure system stability and execution efficiency. It solves the problems of delayed application-layer error handling response, limited coverage, and difficulty in management in distributed scenarios in existing technologies.
Owner:BEIJING VASTDATA TECH

KV cache optimization method and device

The invention relates to a KV cache optimization method and device. The optimization method comprises the following steps: acquiring metadata associated with a cache block in a KV cache; establishing a popularity model of the cache block based on the metadata of the cache block; based on the popularity model of the cache block, executing an optimization strategy for the cache block in response to the trigger condition; and updating a mapping structure associated with the cache block. According to the optimization method provided by the embodiment of the invention, the invalid occupation of the KV cache can be reduced, so that the cache utilization rate, the reasoning efficiency and the system stability are improved.
Owner:MOFFETT AI TECHNOLOGY SHENZHEN CO LTD

Cache optimization method and device for large-model API gateway

This invention relates to the field of cache optimization technology for large language model API gateways, and particularly to a cache optimization method and apparatus for large language model API gateways. The method includes: receiving an original request sent by a client to a large language model API gateway; identifying and removing dynamic noise fields from the original request to obtain a static request context; performing multi-dimensional semantic standardization processing on the static request context to obtain a standardized intermediate semantic representation; extracting intent, entity, and constraint triples from the intermediate semantic representation and combining them with tenant isolation salts to generate an IEC-structured semantic fingerprint; progressively matching the semantic fingerprint through multi-level caches; and obtaining the cached result if a match is found. This invention solves the problems in existing technologies, such as low hit rates due to the inability to reuse caches, high computational power consumption and latency in vector retrieval caching due to GPU dependence, and the inability to effectively handle dynamic request metadata, which easily leads to cache penetration.
Owner:ZHONGHAO XINYING (HANGZHOU) TECH CO LTD

DNS cache optimization system based on multi-node collaboration

The invention discloses a DNS (Domain Name Server) cache optimization system based on multi-node collaboration, which is suitable for public DNS, operator recursion, CDN (Content Delivery Network) scheduling and cloud native scenes. The system is composed of an edge recursion node, a shared second-level cache layer, a cooperative forwarding and near-source rollback module, a distributed active refreshing and lease cooperation module, a cache key generation and routing module and a toughness mechanism module. The domain name and the client network segment are taken as main keys, and the ECS and the two-stage consistent Hash are combined to realize request viscosity and fragmentation stability; when a local miss occurs, preferentially forwarding to a same network segment / same city-same operator / same region collaborative node, and then returning to share cache and authority; a coordination node is elected and refreshed according to domain name fragments, active refreshing of hotspots and immediate entries is driven by lease + priority, and group frightening is avoided through push / pull diffusion; toughness capabilities such as request merging, fusing degradation, health detection and the like are provided, and automatic near-source fallback and removal of abnormities are achieved. According to the method, the hit rate and the first packet delay can be remarkably improved.
Owner:JINAN DIXUN INFORMATION TECH CO LTD

Graph data processing method based on cache optimization

The invention relates to the technical field of graph data processing, and provides a graph data processing method based on cache optimization, which comprises the following steps: constructing a cache optimization representation of graph data, and distributing a bit vector GT-vector with a fixed length of k for each vertex in a graph, each bit representing whether the vertex belongs to a pre-calculated maximum independent set; reserving adjacency list representation of the graph data, and storing the adjacency list representation as reference data in a memory; for an edge query request, reading GT-vectors of two vertexes from a CPU cache, and executing bit and operation; if the bit and the result are non-zero, judging that the edge is non-edge and immediately returning a first Boolean result indicating that the edge does not exist; and if the bit and result is zero, querying the adjacency list in the memory for verification, and returning a second Boolean result indicating that the edge exists or does not exist. According to the method and the device, the hybrid architecture combining the bit vector representation of cache optimization and the adjacency list is constructed, so that the magnitude order improvement of the graph data processing performance is realized on the premise of ensuring that the query result is completely accurate.
Owner:GUANGZHOU UNIVERSITY