Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

550 results about "Video memory" patented technology

Method for automatically drawing OpenGL program by using Vulkan

The invention discloses a method for automatically drawing an OpenGL (Open Graphics Library) program by using Vulkan. The method comprises the following steps of: creating a context used by the Vulkan, initializing each module, processing an OpenGL instruction related to texture and data buffering, and managing storage of texture and data buffering resources in a video memory; a shader program used by the OpenGL is preprocessed into a format acceptable to Vulkan, and an OpenGL shader program instruction is created and destroyed; processing an OpenGL (Open Graphics Library) instruction related to frame buffering to generate structural body information required by Vulkan dynamic rendering; an OpenGL instruction of the sampler is also created; processing an OpenGL (Open Graphics Library) instruction for creating a vertex input format and managing a vertex data buffer area, and maintaining vertex input information, a vertex buffer area and an index buffer area required by Vulkan; and finally, drawing or calculating, distributing and calling Vulkan on the basis of all the instructions.
Owner:ZHEJIANG UNIV +1

End-side multi-mode large model accelerated reasoning method and system

The invention provides an end-side multi-modal large model accelerated reasoning method and system, and the method comprises the steps: carrying out the two-stage screening and rearrangement of visual tokens based on the CLS attention and text-to-visual attention in a visual encoder and pre-filling stage, and constructing a sparse attention and sparse key value cache; in a decoding stage, an important neuron set is judged according to activation gating or historical statistics, only a corresponding feedforward network weight is pulled and calculated, missed weights are loaded on demand through asynchronous I / O, and hot neurons are maintained in a high-speed memory to utilize model sparsity, so that video memory / memory occupancy and calculation overhead are remarkably reduced on an end side; throughput and time delay performance are improved. According to the method, the internal memory and computing resources required by reasoning of the multi-modal large language model are reduced from two dimensions by utilizing the endogenous sparsity of the end-side large language model in input and the model, so that a higher reasoning speed is achieved by utilizing fewer resources on the premise of keeping the size of the model unchanged, and the performance of the whole system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Cache management method and device, storage medium and electronic equipment

The invention provides a cache management method, a cache management device, a computer storage medium and electronic equipment, and relates to the technical field of computers. The method comprises the steps of receiving a reasoning task request and distributing the reasoning task request to a target storage page; key value cache information of the first round of reasoning task is stored in a hard disk cache, and when the second round of reasoning task is executed, key value cache information generated before the second round of reasoning task is preloaded layer by layer from the hard disk cache; when the last round of reasoning task is received, storing first target key value cache information correspondingly generated by the last round of reasoning task into the matched target physical block; and performing hybrid grouping compression on key cache information and value cache information in the first target key value cache information to obtain second target key value cache information after quantization compression. According to the invention, triple balance of video memory-calculation performance-precision can be realized.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Industrial personal computer and multi-graphics card collaborative parallel operation acceleration system

The invention discloses an industrial personal computer and multi-graphics card collaborative parallel computation acceleration system, which relates to the technical field of industrial resource allocation and parallel computation, and comprises a resource monitoring and predicting module, a resource management module and a prediction type resource preparation module, the task splitting and collaborative execution module comprises a task splitting module, a cross-node collaborative module and a collaborative operation engine; the intelligent scheduling and dynamic resource allocation module comprises an intelligent scheduler, a dynamic resource allocation module and a conflict avoidance module. According to the method, the GPU video memory utilization rate, the core utilization rate, the temperature, the video memory fragment rate, the available video memory total amount, the CPU core total utilization rate, the load condition and the idle core number index are collected in real time through a resource monitoring module, and the video memory capacity, the GPU core occupancy rate and the CPU load requirement are predicted in advance before a task is submitted in combination with a gradient boosting decision tree and a neural network prediction model; and resources are reserved, so that the scheduling delay is remarkably reduced, and the scheduling hit rate is improved.
Owner:ZHUHAI SHININGDA TECH CO LTD

Large language model progressive field fine tuning and knowledge fusion method oriented to shield engineering

The invention discloses a large language model progressive field fine tuning and knowledge fusion method for shield engineering. The method comprises the following steps: constructing a layered shield training course containing a wide-area academic theory and a proprietary enterprise construction method; parallelly training a plurality of physically isolated parameter efficient adapters based on the frozen base; performing singular value decomposition on the adapter, extracting a geometric feature subspace representing knowledge distribution, and calculating a conflict correlation degree; based on this, a uniform adaptation mechanism of resource awareness is constructed. The mechanism not only can generate a static fusion model for conflict removal, but also can dynamically activate a specific rank slice of an adapter through a routing network based on real-time hardware resource budget (video memory / FLOPs) and geometry-resource signature. According to the method, multi-source knowledge is reserved, and adaptive dynamic scheduling of edge hardware resources by model reasoning is realized.
Owner:CHINA RAILWAY 14TH BUREAU GRP LARGE SHIELD ENG CO LTD +1

Scene rendering method and system based on three-dimensional Gaussian splashing

The invention discloses a scene rendering method and system based on three-dimensional Gaussian splashing. The method comprises the following steps: organizing and constructing an original 3D Gaussian set to obtain a spatial hierarchical structure; generating corresponding level details for the primitives in the spatial hierarchical structure to obtain a spatial hierarchical structure associated with the level details; traversing the spatial hierarchical structure associated with level details, and executing hierarchical view cone cutting and shielding elimination to obtain a visible node list; traversing the visible node list, calculating a level detail selection standard, and generating a level detail activity Gaussian set; performing optimization sorting on the level detail activity Gaussian set to obtain an activity primitive list; and performing tile-based rasterization on the active primitive list, and performing adaptive processing according to level details of the primitives to obtain a final color value of each pixel. According to the method, the rendering performance can be greatly improved, the occupation of a memory and a video memory is remarkably reduced, the rendering quality is improved, visual flaws are reduced, and the expandability of the 3DGS rendering method is enhanced.
Owner:CHINA ORDNANCE SCI INST

Asynchronous parallel reasoning method, system and equipment for hybrid expert model and medium

The invention discloses an asynchronous parallel reasoning method, system and equipment for a hybrid expert model and a medium, which are corresponding schemes: decoupling synchronization of calculation and communication between GPUs (Graphics Processing Unit) caused by all-to-all set communication in expert parallelism, allowing asynchronous parallelism of model calculation and lexical metadata communication, and solving the problem of asynchronous parallelism of the model calculation and lexical metadata communication. Data communication overhead caused by expert parallelization is fully masked, and synchronization waiting overhead is eliminated; aiming at the phenomenon of uneven cold and heat of experts in reasoning, the hot experts are preferentially placed in the GPU, the cold experts are laterally loaded in the CPU so as to release the video memory space of the GPU, and the calculation efficiency of the GPU can be improved by increasing the batch size during reasoning; efficient resource scheduling is realized by dynamically selecting a computing unit which is most suitable for execution and a cold expert which needs to be loaded; generally speaking, the communication overhead and the waiting overhead during parallel reasoning of experts can be remarkably reduced, meanwhile, the calculation efficiency of the GPU is improved, and the overall throughput performance in the reasoning process is optimized.
Owner:UNIV OF SCI & TECH OF CHINA

Method for improving long text processing efficiency and accuracy

The invention discloses a method for improving long text processing efficiency and accuracy, and relates to the technical field of natural language processing and large language models.According to the method, text word segmentation embedding, sliding block preprocessing, YaRN position code injection, dynamic sparse attention calculation, multi-level attention fusion, graded KV cache management and output generation are sequentially executed; position drift is inhibited through logarithmic scaling, and key contexts are adaptively screened according to the attention activeness, so that the attention calculation complexity is close to linearity; in million-level Token reasoning, the video memory occupation of the method is reduced, the remote dependency recall rate is improved, and the method is suitable for scenes such as document analysis, code auditing and multi-mode streaming understanding.
Owner:BEI JING JING YUE KE JI YOU XIAN GONG SI

High-fidelity lightweight world model construction method for end-to-end automatic driving test

The invention relates to a world model construction method, in particular to a high-fidelity lightweight world model construction method for an end-to-end automatic driving test, which is used for constructing a high-fidelity world model and performing knowledge distillation on the world model to solve the problems of huge parameters, low reasoning efficiency and the like of the world model. Model parameters are reduced on the basis of reserving the world model generation capability, and the reasoning efficiency is improved; a CUDA operator is developed in a user-defined mode for the bottleneck part of world model calculation, video memory allocation is optimized, and the reasoning efficiency of a high-fidelity world model is improved based on a single-device multi-thread scheduling and multi-device cooperative calculation method. According to the method, the high-fidelity lightweight world model for the end-to-end automatic driving test can be constructed, the problems that an existing world model is low in multi-modal information alignment precision, poor in cross-view and cross-frame consistency, low in reasoning efficiency and the like are effectively solved, the confidence coefficient of the end-to-end automatic driving system test process is improved, and the test efficiency of the end-to-end automatic driving system is improved. And the testing efficiency of the end-to-end automatic driving system is greatly improved.
Owner:JILIN UNIVERSITY

Large model batch reasoning and data flow optimization system oriented to MOE architecture

The invention relates to the technical field of project management, in particular to a large-model batch reasoning and data flow optimization system oriented to an MOE architecture. According to the method, a collaborative architecture of the request access module, the environment sensing module, the expert routing engine, the resource scheduling module and the dynamic optimization control module is set, the text length and the subject type are extracted by using the request access module, a basis is provided for accurate routing, and the GPU video memory, the I / O bandwidth and the request queue depth are acquired in real time through the environment sensing module, so that the real-time routing is realized. The system load is comprehensively monitored, meanwhile, an expert sub-network is activated through an expert routing engine according to request features, invalid calculation is avoided, weight loading and resource allocation are managed through a resource scheduling module, the I / O bottleneck is reduced, and finally an optimization strategy is intelligently triggered through a dynamic optimization control module based on routing conflict factors. The problems of large reasoning delay fluctuation and unbalanced resource utilization rate mentioned in the background technology are solved, and stable low-delay response and resource collaborative optimization in a high-concurrency scene is realized.
Owner:VIRTAI TECH BEIJING CO LTD

Lightweight display interface rendering optimization system

The invention discloses a lightweight display interface rendering optimization system, which relates to the technical field of image processing, and comprises an acquisition module, an interface analysis module, a behavior analysis module and an attention analysis and rendering module, dynamically calculating an effective view radius by utilizing a visual tunneling effect; the interface analysis module performs character string fuzzy matching in combination with the context input by the user, identifies the search intention of the user and improves the weight of a corresponding component; the attention analysis module performs multi-modal weighted fusion on the physiological fixation data and the sparse content saliency thermodynamic diagram; according to the method, by recognizing the semantic intention and the physiological fixation point of the user, on the premise that core visual experience continuity is guaranteed, GPU load and video memory bandwidth occupation are remarkably reduced, and balance of high-performance display and low-power-consumption operation is achieved.
Owner:SHENZHEN ZHILINTAI ELECTRONIC TECH CO LTD

Large language model semantic query acceleration method based on sparse KV Cache index

The invention discloses a large language model semantic query acceleration method based on a sparse KV Cache index. The method comprises a KV Cache semantic pruning strategy based on an attention mechanism and a set of asynchronous pipeline reasoning architecture with overlapped calculation and I / O. According to the method, the attention sparsity characteristic of the large language model in the reasoning stage and the asynchronous transmission capacity between the Host memory and the GPU video memory are fully utilized, calculation redundancy and video memory occupation in repetitive semantic query are greatly reduced, and high-throughput, low-delay and high-performance batch semantic data processing service is provided. Through a mechanism for mapping a static text into a compressed semantic index and combining a prefix cache technology in a reasoning process to realize state multiplexing, high-performance reasoning acceleration and high-efficiency storage compression are provided for a data-intensive semantic analysis task in a resource-constrained environment.
Owner:EAST CHINA NORMAL UNIV

OpenGL-based large-scale three-dimensional model object rapid pickup method

The invention relates to an OpenGL (Open Graphics Library)-based large-scale three-dimensional model object quick pickup method. The method comprises the following steps: converting an interactive operation of a user into a three-dimensional space structure; a CPU constructs octree space acceleration structures based on a bounding box of an object in a large-scale three-dimensional model, and leaf nodes, intersecting with a pickup ray or a pickup view pyramid, of an axis alignment bounding box in each octree space acceleration structure are screened out to serve as candidate leaf nodes; collecting triangles contained in all candidate leaf nodes, and constructing a candidate triangle set; dynamically allocating thread groups and threads of the GPU based on the number of candidate triangles in the candidate triangle set; and performing highlight display on the candidate triangles intersected with the pickup ray or the pickup view cone, and feeding back IDs of the corresponding candidate triangles. According to the method, the data transmission quantity and the video memory access pressure are remarkably reduced, the pickup response time is greatly shortened in a large-scale three-dimensional model scene, and the requirements of real-time interaction and visual operation can be met.
Owner:SHENZHEN MAIXI SOFTWARE CO LTD

Video memory management method and device, electronic equipment and storage medium

The invention relates to a video memory management method and device, electronic equipment and a storage medium, the method is applied to a GPU drive in a virtual machine, and the method comprises the steps that when it is detected that the remaining amount of a GPU video memory is lower than a set remaining amount threshold value, cold pages in a GPU page table are recognized; wherein the cold page is a page table item which is not accessed in a recent preset time interval; and under the condition that the cold page is a video memory, replacing a physical address corresponding to the virtual address in the cold page to a physical address in a system memory so as to replace video memory data corresponding to the cold page to the system memory. The embodiment of the invention can improve the overall utilization rate of the video memory.
Owner:MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD

Distributed training system and distributed training method

The embodiment of the invention provides a distributed training system and a distributed training method. The distributed training system and the distributed training method are used for improving the training efficiency of model distributed training. Training nodes in the distributed training system are used for storing check points in video memories of GPUs of the training nodes and sending the check points to memories of CPUs of the training nodes; the training node is also used for sending the check points in the memory of the CPU to the storage system through asynchronous operation; the training node is also used for sending abnormal information to the management node under the condition that the node state information of the training node is abnormal information; the management node is also used for determining whether the training node is a fault node based on the abnormal information, and isolating the fault node under the condition that the training node is the fault node; and the management node is also used for loading a target check point corresponding to the fault node from the storage system to other training nodes under the condition that the training node is the fault node, so that other training nodes complete a model training task of the fault node based on the target check point.
Owner:XFUSION DIGITAL TECH CO LTD

Three-dimensional scene data generation method, analysis method and rendering equipment

The invention relates to the field of three-dimensional scene rendering, in particular to a three-dimensional scene data generation method, an analysis method, rendering equipment and a computer storage medium. According to the method, LOD (Level of Detail) grading processing is performed on three-dimensional scene data, foreground data in the three-dimensional scene data is divided into a plurality of LOD layers, each LOD layer is divided into a plurality of grid nodes, and each grid node is independently stored, so that the data can be loaded as required. Meanwhile, different compression modes are adopted for different data, the GPU can directly analyze and read the compressed data, decompression in a memory is not needed, occupation of the memory and video memory bandwidth is greatly reduced, and rendering efficiency is improved.
Owner:SHENZHEN XGRIDS-INNOVATION CO LTD

Video decoding and unreal engine rendering method based on GPU full-link zero copy

The invention relates to the technical field of image processing, and further relates to a video decoding and unreal engine rendering method based on GPU full-link zero copy, which comprises the following steps of: 1, performing hardware decoding on an input video code stream on a GPU, and constructing a motion vector description buffer region for storing compressed domain motion vector information according to a macro block sequence; step 2, reading the motion vector description buffer area on the GPU through a calculation shader, and generating dynamic special effect control buffer areas in one-to-one correspondence with the macro blocks; and step 3, in a post-processing material of the unreal engine rendering module, taking the shared video texture resource as an input texture, and outputting a video picture. On the premise of not depending on a host processor and not excessively occupying video memory bandwidth, unified management of pixel data and motion data is achieved, delay is remarkably reduced, and special effect stability and direction consistency are improved.
Owner:XIAN IMMERSIVE WONDER FILM TECHNOLOGY CO LTD +1

Large language model-oriented adaptive KV cache compression method and system

The invention relates to the technical field of artificial intelligence and big language model reasoning optimization, and discloses a big language model-oriented adaptive KV cache compression method and system, and the method comprises the steps: constructing a lexical element importance measurement mechanism; analyzing an attention head distribution structure in large language model reasoning, and constructing a plurality of pruning strategies; based on a lexical element importance measurement mechanism and the attention head distribution structure, designing a self-adaptive key value cache compression hybrid strategy set based on a pruning strategy; constructing a static self-adaptive key value cache compression method, and automatically distributing a key value cache compression strategy in a pre-filling stage of large language model reasoning; and in a decoding stage, performing adaptive compression on the key value cache based on the allocated key value cache compression strategy. According to the method, the key value cache can be efficiently compressed on the premise of not depending on explicit attention score calculation, a system-level reasoning optimization framework is compatible, the generation performance is kept, meanwhile, the video memory consumption is remarkably reduced, and the long context reasoning capability is enhanced.
Owner:CENT SOUTH UNIV

GPUBox hardware decoupling system based on Retimer card and PCIeSwitch chip

The invention discloses a GPU Box hardware decoupling system based on a Retimer card and a PCIe Switch chip, and belongs to the technical field of computer hardware architecture and high-speed interconnection. According to the system, a Retimer card and a PCIe Switch chip are integrated in an independent GPU Box, and a decoupling link of a CPU server and a GPU acceleration card is constructed; the Retimer card realizes 30-meter long-distance PCIe signal transmission and breaks through physical distance limitation; the PCIe Switch chip pools GPU resources through a dynamic routing and MRIOV technology, supports flexible allocation of computing power by multiple servers, and realizes Peer-to-Peer direct connection communication between GPUs. Aiming at a large model reasoning scene, the system optimizes KV cache bandwidth allocation and video memory and memory cooperative scheduling, so that the 100B parameter model reasoning throughput is greatly improved; and meanwhile, the usability of the system is greatly improved through fault isolation and hot plug design. According to the method, the problems of physical binding of the CPU and the GPU, limited transmission distance, rigid resource allocation and the like in a traditional architecture are solved, and the method is suitable for large-scale AI calculation and distributed GPU cluster deployment.
Owner:HEFEI FENGZHIYI SEMICON CO LTD

KV cache compression and eviction lexical element recovery method and system for large-scale language model reasoning

The invention relates to the technical field of artificial intelligence and natural language processing, in particular to a KV cache compression and eviction lexical element recovery method and system for large language model reasoning, and the method comprises the steps: for a query vector of an ith lexical element calculated by a current Transform layer, calculating attention scores of the query vector and all key vectors; executing a V cache dynamic updating operation based on the attention score; and for the ith lexical element, after the V cache dynamic updating operation is completely completed, performing attention calculation by using the updated value vector set in the V cache storage pool and the pre-calculated attention score partial product P. According to the technical scheme, the technical problem that in the prior art, collaborative optimization of video memory occupancy and calculation efficiency is difficult is solved, and the method has the advantages that KV cache video memory occupancy is dynamically managed, the model reasoning quality is kept, and the calculation efficiency is improved.
Owner:HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Model reasoning cache management method and device based on software and hardware collaboration, equipment and medium

The embodiment of the invention provides a model reasoning cache management method based on software and hardware collaboration, which comprises the following steps: a user side sends a model access request, executes reasoning calculation, and stores a newly generated key value cache in a video memory cache layer; in the video memory cache layer, updating the access popularity corresponding to each unit-level key value cache, and sorting according to the access popularity from high to low to form a first popularity sequence; and updating the hierarchical access popularity corresponding to each hierarchical key value cache, and sorting according to the hierarchical access popularity from high to low to form a second popularity sequence. And when the space utilization rate of the video memory cache layer reaches a set threshold value, determining and evicting a unit-level key value cache and / or a hierarchical key value cache to be evicted based on the first heat sequence and / or the second heat sequence. The cross-layer data migration efficiency is optimized by utilizing the characteristics of a multi-level memory and through mechanisms such as cache elimination prediction and the like, and efficient utilization of a video memory, reduction of reasoning delay and improvement of the throughput capacity of a system are realized.
Owner:RED BRICK INTELLIGENT MODEL (SHANGHAI) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Multi-view collaborative 3D Gaussian splash optimization method and system

The invention discloses a multi-view collaborative 3D Gaussian splash optimization method and system, and the method comprises the steps: constructing a multi-level heterogeneous video memory pool, and dynamically dividing a video memory in 3D Gaussian reconstruction into a view exclusive memory block and a global shared memory pool; a mixed rendering-gradient pipeline is designed, and hardware-level pipeline parallelism in 3D Gaussian reconstruction is realized through a double-buffer asynchronous switching mechanism based on a CUDA Warp-level parallel primitive fusion forward rendering and back propagation thread group; performing multi-view gradient joint optimization, screening an effective gradient path in 3D Gaussian reconstruction based on the visibility mask matrix, and performing projection error weighted fusion on a multi-view gradient tensor; and implementing a multi-modal densification decision, generating a 3D Gaussian candidate splitting position in 3D Gaussian reconstruction through Monte Carlo sampling, calculating a joint optimization objective function by combining a multi-view projection residual error and a gradient contribution factor, and finally realizing 3D Gaussian reconstruction. According to the invention, high-precision and low-delay large-scale scene real-time rendering and training can be realized.
Owner:ZHEJIANG UNIV

Large language model training method based on space-time tensor division strategy

The invention provides a large language model training method based on a time-space tensor division strategy, and the method comprises the steps: employing a time-space collaborative tensor division strategy to divide tensor operation to a plurality of calculation devices in a time dimension and a space dimension, and enabling the calculation devices to carry out the parallel processing, each computing device only caches matrix data required by the current computing step, complete tensors or redundant copies do not need to be reserved, the problem of repeated storage of data such as activation values and weights in traditional tensor parallel is fundamentally solved, video memory occupation is greatly reduced, and the problem of redundant storage of the tensors is solved; point-to-point data transmission replaces set communication, efficient folding of communication and calculation delay is achieved, communication overhead and video memory occupation in the training process are remarkably reduced, the parallel efficiency and resource utilization rate of large language model training are remarkably improved on the premise that training precision is guaranteed, and the method is suitable for large language model training. And a brand new solution is provided for efficient training of a large language model.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Model performance test method and device, electronic equipment and storage medium

The invention discloses a model performance test method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The theoretical maximum lexical throughput of a target large language model is calculated based on the video memory bandwidth of a graphics processor, the model parameter quantity, the byte number corresponding to the quantization precision and the video memory bandwidth utilization rate; meanwhile, the benchmark performance throughput is obtained, a theoretical corresponding first concurrency number is calculated in combination with the theoretical maximum lexical unit throughput and the concurrency competition loss coefficient, then the model test is executed based on the first concurrency number to obtain the actual maximum lexical unit throughput and a corresponding second concurrency number, and a model performance test result is generated. The problems that in the prior art, due to the fact that manual testing is conducted depending on manual intervention, a continuous approaching attempt mode is adopted, a reasonable test starting point is not deduced in combination with hardware core bottlenecks and key parameters, evaluation is time-consuming and labor-consuming, the result is prone to being affected by artificial factors, and accuracy and consistency are poor can be solved.
Owner:JINAN INSPUR DATA TECH CO LTD

Multi-task parallel processing method for user problems under AI platform

The invention provides a multi-task parallel processing method for user problems under an AI platform, and belongs to the technical field of digital data processing of the AI platform. Task resource requirements are accurately calculated through a video memory pre-estimation function and a memory pre-estimation function, a resource consumption mode of concurrent execution is analyzed through a multi-task resource prediction model; the load capacity of the system is evaluated based on a video memory utilization rate gain index, an optimal task segmentation strategy is determined through a data set splitting degree calculation function, multi-stage video memory sub-pools are constructed to realize differentiated resource management, and a Nash equilibrium point of resource allocation is solved by adopting a data set video memory allocation game model. Collaborative optimization allocation of resources is achieved through the video memory and memory coupling allocation equation set, the task state is monitored in real time in the multi-task concurrent execution process, the resource allocation weight is dynamically adjusted, and the technical problem that the system resource utilization rate is low during AI platform multi-task parallel execution is solved.
Owner:青岛网信信息科技有限公司

Key value cache compression method and device, electronic equipment and storage medium

The invention provides a key value cache compression method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining an attention score between an input current lexical element and a target lexical element set, and the target lexical element set comprises all historical lexical elements from a sequence starting position to a current position and the current lexical element; performing aggregation processing on the plurality of attention scores historically accumulated by each target lexical element to generate an aggregation attention value corresponding to each target lexical element; dynamically determining a cache merging threshold according to the sorting result of the aggregation attention values of the target lexical units; recognizing the lexical element fragments with the aggregation attention values continuously lower than the cache merging threshold value and adjacent positions as target cache fragments; and executing fusion compression on the key value cache corresponding to the target cache fragment and updating the original cache. According to the method and the device, occupation of key value cache on video memory resources can be effectively reduced, and meanwhile, the semantic retention capability is remarkably improved.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Multi-tenant visual large model reasoning resource dynamic allocation and isolation method

The invention provides a multi-tenant visual large model reasoning resource dynamic allocation and isolation method. An integrated system of a multi-tenant environment, visual large model reasoning, dynamic resource allocation and isolation guarantee is constructed. The system controls resources such as model copy number, video memory quota, batch size, queue weight, bandwidth and the like at tenant level fine granularity, and adjusts priority and quota based on closed-loop feedback of real-time indexes (such as queuing length, delay and throughput). Interference suppression among tasks is realized through priority grading, a hard / soft isolation strategy and a tenant-model copy mapping mechanism in combination with GPU MIG, video memory partitioning, queue scheduling and other technologies. Aiming at the characteristics of high video memory, large input and the like of a visual large model, model loading, batch processing and video memory multiplexing strategies are optimized, and differential resource allocation of heterogeneous models is supported. According to the overall scheme, on the premise that the service quality and isolation are guaranteed, the GPU resource utilization rate is remarkably increased, and the operation cost is reduced.
Owner:CHINA COAL TECH & ENG GRP CHONGQING RES INST CO LTD

Large model system based on calculation acceleration chip

The invention discloses a large model system based on a calculation acceleration chip, and relates to the field of large models. Comprising a plurality of calculation acceleration units and a management server, and an inter-chip and off-chip data interaction network is formed and realized; a mixed video memory of the calculation acceleration unit is matched with an SSD to form a multi-source storage mode; a normalized network-on-chip, an interconnection transmission system, a storage control system and a plurality of calculation acceleration cores are arranged in the calculation chip; the normalized network-on-chip can read model parameters of a target position based on the storage control system and send the model parameters to the calculation acceleration core for calculation and recovery; and interacting with the management server and other computing acceleration units based on the interconnection transmission system, and reading and storing external model parameters and inter-chip model parameters. The technical problems of insufficient video memory capacity, too high hardware cost and limited transmission bandwidth in a traditional scheme are effectively solved through collaborative design of a mixed video memory architecture and a normalized network-on-chip in combination with a multi-stage routing control and dynamic configuration mechanism.
Owner:STORAGEX TECH INC

Intelligent grading and excess subscription management system and method for GPU video memory

The invention provides an intelligent grading and excess subscription management system and method for a GPU video memory, and relates to the technical field of GPU video memory management, and the method comprises the steps: firstly building a grading storage system comprising a first performance region and a second performance region; then receiving a video memory allocation request containing task priority and service quality requirements, and monitoring a data access mode and access frequency; the initial placement position of the data block is dynamically determined through the intelligent engine according to the task priority, the data access mode and the frequency; predicting target access probability distribution of each data block by adopting an LSTM network; then migrating data between the first performance area and the second performance area through a swap-in and swap-out mechanism based on the distribution and periodicity characteristics, and ensuring the service quality of the first performance area; finally, the global oversale proportion is dynamically adjusted through an overpurchase safety management module, an independent virtual address space is distributed for each task, and safe overpurchase can be achieved through the process.
Owner:HANHOU (BEIJING) TECH CO LTD

Dynamic risk prediction system

The invention relates to the field of constructional engineering, and discloses a dynamic risk prediction system, which generates space-time alignment input through multi-source data fusion, adopts tensor field modeling to embed contract constraint to construct a risk dynamic model, and solves and outputs a continuous risk field through a partial differential equation; a propagation path is analyzed in combination with asymmetric causal analysis, model parameters are adjusted in real time through a dynamic optimization algorithm, and closed-loop optimization of a risk field is achieved; and finally, through four-dimensional thermodynamic diagram interaction early warning and resource intelligent scheduling, a whole-process closed-loop system of data modeling-causal analysis-dynamic optimization-visual management and control is formed. According to the method, dynamic optimization of risk field parameters is realized through adjoint equation back propagation, and the modeling precision of a complex scene is improved; a four-dimensional space-time thermodynamic diagram rendering technology is innovated to solve the problem of fragmentation of multi-modal information expression, and risk disposal response is accelerated; key task resource supply is guaranteed by combining video memory preemption and containerization scheduling strategies, and the system stability bottleneck in a high-load scene is overcome.
Owner:BEIJING NUO SHICHENG INT ENG PROJECT MANAGEMENT CO LTD