Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

229 results about "Lossless compression" patented technology

Lossless compression is a class of data compression algorithms that allows the original data to be perfectly reconstructed from the compressed data. By contrast, lossy compression permits reconstruction only of an approximation of the original data, though usually with improved compression rates (and therefore reduced media sizes).

Image recognition method based on edge calculation

The invention relates to the technical field of computer vision and image recognition, in particular to an image recognition method based on edge computing, which comprises the following steps: dynamically capturing an original image through a plurality of edge nodes, rejecting redundant regions through a multi-modal perception triggering mechanism, and establishing a cooperative processing group. Illumination equalization, noise filtering and resolution self-adaptive compression tasks are distributed according to dynamic role election, a standardized preprocessed image is generated, a lightweight convolutional neural network is operated in parallel to extract a dual-channel feature vector, and after entropy coding lossless compression and equipment identity tag and time sequence stamp attachment, the dual-channel feature vector is transmitted to a cloud end by adopting a lightweight encryption protocol. The cloud end analyzes the data packet, reconstructs a feature topological graph based on space-time relevance, loads a depth residual error recognition model to execute feature fusion and classification decision, feeds back and updates the weight of an edge node model, solves the problems of low collaborative efficiency and feature distortion, and improves the efficiency and precision of image recognition.
Owner:TUSU AUTOMATION TECH (SHANGHAI) CO LTD

Time series data lossless compression method and system based on dynamic context awareness

The invention discloses a time series data lossless compression method and system based on dynamic context awareness, and relates to the field of data lossless compression, and the method comprises the steps: extracting target data from a device port, and carrying out the preprocessing of the target data through a data processing technology according to the business dimension corresponding to the target data, constructing a structured data set based on a preprocessing result; capturing local time sequence features based on the structured data set, distributing attention weights in combination with a hidden state, generating a feature vector used for sensing a context, and outputting a dimension vector by using a three-layer full-connection network; and according to the dimension vector, analyzing the compression performance score value of each compression algorithm in the resource library, sorting the compression performance score values, and selecting the compression algorithm meeting the target requirement to compress the target data. According to the method, the global periodic features and the local mutation features of the time series data are analyzed through the space-time attention mechanism, and real-time driving algorithm switching is achieved.
Owner:GUODIAN NANJING AUTOMATION

Fine-grained flow lossless compression method combined with multiple threads

The invention relates to the field of traffic storage optimization and traffic data compression, in particular to a multithreading-combined lossless compression method for fine-grained traffic, which comprises the following steps of: acquiring a traffic data file, extracting a header triple based on the traffic data file, generating an identifier according to the header triple, and transmitting the identifier to a server; aggregating the traffic data files with the same identifier to obtain data streams, and generating a sorted data stream size table according to the data volume of the data streams; creating threads according to system computing power, distributing the data flow to the thread with the minimum thread load according to the data flow size table and the thread load condition, and outputting a data flow distribution scheme; and performing fine-grained characterization and serialization on each data stream to obtain an integer sequence and a byte sequence, selecting a compression processing method according to redundancy characteristics of the integer sequence and the byte sequence, obtaining a redundancy-eliminated sequence, and writing the redundancy-eliminated sequence into a compressed file. The invention aims to provide high compression efficiency to reduce the storage cost while ensuring the data precision.
Owner:NORTHEASTERN UNIV CHINA

Parallel adaptive matrix lossless compression method

The invention discloses a parallelizable self-adaptive matrix lossless compression method, which comprises the following steps of: partitioning a first matrix with a larger scale into second matrixes with smaller scales, distributing a splitting process for the second matrixes, splitting all the second matrixes into a plurality of bit matrixes in parallel by the splitting process, and distributing compression threads for the bit matrixes. And adaptively selecting an optimal compression strategy from the compression strategy set by the compression thread, and finally, completing compression of all bit matrixes by the compression thread according to the determined compression strategy in parallel to form a compression matrix and a flag array as final compression data of the first matrix. Through compression process grading and step-by-step parallelization, the memory overhead in the compression process is effectively reduced, the compression speed is improved, an optimal compression strategy is selected through trial compression, self-adaptive compression is achieved, the data compression rate is effectively improved, and lossless compression of data is achieved by constructing a flag array.
Owner:北京麟卓信息科技有限公司

Power grid waveform data lossless compression method, decompression method, equipment, medium and program product

The invention provides a lossless compression method and decompression method for power grid waveform data, equipment, a medium and a program product, and the method comprises the steps: carrying out the periodic feature analysis of the original sampling data of a power grid waveform, so as to determine the periodic structure of the original sampling data, and carrying out the segmentation processing of the original sampling data based on the periodic structure; for the sampling data in each segment, extracting differential information between adjacent sampling points to form a differential data sequence; dynamically determining a corresponding quantization parameter based on the statistical distribution characteristic of the differential data sequence, and performing adaptive quantization processing on the differential data sequence to generate quantized differential data; and entropy coding processing is carried out on the quantized differential data to generate compressed data used for representing the original sampling data, and characteristic parameters used for reconstructing the original sampling data are stored in the compressed data in an associated mode. Unification of high compression ratio, low power consumption and strict lossless reconstruction is realized, and the application requirements of long-term monitoring and remote transmission of the edge side can be met.
Owner:SHANGHAI HOLYSTAR INFORMATION TECH

Large model KV precision lossless compression method and system

The invention provides a large model KV precision lossless compression method and system, and relates to the field of data processing, and the method comprises the steps: obtaining key value cache data generated in a large language model reasoning process; performing high-dimensional feature extraction and similarity analysis on the key value cache data to generate a corresponding hash signature; based on the Hash signature, performing physical reordering on the key value cache data, so that the key value cache data with similar characteristics are adjacently arranged in a physical storage space; transmitting the reordered key value cache data to an intelligent network card through a high-speed interconnection interface; and calling an integrated hardware lossless compression module through the intelligent network card, and performing streaming compression on the reordered key value cache data. According to the technical scheme, the hardware unloading and assembly line technology is utilized, transmission bandwidth occupation and system overall delay are remarkably reduced, and key technical support is provided for efficient reasoning of a large model in a resource-limited edge environment.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Lossless compression storage method and system for heterogeneous data based on credential environment

The invention discloses a lossless compression storage method and system based on heterogeneous data in a credential environment, and the method comprises the steps: firstly, extracting a mapping relation between a file type and a structural feature through scanning a target file, carrying out the preliminary lossless compression after classification and grouping, and further employing a sliding window matching and sequence matching algorithm to optimize compressed data for a redundancy mode, and then generating a metadata structure containing a check code and a compression parameter, packaging the metadata structure into a data packet supporting cross-platform transmission, finally transmitting the data packet to a target storage device through a secure write-in protocol, and verifying data consistency and recovering path feasibility by using the check code. According to the method, the data compression efficiency is remarkably improved, the compatibility of cross-platform transmission and the integrity of stored data are ensured, and efficient and reliable technical support is provided for complex electronic file processing.
Owner:HUNAN YUNDANG INFORMATION TECH CO LTD

Electroencephalogram data transmission method combining lossless and lossy compression

The invention provides an electroencephalogram data transmission method combining lossless and lossy compression, belongs to the field of electroencephalogram data transmission and compression processing, and is used for solving the problems of low transmission efficiency of pure lossless compression and easy loss of key diagnosis information of pure lossy compression in related technologies. According to the method, multi-dimensional features of electroencephalogram data are extracted, a key area and a non-key area are divided, lossless compression and lossy compression processing are adopted respectively, the area division proportion is dynamically adjusted in combination with transmission bandwidth, signal features and application scenes, and corresponding data are transmitted through dual protocols. The transmission efficiency can be improved while the clinical diagnosis marker and the core characteristics are reserved, and the scene adaptability and the transmission reliability are achieved at the same time.
Owner:CLP CLOUD BRAIN (TIANJIN) TECH CO LTD

Compression and decompression method for DNA sequencing data

The invention discloses a compression and decompression method for DNA sequencing data, and belongs to the technical field of biological information. The technical problems that in the prior art, when third-generation sequencing data are compressed, flexibility is poor, and the compression rate is low are solved. According to the compression method for the DNA sequencing data, global compression or block compression can be dynamically selected before compression, different compression methods are adopted for different types of data, and the flexibility of the compression method is improved; when the base sequence is compressed, the adopted self-indexing structure is a lossless compression method, the initial positions of all the sequences can be compressed and restored by using smaller data volume, and the compression rate is improved. Due to the fact that a special design is adopted in a run length coding structure, a character part is removed, an obtained run length segment is shorter, lossless compression of the position and the sequence is guaranteed, and meanwhile a better compression effect is provided. The method is mainly used for compression and decompression of DNA sequencing data.
Owner:HARBIN INST OF TECH

Video lossless compression transmission and storage method and device

The invention relates to a video lossless compression transmission and storage method and device. The method comprises the following steps: analyzing an obtained video frame sequence; executing quality pre-verification, and if yes, determining that the video frame sequence is a first frame sequence; based on the scene detection target, determining a target attention area in the plurality of first sequence frames; selecting a plurality of target feature frames from the plurality of first sequence frames to execute importance evaluation, and determining corresponding feature levels; mapping the plurality of target feature frames into a one-dimensional pixel sequence; dividing the one-dimensional pixel sequence into a plurality of pixel segments, and performing cross-frame aggregation on the plurality of pixel segments according to the feature level to obtain a plurality of feature sequences; and carrying out hierarchical compression coding on the plurality of feature sequences. By adopting the method, the video can be subjected to quality pre-verification, the target concerned area can be adaptively identified, and differential compression and priority transmission are performed, so that the transmission quality and the storage efficiency of the inspection video can be improved in a complex electromagnetic environment and a limited network condition.
Owner:ZHONGJIN PEIKE CONSTR CO LTD

Lossless compression storage method and system based on heterogeneous data under xinxin environment

The application discloses a lossless compression storage method and system based on heterogeneous data in a Xinhua environment, first extracts the file type and structure feature mapping relationship by scanning the target file, classifies and groups, and then carries out preliminary lossless compression, further adopts a sliding window matching and sequence matching algorithm to optimize the compressed data for the redundant mode, then generates a metadata structure containing a check code and compression parameters, encapsulates into a data packet supporting cross-platform transmission, finally transmits to the target storage device through a secure write protocol, and verifies the data consistency and recovery path feasibility by using the check code. The application significantly improves the data compression efficiency, ensures the compatibility of cross-platform transmission and the integrity of storage data, and provides efficient and reliable technical support for complex electronic file processing.
Owner:HUNAN YUNDANG INFORMATION TECH CO LTD

Low earth orbit satellite communication time delay optimization method and system

The invention relates to the technical field of spacecraft satellite-ground communication, and provides a low-earth-orbit satellite communication time delay optimization method and system aiming at the defects that a traditional low-earth-orbit satellite communication satellite system and a satellite communication frame have high time delay and poor real-time communication function. The core of the invention comprises three parts: firstly, an efficient receiving and transmitting end coding and decoding technology: a transmitting end adopts a subblock coding technology of a Turbo code; and a receiving end uses an MAX-Log-MAP algorithm to carry out rapid iterative decoding on the word block code of the Turbo code. And secondly, compressing the user data: carrying out lossless compression on the user data and a satellite communication frame header by using Huffman coding, reducing the transmission data volume, improving the data processing speed, and thus reducing the processing time delay. And finally, simplifying the frame header design: optimizing the frame header of the communication frame so as to reduce the total length of each frame. The method is particularly suitable for communication satellites (the orbit height is less than or equal to 1500km) in a low earth orbit.
Owner:HARBIN GONGDA SATELLITE TECH CO LTD

Deep learning data compression using multiple hardware accelerator architectures

Deep learning data compression using multiple hardware accelerator architectures is provided herein. A system includes a computing device and first and second hardware accelerators coupled thereto. The first and second hardware accelerators may be of different types, such as a tensor streaming processor and a field programmable gate array. The first and second hardware accelerators may be directly connected to one another, such as by a chip-to-chip connection. The first and second accelerators may implement different stages of a data pipeline, such as lossless and lossy compression stages of a learned image compression.
Owner:GROQ UK LTD

Lossless compression and decoding method of depth image

This invention provides a lossless compression and decoding method for depth images. Lossless compression includes: dividing sequential data into multiple independent image blocks according to predetermined partitioning rules; selecting a corresponding thread model based on the characteristics of the depth image; calling a thread based on the thread model and using a reversible encoding algorithm to independently encode each image block; wherein, at the start of encoding for each image block, the encoder state is reset so that the encoding of each image block does not depend on the data of the preceding image block; generating metadata containing position index and compressed data length information and assembling it with the encoded image block for output. Lossless decoding includes: reading the metadata and decoding at least one image block based on the metadata. Through the above methods, parallel encoding and decoding are achieved; precise positioning via physical offsets and random access and on-demand region decoding are supported, improving the processing response of high frame rate depth streams under low computing power environments while ensuring lossless restoration.
Owner:WISDOM CORNERSTONE (SHANGHAI) TECHNOLOGY CO LTD

Binary Image Compression Encoding and Decoding Method Based on Base62 Encoding

This invention relates to a binary image compression and decoding method based on Base62 encoding. The method involves converting the original binary image to Base32 to generate a Base32 two-dimensional array; setting the number of compressed bytes; compressing the Base32 array in blocks according to the number of compressed bytes to create an index matrix; performing a first compression on the Base32 two-dimensional array based on the index matrix; and performing a second compression on consecutive empty data segments appearing in the first compression result, based on the characteristics of numerous white blocks and similar adjacent rows in printed labels, using run-length encoding and character substitution algorithms. The resulting binary image compression is lossless, achieving a high compression ratio. Furthermore, being based on Base62 encoding, it is suitable for network transmission. The Base62 encoding table is fixed, simplifying decoding and meeting the needs of narrowband transmission in the Internet of Things (IoT) and decompression requirements of low-performance terminals.
Owner:SHANGHAI SEARI INTELLIGENT SYST CO LTD

Lossless compression with probabilistic circuits

Devices, methods, and systems for lossless compression utilizing probabilistic circuits (PCs) are provided. In one embodiments, a method for lossless compression using PCs is provided, the method comprising: receiving image data comprising a plurality of pixels, wherein each of the plurality of pixels is represented by a variable; sequentially compressing the variables one-by-one using conditional probabilities, wherein the conditional probabilities are computed by: calculating at least one marginal; initially setting to 1 a probability p(n) for every PC unit n; defining an evali for a set of PC units n that need to be evaluated in an ith iteration; and for i=1 to a dataset D, evaluating PC units n in evali using a bottom-up process and computing a target probability; and generating a bitstream using a streaming code for the compressed variables using the conditional probabilities.
Owner:RGT UNIV OF CALIFORNIA

Encoding a binary image

In various examples there is a method for encoding a binary image, the method comprising receiving the binary image; performing run-length encoding on pixel data of the binary image to produce run-length encoded data; performing differential encoding on the run-length encoded data to produce differential encoded data; performing variable length encoding on the differential encoded data to produce variable length encoded data; and applying a lossless compressor to the variable length encoded data to produce compressed data, the compressed data being an encoded binary image.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Three-dimensional medical image lossless compression method and system based on transformer decoder

The application discloses a three-dimensional medical image lossless compression method and system based on a Transformer decoder, which comprises the following steps: pre-processing a three-dimensional medical image to obtain a voxel sequence, segmenting the voxel sequence to obtain a multi-voxel sequence block, and inputting the multi-voxel sequence block into a double-layer decoder model; a first layer decoder performs rough prediction on the multi-voxel block sequence; a second layer decoder takes the rough prediction result as prior features to predict each voxel and obtain a prediction probability distribution of the next voxel; the prediction probability distribution of all voxels and each voxel are input into an encoder to obtain a code stream; the code stream is decoded by using the same prediction mode and sequence as the double-layer decoder model to obtain a decoded voxel sequence; and inverse pre-processing is performed on the decoded voxel sequence to obtain a lossless compressed three-dimensional medical image. The application models voxels from rough to fine, reduces the calculation time of each voxel, realizes long-distance modeling dependence, and realizes high-performance lossless compression.
Owner:SHANGHAI JIAOTONG UNIV

Information processing device and communication system

An information processing device (UPF) according to one embodiment of the present invention constitutes a node located in an interface segment between a user terminal (UE) and a data network (DN) within a core network, wherein: via an N6 segment that is an interface segment between the information processing device and the DN, at least (1) instance of processing is performed to transmit a protocol data unit (PDU) with a reversibly compressed payload to the DN by uplink communication or (2) processing to receive a PDU with a reversibly compressed payload from the DN by downlink communication; and such reversible compression is not performed on a PDU at a stage to be transmitted by the UE and a PDU at a stage to be received by the PDU.
Owner:SOFTBANK CORPORATION

Radar data compression and transmission method based on wavelet lifting-DCT (Discrete Cosine Transformation)-mixed entropy coding

PendingCN121417907ACode conversionComplex mathematical operationsBiorthogonal wavelet transformRadar
The invention discloses a radar data compression and transmission method based on wavelet lifting-DCT (Discrete Cosine Transformation)-mixed entropy coding. The method comprises the following steps: processing data into a two-dimensional matrix format; the network condition is monitored in real time, and the network is selected according to the availability and bandwidth characteristics of the network; setting a dynamic compression ratio according to a network environment; performing biorthogonal wavelet transformation on the two-dimensional data in the row direction, performing biorthogonal wavelet transformation on the transformed result in the column direction, and removing redundant information; a block-shaped DCT method is adopted, DCT is further conducted on the low-frequency part and the high-frequency part after wavelet transformation, coefficients are effectively concentrated to the upper left corner of a matrix, energy distribution is concentrated, the compression efficiency is improved, and meanwhile the block size is dynamically determined by monitoring the network state according to the parameters and a dynamic compression ratio setting rule; a mixed entropy coding method is adopted, and lossless compression is carried out by using data statistical characteristics; and transmitting the compressed data.
Owner:NANJING NRIET IND CORP

Data hierarchical compression storage method, energy storage management system, device and storage medium

The application discloses a data hierarchical compression storage method and an energy storage management system, device and storage medium, and comprises the following steps: acquiring operation data of a cascaded energy storage system; classifying the operation data into first data, second data and third data according to the reading frequency and the importance of the data; compressing the second data by using a second lossless compression step to form second storage data, wherein the compression rate of the first lossless compression step is lower than that of the second lossless compression step; screening the third data by using a decision tree screening step to form third storage data, wherein the decision tree screening step is provided with a screening condition, part of the content in the third data is deleted based on the screening condition, and the remaining content in the third data forms the third storage data; and storing the first storage data, the second storage data and the third storage data in a local storage, so that the classified storage improves the flexibility of storage reading, reduces the storage load of the storage and improves the reliability of data storage.
Owner:GUANGDONG MINGYANG LONGYUAN POWER ELECTRONICS

A Parallelizable Adaptive Matrix Lossless Compression Method

This invention discloses a parallelizable adaptive lossless matrix compression method. It involves dividing a large first matrix into smaller second matrices, assigning a splitting process to each second matrix, and having these splitting processes divide all second matrices into multiple bit matrices in parallel. Compression threads are then assigned to each bit matrix, and these compression threads adaptively select the optimal compression strategy from a set of compression strategies. Finally, the compression threads compress all bit matrices in parallel according to the determined compression strategy, forming a compressed matrix and a flag array as the final compressed data of the first matrix. By hierarchically and parallelizing the compression process, the memory overhead of the compression process is effectively reduced and the compression speed is improved. Adaptive compression is achieved by selecting the optimal compression strategy through trial compression, effectively improving the data compression ratio. Lossless data compression is achieved by constructing a flag array.
Owner:北京麟卓信息科技有限公司

Real-time lossless compression method for linear scanning panoramic infrared image data

The invention relates to the technical field of data compression, in particular to a real-time lossless compression method for linear scanning panoramic infrared image data. The method comprises the following steps: receiving current column original pixel data output by a linear column scanning infrared sensor; taking the original pixel data of the current column and the coded pixel data of the adjacent column in one frame as reference columns, and performing backward inter-column difference prediction to obtain predicted residual data of the current column; mapping the predicted residual data into a symbol value; carrying out entropy coding on the symbol value by using a static Huffman coding table which is pre-generated and solidified based on the statistical characteristics of the linear scanning infrared image, and outputting a compressed code stream; according to the scheme, the technical problem that high real-time performance and high compression ratio are difficult to consider when real-time lossless compression is carried out on linear scanning panoramic infrared image data in the prior art is solved.
Owner:CHONGQING GAUSS INTELLIGENT COMPUTING TECHNOLOGY CO LTD

Large language model weight compression method, inference operation method, system and medium

The application discloses a large language model weight compression method, an inference operation method, a system and a medium, and belongs to the technical field of large language model weight compression and inference operation optimization. The method comprises the following steps: acquiring a weight matrix of a large language model and constructing an exponential high-frequency window and a benchmark value, classifying and encoding each weight data into a fixed-length code word and splitting the fixed-length code word into an independent bitmap, synchronously generating a compact value stream and a rollback value stream, and packing the compact value stream and the rollback value stream into weight compression data; loading the weight compression data, generating a state mask and a channel prefix mask in a register based on the bitmap, calculating a reading offset by using a population count instruction, reading a compressed data segment from a corresponding data stream, reconstructing weight data in the register in combination with the code word, and finally directly inputting the weight data into a matrix multiplication unit for operation; therefore, by implementing the application, the problem of efficiency reduction caused by control flow divergence and redundant memory access in GPU inference of existing variable-length encoding can be solved, and the model inference efficiency under lossless compression can be improved.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU) +1

Lossy frame buffer compression

Frame buffer compression schemes used for image compression are oftentimes lossless so that the image can be decompressed as close as possible back to its original state. However, lossless compression schemes require that any image data that cannot be successfully compressed (i.e. without losing significant data) be kept in a non-compressed state for transmission and storage. As a result, lossless compression can reduce bandwidth requirements but not memory requirements. The more recently introduced lossy frame buffer compression schemes do allow for some data loss and therefore can save both bandwidth and memory, however, lossy frame buffer compression schemes are limited particularly in the amount by which image data can practically be reduced. The present disclosure provides lossy frame buffer compression which involves an additional compression step, thereby allowing image data to be compressed to a lower rate. This lossy frame buffer compression can reduce both bandwidth and memory usage.
Owner:NVIDIA CORP

Dot matrix font library compression method and device, electronic equipment and storage medium

The invention discloses a dot matrix font library compression method and device, electronic equipment and a storage medium. The method comprises the following steps: in response to triggering of a dot matrix font library compression event, determining zero data in a to-be-compressed dot matrix font library; wherein the zero data is dot matrix data which is not depicted; deleting each piece of zero data from the dot matrix font library to be compressed, and respectively determining an original initial address of each piece of zero data and an accumulated zero removal length before each piece of zero data; generating an initial zero removal table based on the original initial address and the accumulated zero removal length corresponding to each piece of zero data; and generating a target dot matrix font library based on the initial zero removal table and the to-be-compressed dot matrix font library after the zero data is deleted. Through the technical scheme provided by the embodiment of the invention, lossless compression can be performed on the dot matrix font library, and the storage space occupied by the dot matrix font library is effectively reduced.
Owner:ZHEJIANG UNIVIEW TECH CO LTD

A semi-automatic bootstrap medical image sequence labeling method

PendingCN122473209ADICOMMultiple frame
The application discloses a kind of semi-automatic bootstrap medical image sequence marking method, comprising the following steps: facing single frame or multiple frames DICOM and its derived sequence, frame analysis and export lossless image sequence, establish sequence index to support hierarchical batch access, sampling and cross-frame operation, estimate imaging window (sound window) and crop by interframe pixel variation statistics, support manual frame selection, reuse history crop frame and skip abnormal sequence, when marking, based on prompt point set and neighborhood brightness distribution adaptive threshold determination, connected extraction generates effusion target mask;Magnetic tracing is provided to highlight boundary and bone surface boundary is generated along texture fitting, optionally reuse last frame prompt point / mask or optical flow propagation to subsequent frame, and support cross-frame interpolation, and can be integrated segmentation model automatic marking, correction and incremental training to form closed loop;Mask is saved with file tail additional lossless compression section, supports cover update, does not affect general reading and can be restored.
Owner:SOUTH CHINA UNIV OF TECH

An end-side large model inference acceleration method and device based on lossless compression

PendingCN122363909ABatch processingAlgorithm
The application discloses an end-side large model inference acceleration method and equipment based on lossless compression. In the offline preprocessing stage, firstly, the large model weight parameters are decomposed by bits, and the exponential bit data and the decimal bit data are separated; then, the exponential bit data is losslessly compressed, and the compressed exponential bit result and the uncompressed decimal bit data are stored in a solid state disk. In the online inference stage, the GPU memory allocation budget is sensed in real time, and the tensor cache strategy is dynamically adjusted accordingly; when the parameter loading is performed, the scheduler optimally schedules the calculation and I / O workflow related to the tensor recovery according to the current cache state. Under the premise of ensuring zero precision loss of the end-side inference of the large model, the first word delay and the word delay of the model inference are significantly reduced, and the overall throughput in the batch processing scene is greatly improved.
Owner:NANJING UNIV