Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

172 results about "Mixed precision" patented technology

Heterogeneous computing power cooperative scheduling system and method for mixed precision training

The invention discloses a heterogeneous computing power cooperative scheduling system and method for mixed precision training, and belongs to the technical field of artificial intelligence computing. The system comprises a computational graph analysis and operator portrait module which is used for analyzing and dividing a model computational graph and extracting operator features; the heterogeneous hardware capability sensing and matching module is used for managing performance files and real-time states of heterogeneous hardware in the cluster and matching optimal execution hardware for each calculation partition; and the data flow coordination and pipeline parallel controller is used for generating a global execution plan, managing cross-device data dependence and communication and calculating overlapping optimization execution efficiency through communication. According to the method, the problem of low scheduling efficiency of mixed precision training in a heterogeneous environment is solved, automatic and accurate mapping from a calculation task to heterogeneous hardware is realized, the training speed is remarkably improved, the training cost is reduced, and the overall resource utilization rate of a cluster is improved.
Owner:HANHOU (BEIJING) TECH CO LTD

Low-bit-width high-energy-efficiency floating point storage and calculation integrated circuit based on partial pre-alignment architecture

The invention belongs to the technical field of storage and calculation integration, and particularly relates to a low-bit-width and high-energy-efficiency floating point storage and calculation integrated circuit based on a partial pre-alignment framework. The circuit comprises a memory array, a pre-calculation unit, an adder tree, a configurable arithmetic unit and a normalization unit, and supports mixed precision operation of FP8MACFP4 and FP8MACFP8. The method is characterized in that a partial pre-alignment strategy dominated by an activation value is adopted, the maximum index of the activation value is dynamically counted, the mantissa of the maximum index is aligned, and multiple partial pre-alignment intermediate results are pre-calculated and latched for reuse; in combination with a customized lookup table and a multiplexer, a pre-calculation result is directly selected to replace real-time multiplication and displacement; and through the reconfigurable hardware, the FP8MACFP8 high-precision operation is realized by utilizing the FP8MACFP4 unit combination. According to the method, complete online floating point multiplication and addition operation is realized, and excellent energy efficiency ratio and operation speed are obtained while high precision is kept.
Owner:FUDAN UNIVERSITY

Mixed-precision quantization method and apparatus, and device

PCT designated stageWO2025199906A1Neural learning methodsGoal nodeMixed precision
The present application relates to the technical field of computers. Provided are a mixed-precision quantization method and apparatus, and a device. The method is applied to each quantization layer comprised in a quantization model, and comprises: acquiring a plurality of parameter files, wherein each parameter file at least comprises a preset quantization bit width and a preset number of iterations; on the basis of each parameter file, generating an initial quantization layer to obtain a plurality of initial quantization layers; on the basis of a target node comprised in each initial quantization layer, constructing an initial node link to obtain a plurality of initial node links, and on the basis of a quantization result of each initial node link, determining a quantization loss value; and on the basis of the quantization loss value corresponding to each initial node link, determining a target node link, and determining the preset quantization bit width corresponding to the target node link to be a target bit width. The method in the present application improves the automation degree and accuracy in determining the quantization bit width of a quantization model.
Owner:ECARX (HUBEI) TECHCO LTD

Model reasoning method and device, electronic equipment, storage medium and program product

The invention provides a model reasoning method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: performing attention calculation on inferred tokens by adopting first numerical precision through a large language model to obtain a first attention score corresponding to each inferred token; screening a target tokens from the inferred tokens based on the first attention score; reasoning the input sequence by adopting second numerical precision through the large language model to obtain a reasoning result output by the large language model; the input sequence comprises a target token and tokens to be input corresponding to the token to be inferred; the tokens to be input are tokens pre-selected from the inferred tokens according to a preset rule; the first numerical precision is lower than the second numerical precision. According to the method, the sparse processing of the token is realized by adopting mixed precision calculation, so that the reasoning efficiency of the large language model is improved.
Owner:NANJING ILUVATAR COREX TECH CO LTD (DBA ILUVATAR COREX INC NANJING)

Quantization method and system of large language model and electronic equipment

The invention provides a quantification method for a large language model, and the method comprises the steps: for each linear layer, calculating the weight quantification sensitivity of the linear layer, and determining the high-precision bit width ratio of each linear layer according to the weight quantification sensitivity of each linear layer and a target precision mixing ratio; for a plurality of channels of each linear layer, based on the input activation of the linear layer, obtaining the channel quantization sensitivity of each channel in the linear layer; sequencing a plurality of channels in the linear layer according to the channel quantization sensitivity, clustering the sequenced channels in combination with the high-precision bit width ratio of the linear layer, and distributing quantization bit widths with different precisions for weight parameters corresponding to the channels of different clusters; and rearranging a plurality of channels of the same cluster according to distribution similarity indexes, grouping the rearranged channels, and synchronously quantizing weight parameters of each group of channels according to the allocated quantization bit width to obtain a quantized large language model.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Mixed-Precision Model Quantization Method and System for a Residual Connection of a Trained Model

A mixed-precision model quantization method includes loading a trained model, and quantizing the trained model with a mixed-precision setting to generate a quantized model for inference. The trained model includes a plurality of residual connections. In each residual connection, a first activation bypasses at least one operator and is added to a second activation to generate a fourth activation. The second activation is the output of the first activation after being processed by the at least one operator, The mixed-precision setting includes (a) the first activation, the second activation, and the fourth activation in at least one residual connection of the plurality of residual connections being assigned a first precision, and (b) third activations in all operators bypassed by the at least one residual connection being assigned a second precision. The third activations are generated by the bypassed operators and processed within the bypassed operators.
Owner:MEDIATEK INC

Lightweight large model operation method based on end side deployment

The invention provides a lightweight large model operation method based on end-side deployment. The lightweight large model operation method comprises the following steps: acquiring equipment operation data and storing a model file in a mixed precision format according to quantization precision supported by terminal equipment; determining a mixing precision quantification strategy and carrying out strategy analysis on the equipment operation data; a Key-Value cache file obtained through model reasoning in the historical dialogue is reserved; calculating a semantic embedding vector of the second user request, calculating a cosine similarity between the semantic embedding vector and the first V vector, and judging whether the second user request has a similar intention or not according to a calculation result; and if yes, performing incremental reasoning on the second user request by multiplexing the Key-Value cache file. By monitoring the state of the terminal equipment in real time, dynamically adjusting the quantization level, multiplexing the cache data of the semantic related requests and merging batch request reasoning, the effects of balancing the energy consumption and performance of the terminal equipment, reducing repeated calculation and improving the service time of the equipment and the user experience are achieved.
Owner:GUANGDONG GUOLI EDUCATION TECH CO LTD

Expert parallelism processing method and system of large language model based on MoE

The invention belongs to the field of machine learning, discloses an expert parallelism processing method and system for a large language model based on MoE, and realizes efficient parallel processing of the MoE model by dynamically distributing expert quantization bit width and sparse mode, predicting and prefetching to-be-activated expert parameters, grouping tokens to generate task queues and dynamically configuring hardware accelerators. Firstly, an importance score is calculated based on expert historical activation frequency, weight distribution and a model structure, so that quantization precision and a sparse proportion are adaptively allocated, and resource waste and precision loss of a unified strategy are avoided; secondly, predicting an expert to be activated by using a current layer hidden state and a historical activation sequence, loading parameters to a special cache in advance, reducing high-bandwidth memory access, and relieving bandwidth peak scrambling; moreover, tokens are grouped through a token-expert mapping table, a task queue is created, and dynamic configuration of a systolic array is combined, so that the problem of expert heterogeneity after compression is solved, and efficient and parallel hybrid precision matrix operation is ensured.
Owner:NINGBO ORIENTAL UNIVERSITY OF TECHNOLOGY +2

Method and system for realizing mixed-precision low-resource digital control oscillator based on FPGA (Field Programmable Gate Array)

The invention relates to the technical field of digital signal processing, and particularly discloses a mixed precision low-resource digital control oscillator implementation method and system based on an FPGA (Field Programmable Gate Array), and the method comprises the steps: dividing a 32-bit phase accumulator into a high-bit 24-bit dominant frequency control section and a low-bit 8-bit dynamic compensation section, and carrying out the phase accumulation in parallel through double carry chains, the high-order section receives high 24 bits of a frequency tuning word FTW to control a dominant frequency, and the low-order section receives low 8 bits of the FTW and overlaps an overflow remainder to dynamically compensate for a phase error. According to the implementation method of the mixed precision low-resource digital control oscillator based on the FPGA disclosed by the embodiment of the invention, through collaborative innovation of mixed precision phase accumulation, hierarchical mixed storage, high-order interpolation reconstruction and dynamic phase compensation, the core problems of a traditional NCO in resource occupation, phase continuity, waveform precision and time sequence performance are systematically solved.
Owner:ANHUI XICHENG TECH CO LTD

High-energy-efficiency mixed-precision charge domain in-memory computing architecture and working method thereof

The invention belongs to the field of storage, and discloses a high-energy-efficiency mixed-precision charge domain in-storage computing architecture and a working method thereof. Comprising a single-slope analog-to-digital converter, a sparsity perception input alignment module, a multi-bit input accumulation module, a controller and an accumulation module, wherein the single-slope analog-to-digital converter consists of an index calculation array, a mantissa calculation array, a shared ramp voltage generator and a bidirectional counter. According to the invention, serial input binary coding based on capacitor voltage is suitable for floating point and integer multiply-accumulate operation with flexible bit width; according to the invention, a shared single-slope analog-to-digital converter (SS-ADC) is introduced to realize maximum index search and index difference calculation; according to the method, a sparsity perception calculation scheme is provided, low-importance input-weight pairs are filtered out through an adjustable threshold value, and invalid power consumption is reduced; according to the method, a multi-bit input accumulation method is further combined, and an ADC redundancy optimization quantization and normalization process is utilized, so that the overall energy efficiency is improved.
Owner:ZHEJIANG UNIV

Self-adaptive mixing precision quantification method, device, equipment and medium

The invention relates to a self-adaptive mixing precision quantification method and device, equipment and a medium, and the method comprises the steps: carrying out the comprehensive sorting through the product of the cosine similarity difference value of adjacent layers and the sensitivity weight of each layer, so as to guarantee that bottleneck layers which are liable to be influenced by quantification and are crucial to the final precision can be accurately recognized; therefore, precise positioning of a protection target is realized, an iterative optimization loop is introduced, an optimal solution meeting a preset performance target can be spontaneously found finally by continuously evaluating a time-precision balance point of an overall model and automatically adjusting configuration, and a suboptimal result caused by improper primary configuration is avoided; and the input of the user is simplified into a visual final performance target, and the complicated layer sorting and selection process is automatically processed in the system, so that the use threshold of the technology is reduced.
Owner:HUNAN GREAT WALL GALAXY TECH CO LTD

Deep learning acceleration with mixed precision

A device for deep learning acceleration with mixed precision may include matrix-vector (MV) components that each include vector-vector (VV) components that are each configured to generate a respective VV output based on an input precision mode, an output precision mode, and an accumulation of products. The accumulation of products may be calculated by adding products based on the input precision mode. Each product may be calculated by multiplying, based on the input precision mode, a map data segment and a kernel data segment. Each MV component may include one or more components configured to concatenate VV outputs to generate a concatenated VV output. The device may include activation function components that are each configured to receive a corresponding concatenated VV output, generate an activation function output based on the corresponding concatenated VV output and the output precision mode, and output the activation function output.
Owner:MICRON TECHNOLOGY INC

Hybrid precision parallel compression method and system for optimizing large model key value cache

The invention discloses a mixed precision parallel compression method for optimizing large model key value cache, which combines the advantages of mixed precision key value cache compression with an advanced system optimization technology. Based on the characteristic that key value pairs needing high-precision retention in mixed precision compression are the same as key value pairs used for attention calculation in a prefetching strategy, low-precision KV caches are stored in a GPU memory, and meanwhile predicted high-precision important KV caches are dynamically prefetched from a CPU memory according to needs. The technical problems that an existing multi-head attention mechanism-based method is incompatible with an existing pre-training model and cannot be directly applied to a closed source or a fine-tuned large model, and generalization of the method is reduced can be solved, and the situation that an existing pruning-based method is prone to deleting unimportant marks in the current stage can be solved. And contextual information loss is caused.
Owner:HUNAN UNIV

Data layout optimization method and device applied to NPU code compiling and medium

The invention discloses a data layout optimization method and device applied to NPU code compilation and a medium. The method comprises the steps that an intermediate representation IR input by an upper layer is split into an operation type OP and an operand, and the operand input by the upper layer is divided into logic data and a logic mask; converting the logic data of the upper layer into data vector representation of the abstraction layer, and converting the logic mask of the upper layer into mask vector representation of the abstraction layer; obtaining a corresponding bottom hardware instruction capability according to an operation type OP input by an upper layer, and performing legalization operation on the operation type OP; and performing legalization processing on the converted abstract data vector type representation and mask vector representation according to the target hardware capability, performing instruction mapping on the processed abstract instruction according to the underlying hardware pair operation type OP, and generating a target LLVM instruction compatible with the underlying hardware. Register layout abstraction and conversion can be carried out on the NPU code under mixed precision input, and the utilization efficiency of the NPU bottom layer register is improved.
Owner:SOUTH CHINA UNIV OF TECH

Large model fine-tuning optimization method based on multi-strategy fusion

The invention discloses a large model fine tuning optimization method based on multi-strategy fusion, which comprises the following steps: designing a dynamic parameter selection mechanism, adaptively determining a parameter subset needing fine tuning according to a task demand and a model structure, and reducing unnecessary parameter updating calculation; constructing a dynamic low-rank decomposition framework, dynamically adjusting the rank of a low-rank matrix according to a model training state and data characteristics, and keeping key information while compressing a parameter scale; a self-adaptive task sensing mechanism is introduced, a fine adjustment strategy is automatically adjusted according to different task characteristics, and the adaptability of the model to various tasks is improved; and a mixed precision training method is adopted, so that the calculation complexity and the memory occupation are reduced on the premise of ensuring the model precision. According to the method, a parameter efficient fine tuning technology and a dynamic low-rank decomposition strategy are innovatively combined, and an adaptive task perception mechanism and a mixed precision training technology are introduced, so that the operand and resource requirements of model training are effectively reduced, and the fine tuning efficiency and the model performance are improved.
Owner:JIANGSU JIYUAN MEDICAL TECH CO LTD

High-efficiency INT6 quantification method, device and equipment for large language model

The invention discloses an efficient INT6 quantification method, device and equipment for a large language model, and the method comprises the steps: carrying out the mixing precision quantification of the large language model, and obtaining a quantized large language model; performing bit-level data packaging on the weight and the activation value in the quantized large language model to obtain bit-level data; loading the bit-level data to a register of the GPU, and carrying out matrix product accumulation operation and weighted summation by utilizing BTC to obtain output data; and storing the output data back to a global memory of the GPU so as to complete the quantitative reasoning process of the large language model. According to the method, mixed precision quantification is carried out on a large language model by utilizing different precisions, the reasoning speed is improved through a scheduling strategy of the GPU while relatively high quantification precision is achieved, and the reasoning potential of the GPU is fully mined, so that all potential of 6-bit quantification is released.
Owner:XIDIAN UNIV

Computing device and method based on RISC-V extension instruction

The invention provides a computing device and method based on RISC-V extension instructions, the computing device supports approximate computation of mixed precision according to approximate computation instructions in an extended approximate computation instruction set, and the computing device comprises an out-of-order scheduling and register reading module used for executing instruction dependency analysis and operand preloading, scheduling the non-approximate calculation instruction and the approximate calculation instruction to different transmitting queues respectively; the first instruction transmitting queue is used for temporarily storing a to-be-transmitted non-approximate calculation instruction; the second instruction transmitting queue is used for temporarily storing approximate calculation instructions to be transmitted; the precise calculation module is used for completing precise calculation related tasks according to the instruction from the first instruction transmitting queue; and the approximate calculation module is used for completing approximate calculation related tasks according to the instructions from the second instruction transmitting queue, and supports approximate calculation of various precisions.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

MoE fine tuning training method and device, electronic equipment and storage medium

The invention provides a fine tuning training method and device of MoE, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of distributed training and optimization of a large-scale machine learning model, compression and acceleration of a deep learning model and the like. According to the implementation scheme, model structure analysis is conducted on a MoE model to be finely adjusted, and an expert module and a non-MoE module in the MoE model are recognized; quantizing the weight parameter of the expert module by adopting first low-bit precision, and quantizing the weight parameter of the non-MoE module by adopting second low-bit precision; performing model forward calculation and back propagation calculation based on the quantized weight by using a mixed precision calculation operator; and based on an LoRA fine tuning technology, performing fine tuning training on the quantized MoE model by using a distributed training framework to obtain a target MoE model. The scheme can significantly reduce video memory occupation, effectively improve training efficiency, break through model scale limitation, and maintain excellent model performance.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Transform network-oriented mixed precision multiplier

The invention discloses a Tranformer network-oriented hybrid precision multiplier, which is applied to an FPGA (Field Programmable Gate Array), and is used for allocating a target high-resolution image according to the weight precision of the weight data of the current layer of the target Tranformer network, the feature map precision of the feature map data of the target high-resolution image and a preset data allocation rule. Determining a first quantity of weight data and a second quantity of feature map data of a single calculation cycle; according to the first quantity of weight data and the second quantity of feature map data, recoding the weight data and the feature map data after determining a target recoding algorithm to obtain weight coded data and feature map coded data; performing multiplication calculation on the weight coding data and the feature map coding data in a single calculation period to obtain an initial product result of the single calculation period; and decoding the initial product result, and merging the multiple sub-feature map data of all the calculation periods corresponding to the current layer of the target Transform network to obtain target feature map data. According to the method, the resource utilization rate is improved on the basis of ensuring the model precision.
Owner:XIDIAN UNIV

GEMM load-oriented GPU modeling method

A GEMM load-oriented GPU modeling method is characterized in that through a multi-stage collaborative modeling mechanism, cache behaviors, instruction overhead and calculation intensity are deeply coupled, accurate performance prediction of GPU execution GEMM operators is realized, the method can be widely applied to scheduling optimization of GPU intensive scenes such as AI training and scientific calculation, firstly, a three-stage cache weight distribution mechanism is established, and then, a three-stage cache weight distribution mechanism is established; quantifying the contribution of the L1 / L2 cache hit rate and the DRAM bandwidth degradation factor to the effective bandwidth; secondly, an instruction-level memory access overhead correction mechanism is introduced, and the mixing precision and the real calculation strength of a sparse calculation scene are captured through dynamic parameter adjustment and optimization; then combining the calculation force peak value and the bandwidth upper limit to construct a double-boundary constraint model, and generating a theoretical performance critical value; further predicting a stream multiprocessor utilization rate based on a neural network, and quantifying efficiency loss caused by hardware resource contention through a multi-layer perceptron structure; and finally, the integration module outputs task execution time to realize end-to-end performance prediction.
Owner:BEIHANG UNIV

Mixed-precision matrix multiplication

Systems and techniques for providing mixed-precision matrix multiplication in multi-chiplet processors recognize different precision formats of matrices to be multiplied based on, e.g., parameters provided with instructions or start and end memory locations of the matrices. A plurality of different multiplication chains are provided for different formats such that mixed-precision matrix multiplication can be performed using multiplication chains configured to handle multiplication of different precision formats. The multiplication chains are automatically selected based on the precision formats of the matrices to be multiplied, enabling programmers to utilize the chains without having to directly access the individual multiplication chains.
Owner:ADVANCED MICRO DEVICES INC

Self-adaptive mixing precision calculation method, system and equipment and storage medium

The invention relates to the field of power system simulation, and provides a self-adaptive mixing precision calculation method, system and device and a storage medium, applied to a multi-type programmable chip cooperation system, the method comprises the following steps: receiving runtime monitoring information returned by a multi-chip cooperation controller and a precision gateway through an intelligent precision manager, updating precision configuration information and a gateway conversion strategy according to the monitoring information during operation; decomposing the target calculation task into a plurality of sub-tasks through the multi-chip cooperative controller according to the precision configuration information, and mapping each sub-task to a target programmable chip corresponding to each precision partition; wherein when cross-precision transmission of data exists, data conversion and error control are carried out through the precision gateway according to the gateway conversion strategy, and an output result is obtained. According to the method, fine-grained identification and precision configuration can be realized, standardized conversion and error suppression are realized in a cross-precision boundary, closed-loop cooperation is carried out, and self-adaptive cooperative calculation of mixed precision is realized.
Owner:ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1

High-speed rail platform safety determination method and system based on mixed precision inference

PendingCN122333241ARelation graphAlgorithm
The application relates to the technical field of intelligent reasoning and safety judgment, and discloses a high-speed rail platform safety judgment method and system based on mixed-precision reasoning, which comprises the following steps: acquiring multi-source semantic observation records and generating a safety observation element set; constructing a platform safety relation graph; determining relation conflict density, closed residual error, cross-source divergence degree and reasoning difficulty level; generating a mixed-precision bit width scheduling table; performing graph relation reasoning to obtain an initial safety judgment vector and an initial judgment boundary quantity; in step 6, a final judgment vector is determined; and in step 7, a locked safety level is determined and a safety judgment package is output. The application realizes mixed-precision safety judgment and locked output driven by multi-source semantic relation of a high-speed rail platform.
Owner:XIAMEN SILICON TECHNOLOGY CO LTD

Decision tree training and inference with mixed precision

A method, system, and computer program product perform machine-learning inferences with a tree-based model. The tree-based model includes a decision tree that was trained on a first system, which is configured to perform computations with a first arithmetic precision. The inferences are performed with the tree-based model on a second system, which is configured to perform computations with an arithmetic precision that is lower than the first arithmetic precision. Performing the inference includes determining that an input feature value is equal to a threshold value of a corresponding node and, in response, using a majority voting to select a left or right path of the decision tree. The majority voting is based on historical statistical data that includes tree-path statistics.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

A method for generating an adaptive mixed-precision quantization network based on self-learning

The application relates to the technical field of model quantization, and discloses a generation method of a self-adaptive mixed-precision quantization network based on self-learning, which comprises the following steps: obtaining a teacher bit width set and a student meta-network to be trained, the teacher bit width set comprising multiple candidate teachers with high bit widths, and the student meta-network sharing full-precision weights and supporting multiple bit width configurations; for any bit width configuration, a corresponding target teacher is determined according to the interlayer Manhattan distance between the student meta-network and each candidate teacher and the entropy value of the prediction probability distribution of each candidate teacher; all bit width configurations of the student meta-network are jointly trained to obtain a self-adaptive bit width student meta-network; based on the interlayer importance evaluation result, a bit width search strategy is used to determine the target bit width of each layer network; and based on the target bit width of each layer network and the self-adaptive bit width student meta-network, a target mixed-precision quantization network is generated, thereby solving the technical problem of low diagnosis accuracy of related models in power equipment state monitoring.
Owner:WENZHOU ELECTRIC POWER BUREAU

Space-air-ground integrated multi-agent collaborative large model online training method and system

The invention discloses an air-space-ground integrated multi-agent collaborative large model online training method and system. According to the method, a distributed mixing precision quantitative configuration table is dynamically generated based on multi-node resource heterogeneity, and appropriate parameter precision is allocated to different nodes; a cross-node gradient residual error distribution and local compensation strategy is adopted, only compressed gradient residual errors are transmitted between nodes, errors are locally accumulated, and high-precision compensation updating is carried out regularly; elastic division and dynamic scheduling of model parameters and learning tasks among multiple nodes are supported, and task allocation and model segmentation are adjusted in real time according to the performance of each node and the network condition; when the nodes are lost or the calculation load is uneven, the system has fault-tolerant and adaptive scheduling capabilities, and can automatically redistribute model fragments or adjust the training process. According to the method, the computing power resources of the heterogeneous multi-agent are fully utilized, and efficient cooperative training of the super-large-scale model under the scene that the bandwidth is limited and the nodes dynamically change is realized.
Owner:SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT) +1

Method and apparatus for controlling input / output operation of vector processor in mixed precision environment

A The present invention relates to a technique for controlling input / output operation of a vector processor, which is designed to optimize vector operation in a mixed precision environment, and to a technique for maximizing data processing performance while minimizing waste of operation resources by improving the data conversion process between the memory and the vector processor.
Owner:SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION

Novel translation model reasoning method and system based on rwkv

PendingCN122287660AImplement reasoning methodsscale upComputation complexityTheoretical computer science
The RWKV-based novel translation model inference method and system includes the following steps: 1) Collecting novel translations and extracting parallel corpora, and using dynamic MicroBatch concatenation technology for sequence compression; 2) Introducing a lightweight group query attention mechanism on the basis of the RWKV architecture, directly obtaining KV information from the Embedding layer to build the model; 3) Employing a sublinear complexity hybrid parallel training mode, combined with a global scalar scaling FP16 mixed precision strategy for training; 4) Applying a hierarchical distributed heterogeneous architecture, offloading the optimizer to low-performance devices and performing gradient compression transmission; 5) Outputting the translation using a joint decoder and dynamic batch inference technology. This invention, through the above method and system, effectively reduces the computational complexity and memory usage in the long text translation process, improves model training efficiency and inference throughput, and significantly improves the translation efficiency and contextual coherence of ultra-long texts.
Owner:LIAONING UNIVERSITY

Transform hardware acceleration method and accelerator based on hybrid precision quantization and huffman coding

This invention discloses a hardware acceleration method and accelerator for Transformer based on mixed-precision quantization and Huffman coding. The acceleration method includes: using a genetic algorithm to obtain several configuration schemes for mixed-precision quantization of Transformer network layers; performing mixed-precision quantization on each Transformer network layer based on each quantization configuration scheme to obtain a corresponding KL divergence; training a multilayer perceptron to obtain a quantization configuration prediction network using the quantization configuration scheme and the corresponding KL divergence as the output label and input feature, respectively; receiving a user-set target KL divergence value, using the quantization configuration prediction network to obtain the corresponding quantization configuration scheme, and performing mixed-precision quantization on each network layer based on the quantization scheme; and using Huffman coding to encode and compress all quantization weights before on-chip storage. This invention can reduce storage and computational overhead while maintaining model accuracy.
Owner:HUNAN NORMAL UNIVERSITY