Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

101 results about "Approximate computing" patented technology

Approximate computing is a computation technique which returns a possibly inaccurate result rather than a guaranteed accurate result, and can be used for applications where an approximate result is sufficient for its purpose. One example of such situation is for a search engine where no exact answer may exist for a certain search query and hence, many answers may be acceptable. Similarly, occasional dropping of some frames in a video application can go undetected due to perceptual limitations of humans. Approximate computing is based on the observation that in many scenarios, although performing exact computation requires large amount of resources, allowing bounded approximation can provide disproportionate gains in performance and energy, while still achieving acceptable result accuracy. For example, in k-means clustering algorithm, allowing only 5% loss in classification accuracy can provide 50 times energy saving compared to the fully accurate classification.

Graph neural network execution on neural processing unit

Workloads for executing a graph neural network (GNN) may be divided among various processing units, such as a central processing unit (CPU) and a neural processing unit (NPU). The NPU may include a data processing unit (DPU) and a digital signal processor (DSP). The CPU may perform precomputation, model optimization, hardware optimization, and compilation. For example, the CPU may precompute a parameter matrix and use the parameter matrix as internal parameters of a GNN. The CPU may also perform node padding, approximation computation, or transfer of DSP operations to DPU to optimize the GNN. The CPU may also perform sparsity data compute and storage, vertical fusion of DSP operations and DPU operations, or data quantization to optimize performance of the NPU. The compiled GNN may be provided to the NPU, and the DPU and DSP may perform the operations in the compiled GNN to produce a prediction of the GNN.
Owner:INTEL CORP

A Low-Complexity Decoding Method for LDPC-Hadamard Codes Based on Prototype Diagrams

This invention discloses a low-complexity decoding method for LDPC-Hadamard codes based on a prototype diagram, belonging to the field of digital signal processing. The implementation method is as follows: information bits are LDPC encoded, and then the constraint relationship of the check nodes of the LDPC code is encoded into Hadamard code to obtain the prototype diagram of the LDPC-Hadamard code; the encoded codewords are transmitted and received through a channel, and the PVN and HCN nodes are initialized with the received information; Hadamard decoding is performed using a symbol-by-symbol maximum a posteriori probability decoding algorithm, and DFHT approximation is used to reduce decoding complexity; the decoded information is transmitted to the PVN nodes; the updated HCN node information is summed with the information from the previous iteration, and the summation result is used to determine the decoding; a decoding decision is made on the PVN information, and decoding is performed. This invention combines LDPC encoding, Hadamard encoding, and Max-Log approximation calculation methods to reduce decoding complexity and enhance the stability and reliability of the communication system.
Owner:BEIJING INST OF TECH

Big data real-time monitoring system and method

PendingCN121166478AFault responseSemantic analysisComplex event processingTimestamp
The invention relates to a big data real-time monitoring system and method. The big data real-time monitoring system comprises a data access module, a preprocessing standardization module, a streaming computation module, a rule model engine module, an alarm arrangement and disposal module, a storage query module and a visualization and operation and maintenance module. The data access module is used for collecting multi-source heterogeneous data from a message queue, a log agent or an Internet of Things protocol, and labeling an event timestamp and a source identifier for each piece of data; the preprocessing standardization module is used for performing deduplication, cleaning, desensitization and dimension table association on the collected data and outputting a unified event structure; the stream-oriented computing module is used for executing window aggregation, state updating and complex event processing under event time semantics and supporting water level management and just one-time semantic processing, high-timeliness and consistent monitoring is achieved through self-adaptive water level, double-layer windows and dynamic threshold gating fusion, versioning playback and approximate computing cost reduction are supported, and the method is suitable for large-scale popularization and application. And the system has automatic degradation and self-healing capabilities.
Owner:SHANGHAI QUANTITATIVE FOREST TECH CO LTD

Computing device and method based on RISC-V extension instruction

The invention provides a computing device and method based on RISC-V extension instructions, the computing device supports approximate computation of mixed precision according to approximate computation instructions in an extended approximate computation instruction set, and the computing device comprises an out-of-order scheduling and register reading module used for executing instruction dependency analysis and operand preloading, scheduling the non-approximate calculation instruction and the approximate calculation instruction to different transmitting queues respectively; the first instruction transmitting queue is used for temporarily storing a to-be-transmitted non-approximate calculation instruction; the second instruction transmitting queue is used for temporarily storing approximate calculation instructions to be transmitted; the precise calculation module is used for completing precise calculation related tasks according to the instruction from the first instruction transmitting queue; and the approximate calculation module is used for completing approximate calculation related tasks according to the instructions from the second instruction transmitting queue, and supports approximate calculation of various precisions.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Real-valued super-resolution direction-of-arrival estimation method for high-order acoustic field sensor array

The application relates to a real-valued super-resolution direction estimation method of a high-order acoustic field sensor array, which uses a received signal covariance matrix to approximately calculate an estimated value of noise power, avoids the iterative operation process of noise power in the prior art, improves the calculation efficiency, and reduces the floating-point operation amount. By constructing an augmented matrix, the array receiving data matrix and the array manifold matrix with a multi-dimensional structure of an element are made into Hermitian matrices, so that the unitary transformation processing of the related parameters of the high-order acoustic field sensor array is realized. On the basis of obtaining the array receiving data matrix and the array manifold matrix in the real number domain, the array receiving data matrix and the array manifold matrix are applied to a sparse approximate minimum variance method with a variable exponential factor, and the completely real-valued sparse approximate minimum variance direction estimation with a variable exponential factor is realized.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

A low-power approximate equalization circuit and method based on frequency domain error property mapping

This invention relates to the fields of high-speed serial communication and digital signal processing, and discloses a low-power approximate equalization circuit and method based on frequency domain error attribute mapping. The circuit includes: a Fast Fourier Transform (FFT) module for converting the time-domain signal to be equalized into a frequency-domain signal; a frequency-domain equalization processing module for equalizing the frequency-domain signal; an Inverse Fast Fourier Transform (IFT) module for converting the equalized frequency-domain signal into a time-domain output signal; and a mapping control module for constructing a three-dimensional error attribute mapping relationship with frequency domain sub-bands as the first dimension, approximate calculation accuracy level as the second dimension, and the error polarity of the approximate arithmetic unit as the third dimension. This invention simultaneously considers frequency domain position, approximate accuracy, and error polarity characteristics in frequency-domain equalization, and achieves system-level error balance without introducing additional compensation circuits, thereby achieving low-power, high-reliability frequency-domain equalization.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

A heterogeneous platform approximate computing task optimization mapping method based on DVFS and DPM

The application discloses a heterogeneous platform approximate computing task optimization mapping method based on DVFS and DPM. Firstly, real-time tasks with correlation are modeled as an approximate computing task model, so that a task directed acyclic graph (DAG), a task correlation matrix and a six-tuple representing task characteristics can be obtained; then, a mechanism of DVFS and DPM combination is introduced based on a heterogeneous multi-core platform; a problem description of task mapping based on QoS and energy joint optimization is constructed; a variable substitution method and a Big-M reconstruction method are used to process nonlinear terms in the problem, the task mapping problem is linearized, and an optimal solution is obtained through a Gurobi solver; a task layering method and a greedy algorithm are used to design a low-complexity heuristic algorithm, and the scalability of the mapping method is improved. The method of the application adopts the mechanism of DVFS+DPM joint optimization under the premise of meeting the system real-time, energy efficiency and reliability constraints, and improves the QoS of the system.
Owner:SOUTHEAST UNIV

Layered multi-objective optimization method and system for approximate calculation parameter optimization

The invention relates to a hierarchical multi-objective optimization method and system for approximate calculation parameter optimization. The method comprises the following steps: constructing a hierarchical search model for approximate calculation commitment optimization; and in the category-type decision-making layer, taking each approximation strategy as an arm, maintaining the arms by adopting a multi-arm tiger machine algorithm, obtaining random sampling values in each round of iteration according to the obtained Beta distribution, and evaluating the target arm corresponding to the maximum value so as to obtain an arm to be optimized. And inputting the parameter point to be evaluated into a numerical decision-making layer to perform parameter iterative optimization, and generating a next parameter point to be evaluated by maximizing a qNEHVI acquisition function, so that a program, corresponding to the parameter point, of the arm to be optimized is executed through a multi-target Bayesian optimization engine, and an optimization evaluation result is obtained. And obtaining the most approximate calculation configuration according to the result and the preset evaluation budget. The method can face mixed and hierarchically dependent variable space, and can improve the tuning efficiency and reduce the cost.
Owner:NAT UNIV OF DEFENSE TECH

An aging-aware precision reconfigurable logic synthesis method

PendingCN122334124AHemt circuitsLogisim
The application discloses an aging-aware precision reconfigurable logic synthesis method and system, and belongs to the technical field of integrated circuit EDA. The application aims to solve the problems of limited optimization space, introduced error from the beginning and lack of aging modeling in the existing aging-aware synthesis. The application constructs an aging standard cell library, iteratively applies constant replacement to the original circuit to generate multiple approximate candidate circuits, and then constructs a precision reconfigurable circuit composed of reconfigurable modules, which dynamically switches between accurate mode and approximate mode. Through aging-aware timing modeling, a power function model of gate delay with running time is established, and the aging continuity during mode switching is processed. The accurate mode life and the approximate mode remaining life are predicted, and the circuit with the highest score is automatically selected according to the score function. The application can improve the circuit life by about 9.5 times with an additional area overhead of about 3.72% under the premise of meeting the error constraint, which significantly enhances the long-term reliability of digital circuits.
Owner:SHANGHAI JIAOTONG UNIV

A method for predicting the centering state of a permanent magnet coupling based on a time sequence signal

The application belongs to the technical field of permanent magnet magnetic drive, and provides a permanent magnet coupling centering state prediction method based on time sequence signals. First, two mutually orthogonal state vectors are introduced to represent the eccentricity and position angle of the inner and outer rotors of the permanent magnet coupling, and the axial section of the inner and outer rotors of the permanent magnet coupling is theoretically modeled; in combination with the approximate calculation model of the Monte Carlo algorithm, the measured results are referenced to the calculation to reevaluate the weight of the data set, and the fast prediction and evaluation of the centering state of the permanent magnet coupling are realized. The application can realize the prediction and evaluation of the centering state of different types of permanent magnet couplings, and has strong adaptability; in addition, the measured data is reevaluated by the statistical probability method, the calculation result is corrected, the accuracy of the prediction value is ensured, and important technical support is provided for the multi-parameter state evaluation and prediction of the permanent magnet coupling. In the aspect of actual engineering application, it has good actual application performance, simple operation and small calculation amount.
Owner:DALIAN UNIV OF TECH

High-accuracy calibration of an analog quantum simulator

One example aspect of the present disclosure is directed to a method for operating a quantum computing system (QCS). The QCS includes a quantum processor device that includes a set of qubits and a set of qubit couplers. The method includes determining a set of dressed qubit frequencies. Each dressed qubit frequency of the set of dressed qubit frequencies corresponds to a separate qubit pair of the set of qubits. A set of bare qubit frequencies is generated based on the set of dressed qubit frequencies. A spin-Hamiltonian for the set of qubits is approximated based on the set of bare qubit frequencies
Owner:GOOGLE LLC

Method and device for evaluating reliability of engineering production system

The invention relates to the technical field of reliability engineering, in particular to a reliability evaluation method and evaluation device for an engineering production system, and the method comprises the steps: obtaining a reliability block diagram of the engineering production system, determining the metadata of each function block, generating a reliability function network, marking a typical reliability structure unit, and obtaining a reliability function network; the method comprises the following steps: determining a function subnet, configuring a time domain strategy to generate a self-adaptive time grid, solving the equivalent reliability of the function subnet at a self-adaptive time point of the self-adaptive time grid to obtain a reliability sampling sequence of the engineering production system, and generating a reliability evaluation result. Therefore, the problems that in the related technology, due to the fact that semantic information corresponding to a reliability block diagram depends on manual disassembly, and a reliability evaluation result is calculated through fixed step size discrete approximation, missed judgment and misjudged are likely to occur due to complex topology, the calculation precision is poor in a high-risk time period, the calculation efficiency is low in a stable time period, and engineering decision is difficult to support are solved.
Owner:TSINGHUA UNIVERSITY

A mechanical arm obstacle avoidance path planning method based on an improved ant colony algorithm

PendingCN122323217ARobotic armMetric tensor
This invention proposes a robotic arm obstacle avoidance path planning method based on an improved ant colony algorithm, relating to the fields of robot path planning and intelligent optimization algorithms. This method endows the joint configuration space with a non-uniform metric structure induced by a metric tensor G(q), where G(q) is derived from the normalized joint inertia matrix W. I The algorithm consists of three parts: the singularity gradient outer product term and the obstacle spacing gradient outer product term. Local geodesic distances are approximated using the mean of the metric tensors at both ends of the node. All edge weights are pre-calculated and cached during the PRM graph construction phase. The ant colony uses the reciprocal of the geodesic distance as a heuristic function, drives non-uniform pheromone evaporation using the normalized value of the metric tensor trace increment, and uses the weighted sum of geodesic length cost and inertia-weighted velocity mutation penalty as the comprehensive path cost. The beetle whisker algorithm performs pre-search and completes non-uniform pheromone initialization under geodesic metrics. This method effectively improves path safety, continuity, and dynamic adaptability.
Owner:LUDONG UNIVERSITY

Approximate computing digital circuit for post-quantum cryptography applications

A digital circuit for including a scalar product between two vectors (a0, a1, . . . , ai, . . . , aN-1) and (s0, s1, . . . , si, . . . , sN-1). The digital circuit includes a multiplier, an accumulator including at least one adder and a register, as well as a control circuit of the accumulator. At a clock tick of index i, the multiplier is configured to compute the result ri of the multiplication ai×si, and the accumulator is configured to add ri with the current value of the register. Afterwards, the result of the addition is memorised in the register. The control circuit is configured to control the accumulator so as to perform the addition in an approximate manner for at least one addition amongst the N additions of the computation of the scalar product. In particular, the digital circuit is intended to be used in an electronic device implementing a cryptographic algorithm based on a “Learning With Errors” (LWE) technology.
Owner:COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES

Homomorphic ciphertext-based maximum and minimum function processing method and system, and storage medium

The application provides a maximum and minimum value function processing method and system and a storage medium, which are characterized by the following steps: S1-1, obtaining a maximum and minimum value function and first and second ciphertexts to be compared; S1-2, summing the first and second ciphertexts based on a homomorphic encryption algorithm to obtain a first result; S1-3, calculating a second result obtained by dividing the first result by 2 based on a predetermined homomorphic division operation rule; S1-4, performing square root on the square of the difference between the first and second ciphertexts based on a predetermined square root approximation operation rule to obtain an approximate third result; S1-5, calculating a fourth result obtained by dividing the third result by 2 based on the predetermined homomorphic division operation rule; S1-6, performing corresponding homomorphic encryption processing on the second and fourth results according to the type of the maximum and minimum value function to obtain a fifth result; and S1-7, decrypting to obtain the processing result of the maximum and minimum value function.
Owner:SHANGHAI HOMO STATE INFORMATION TECH CO LTD

Analyzing computing products using computer-vision

A method of analyzing a particular computing product, including: receiving a plurality of images of a particular layout of the particular computing product; segmenting, using a classification model, the plurality of images to identify computing components of the particular computing product; analyzing the particular layout, including, for each computing component of the particular layout: approximating a physical size of the computing component; identifying a predetermined layout weight of the computing component; determining a proximity of the component to each other computing component; calculating a computing component score for the computing component for the particular layout based on i) the physical size of the computing component, ii) the predetermined layout weight of the computing component, and iii) the proximity of the computing component to each other computing component; and determining a layout score of the particular layout based on the computing component score of each of the computing components.
Owner:DELL PROD LP

Improved model-free current prediction control method, device and system based on forgetting factor

The application discloses an improved model-free current prediction control method and device and system based on a forgetting factor. The method does not need to acquire system parameters in advance, only needs to collect a load current, and then approximately calculates a current increment from historical current data, and then outputs an optimal inverter switch state from a minimized current value function, and meanwhile, a forgetting factor is introduced to weaken the adverse effect of a sampling error on the system, so that the model-free current prediction control is realized. On the basis of guaranteeing the realization of the parameter-free control, the forgetting factor is introduced to weaken the adverse effect of the sampling error on the system, the current quality is improved while the response speed is faster, and the robustness of the control is improved. The application is suitable for different types of power electronic topological structures, and does not need to separately perform mathematical modeling on different types of power electronic topological structures, has a reference significance for the model-free prediction current control of the power electronic converter, and has a wide application prospect.
Owner:TONGJI UNIV

Nonlinear operator approximate calculation device and method, neural network processor and medium

The embodiment of the invention provides a nonlinear operator approximate calculation device and method, a neural network processor and a medium, and belongs to the technical field of neural networks. The device comprises: a floating point number evaluation unit receiving original floating point data of an original neural network operator, and performing compensation interval evaluation on the original floating point data according to a preset floating point value domain range to obtain compensation interval evaluation information; if the compensation interval evaluation information represents that the original floating point data is not in the preset floating point value domain range, the index splitting unit splits the original floating point data into a first floating point number and a second floating point number; the operation compensation unit performs fitting compensation on the first floating-point number to obtain first output data; and the splicing unit splices the first output data and the second floating-point number passing through the original neural network operator to obtain target output data. According to the invention, the computing resources and time of a computer system can be reduced, the full-value-domain compensation of the neural network operator is completed with few hardware resources, the computing precision is improved, and the reasonability of network reasoning is ensured.
Owner:SHENZHEN WEIXUN TECH CO LTD

A 2-group signed tensor computing circuit structure based on 6-bit approximate full adder

The application discloses a 2-group signed tensor calculation circuit structure based on a 6-bit approximate full adder, relates to the field of neural network hardware acceleration, and comprises a 6-bit approximate full adder module, a signed 8*8 approximate multiplier circuit structure based on the 6-bit approximate full adder and the 2-group signed tensor calculation circuit structure based on the 6-bit approximate full adder. Since the neural network accelerator can sacrifice part of the accuracy of data in exchange for the optimization of delay, area and power consumption of a circuit structure, the 6-bit approximate full adder module ignores part of the carry, thereby reducing the circuit area and lowering the circuit power consumption, the signed 8*8 multiplier calculation process and the 2-group signed tensor calculation process are optimized by using the 6-bit approximate full adder module, approximate calculation is introduced at some positions, and in exchange for the improvement of area and power consumption of the circuit structure, part of the accuracy is lost.
Owner:SOUTHEAST UNIV

Fast approximate analysis method for mutual coupling-containing pattern of a large-scale ultra-wideband heterogeneous array

The present disclosure discloses a fast approximate analysis method for mutual coupling-containing pattern of a large-scale ultra-wideband heterogeneous array, comprises: dividing the entire large-scale ultra-wideband heterogeneous array into a plurality of sub-arrays according to a type of element used by a heterogeneous array, ensuring that each sub-array contains elements of the same type; classifying the elements in the sub-arrays; selecting representative elements for representing environmentally similar elements, removing rows and columns that do not have the representative elements, and retaining all key features of the heterogeneous array; performing full-wave simulation on a constructed compact representative array, extracting AEPs of all the representative elements, and storing same; and replacing AEPs of the environmentally similar elements with the AEPs of the representative elements, performing approximate computing to obtain patterns of all the sub-arrays, and superimposing the obtained results to obtain a mutual-coupling-containing pattern of the heterogeneous array.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Privacy protection method for two-party radial basis kernel function calculation

The invention discloses a privacy protection method for two-party radial basis kernel function calculation, and belongs to the field of multi-party security calculation. The method comprises the following steps: combining a Trunc protocol, a CrossTerm protocol, an MCrossTerm protocol and an MUX protocol in the prior art, constructing secret sharing of a radial basis kernel function by sequentially utilizing operations such as amplification, dot product, two-norm and zooming on secret sharing under a framework of multi-party security calculation, and realizing approximate calculation under a ciphertext equivalent to plaintext radial basis kernel function calculation. According to the method disclosed by the invention, the function of a million-opercuity protocol can be realized only by adopting one-time casual transmission, so that repeated calling and casual transmission in the prior art are avoided, the communication and calculation overhead can be greatly reduced, and the efficiency is improved; the method is applied to a machine learning model training task, data safety of a user can be guaranteed, and efficient model training is achieved.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

Chaotic sequence prediction network training and use method, interference system and medium

The present application relates to a kind of chaos sequence prediction network training and use method, interference system and medium, method includes: obtaining the derivative of several chaos sequence data of pre-set chaos system;Approximate calculation of the derivative of several chaos sequence data is based on pre-set numerical difference equation, determine the trend data of chaos sequence;Obtain initial KANs, the network parameter of initial KANs is trained based on the trend data and the chaos sequence data, obtain initial prediction network;Obtain the prediction fixed point of initial prediction network, the training state of initial prediction network is evaluated based on the prediction fixed point, when the training state is training completion, the initial prediction network is output as chaos sequence prediction network.On the basis of high prediction accuracy, training speed is improved, and then the real-time, accurate prediction of chaos sequence in intelligent interference system is realized.
Owner:成都流体动力创新中心

Homomorphic encryption maximum value calculation method based on multivariate symmetric polynomial approximation

The invention discloses a homomorphic encryption maximum value calculation method based on multivariate symmetric polynomial approximation, and belongs to the field of information security. The method comprises the following steps: firstly, setting a polynomial maximum number of times for fitting a maximum value function according to precision and depth budget; secondly, mapping an input variable into a finite dimension moment variable, and training a fitting polynomial based on the original data set; through rearrangement and layer number analysis on monomials in a polynomial, identifying and reconstructing border-crossing items exceeding depth budget, and converting the border-crossing items into a tree multiplication structure conforming to depth constraint; and finally, generating a final low-depth polynomial expression by caching and multiplexing the intermediate variable. Through the moment variable dimension reduction and low-depth reconstruction technology, the problem of combinatorial explosion in a traditional method is effectively avoided, the multiplication depth and multiplication times of ciphertext calculation are remarkably reduced, efficient and low-delay maximum value approximate calculation is achieved in homomorphic encryption schemes such as CKKS, BGV and BFV, and the throughput rate and deployability of ciphertext reasoning are improved.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

High-precision random calculation method based on binary partial product and multiplier

The invention belongs to the field of heterogeneous approximate calculation, and relates to a high-precision random calculation method based on a binary partial product and a multiplier. The high-precision random calculation method based on the binary partial product directly uses the weight of a bit corresponding to a binary number to generate a random bit stream. Bit streams are combined in a uniform distribution mode, phase and sum are carried out bit by bit, and all phase and results are added to obtain the random calculation multiplier with the optimal precision. According to the method, binary weights are combined, a shifting and splicing mode is used for replacing addition, partial product addition is used for achieving random calculation, and an approximate parallel counter is adopted for optimization. The multiplier provided by the invention fully combines the advantages of random calculation and binary calculation, the hardware cost is less than half of that of a binary multiplier, and the delay is lower. When the provided multiplier is applied to neural network reasoning, the precision of model loss can be almost ignored.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Method and device for enhancing NR channel estimation frequency offset measurement

The invention discloses a method and a device for enhancing NR channel estimation frequency offset measurement, which relate to the technical field of communication, convert nonlinear operation into linear operation, and comprise the following steps of: performing complex conjugate multiplication and average processing on channel estimation values of adjacent pilot symbols to obtain an average complex value; mapping the complex value phase angle to a first quadrant, and determining a Taylor expansion point in a preset equally-spaced division interval; a phase angle is approximately calculated by using a first-order Taylor expansion formula, and a final phase angle is obtained by combining quadrant judgment, so that a high-precision frequency offset value is calculated. According to the method, through theoretical innovation, the calculation complexity is reduced from a nonlinear high order to a linear low order, and the required parameter storage space is remarkably reduced; experimental data show that under the same resource condition, the method can realize about 17dB frequency offset estimation precision improvement; when similar precision is achieved, required resource consumption can be reduced to 1 / 50 or below of that of a traditional method, and operation efficiency and engineering applicability are greatly improved.
Owner:BEIJING CHANGKUN TECHNOLOGY LTD

A matrix multiplication approximate calculation method based on multi-hash voting mechanism

This application relates to a matrix multiplication approximation calculation method based on a multi-hash voting mechanism. The method includes: a source computing node in the inference system obtains the token activation value matrix of the data to be inferred, which is to be sent to the target computing node where the target expert resides; the source and target computing nodes perform hash compression and voting operations on the token activation value matrix and the expert weight matrix respectively based on multiple preset hash functions to generate corresponding sketch sets; during the All-to-All communication process at the MoE layer, the source computing node sends the sketch set of the token activation value matrix to the target computing node; the target computing node performs approximate calculations based on the rows selected from the intersection of the two sketch sets to obtain the inference result. This method improves the inference efficiency of the inference system by introducing multiple independent hash functions for collaborative sampling and voting, thereby reducing computation, memory usage, network load, and GPU idle waiting time.
Owner:GREEN IND INNOVATION RES INST OF ANHUI UNIV

Unstructured data processing method, system and equipment based on online aggregation of large language model and medium

The invention relates to an unstructured data processing method, system, equipment and medium based on online aggregation of a large language model, and the method comprises the steps: carrying out the dynamic stratified sampling of a to-be-processed unstructured data field, and converting the sampling samples of each layer into a structured data stream in a unified manner in combination with the large language model; and performing online aggregation operation on the structured data stream according to the sampling sequence, gradually calculating an aggregation result and a corresponding confidence interval, and obtaining a user query result based on the aggregation result and the confidence interval. According to the method, the powerful semantic understanding and processing capacity of a large language model and the real-time advantage of an online aggregation technology are creatively fused, progressive approximate calculation of text analysis can be achieved, and the efficiency and flexibility of unstructured data processing are effectively improved. Therefore, the method can be widely applied to the fields of artificial intelligence and data processing.
Owner:RENMIN UNIVERSITY OF CHINA

Method and apparatus for approximating computation of a softmax function

The application discloses a method and device for approximating calculation of a softmax function. The method simplifies the exponential operation of e into twice constant multiplication and once exponential operation of e with the output range of (1, 2), which is mainly to limit the range of e exponent to utilize the characteristics of small range and high precision of the cordic algorithm; and simplifies N times division operation into once reciprocal operation with fixed input range, N times multiplication and shift operation. The device of the application comprises a constant multiplication unit, an exponential calculation unit, a floating point addition unit, a mantissa calculation unit, a subtraction array unit and a multiplication unit, adopts a carry-save adder instead of a traditional adder, and further shortens the critical path. The scheme of the application greatly improves the calculation speed while maintaining high precision, and reduces the consumption of calculation resources.
Owner:NANJING UNIV

Vision transformer neural network acceleration system and method based on cpu and fpga

This invention proposes a visual Transformer neural network acceleration system and method based on CPU and FPGA. The method's implementation steps are as follows: an embedding module constructs a feature matrix; a preprocessing module preprocesses the feature matrix and weight matrix; the preprocessed feature matrix and weight matrix are moved and cached; an acceleration unit constructs a self-attention matrix; the self-attention matrix is ​​cached and moved; a post-processing unit constructs a global feature matrix; and an MLP Head module obtains the classification result. The preprocessing module in the CPU reduces the additional data processing time of the FPGA acceleration unit by preprocessing the feature matrix and weight matrix, effectively improving the neural network's computation speed. Simultaneously, the normalization operation module in the acceleration unit utilizes exponential and logarithmic approximation results during approximation calculations, eliminating floating-point exponentiation and division operations, effectively reducing hardware resource consumption.
Owner:XIDIAN UNIV