Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2 results about "Linear operators" patented technology

Computing systems, data processing methods, apparatus, and media for high-bandwidth inference

PendingCN122263994AReduce computing loadReduce handling bandwidth requirementsDigital storageInference methodsIntegrated circuitNonlinear operators
The disclosure provides a computing system, a data processing method, equipment and a medium for high-bandwidth inference, and relates to the technical field of integrated circuits. The computing system is used for performing a decoding stage of a Transformer-based model inference, and comprises a host processor for performing a decoding stage nonlinear operator, an offload subsystem for performing at least a part of a decoding stage linear operator, and a standard high-speed interface module for transmitting an input activation vector and an output result vector between the host processor and the offload subsystem. The offload subsystem comprises a weight lock storage array for storing a weight matrix in a static residence manner, an input vector streaming interface for streamingly receiving the input activation vector, a matrix vector multiplication calculation unit for performing a matrix vector multiplication operation on the input activation vector and the weight matrix, and a result processing module for reducing or arranging the operation result to obtain the output result vector.
Owner:ICY TECHNOLOGY (BEIJING) CO LTD

Conversion method of artificial neural network model, storage medium, and program product

The embodiment of the application provides a conversion method, a storage medium and a program product of an artificial neural network model, relates to the technical field of artificial intelligence, and the method comprises the following steps: after obtaining a pre-trained artificial neural network model, converting each nonlinear operator in the artificial neural network model into a corresponding pulse module, wherein the pulse module comprises a differential expectation compensation module, the differential expectation compensation module is used for calculating an output increment according to cumulative membrane potential, a differential pulse neuron is inserted in each pulse module, the differential pulse neuron updates an encoding activation value when a pulse is emitted, otherwise the encoding activation value remains unchanged, a bias term of a linear operator located in a previous layer of each nonlinear operator is removed, and an initial membrane potential of the differential pulse neuron inserted in the pulse module corresponding to the nonlinear operator is set as the bias term, and the encoding activation value is only updated when a pulse is emitted, so that the loss caused by updating the encoding activation value regardless of whether a pulse is emitted or not is avoided, and energy consumption is significantly reduced.
Owner:PEKING UNIV