Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

47results about "Operational speed enhancement" patented technology

Pulse neural network hardware accelerator and data processing method

PendingCN121635840AOperational speed enhancementDigital computer detailsCoprocessorAlgorithm
The invention discloses a pulse neural network hardware accelerator and a data processing method, and the accelerator is characterized in that a low-power-consumption three-stage pipeline CPU module is used for receiving input data, scheduling an SNN network acceleration instruction, and sending the input data to an asynchronous edge SNN hardware accelerator module through a coprocessor interface; the asynchronous edge SNN hardware accelerator module comprises a pulse data encoding and decoding module, L neuromorphic kernels and an on-chip network, the pulse data encoding and decoding module encodes input data into a pulse form and sends the pulse form into the neuromorphic kernels, and the neuromorphic kernels are used for performing calculation based on the data in the pulse form; the network-on-chip is used for communication between the neuromorphic kernels, and the connection between neurons before and after synapses in the neuromorphic kernels is realized by adopting a synaptic cross array. According to the invention, the data processing acceleration performance can be greatly improved.
Owner:WUHAN UNIV +1

Method and apparatus for controlling hardware accelerator

ActiveUS12591456B2Operational speed enhancementProgram initiation/switchingSleep stateEmbedded system
Disclosed are a method and apparatus for controlling a hardware accelerator. The method includes setting a delay time related to task performance of a specific hardware accelerator, switching to a sleep state during the set delay time from a request time point when making a request, to a central processing unit (CPU), for requesting the specific hardware accelerator to accelerate a task, polling the state of the specific hardware accelerator to the CPU by switching from the sleep state to an operating state after the delay time, not polling a state of the specific hardware accelerator to the CPU during the delay time when switching to the sleep state, and receiving result information corresponding to the polling from the CPU in response to the polling.
Owner:MOBILINT INC

System and architecture neural network accelerator including filter circuit

ActiveUS12632257B2Operational speed enhancementConcurrent instruction executionMemory addressInternal memory
A system and an accelerator circuit includes an internal memory to store data received a memory associated with a processor and a filter circuit block comprising a plurality of circuit stripes, each circuit stripe including a filter processor, a plurality of filter circuits, and a slice of the internal memory assigned to the plurality of filter circuits, where the filter processor is to execute a filter instruction to read data values from the internal memory based on a first memory address, for each of the plurality of circuit stripes: load the data values in weight registers and input registers associated with the plurality of filter circuits of the circuit stripe to generate a plurality of filter results, and write a result generated using the plurality of filter circuits in the internal memory at a second memory address.
Owner:OPTIMUM SEMICON TECH

Computing and archiving research environment using multiple data integration

PendingCN122070534AOperational speed enhancementResource allocationPerformance computingEngineering
The present invention relates to a computing and archiving research environment that uses multi-data integration between different information or data with high requirements for data storage and high performance computing. The invention provides a platform for easy storage, analysis and sharing of different scientific data by providing different service combinations such as high performance computing (HPC), scientific clouds and data archiving.
Owner:DEPT OF SCI & TECH ADVANCED SCI & TECH INST

Data processing apparatus, method, device and chip

ActiveCN119987855BOperational speed enhancementEnergy efficient computingComputer hardwareHardware structure
This application relates to the field of data processing, and in particular to a data processing apparatus, method, device, and chip. In this apparatus, a controller module receives a task selection signal from a task request terminal and determines the target task to be executed by the task execution unit based on the task selection signal. The task execution unit selects the corresponding hardware processing module based on the target task to construct the corresponding target task path. The task execution unit receives data to be processed and inputs the data into the target task path to start the target task and obtain the task execution result. The result reading module outputs the task execution result to the task request terminal. This apparatus adapts to the computational needs of various polynomial multiplication tasks through task selection signals and dynamically combined hardware processing modules, improving the scalability of the hardware structure and increasing resource utilization.
Owner:SHANGHAI TAIZE SEMICONDUCTOR CO LTD

Model compiling method and device, computer equipment and storage medium

PendingCN121900762AOperational speed enhancementResource allocationAlgorithmTheoretical computer science
The invention relates to a model compiling method and device, computer equipment and a storage medium. The method comprises the following steps: performing analysis processing on a target ONNX model to obtain model feature data, and performing classification processing on each operator type according to operator capability and a constraint library to obtain a corresponding first target operator type and a corresponding second target operator type; performing rewriting processing on each second target operator type according to the rewriting strategy mode to obtain a corresponding rewritten second target operator type, and performing verification and constraint check processing on each rewritten second target operator type to obtain a corresponding check result; and according to the operators corresponding to the first target operator types in the target ONNX model and the operators corresponding to the third target operator types in the target ONNX model, generating execution subgraph fragments of the deep learning accelerators, and issuing the execution subgraph fragments to the corresponding deep learning accelerators for compiling processing. By adopting the method, the compiling efficiency can be improved.
Owner:BEIJING XIAOMA YIYI TECH CO LTD

Data processing method and device, electronic equipment, storage medium and program product

ActiveCN121900974AOperational speed enhancementProgram initiation/switchingComputer hardwareGraphics
The invention provides a data processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of large models.The method comprises the steps that an input matrix is divided into a plurality of sub-matrixes, and the number and the position index of each sub-matrix are determined based on a grouping staggering method; executing corresponding sub-matrix operation based on at least one calculation thread block in the thread block network, and transmitting an operation result of the corresponding sub-matrix based on at least one communication thread block corresponding to the at least one calculation thread block; based on the serial number sequence corresponding to each sub-matrix, storing the operation result of each sub-matrix in a cache, and based on the serial number and the position index corresponding to each sub-matrix and the operation results of the sub-matrixes sequentially stored in the cache, determining an output matrix; therefore, the communication thread blocks are arranged between the calculation thread blocks, so that the calculation thread blocks can synchronously perform sub-matrix operation in the communication process of the communication thread blocks, and the resource utilization rate of the graphics processor is improved.
Owner:INSPUR (SHANDONG) COMPUTER TECH CO LTD

Convolutional neural network software and hardware collaborative accelerator based on ARM and FPGA

PendingCN121787475AOperational speed enhancementPhysical realisationComputer hardwareNerve network
The invention relates to a convolutional neural network software and hardware collaborative accelerator based on an ARM and an FPGA, which is characterized in that a convolutional neural network is disassembled according to a network layer based on the essential characteristics of mathematical operation, a convolutional layer and a full connection layer which are essentially multiplied and added operation are deployed on the FPGA, a Softmax layer which is essentially exponential and division operation is deployed on the ARM, and the convolutional neural network software and hardware collaborative accelerator based on the ARM and the FPGA is obtained. The pooling layer is deployed on the FPGA; parameters on a convolution layer, a pooling layer and a full-connection layer deployed on an FPGA are quantized into a fixed-point number format and stored in an on-chip ROM of the FPGA, and a Softmax layer deployed on an ARM adopts a floating-point number operation mode; a channel parallel strategy is adopted for a convolution layer and a pooling layer deployed on the FPGA, and convolution, pooling and full-connection hardware operation units are designed respectively.
Owner:HEBEI UNIV OF TECH

Information processing apparatus, information processing method, and program

ActiveEP3905073B1Operational speed enhancementMathematical models
A method includes: acquiring, from a search node configured to perform a search for a ground state represented by plural state variables included in an energy function by using plural temperature values and hold a value of the energy function for the plural state variables, a value of the energy function obtained for the plural state variables at a first temperature value among the plural temperature values; determining whether the value acquired is smaller than a smallest value of the energy function obtained for the plural state variables before reaching the first temperature value; recording update information indicating that the smallest value has been updated at the first temperature value in a case where the value is smaller than the smallest value; and outputting a second temperature value based on the first temperature value at which the update information has been recorded among the plural temperature values.
Owner:FUJITSU LTD

Application Programming Interface for Monitoring Resource Usage

ActiveJP7822930B2Operational speed enhancementSingle instruction multiple data multiprocessors
Apparatus, systems, and techniques for generating one or more data structures to be used to monitor use of information by a computer program. In at least one embodiment, the one or more data structures to be used to monitor use of information by a computer program are generated based on, for example, CUDA or other parallel computing platform code.
Owner:NVIDIA CORP

Data online compression method and device of neural network hardware accelerator

ActiveCN115660056BOperational speed enhancementEnergy efficient computingAlgorithmNeural network hardware
The application discloses a data online compression method and device of a neural network hardware accelerator. The method comprises converting a first activation value output by a neural network to obtain a first activation mask; dividing the first activation mask into at least two groups of activation sub-masks, and sequentially performing accumulation processing on each group of activation sub-masks in a preset order to obtain an activation position mask; calculating an activation selection mask based on the first activation mask, the activation position mask and a weight value output by the neural network; performing screening processing on the first activation value according to the activation selection mask to obtain a target activation value, and generating a second activation mask based on the target activation value. Through the online mask setting of the activation value and the offline compression of the weight value, the adaptability to different neural network compressions is high, the data moving efficiency can be improved, the power consumption is reduced, and the throughput is ensured.
Owner:JIANGNAN INST OF COMPUTING TECH

Ai accelerator construction method and related apparatus

ActiveCN120994247BOperational speed enhancementResource allocationAlgorithmTheoretical computer science
The application relates to the field of artificial intelligence, in particular to an AI accelerator construction method and related device. The method comprises the following steps: importing machine learning models of different AI frameworks through a multi-framework adaptation interface, and converting the machine learning models into linear algebra dialect representations by using an MLIR dialect conversion chain; performing target-independent optimization processing, retaining core calculation semantics in the linear algebra dialect, and reconstructing the linear algebra dialect representations into optimized intermediate representations with efficient data flow structures; identifying target hardware characteristics of a target hardware; adaptively performing hardware optimization processing matched with the target hardware characteristics on the optimized intermediate representations; converting the target intermediate representations into hardware operation units executable by the target hardware, loading the hardware operation units into the target hardware to obtain an AI accelerator, so that the hardware operation units are run on the target hardware through the AI accelerator, the corresponding calculation resources in the target hardware are scheduled, and hardware acceleration of the machine learning models is completed.
Owner:ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD

Graphics processing unit processing and caching improvements

PendingUS20260170600A1Memory architecture accessing/allocationOperational speed enhancement
Embodiments described herein are generally directed to improvements relating to power, latency, bandwidth and / or performance issues relating to GPU processing / caching. According to one embodiment, a system includes a producer intellectual property (IP) (e.g., a media IP), a compute core (e.g., a GPU or an AI-specific core of the GPU), a streaming buffer logically interposed between the producer IP and the compute core. The producer IP is operable to consume data from memory and output results to the streaming buffer. The compute core is operable to perform AI inference processing based on data consumed from the streaming buffer and output AI inference processing results to the memory.
Owner:INTEL CORP

Data processing method, data processing device, and electronic device including data processing device

ActiveJP7848950B2Operational speed enhancementResource allocation
To provide a data processing method, a data processing device and an electronic device including the data processing device, and an accelerator system.SOLUTION: A data processing method includes steps of: receiving a request to execute a model of a neural network by an accelerator; generating a plurality of candidate kernels with respect to each of a plurality of layers contained in the model; and allocating either one candidate kernel, selected based on corresponding kernel information and state information of the accelerator out of the plurality of candidate kernels with respect to the layer executed by the accelerator, to the accelerator.SELECTED DRAWING: Figure 1
Owner:SAMSUNG ELECTRONICS CO LTD

Latency and throughput centric reconfigurable storage device

ActiveCN113253916BOperational speed enhancementInput/output to record carriersComputer architectureControl store
A storage device includes a storage controller that receives data from a host device and stores the data in a storage memory, a reconfigurable integrated circuit that is communicatively connected to the storage controller and accelerates logical operations performed on the data stored in the storage memory, the reconfigurable integrated circuit including a first logic block that performs static logical operations among the logical operations, a second logic block that performs one or more dynamic logical operations among the logical operations, and a plurality of memory buffers configured to store inputs and outputs of the first logic block and the second logic block.
Owner:SAMSUNG ELECTRONICS CO LTD

Computing and archiving research environment using multiple data integration

PCT designated stageWO2026063793A1Operational speed enhancementResource allocationEngineeringResearch environment
The present invention relates to a computing and archiving research environment using multiple data integration between different information or data that have a high requirement for data storage and a high-performance computation. The present invention provides a platform for an easy storage, analysis, and sharing of different scientific data by providing different combination of services such as a high-performing computing (HPC), science cloud, and data archiving.
Owner:DEPT OF SCI & TECH ADVANCED SCI & TECH INST

Reducing latency in highly scalable hpc applications via accelerator-resident runtime management

PendingEP4430473A4Operational speed enhancementProgram initiation/switchingOperating systemComputer engineering
Methods and systems for runtime management by an accelerator-resident manager. Techniques include receiving, by the manager, a representation of a processing flow of an application, including a plurality of kernels and respective dependencies. The manager, then, assigns the plurality of kernels to one or more APUs managed it and launches the plurality of kernels on their assigned APUs to run in an iteration according to the respective dependencies.
Owner:ADVANCED MICRO DEVICES INC

Data processing method and apparatus

ActiveCN114386560BOperational speed enhancementKernel methodsAlgorithmTheoretical computer science
Data processing methods and apparatus are disclosed. A processor-implemented data processing method includes receiving a request for executing a neural network model on an accelerator, generating a plurality of candidate kernels for each of a plurality of layers included in the model, and assigning a single candidate kernel to the accelerator, the single candidate kernel being selected from the plurality of candidate kernels generated for a layer to be run on the accelerator based on corresponding kernel information and state information of the accelerator.
Owner:SAMSUNG ELECTRONICS CO LTD

Memory expander, heterogeneous computing device using memory expander, and operation method of heterogenous computing

ActiveUS12572300B2Operational speed enhancementInput/output to record carriersEngineeringController (computing)
A memory expander includes a memory device that stores a plurality of task data. A controller controls the memory device. The controller receives metadata and a management request from an external central processing unit (CPU) through a compute express link (CXL) interface and operates in a management mode in response to the management request. In the management mode, the controller receives a read request and a first address from an accelerator through the CXL interface and transmits one of the plurality of task data to the accelerator based on the metadata in response to the read request.
Owner:SAMSUNG ELECTRONICS CO LTD

Neural network intelligent power quality data interpolation method and device based on NPU acceleration

The invention relates to the technical field of artificial intelligence, in particular to a neural network intelligent power quality data interpolation method and device based on NPU acceleration, and the method comprises the steps: obtaining a power quality data interpolation network model obtained through the training of a power quality training data set; performing quantization and format conversion on the trained power quality data interpolation network model, and then deploying the power quality data interpolation network model to a neural network processing unit (NPU); inputting the collected real-time electric energy quality data into the NPU so as to carry out electric energy quality data interpolation calculation through an electric energy quality data interpolation network model to obtain an electric energy quality data interpolation result; and outputting an electric energy quality data interpolation result to a central processing unit (CPU). Therefore, by inputting the data, the NPU quickly completes power quality data interpolation and outputs the result to the CPU, so that the system can efficiently execute various tasks, and computing resource conflicts and performance bottlenecks are effectively avoided.
Owner:SHENZHEN RENERGY TECH

Vector Extraction and Merge Instruction

PendingJP2025522516A5Operational speed enhancementRegister arrangements
Apparatus, method, and medium are provided. The apparatus includes a decoder circuit that generates a control signal in response to a vector extraction and merge instruction that specifies control parameters, a first vector register, a second vector register, and a destination vector register. The apparatus includes a processing circuit that executes processing of a plurality of beats in response to the control signal, and each beat includes processing corresponding to at least a part of the first vector register and the destination vector register. The processing for the K 番目 beats includes extracting bits specified by the control parameters from the K 番目 portion of the first vector register, concatenating the bits with further bits, and storing the result in the K 番目 portion of the destination register. The further bits are, for the first portion, extracted from the first portion of the second vector register, and otherwise, from the (K - 1) 番目 portion of the first vector register.
Owner:ARM LTD

MATRIX-MULTIPLIKATIONSMASCHINE

PendingDE102025133499A1Operational speed enhancementComplex mathematical operationsOperandMechanical engineering
A matrix multiplication engine can include a first operand buffer and a second operand buffer, each capable of storing multiple operand elements arranged in rows and columns. A cell array can be formed from cells, each cell comprising a memory and accumulator circuit to receive operand elements column-wise from both the first and second operand buffers, to compute a dot product of the received operand elements, and to accumulate the dot product in a corresponding tile state element in memory. Matrix elements of the operand matrices to be multiplied can be loaded row-wise into rows of the operand buffers and read column-wise into the cells. The number of elements for which a dot product is calculated can be selected depending on the operand element width.
Owner:SIFIVE INC

Large model reasoning acceleration processing method and device, electronic equipment and storage medium

PendingCN121860068AOperational speed enhancementBiological modelsConcurrent computationAlgorithm
The invention provides a large model reasoning acceleration processing method and device, electronic equipment and a storage medium, and belongs to the technical field of large model computation.The method comprises the steps that a query sequence is divided into a plurality of query groups, and a parallel computing unit is allocated to each query group, pre-loading the required key value cache in the local storage space of the parallel computing unit; in each parallel computing unit, a plurality of thread bundles are used for carrying out parallel computing on the query of the current token and the dot product of all historical keys, and a similarity vector is generated; performing normalization processing on each similarity vector, and outputting a plurality of attention weights; performing weighted summation on the attention weights and values to obtain a weighted sum, and temporarily storing the weighted sum in a local storage space of the parallel computing unit; under the condition that all the parallel computing units complete computing, the weighted sum of all the parallel computing units is merged, and a complete output result is written back into the global storage space. The large model reasoning delay is reduced, and the throughput and the bandwidth utilization rate are improved.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Apparatus, NPU and chipset implemented for fusion neural network

ActiveUS12594965B2Operational speed enhancementDigital data processing detailsSpecial function unitChipset
A neural processing unit (NPU) includes a controller including a scheduler, the controller configured to receive from a compiler a machine code of an artificial neural network (ANN) including a fusion ANN, the machine code including data locality information of the fusion ANN, and receive heterogeneous sensor data from a plurality of sensors corresponding to the fusion ANN; at least one processing element configured to perform fusion operations of the fusion ANN including a convolution operation and at least one special function operation; a special function unit (SFU) configured to perform a special function operation of the fusion ANN; and an on-chip memory configured to store operation data of the fusion ANN, wherein the scheduler is configured to control the at least one processing element and the on-chip memory such that all operations of the fusion ANN are processed in a predetermined sequence according to the data locality information.
Owner:DEEPX CO LTD

The method implemented by the first computing node, the first computing node, and the readable medium

ActiveCN115335804BOperational speed enhancementTransmissionData packEngineering
In distributed training, to avoid network congestion, a first computing node can determine, according to a node-aware halving doubling algorithm, a rendezvous identifier for sending a data packet from a first process to a second process, the first process and the second process belonging to different nodes connected to different tile switches under a particular network topology. The first computing node can then send the data packet from the first process to the second process through a rendezvous switch corresponding to the rendezvous identifier.
Owner:ALIBABA GROUP HOLDING LTD

Stacked dies for machine learning accelerator

An apparatus is disclosed. The apparatus includes a machine learning die including a memory and one or more machine learning accelerators; and a processing core die stacked with the machine learning die, the processing core die configured to execute a shader program to control operations on the machine learning die, wherein the memory is configurable as either or both of a cache and directly accessible memory.
Owner:ADVANCED MICRO DEVICES INC

Apparatus and system for processing fusion of heterogeneous data including synchronized infrared images

PendingUS20260116431A1Operational speed enhancementDigital data processing detailsSequence processingAlgorithm
A neural processing unit (NPU) includes a controller including a scheduler, the controller configured to receive from a compiler a machine code of an artificial neural network (ANN) including a fusion ANN, the machine code including data locality information of the fusion ANN, and receive heterogeneous sensor data from a plurality of sensors corresponding to the fusion ANN; at least one processing element configured to perform fusion operations of the fusion ANN including a convolution operation and at least one special function operation; a special function unit (SFU) configured to perform a special function operation of the fusion ANN; and an on-chip memory configured to store operation data of the fusion ANN, wherein the schedular is configured to control the at least one processing element and the on-chip memory such that all operations of the fusion ANN are processed in a predetermined sequence according to the data locality information.
Owner:DEEPX CO LTD