Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

122results about "Operational speed enhancement" patented technology

Cluster-oriented large model parallel method and device and electronic device

The invention relates to a cluster-oriented large model parallelization method and device and an electronic device.The method is applied to the field of large models.The method comprises the steps that operator information of a cluster-oriented large model in a preset microprocessing batch and a preset parallelization mode is obtained, the operator information comprises operator time information and operator memory information of operators in the cluster-oriented large model; the cluster comprises one or more types of accelerators; determining an initial operator parallel configuration strategy of a plurality of assembly lines in the large model based on the operator information, model memory information required by the large model and a memory extreme value of an accelerator; and performing recursion processing on the initial operator parallel configuration strategy according to a preset load balancing mode to obtain a target operator parallel configuration strategy, and running the cluster-oriented large model based on the target operator parallel configuration strategy. According to the method and the device, the utilization efficiency of chip calculation performance during large model parallel configuration is improved, and high efficiency and wide application range of large model training are realized.
Owner:ZHEJIANG LAB

AI accelerator construction method and related device

The invention relates to the field of artificial intelligence, in particular to an AI accelerator construction method and a related device. The method comprises the following steps: importing machine learning models of different AI frames through a multi-frame adaptation interface, and converting the machine learning models into linear algebraic dialect representation by adopting an MMIR dialect conversion chain; through target-independent optimization processing, core calculation semantics in linear algebraic dialects are reserved, and linear algebraic dialect representation is reconstructed into optimization intermediate representation with an efficient data flow structure; identifying target hardware characteristics of the target hardware; adaptively executing hardware optimization processing matched with target hardware characteristics on the optimization intermediate representation; and converting the target intermediate representation into a hardware operation unit executable by the target hardware, loading the hardware operation unit into the target hardware to obtain an AI accelerator, thereby running the hardware operation unit on the target hardware through the AI accelerator, scheduling corresponding computing resources in the target hardware, and completing hardware acceleration of the machine learning model.
Owner:ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD

Pulse neural network hardware accelerator and data processing method

The invention discloses a pulse neural network hardware accelerator and a data processing method, and the accelerator is characterized in that a low-power-consumption three-stage pipeline CPU module is used for receiving input data, scheduling an SNN network acceleration instruction, and sending the input data to an asynchronous edge SNN hardware accelerator module through a coprocessor interface; the asynchronous edge SNN hardware accelerator module comprises a pulse data encoding and decoding module, L neuromorphic kernels and an on-chip network, the pulse data encoding and decoding module encodes input data into a pulse form and sends the pulse form into the neuromorphic kernels, and the neuromorphic kernels are used for performing calculation based on the data in the pulse form; the network-on-chip is used for communication between the neuromorphic kernels, and the connection between neurons before and after synapses in the neuromorphic kernels is realized by adopting a synaptic cross array. According to the invention, the data processing acceleration performance can be greatly improved.
Owner:WUHAN UNIV +1

GIS buffer zone parameter inference method and system based on deep learning

The invention relates to the technical field of geographic information systems, in particular to a GIS buffer zone parameter inference method based on deep learning, and the method comprises the following steps: obtaining a natural language instruction input by a user; analyzing the natural language instruction to extract a geographic entity keyword and a fuzzy description word; based on the fuzzy description word, determining an initial distance parameter through a pre-trained language-space mapping model; querying a spatial database based on the geographic entity keyword to obtain geographic entity attributes; according to the geographic entity attribute, correcting the initial distance parameter to obtain a corrected distance parameter; a field programmable gate array (FPGA) is used for carrying out parallel calculation on the spatial relation matrix to generate an optimized buffer area; and outputting the optimized buffer area. According to the method, the FPGA is introduced to carry out parallel calculation of the spatial relation matrix, so that the data processing speed and the energy efficiency ratio are remarkably improved, and the real-time generation of the buffer area is realized.
Owner:MUCHENG SURVEYING & MAPPING (BEIJING) CO LTD

Data reduction method and apparatus, and device and computer-readable storage medium

The present application belongs to the technical field of communications. Disclosed are a data reduction method and apparatus, and a device and a computer-readable storage medium. The method is applied to at least one root node in a multi-dimensional direct-connect topology, wherein nodes in the multi-dimensional direct-connect topology are used for model training. The method comprises: receiving first data, wherein the first data is part of transmitted data, which is data transmitted between nodes in a model training process; obtaining second data, wherein the second data is part of first generated data, which is data generated by a root node in the model training process; performing reduction on the first data and the second data to obtain third data, wherein the third data is used by nodes in a multi-dimensional direct-connect topology to execute subsequent model training; and sending the third data to a first node in the multi-dimensional direct-connect topology, wherein the first node is a node in the multi-dimensional direct-connect topology other than the root node. The method can improve the communication efficiency in a data reduction process.
Owner:HUAWEI TECH CO LTD

Systolic arithmetic on sparse data

Embodiments described herein provided for an instruction and associated logic to enable a processing resource including a tensor accelerator to perform optimized computation of sparse submatrix operations. One embodiment provides a parallel processor comprising a processing cluster coupled with the cache memory. The processing cluster includes a plurality of multiprocessors coupled with a data interconnect, where a multiprocessor of the plurality of multiprocessors includes a tensor core configured to load tensor data and metadata associated with the tensor data from the cache memory, wherein the metadata indicates a first numerical transform applied to the tensor data, perform an inverse transform of the first numerical transform, perform a tensor operation on the tensor data after the inverse transform is performed, and write output of the tensor operation to a memory coupled with the processing cluster.
Owner:INTEL CORP

Data reduction method, device and equipment and computer readable storage medium

The invention discloses a data reduction method, device and equipment and a computer readable storage medium, and belongs to the technical field of communication. The method is applied to at least one root node in a multi-dimensional direct connection topology, the nodes in the multi-dimensional direct connection topology are used for model training, and the method comprises the steps that first data are received, the first data are partial data of transmission data, and the transmission data are data transmitted among the nodes in the model training process; second data is obtained, the second data is part of first generated data, and the first generated data is data generated by the root node in the model training process; the first data and the second data are reduced to obtain third data, and the third data are used for nodes in the multi-dimensional direct connection topology to execute subsequent model training; and sending the third data to a first node in the multi-dimensional direct connection topology, wherein the first node is a node except the root node in the multi-dimensional direct connection topology. According to the method, the communication efficiency in the data reduction process can be improved.
Owner:HUAWEI TECH CO LTD

Instruction distribution device and method, processor and electronic equipment

The embodiment of the invention provides an instruction distribution device and method, a processor and electronic equipment. The instruction distribution device comprises a write selection module, a plurality of micro instruction queues and a read distribution module. The write selection module is configured to write microinstructions into a plurality of microinstruction queues; the plurality of microinstruction queues are configured to cache written microinstructions; and the read distribution module is configured to read the microinstructions from the plurality of microinstruction queues, obtain a target microinstruction group to be distributed at this time based on the read microinstructions, and execute a distribution operation. The instruction distribution device improves the bandwidth of instruction distribution by setting a plurality of microinstruction queues and a scheduling method of the plurality of microinstruction queues, thereby promoting the improvement of the performance of a processor.
Owner:HYGON INFORMATION TECH CO LTD

Load balancing system for the execution of applications on reconfigurable processors

A data processing system is presented in a client-server configuration for executing first and second applications that a client in the client-server configuration can offload for execution onto the data processing system. The data processing system includes a server and a pool of reconfigurable data flow resources that is configured to execute the first application in a first runtime context and the second application in a second runtime context. The server is configured to establish a session with the client, receive first and second execution requests for executing the first application and the second application from the client, start respective first and second execution of the first and second applications in the respective first and second runtime contexts in response to receiving the first and second execution requests, and balance a first load from the first execution with a second load from the second execution.
Owner:SAMBANOVA SYSTEMS INC

Method, apparatus, device, medium and product for accelerating multiplication of double sparse matrices

The present invention discloses a method, apparatus, device, medium and product for accelerating the multiplication of double sparse matrices. The method includes: locating a set of compressed matrices matching a first sparse matrix and a second sparse matrix in the global memory of a computing chip according to the sparse matrix multiplication requirement; sequentially transferring the set of compressed matrices and the second sparse matrix from the global memory to the hardware registers of the computing chip in the form of data blocks; and gradually calculating the multiplication result of the first sparse matrix and the second sparse matrix by a sparse computing unit of the computing chip according to the data loaded in batches in the hardware registers. The technical solution of the embodiments of the present invention can give full play to the hardware acceleration performance of the sparse computing unit in the computing chip, and thus greatly optimize the overhead of computing, bandwidth and storage resources in the process of multiplying double sparse matrices, and is particularly applicable to the model calculation scenario of a mixture-of-experts model.
Owner:SHANGHAI SUIYUAN TECH CO LTD

Method and apparatus for controlling hardware accelerator

Disclosed are a method and apparatus for controlling a hardware accelerator. The method includes setting a delay time related to task performance of a specific hardware accelerator, switching to a sleep state during the set delay time from a request time point when making a request, to a central processing unit (CPU), for requesting the specific hardware accelerator to accelerate a task, polling the state of the specific hardware accelerator to the CPU by switching from the sleep state to an operating state after the delay time, not polling a state of the specific hardware accelerator to the CPU during the delay time when switching to the sleep state, and receiving result information corresponding to the polling from the CPU in response to the polling.
Owner:MOBILINT INC

Computational storage device and method of operating the same

An operating method of a computational storage device includes: setting a first computing namespace, including a first queue and a first accelerator and having a first value as its first ID, per instructions from a first host; setting a second computing namespace, including a second queue and a second accelerator and having a second value as its first ID, per instructions from a second host; loading a first program from the first host in the first computing namespace; loading a second program from the second host in the second computing namespace; setting a second ID of the first computing namespace to a third value based on an ID of the first program per instructions to activate the first program; and setting the second ID of the second computing namespace to a fourth value based on an ID of the second program per instructions to activate the second program.
Owner:SAMSUNG ELECTRONICS CO LTD

chip

A chip is disclosed, belonging to the field of electronic technology. The chip includes a dedicated scheduler, a general scheduler, and a plurality of hardware accelerators. The plurality of hardware accelerators are connected by soft connections, at least one hardware accelerator is connected to the dedicated scheduler, and at least one hardware accelerator is connected to the general scheduler. The processing flow of the dedicated scheduler is relatively fixed and the processing efficiency is relatively high. The processing flow of the general scheduler is relatively flexible and can overcome hardware design defects and improve performance through software programming. Therefore, it can provide processing flexibility while ensuring the processing efficiency of the chip.
Owner:HUAWEI TECH CO LTD

System and method for directly processing data in storage medium

The invention relates to the technical field of in-storage computing, and discloses a system and method for directly processing data in a storage medium, the storage medium is used for providing a hardware acceleration assembly specially used for data processing, and the hardware acceleration assembly comprises a mixed data loading component and a data processing component, the data directional division component divides data to be calculated and weights into different data temporary storage blocks; the task scheduling component schedules the divided data to a data calculation matrix unit; the data calculation matrix unit executes specific feature extraction operation to obtain an operation result; according to the method, an in-storage calculation architecture is adopted, a hardware acceleration component capable of efficiently carrying out parallel convolution operation is deployed, all calculation tasks are unloaded into the storage medium, feature extraction operation is executed in the storage medium, and a final result is returned. And a large amount of data movement overhead is saved.
Owner:SHENZHEN UNIV

System and architecture neural network accelerator including filter circuit

A system and an accelerator circuit includes an internal memory to store data received a memory associated with a processor and a filter circuit block comprising a plurality of circuit stripes, each circuit stripe including a filter processor, a plurality of filter circuits, and a slice of the internal memory assigned to the plurality of filter circuits, where the filter processor is to execute a filter instruction to read data values from the internal memory based on a first memory address, for each of the plurality of circuit stripes: load the data values in weight registers and input registers associated with the plurality of filter circuits of the circuit stripe to generate a plurality of filter results, and write a result generated using the plurality of filter circuits in the internal memory at a second memory address.
Owner:OPTIMUM SEMICON TECH

Computing and archiving research environment using multiple data integration

The present invention relates to a computing and archiving research environment that uses multi-data integration between different information or data with high requirements for data storage and high performance computing. The invention provides a platform for easy storage, analysis and sharing of different scientific data by providing different service combinations such as high performance computing (HPC), scientific clouds and data archiving.
Owner:DEPT OF SCI & TECH ADVANCED SCI & TECH INST

On-chip code breakpoint debugging method, on-chip processor, and chip breakpoint debugging system

The present invention relates to an on-chip code breakpoint debugging method, an on-chip processor, and a chip breakpoint debugging system. The method comprises: the on-chip processor starts and executes an on-chip code, and an output function is set at a breakpoint position of the on-chip code; the on-chip processor obtains output information of the output function, the output information is output information of the output function when the on-chip code is executed to the output function; the on-chip processor stores the output information into an off-chip memory. In the embodiments of the present invention, according to the output information, which is stored in the off-chip memory, of the output function, the on-chip processor can obtain execution conditions of breakpoints of the on-chip code in real time, can achieve the purpose of debugging multiple breakpoints in the on-chip code at the same time, and debugging efficiency of the on-chip code is improved.
Owner:SHANGHAI CAMBRICON INFORMATION TECH CO LTD

Electronic apparatus and control method thereof

An electronic apparatus may include a memory, and a processor configured to perform an operation between a first index data and a second index data corresponding to a first data set and a second data set, respectively, to acquire first output data; identify a sparsity ratio of the first output data; based on the sparsity ratio exceeding the threshold value, determine a first method of processing the first data set and the second data set using the first data set and the second data set, as an operation method of the electronic apparatus; and based on the sparsity ratio being equal to or less than the threshold value, identify valid data included in the first index data and the second index data, and determine a second method of processing the first data set and the second data set using the valid data, as the operation method of the electronic apparatus.
Owner:SAMSUNG ELECTRONICS CO LTD

Method and system for optimizing computing performance of neural network based on on-chip storage

The invention discloses a method for optimizing the computing performance of a neural network based on on-chip storage, which comprises the following steps: firstly, carrying out SPM space division according to the access characteristics of pulse input data, internal weight and external weight so as to improve the data utilization efficiency, then, carrying out batch processing on the computing process in the time step dimension so as to enable the task scheduling to be more balanced, and then, carrying out optimization on the computing performance of the neural network. According to the method, repeated transmission is reduced through one-time loading of internal weights, storage pressure is avoided through blocking loading of external weights according to needs, and finally, a continuous and stable neural network calculation process is constructed in combination with membrane potential updating and pulse output management. The method can be widely applied to scenes sensitive to computing resources and storage bandwidth, such as embedded AI chips and edge computing equipment. According to the method, the technical problems of unreasonable SPM space division and serious access conflict existing in an existing neural network calculation performance optimization method based on a hardware accelerator can be solved.
Owner:HUNAN UNIV

Data processing apparatus, method, device and chip

This application relates to the field of data processing, and in particular to a data processing apparatus, method, device, and chip. In this apparatus, a controller module receives a task selection signal from a task request terminal and determines the target task to be executed by the task execution unit based on the task selection signal. The task execution unit selects the corresponding hardware processing module based on the target task to construct the corresponding target task path. The task execution unit receives data to be processed and inputs the data into the target task path to start the target task and obtain the task execution result. The result reading module outputs the task execution result to the task request terminal. This apparatus adapts to the computational needs of various polynomial multiplication tasks through task selection signals and dynamically combined hardware processing modules, improving the scalability of the hardware structure and increasing resource utilization.
Owner:SHANGHAI TAIZE SEMICONDUCTOR CO LTD

Model compiling method and device, computer equipment and storage medium

The invention relates to a model compiling method and device, computer equipment and a storage medium. The method comprises the following steps: performing analysis processing on a target ONNX model to obtain model feature data, and performing classification processing on each operator type according to operator capability and a constraint library to obtain a corresponding first target operator type and a corresponding second target operator type; performing rewriting processing on each second target operator type according to the rewriting strategy mode to obtain a corresponding rewritten second target operator type, and performing verification and constraint check processing on each rewritten second target operator type to obtain a corresponding check result; and according to the operators corresponding to the first target operator types in the target ONNX model and the operators corresponding to the third target operator types in the target ONNX model, generating execution subgraph fragments of the deep learning accelerators, and issuing the execution subgraph fragments to the corresponding deep learning accelerators for compiling processing. By adopting the method, the compiling efficiency can be improved.
Owner:BEIJING XIAOMA YIYI TECH CO LTD

Convolution Input Data Determination, Convolution Operation, Image Denoising, Method and System

The present application discloses a method and system for determining convolution input data, performing convolution operations, image denoising, an accelerator, and an image processing chip. The method for determining convolution input data includes: determining convolution operation input data corresponding to a current operation result; storing k-s rows of convolution results before the current operation result into BT; and storing k-s columns of convolution results before the current operation result in the BT into BL, so as to determine convolution input data for the next layer of convolution operation according to the BT and the BL. The present application can avoid repeated calculations of the neural network depth level when calculating pixel outputs at different positions of the output layer, can implement on-chip data caching of neural network data, avoid frequent off-chip data interactions when obtaining upper-layer intermediate results, improve the operation efficiency of the accelerator, and reduce operation power consumption.
Owner:上海为旌科技有限公司

A low-bit neural network inference method and system for microcontrollers

This invention discloses a low-bit neural network inference method and system for microcontrollers, relating to the field of artificial intelligence technology. The method involves acquiring convolutional neural network weight data and input data, and performing sub-byte quantization. Within the microcontroller firmware, an inference computation library interface is called to pass the quantized weight data and input data into the inference computation library for execution. Specifically, the quantized weight data and input data are input to the neural network function, and a neural network auxiliary function is called to preprocess them, obtaining a weight vector and an input vector. The neural network function then performs parallel convolution calculations based on the weight vector and input vector, obtaining the convolution result. Finally, the neural network auxiliary function is called to compress the convolution result, obtaining and outputting the compressed convolution result. This method effectively improves inference execution efficiency and reduces system resource consumption while maintaining essentially no reduction in model accuracy.
Owner:SHANDONG NORMAL UNIV

Data processing method and device, electronic equipment, storage medium and program product

The invention provides a data processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of large models.The method comprises the steps that an input matrix is divided into a plurality of sub-matrixes, and the number and the position index of each sub-matrix are determined based on a grouping staggering method; executing corresponding sub-matrix operation based on at least one calculation thread block in the thread block network, and transmitting an operation result of the corresponding sub-matrix based on at least one communication thread block corresponding to the at least one calculation thread block; based on the serial number sequence corresponding to each sub-matrix, storing the operation result of each sub-matrix in a cache, and based on the serial number and the position index corresponding to each sub-matrix and the operation results of the sub-matrixes sequentially stored in the cache, determining an output matrix; therefore, the communication thread blocks are arranged between the calculation thread blocks, so that the calculation thread blocks can synchronously perform sub-matrix operation in the communication process of the communication thread blocks, and the resource utilization rate of the graphics processor is improved.
Owner:INSPUR (SHANDONG) COMPUTER TECH CO LTD

Convolutional neural network software and hardware collaborative accelerator based on ARM and FPGA

The invention relates to a convolutional neural network software and hardware collaborative accelerator based on an ARM and an FPGA, which is characterized in that a convolutional neural network is disassembled according to a network layer based on the essential characteristics of mathematical operation, a convolutional layer and a full connection layer which are essentially multiplied and added operation are deployed on the FPGA, a Softmax layer which is essentially exponential and division operation is deployed on the ARM, and the convolutional neural network software and hardware collaborative accelerator based on the ARM and the FPGA is obtained. The pooling layer is deployed on the FPGA; parameters on a convolution layer, a pooling layer and a full-connection layer deployed on an FPGA are quantized into a fixed-point number format and stored in an on-chip ROM of the FPGA, and a Softmax layer deployed on an ARM adopts a floating-point number operation mode; a channel parallel strategy is adopted for a convolution layer and a pooling layer deployed on the FPGA, and convolution, pooling and full-connection hardware operation units are designed respectively.
Owner:HEBEI UNIV OF TECH

Acceleration assembly and electronic equipment

The invention relates to an acceleration assembly and electronic equipment. Wherein the acceleration unit can be included in a combined processing device, and the combined processing device can further comprise an interconnection interface and other processing devices. And the acceleration unit interacts with other processing devices to jointly complete calculation operation specified by a user. The combined processing device can further comprise a storage device, and the storage device is connected with the acceleration unit and the other processing devices and used for data services of the acceleration unit and the other processing devices. By means of the content disclosed by the invention, high-speed processing of mass data is realized.
Owner:ANHUI CAMBRICON INFORMATION TECH CO LTD