Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

90results about "Operational speed enhancement" patented technology

AI accelerator construction method and related device

The invention relates to the field of artificial intelligence, in particular to an AI accelerator construction method and a related device. The method comprises the following steps: importing machine learning models of different AI frames through a multi-frame adaptation interface, and converting the machine learning models into linear algebraic dialect representation by adopting an MMIR dialect conversion chain; through target-independent optimization processing, core calculation semantics in linear algebraic dialects are reserved, and linear algebraic dialect representation is reconstructed into optimization intermediate representation with an efficient data flow structure; identifying target hardware characteristics of the target hardware; adaptively executing hardware optimization processing matched with target hardware characteristics on the optimization intermediate representation; and converting the target intermediate representation into a hardware operation unit executable by the target hardware, loading the hardware operation unit into the target hardware to obtain an AI accelerator, thereby running the hardware operation unit on the target hardware through the AI accelerator, scheduling corresponding computing resources in the target hardware, and completing hardware acceleration of the machine learning model.
Owner:ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD

Pulse neural network hardware accelerator and data processing method

The invention discloses a pulse neural network hardware accelerator and a data processing method, and the accelerator is characterized in that a low-power-consumption three-stage pipeline CPU module is used for receiving input data, scheduling an SNN network acceleration instruction, and sending the input data to an asynchronous edge SNN hardware accelerator module through a coprocessor interface; the asynchronous edge SNN hardware accelerator module comprises a pulse data encoding and decoding module, L neuromorphic kernels and an on-chip network, the pulse data encoding and decoding module encodes input data into a pulse form and sends the pulse form into the neuromorphic kernels, and the neuromorphic kernels are used for performing calculation based on the data in the pulse form; the network-on-chip is used for communication between the neuromorphic kernels, and the connection between neurons before and after synapses in the neuromorphic kernels is realized by adopting a synaptic cross array. According to the invention, the data processing acceleration performance can be greatly improved.
Owner:WUHAN UNIV +1

GIS buffer zone parameter inference method and system based on deep learning

The invention relates to the technical field of geographic information systems, in particular to a GIS buffer zone parameter inference method based on deep learning, and the method comprises the following steps: obtaining a natural language instruction input by a user; analyzing the natural language instruction to extract a geographic entity keyword and a fuzzy description word; based on the fuzzy description word, determining an initial distance parameter through a pre-trained language-space mapping model; querying a spatial database based on the geographic entity keyword to obtain geographic entity attributes; according to the geographic entity attribute, correcting the initial distance parameter to obtain a corrected distance parameter; a field programmable gate array (FPGA) is used for carrying out parallel calculation on the spatial relation matrix to generate an optimized buffer area; and outputting the optimized buffer area. According to the method, the FPGA is introduced to carry out parallel calculation of the spatial relation matrix, so that the data processing speed and the energy efficiency ratio are remarkably improved, and the real-time generation of the buffer area is realized.
Owner:MUCHENG SURVEYING & MAPPING (BEIJING) CO LTD

Data reduction method and apparatus, and device and computer-readable storage medium

The present application belongs to the technical field of communications. Disclosed are a data reduction method and apparatus, and a device and a computer-readable storage medium. The method is applied to at least one root node in a multi-dimensional direct-connect topology, wherein nodes in the multi-dimensional direct-connect topology are used for model training. The method comprises: receiving first data, wherein the first data is part of transmitted data, which is data transmitted between nodes in a model training process; obtaining second data, wherein the second data is part of first generated data, which is data generated by a root node in the model training process; performing reduction on the first data and the second data to obtain third data, wherein the third data is used by nodes in a multi-dimensional direct-connect topology to execute subsequent model training; and sending the third data to a first node in the multi-dimensional direct-connect topology, wherein the first node is a node in the multi-dimensional direct-connect topology other than the root node. The method can improve the communication efficiency in a data reduction process.
Owner:HUAWEI TECH CO LTD

Data reduction method, device and equipment and computer readable storage medium

The invention discloses a data reduction method, device and equipment and a computer readable storage medium, and belongs to the technical field of communication. The method is applied to at least one root node in a multi-dimensional direct connection topology, the nodes in the multi-dimensional direct connection topology are used for model training, and the method comprises the steps that first data are received, the first data are partial data of transmission data, and the transmission data are data transmitted among the nodes in the model training process; second data is obtained, the second data is part of first generated data, and the first generated data is data generated by the root node in the model training process; the first data and the second data are reduced to obtain third data, and the third data are used for nodes in the multi-dimensional direct connection topology to execute subsequent model training; and sending the third data to a first node in the multi-dimensional direct connection topology, wherein the first node is a node except the root node in the multi-dimensional direct connection topology. According to the method, the communication efficiency in the data reduction process can be improved.
Owner:HUAWEI TECH CO LTD

Instruction distribution device and method, processor and electronic equipment

The embodiment of the invention provides an instruction distribution device and method, a processor and electronic equipment. The instruction distribution device comprises a write selection module, a plurality of micro instruction queues and a read distribution module. The write selection module is configured to write microinstructions into a plurality of microinstruction queues; the plurality of microinstruction queues are configured to cache written microinstructions; and the read distribution module is configured to read the microinstructions from the plurality of microinstruction queues, obtain a target microinstruction group to be distributed at this time based on the read microinstructions, and execute a distribution operation. The instruction distribution device improves the bandwidth of instruction distribution by setting a plurality of microinstruction queues and a scheduling method of the plurality of microinstruction queues, thereby promoting the improvement of the performance of a processor.
Owner:HYGON INFORMATION TECH CO LTD

Method and apparatus for controlling hardware accelerator

Disclosed are a method and apparatus for controlling a hardware accelerator. The method includes setting a delay time related to task performance of a specific hardware accelerator, switching to a sleep state during the set delay time from a request time point when making a request, to a central processing unit (CPU), for requesting the specific hardware accelerator to accelerate a task, polling the state of the specific hardware accelerator to the CPU by switching from the sleep state to an operating state after the delay time, not polling a state of the specific hardware accelerator to the CPU during the delay time when switching to the sleep state, and receiving result information corresponding to the polling from the CPU in response to the polling.
Owner:MOBILINT INC

System and method for directly processing data in storage medium

The invention relates to the technical field of in-storage computing, and discloses a system and method for directly processing data in a storage medium, the storage medium is used for providing a hardware acceleration assembly specially used for data processing, and the hardware acceleration assembly comprises a mixed data loading component and a data processing component, the data directional division component divides data to be calculated and weights into different data temporary storage blocks; the task scheduling component schedules the divided data to a data calculation matrix unit; the data calculation matrix unit executes specific feature extraction operation to obtain an operation result; according to the method, an in-storage calculation architecture is adopted, a hardware acceleration component capable of efficiently carrying out parallel convolution operation is deployed, all calculation tasks are unloaded into the storage medium, feature extraction operation is executed in the storage medium, and a final result is returned. And a large amount of data movement overhead is saved.
Owner:SHENZHEN UNIV

System and architecture neural network accelerator including filter circuit

A system and an accelerator circuit includes an internal memory to store data received a memory associated with a processor and a filter circuit block comprising a plurality of circuit stripes, each circuit stripe including a filter processor, a plurality of filter circuits, and a slice of the internal memory assigned to the plurality of filter circuits, where the filter processor is to execute a filter instruction to read data values from the internal memory based on a first memory address, for each of the plurality of circuit stripes: load the data values in weight registers and input registers associated with the plurality of filter circuits of the circuit stripe to generate a plurality of filter results, and write a result generated using the plurality of filter circuits in the internal memory at a second memory address.
Owner:OPTIMUM SEMICON TECH

Computing and archiving research environment using multiple data integration

The present invention relates to a computing and archiving research environment that uses multi-data integration between different information or data with high requirements for data storage and high performance computing. The invention provides a platform for easy storage, analysis and sharing of different scientific data by providing different service combinations such as high performance computing (HPC), scientific clouds and data archiving.
Owner:DEPT OF SCI & TECH ADVANCED SCI & TECH INST

Method and system for optimizing computing performance of neural network based on on-chip storage

The invention discloses a method for optimizing the computing performance of a neural network based on on-chip storage, which comprises the following steps: firstly, carrying out SPM space division according to the access characteristics of pulse input data, internal weight and external weight so as to improve the data utilization efficiency, then, carrying out batch processing on the computing process in the time step dimension so as to enable the task scheduling to be more balanced, and then, carrying out optimization on the computing performance of the neural network. According to the method, repeated transmission is reduced through one-time loading of internal weights, storage pressure is avoided through blocking loading of external weights according to needs, and finally, a continuous and stable neural network calculation process is constructed in combination with membrane potential updating and pulse output management. The method can be widely applied to scenes sensitive to computing resources and storage bandwidth, such as embedded AI chips and edge computing equipment. According to the method, the technical problems of unreasonable SPM space division and serious access conflict existing in an existing neural network calculation performance optimization method based on a hardware accelerator can be solved.
Owner:HUNAN UNIV

Data processing apparatus, method, device and chip

This application relates to the field of data processing, and in particular to a data processing apparatus, method, device, and chip. In this apparatus, a controller module receives a task selection signal from a task request terminal and determines the target task to be executed by the task execution unit based on the task selection signal. The task execution unit selects the corresponding hardware processing module based on the target task to construct the corresponding target task path. The task execution unit receives data to be processed and inputs the data into the target task path to start the target task and obtain the task execution result. The result reading module outputs the task execution result to the task request terminal. This apparatus adapts to the computational needs of various polynomial multiplication tasks through task selection signals and dynamically combined hardware processing modules, improving the scalability of the hardware structure and increasing resource utilization.
Owner:SHANGHAI TAIZE SEMICONDUCTOR CO LTD

Model compiling method and device, computer equipment and storage medium

The invention relates to a model compiling method and device, computer equipment and a storage medium. The method comprises the following steps: performing analysis processing on a target ONNX model to obtain model feature data, and performing classification processing on each operator type according to operator capability and a constraint library to obtain a corresponding first target operator type and a corresponding second target operator type; performing rewriting processing on each second target operator type according to the rewriting strategy mode to obtain a corresponding rewritten second target operator type, and performing verification and constraint check processing on each rewritten second target operator type to obtain a corresponding check result; and according to the operators corresponding to the first target operator types in the target ONNX model and the operators corresponding to the third target operator types in the target ONNX model, generating execution subgraph fragments of the deep learning accelerators, and issuing the execution subgraph fragments to the corresponding deep learning accelerators for compiling processing. By adopting the method, the compiling efficiency can be improved.
Owner:BEIJING XIAOMA YIYI TECH CO LTD

A low-bit neural network inference method and system for microcontrollers

This invention discloses a low-bit neural network inference method and system for microcontrollers, relating to the field of artificial intelligence technology. The method involves acquiring convolutional neural network weight data and input data, and performing sub-byte quantization. Within the microcontroller firmware, an inference computation library interface is called to pass the quantized weight data and input data into the inference computation library for execution. Specifically, the quantized weight data and input data are input to the neural network function, and a neural network auxiliary function is called to preprocess them, obtaining a weight vector and an input vector. The neural network function then performs parallel convolution calculations based on the weight vector and input vector, obtaining the convolution result. Finally, the neural network auxiliary function is called to compress the convolution result, obtaining and outputting the compressed convolution result. This method effectively improves inference execution efficiency and reduces system resource consumption while maintaining essentially no reduction in model accuracy.
Owner:SHANDONG NORMAL UNIV

Data processing method and device, electronic equipment, storage medium and program product

The invention provides a data processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of large models.The method comprises the steps that an input matrix is divided into a plurality of sub-matrixes, and the number and the position index of each sub-matrix are determined based on a grouping staggering method; executing corresponding sub-matrix operation based on at least one calculation thread block in the thread block network, and transmitting an operation result of the corresponding sub-matrix based on at least one communication thread block corresponding to the at least one calculation thread block; based on the serial number sequence corresponding to each sub-matrix, storing the operation result of each sub-matrix in a cache, and based on the serial number and the position index corresponding to each sub-matrix and the operation results of the sub-matrixes sequentially stored in the cache, determining an output matrix; therefore, the communication thread blocks are arranged between the calculation thread blocks, so that the calculation thread blocks can synchronously perform sub-matrix operation in the communication process of the communication thread blocks, and the resource utilization rate of the graphics processor is improved.
Owner:INSPUR (SHANDONG) COMPUTER TECH CO LTD

Convolutional neural network software and hardware collaborative accelerator based on ARM and FPGA

The invention relates to a convolutional neural network software and hardware collaborative accelerator based on an ARM and an FPGA, which is characterized in that a convolutional neural network is disassembled according to a network layer based on the essential characteristics of mathematical operation, a convolutional layer and a full connection layer which are essentially multiplied and added operation are deployed on the FPGA, a Softmax layer which is essentially exponential and division operation is deployed on the ARM, and the convolutional neural network software and hardware collaborative accelerator based on the ARM and the FPGA is obtained. The pooling layer is deployed on the FPGA; parameters on a convolution layer, a pooling layer and a full-connection layer deployed on an FPGA are quantized into a fixed-point number format and stored in an on-chip ROM of the FPGA, and a Softmax layer deployed on an ARM adopts a floating-point number operation mode; a channel parallel strategy is adopted for a convolution layer and a pooling layer deployed on the FPGA, and convolution, pooling and full-connection hardware operation units are designed respectively.
Owner:HEBEI UNIV OF TECH

Acceleration assembly and electronic equipment

The invention relates to an acceleration assembly and electronic equipment. Wherein the acceleration unit can be included in a combined processing device, and the combined processing device can further comprise an interconnection interface and other processing devices. And the acceleration unit interacts with other processing devices to jointly complete calculation operation specified by a user. The combined processing device can further comprise a storage device, and the storage device is connected with the acceleration unit and the other processing devices and used for data services of the acceleration unit and the other processing devices. By means of the content disclosed by the invention, high-speed processing of mass data is realized.
Owner:ANHUI CAMBRICON INFORMATION TECH CO LTD

Information processing apparatus, information processing method, and program

A method includes: acquiring, from a search node configured to perform a search for a ground state represented by plural state variables included in an energy function by using plural temperature values and hold a value of the energy function for the plural state variables, a value of the energy function obtained for the plural state variables at a first temperature value among the plural temperature values; determining whether the value acquired is smaller than a smallest value of the energy function obtained for the plural state variables before reaching the first temperature value; recording update information indicating that the smallest value has been updated at the first temperature value in a case where the value is smaller than the smallest value; and outputting a second temperature value based on the first temperature value at which the update information has been recorded among the plural temperature values.
Owner:FUJITSU LTD

Fully configurable floating-point format

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets connected to the plurality of chiplet sockets. At least one of the plurality of chiplets comprises a graphics processing cluster including a plurality of processing resources. At least one processing resource of the plurality of processing resources includes a dynamic precision floating-point unit having floating-point circuitry configured to process an input in a configurable floating-point format with a variable number of exponent and mantissa bits.
Owner:INTEL CORP

Accelerator for mathematical operations based on analog computing

A computing device comprising input device terminals (x0-x4) for receiving respective analog input signals of the device; output device terminals (p0-p7) for receiving respective analog output signals of the device; rows of analog cells (C00-C01), wherein an analog cell (C00) comprises an input cell terminal and an output cell terminal, wherein an analog cell is configured to generate at the output cell terminal an output analog signal whose amplitude is the product of a multiplication coefficient by the amplitude of the input analog signal received at the input cell terminal, wherein all input terminals of the cells in a row are connected to a same input device terminal (x0-x4); a network of switches (00-33) for selectively interconnecting the output cell terminals of the analog cells and selectively connecting the output cell terminal of each of the analog cells to an output device terminal (p0-p7).
Owner:NOKIA TECHNOLOGIES OY

Techniques for providing shared memory for accelerator boards

Techniques for providing shared memory for accelerator boards include an accelerator board to receive a memory access request from an accelerator device to access a memory region with a memory controller. The request is to identify the memory region with a logical address. In addition, the accelerator board is to determine a physical address associated with the memory region from a mapping of logical addresses and associated physical addresses. In addition, the accelerator board is to route the memory access request to a memory device associated with the determined physical address.
Owner:INTEL CORP

Facilitating secure execution of external workflows for genomic sequencing diagnostics

This disclosure describes methods, non-transitory computer readable media, and systems that can facilitate execution of external workflows for diagnostic analysis of nucleotide sequencing data utilizing a container orchestration engine. For example, the disclosed systems can utilize a container orchestration engine to allow external systems (e.g., third-party systems) to generate and implement workflows for analyzing sequencing data. In executing individual workflow containers of a sequencing diagnostic workflow, the disclosed systems can isolate the workflow containers to prevent access to, or corruption of, other data while also orchestrating allocation of computing resources available at a genomic sequence processing device to execute the workflow containers.
Owner:ILLUMINA INC

Method and apparatus for implementing out-of-order pipeline execution of statically mapped workloads

The present disclosure relates to methods and apparatus for implementing statically mapped out-of-order pipeline execution of a workload. Disclosed are methods, apparatus, systems, and articles of manufacture for implementing statically mapped out-of-order pipeline execution of a workload to one or more compute building blocks of an accelerator. An example apparatus includes an interface for loading a first number of credits into a memory; a comparator for comparing the first number of credits to a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and a dispatcher for selecting a workload node in the workload to execute at a first compute building block in one or more compute building blocks when the first number of credits satisfies the threshold number of credits.
Owner:INTEL CORP

Thread allocation method, thread allocation device, and computer readable recording medium

There is provided a thread allocation method configured to generate a plurality of thread groups based on a number of active processing cores; determine a number of threads to be allocated to each thread group among a plurality of threads based on a computation capacity per each thread; allocate at least one thread to the respective thread groups based on a priority of each threads and the determined number of threads; and allocate each thread group to each active processing core.
Owner:UNIST (ULSAN NAT INST OF SCI & TECH)

Operation method of host processor and accelerator, and electronic device including the same

An operation method includes: dividing a model to be executed in an accelerator into a plurality of stages; determining, for each of the stages, a maximum batch size processible in an on-chip memory of the accelerator; determining the determined maximum batch sizes to each be a candidate batch size to be applied to the model; and determining, to be a final batch size to be applied to the model, one of the determined candidate batch sizes that minimizes a sum of a computation cost of executing the model in the accelerator and a memory access cost.
Owner:SAMSUNG ELECTRONICS CO LTD

On-chip code breakpoint debugging method, on-chip processor, and chip breakpoint debugging system

The present invention relates to an on-chip code breakpoint debugging method, an on-chip processor, and a chip breakpoint debugging system. The method comprises: the on-chip processor starts and executes an on-chip code, and an output function is set at a breakpoint position of the on-chip code; the on-chip processor obtains output information of the output function, the output information is output information of the output function when the on-chip code is executed to the output function; the on-chip processor stores the output information into an off-chip memory. In the embodiments of the present invention, according to the output information, which is stored in the off-chip memory, of the output function, the on-chip processor can obtain execution conditions of breakpoints of the on-chip code in real time, can achieve the purpose of debugging multiple breakpoints in the on-chip code at the same time, and debugging efficiency of the on-chip code is improved.
Owner:SHANGHAI CAMBRICON INFORMATION TECH CO LTD

Data online compression method and device of neural network hardware accelerator

The application discloses a data online compression method and device of a neural network hardware accelerator. The method comprises converting a first activation value output by a neural network to obtain a first activation mask; dividing the first activation mask into at least two groups of activation sub-masks, and sequentially performing accumulation processing on each group of activation sub-masks in a preset order to obtain an activation position mask; calculating an activation selection mask based on the first activation mask, the activation position mask and a weight value output by the neural network; performing screening processing on the first activation value according to the activation selection mask to obtain a target activation value, and generating a second activation mask based on the target activation value. Through the online mask setting of the activation value and the offline compression of the weight value, the adaptability to different neural network compressions is high, the data moving efficiency can be improved, the power consumption is reduced, and the throughput is ensured.
Owner:JIANGNAN INST OF COMPUTING TECH