Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

521results about "Architecture with multiple processing units" patented technology

Programmable in-memory computing accelerator for low-precision deep neural network inference

A programmable in-memory computing (IMC) accelerator for low-precision deep neural network inference, also referred to as PIMCA, is provided. Embodiments of the PIMCA integrate a large number of capacitive-coupling-based IMC static random-access memory (SRAM) macros and demonstrate large-scale integration of IMC SRAM macros. For example, a 28 nm prototype integrates 108 capacitive-coupling-based IMC SRAM macros of a total size of 3.4 megabytes (Mb), demonstrating one of the largest IMC hardware to date. In addition, a custom instruction set architecture (ISA) is developed featuring IMC and single-instruction-multiple-data (SIMD) functional units with hardware loop to support a range of deep neural network (DNN) layer types. The 28 nm prototype chip achieves a peak throughput of 4.9 tera operations per second (TOPS) and system-level peak energy-efficiency of 437 TOPS per watt (TOPS / W) at 40 megahertz (MHz) with a 1 volt (V) supply.
Owner:THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK +1

Caching method and system for access unit of superscalar processor

The invention belongs to the field of integrated circuits and computer system structures, and provides a caching method and system for a memory access unit of a superscalar processor, and the method comprises the steps: receiving a plurality of memory access instructions in the same period, and determining a corresponding Bank in a to-be-accessed cache through the memory access instructions; after the memory access instruction obtains a cache access permission, if cache line missing occurs, generating a missing request, merging all the missing requests, and performing parallel prefetching training on the merged missing requests by utilizing a mode of fusing a constant step length prefetching mode and a complex step length prefetching mode to obtain a prefetching request and a prefetching cache address corresponding to the prefetching request; requesting a missing cache line from the first-level cache to the second-level cache based on the missing queue, and writing the missing cache line back to the cache line of the corresponding data cache in the first-level cache; and storing the bus consistency request by using the sniffing queue, judging whether the data in the multi-core cache are consistent or not by using the consistency request, and performing consistency modification according to a judgment result. The cache hit rate and the bandwidth utilization rate are improved.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Large language model reasoning system and method based on multi-chip parallel computing

The invention provides an inference system and method of a large language model based on multi-chip parallel computing, and relates to the technical field of artificial intelligence. The system comprises a pre-calculation module used for processing input instruction information to generate to-be-reasoned data, and the to-be-reasoned data is in a matrix form; the expert parallel module is used for sending the to-be-reasoned data to accelerator chips in the expert parallel module and determining sub-reasoning data processed by the activation expert units corresponding to the accelerator chips respectively, so that the activation expert units carry out calculation based on the corresponding sub-reasoning data and complete parallel calculation result data is determined. The input data is broadcasted to all the accelerator chips, each accelerator chip selects the corresponding input data for calculation according to the set activation expert unit, the same complete calculation result is obtained through global protocol operation among all the accelerator chips, and the overall operation performance and efficiency are improved.
Owner:SHENZHEN CORERAIN TECH CO LTD

SYSTEMS AND METHODS IN THE FIELD OF SELF-SUPERVISED DETECTION OF FACIAL FLAGSHIPS

SYSTEMS AND METHODS IN THE FIELD OF SELF-SUPERVISED DETECTION OF FACIAL BOUNDARIES. Systems and methods for self-supervised learning (SSL) in facial detection networks are proposed. In one embodiment, a facial detection network comprises encoder components configured to encode facial features, the encoder components including trained components of a masked image modeling (MIM) network configured to process non-overlapping patches determined from the input image, the MIM network trained with an SSL objective; and decoder components configured by training to determine local matches between features to determine estimates for facial landmarks. In one embodiment, the MIM network is an MAE network.In one embodiment, the decoder components are derived from those of a second trained network comprising the encoder components as trained but fixed, wherein the decoder components of the second network are trained using locality constraint repulsion loss (LCR). Methods are proposed for SSL training of the encoder and decoder components. Figure for abstract: none.
Owner:LOREAL SA

Methods and circuits for streaming data to processing elements in stacked processor-plus-memory architecture

A stacked processor-plus-memory device includes a processing die with an array of processing elements of an artificial neural network. Each processing element multiplies a first operand—e.g. a weight—by a second operand to produce a partial result to a subsequent processing element. To prepare for these computations, a sequencer loads the weights into the processing elements as a sequence of operands that step through the processing elements, each operand stored in the corresponding processing element. The operands can be sequenced directly from memory to the processing elements or can be stored first in cache. The processing elements include streaming logic that disregards interruptions in the stream of operands.
Owner:RAMBUS INC

Multi-core processor chip and storage access method and device of multi-core processor chip

The invention provides a multi-core processor chip and a storage access method and device of the multi-core processor chip, and relates to the technical field of computer processors, the multi-core processor chip comprises at least one processor core, the processor core comprises a general processor core, a first address space mapping module and a first core interconnection interface, the universal processor core supports high-speed cache consistency; at least one input / output core grain, wherein the input / output core grain comprises a memory interface, a directory, a second address space mapping module and a second core grain interconnection interface; both the first address space mapping module and the second address space mapping module store an address space mapping table, and the address space mapping table stores a mapping relationship between a memory address accessed by a processor core grain and a memory interface of an input / output core grain; the directory stores cache consistency information of the cache blocks corresponding to the memory interfaces of the input / output core grain and other input / output core grains corresponding to the directory.
Owner:BEIJING VCORE TECH CO LTD

Ultra-wide RISC-V long vector processor

The invention relates to the technical field of processor hardware design, and discloses an ultra-wide RISC-V long vector processor, which comprises a 64-bit scalar RISC-V core and a plurality of vector clusters, wherein each vector cluster comprises an instruction dispatcher and an instruction scheduler, the instruction dispatcher receives vector instructions sent from the scalar core and distributes the instructions to idle processing channels, and the instruction scheduler controls the execution sequence of the instructions in time; each channel is connected with the mask unit, the sliding unit and the vector read-write unit through the full-interconnection crossbar switch, and the full-interconnection crossbar switch and the mask unit carry out conditional execution on elements in a vector instruction based on the mask register. According to the method, the long vector can be quickly and efficiently calculated, the expandability problem of a full interconnection structure is solved by adopting a special layered pipeline interconnection structure, and then long vector support can be carried out on an extended vector processor architecture.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Vector mask buffers in a vector instruction execution pipeline

Systems and methods related to vector mask buffers in a vector instruction execution pipeline are disclosed herein. The vector instruction execution pipeline may include several lanes. Each lane may include a vector register file, a vector mask buffer, and a functional processing unit. The vector register file may store operand data and the vector mask buffer may store a vector mask associated with the operand data. In a lane, the operand data may be read from the register file into a functional processing unit, and the vector mask may be read from the vector mask buffer to the functional processing unit. The functional processing unit may process the operand data based on the vector mask. The lane-specific vector mask buffers improve the efficiency of the vector instruction execution pipeline by storing the vector masks proximate to where the vector masks will be used.
Owner:TENSTORRENT USA INC

Processor, tensor processing method, and device

The present application relates to the technical field of data processing, and is used for realizing universal and efficient processing of a multi-dimensional tensor. Provided are a processor, a tensor processing method, and a device. A tensor processing unit in the processor comprises: a preprocessing circuit, which is used for adjusting first dimension information of a multi-dimensional tensor, so as to obtain second dimension information, wherein the product of the size of any dimension in the second dimension information and the capacity of a unit access storage space is smaller than or equal to the capacity of a data cache; an address generation circuit, which is used for generating, on the basis of the second dimension information, a plurality of source addresses corresponding to a plurality of pieces of data in the multi-dimensional tensor; an access control circuit, which is used for determining a plurality of pieces of cache mapping information corresponding to the plurality of source addresses, and sending a plurality of access requests on the basis of the plurality of source addresses, and is used for acquiring from a first memory the plurality of pieces of data of the multi-dimensional tensor; and the data cache, which is used for caching the plurality of pieces of data on the basis of the plurality of pieces of cache mapping information.
Owner:HUAWEI TECH CO LTD

Tensor dimension recombination method and device for tensor processing unit, and chip

The invention relates to the field of tensor processing, and provides a tensor dimension recombination method and device for a tensor processing unit, and a chip. The method comprises the following steps: checking whether the total number of elements of an input tensor is equal to that of elements of a target output tensor; based on the shape parameters of the input tensor and the target output tensor, storage layout information and hardware architecture characteristics of a tensor processing unit, performing mode judgment on the current dimension recombination operation to judge whether the current dimension recombination operation is matched with a preset high-frequency special judgment mode or not; when the current dimension recombination operation is matched with the high-frequency special judgment mode, a hardware acceleration execution mode corresponding to the high-frequency special judgment mode is adopted, and tensor dimension recombination is completed in an on-chip storage range; when the current dimension recombination operation is not matched with the high-frequency special judgment mode, tensor dimension recombination is completed through parallel computing and cooperative processing in a general recombination execution mode oriented to a hierarchical storage structure. The respe execution efficiency can be improved, occupation of on-chip memory resources is reduced, and waste of bandwidth resources is avoided.
Owner:ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD

Master control election method for multiple micro-control units and related device

The invention discloses a main control election method for multiple micro-control units and a related device, and relates to the technical field of information analysis, and the method comprises the steps: calculating a priority value of each micro-control unit based on a starting timestamp, an internal temperature, a central processing unit utilization rate and voltage stability; each micro-control unit broadcasts a current voting round and proposes a master control identity identification number and a priority value so as to perform priority evaluation on each micro-control unit by utilizing a triple comparison rule; performing proposal updating based on a priority evaluation result to obtain proposal updating information; determining a main control unit and a plurality of standby control units by using a preset agreement mechanism based on the proposal update information; and the main control unit broadcasts a heartbeat frame, each standby control unit judges whether the main control unit has a fault based on the heartbeat frame, and if the main control unit is judged to have the fault, the main control election is performed again. According to the method, the consistency and reliability of the master control election process are guaranteed, and meanwhile self-judgment and self-recovery after the master control fails are achieved.
Owner:SHANGHAI FUKUN AVIATION TECH CO LTD

Quantization prediction for block data

A scalar processor associated with a vector processor reduces the quantization error for blocked data with a relatively small register size by predicting adjustments for shared scalars used in runtime quantization. The scalar processor provides a recommended scale value to the vector processor for scaling a block of data from a wide data type format to a narrow data type format. The scalar processor and the vector processor share a register at which the scalar processor stores the recommended scale value and from which the vector processor accesses the recommended scale value. The vector processor performs an operation to quantize at least a portion of the block of data by applying a scale value that is based on the recommended scale value.
Owner:ADVANCED MICRO DEVICES INC +1

Heterogeneous sensing adaptive low-bit neural network deployment method

The invention provides a heterogeneous perception adaptive low-bit neural network deployment method, and relates to the technical field of artificial intelligence and heterogeneous computing, and the method comprises the steps: firstly obtaining the hierarchical computing feature information of each layer of a neural network and the dynamic feature parameter information of heterogeneous hardware, and forming a multi-level basic information set; and then a hierarchical efficiency association model is constructed to describe the association relationship among the calculation precision, the hardware dynamic characteristics and the layer calculation efficiency. During online operation, a hardware real-time load state and an energy efficiency constraint condition are tracked to generate a dynamic state monitoring result, and the dynamic state monitoring result and the dynamic state monitoring result are combined to generate a hierarchical deployment configuration scheme by adopting a multi-objective optimization algorithm. And finally, a lightweight runtime scheduler is called to allocate calculation tasks according to the scheme, and interlayer dependency data is loaded and executed, so that dynamic neural network deployment across hardware equipment is realized, and the deployment effect and the operation efficiency are improved.
Owner:XINGFAN XINGQI (CHENGDU) TECH CO LTD

Configurable processor element arrays for implementing convolutional neural networks

PendingUS20260154525A1Neural architecturesPhysical realisationData streamProcessor element
Example apparatus disclosed herein include an array of processor elements, the array including rows each having a first number of processor elements and columns each having a second number of processor elements. Disclosed example apparatus also include configuration registers to store descriptors to configure the array to implement a layer of a convolutional neural network based on a dataflow schedule corresponding to one of multiple tensor processing templates, ones of the processor elements to be configured based on the descriptors to implement the one of the tensor processing templates to operate on input activation data and filter data associated with the layer of the convolutional neural network to produce output activation data associated with the layer of the convolutional neural network. Disclosed example apparatus further include memory to store the input activation data, the filter data and the output activation data associated with the layer of the convolutional neural network.
Owner:INTEL CORP

System

A system is provided.SOLUTION: A system comprising: means for registering a face image of a user; means for extracting feature points from the registered face image; means for generating a 3D model of a face based on the feature points; means for automatically generating an angle, a pose, and an expression to be combined with a background picture; means for representing the face image by the 3D model according to the generated angle, pose, and expression and naturally combining the face image with the background picture; and means for presenting the combined image to the user and storing the combined image after checking.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Image processor and method for processing an image

The present application relates to image processors and methods for processing images. A method of computing a warp result is disclosed that can include performing, for each target pixel in a set of target pixels, a warp computation process that includes receiving, by a first set of processing units in an array of processing units, a first weight and a second weight associated with the target pixel; receiving, by a second set of processing units in the array, values of neighboring source pixels associated with the target pixel; computing, by the second set, a warp result responsive to the values of the neighboring source pixels and the pair of weights; and providing the warp result to a memory module.
Owner:MOBILEYE VISION TECH LTD

Universal computing unit and instruction scheduling method

The invention provides a general purpose computing unit and an instruction scheduling method for the general purpose computing unit. The general-purpose computing unit comprises a plurality of execution units, wherein each execution unit is used for executing an operation instruction by taking a thread bundle as a unit; and an instruction scheduler for determining a priority of the plurality of operation instructions based on whether the multiplexing flag exists, and scheduling each operation instruction to one of the plurality of execution units based on the priority; wherein each execution unit comprises one or more reuse registers, each reuse register is used for registering an operand of a previous operation instruction and a label, and the label is used for indicating an address and a thread bundle of the operand; and the instruction decoding unit is used for decoding the received operation instruction to determine the reading position of the operand of the operation instruction. The instruction is preferentially scheduled by adding a reuse mark to the whole operation instruction, so that the hit rate of the operand is improved, and excessive instruction bit fields do not need to be occupied.
Owner:SHANGHAI BIREN TECH CO LTD

Multi-transmitting channel architecture optimization method and system based on superscalar processor

The invention provides a multi-transmitting-channel architecture optimization method and system based on a superscalar processor, and belongs to the technical field of processors, and the method comprises the steps: reading an instruction by adopting a data capturing and transmitting mechanism, judging whether a register is ready, introducing a feedforward data cache region, optimizing a data path, reading and obtaining a value of the register, and obtaining a multi-transmitting-channel architecture of the superscalar processor; the read values and instructions are stored in a transmitting queue; loading storage instruction transmission optimization is carried out on the instruction, a storage instruction is divided into a storage address and storage data, the storage data is allowed to be transmitted in advance, and table items are established in a loading storage unit; and reconstructing the transmitting queues, reconstructing the LSIQ transmitting queues into a Load instruction queue and a Store instruction queue, arbitrating the transmitting queues respectively, and waiting for execution. According to the method, the access pressure of the register file can be effectively relieved, the data feedforward efficiency is improved, RAW address conflicts are reduced, and the instruction transmitting efficiency and the overall processor performance are improved.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Floating point data parallel computing method and device of vector processor and vector processor

The invention provides a floating point data parallel computing method and device based on a vector processor and the vector processor, and the method comprises the steps: carrying out vectorization and parallel absolute value calculation on a plurality of pieces of input floating point data, and dividing each piece of data into different computing intervals according to the calculated absolute value; for data located in different calculation intervals, calculating branches in different intervals are combined, branch differences are uniformly represented through symbol transformation variables, and fitting intermediate variables corresponding to all the data are obtained through uniform vector operation instruction parallel calculation; and performing parallel calculation on the basis of the intermediate variable for fitting to obtain a preliminary result corresponding to each data, and correcting the preliminary result according to a symbol of original input data to output a final calculation result in a vector form. According to the method, the problems of low hardware resource utilization rate and poor calculation efficiency of high-precision floating point function calculation realized by adopting a scalar serial processing mode in the prior art are solved.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Computing core particle, computing chip and computing system

A computing core particle, a computing chip, and a computing system are provided. Provided is a computing core particle, characterized in that the computing core particle comprises: a storage module, the storage module comprising a plurality of storage layers; the logic module comprises one or more logic layers, each logic layer in the one or more logic layers comprises a control core and a plurality of computing cores, and the control core of each logic layer is connected with the plurality of computing cores of the same logic layer; each computing core of each logic layer is connected with the adjacent computing core; wherein the plurality of storage layers and the one or more logic layers are vertically stacked, and each of the plurality of computing cores of each logic layer is connected to one or more storage layers of the plurality of storage layers.
Owner:张江国家实验室

High-parallelism-degree data prefetching implementation method for target tracking hardware accelerator

The invention provides a high-parallelism-degree data prefetching implementation method for a target tracking hardware accelerator. A controller, an input cache module, a weight cache module, a matrix processing unit, an output cache module, a write-back module and an external memory are included. The controller controls the calculated data flow by controlling the input cache module, the matrix processing unit and the output cache module; the input cache module is responsible for completing pre-fetching and pre-processing operations of input data; the weight caching module is used for completing prefetching of weight data; the matrix processing unit is used for receiving output data of the input cache module and the weight cache module, carrying out parallel operation and outputting a result to the output cache module, and the output of the output cache module is written back to an external memory through the write-back module. The configurable accelerator architecture based on the vector processing unit (VPU) is combined with the optimized sliding window design and the multi-level parallel strategy, so that the calculation performance and the energy efficiency ratio are effectively improved.
Owner:SHANGHAI JIAOTONG UNIV

Configuring a tensor operation pipeline in a hardware accelerator

A computing method is provided for configuring a tensor operation pipeline. In one example implementation, the method includes receiving a tensor operation pipeline definition and tensor data from a processor, at a configurable pipeline processing element array of a hardware accelerator. The method further includes, in each of a plurality of processing elements of the array, processing the tensor data by implementing a configurable tensor operation pipeline including one or more of the fixed tensor operation logic units according to the tensor operation pipeline definition. The method further includes outputting a tensor operation pipeline result based on the processing of the tensor data by each tensor operation pipeline in each processing element.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Autonomous resonance network for global energy reduction using novel quantum cryptanalysis

System for reducing energy consumption when executing a Deep-Q Network (DQN) on a graphics processing unit (GPU), comprising:
Owner:TRAUTH STEFAN

Apparatus and method for dynamic core management

An apparatus and method are described for intelligently scheduling threads across a plurality of logical processors. For example, one embodiment of a processor comprises: a plurality cores and power management circuitry to associate a plurality of performance values and a plurality of efficiency values with the plurality of cores. In some implementations, each core is associated with at least one performance value and at least one efficiency value. The performance values and efficiency values are used by a scheduler for scheduling threads on the plurality of cores. Some implementations include dynamic core configuration hardware logic coupled to or integral to the power management circuitry to resolve a plurality of configuration hints into a consolidated hint for updating one or more performance values of the plurality of performance values and / or one or more efficiency values of the plurality of efficiency values.
Owner:INTEL CORP

Information processing system and information processing program

To provide an information processing system and information processing program for making a prediction on application of color gradation correction when printing.SOLUTION: An estimation unit 106 inputs new input information, including new print information and information regarding generation of the new print information, to a trained model pre-trained to output application information regarding application of color gradation correction to print information upon input of input information including the print information and information regarding generation of the print information, so as to output application information corresponding to the new input information.SELECTED DRAWING: Figure 4
Owner:FUJIFILM BUSINESS INNOVATION CORP

Learning data generation device, learning data generation method, and learning data generation program

This enables the generation of effective image data to improve detection accuracy without human intervention. [Solution] The image conversion unit 21 receives image data, which is training data 31, as input to the image conversion model 41 that converts image data, and uses this training data to train an object detection model that detects objects with the attributes of the target object from the image data. The image conversion model 41 then converts the image data to obtain a converted image. The training data addition unit 22 adds the converted image obtained by the image conversion unit 21 to the training data 31.
Owner:MITSUBISHI ELECTRIC DIGITAL INNOVATION CORP

Image processing device, learning method, and program

To simultaneously train a feature extractor and a restorer which restores an ideal image from a feature quantity outputted from the feature extractor, so that they are totally optimized.SOLUTION: An image processing apparatus includes: acquisition means which acquires a first image including an object, a second image which includes the object and is different from the first image in condition relating to photographing, and identification information for identifying the object; extraction means which extracts a feature quantity from the second image; classification means which uses the feature quantity extracted by the extraction means to classify the object; restoration means which uses the feature quantity extracted by the extraction means to generate a restored image brought closer to the first image from the second image; and training means which trains the extraction means and the restoration means so that a value corresponding to a first difference relating to a relation between a classification result of the object obtained by the classification means and the identification information and a second difference relating to a relation between the restored image generated by the restoration means and the first image is reduced.SELECTED DRAWING: Figure 4
Owner:CANON KK

Image processing system and method for generating decorative character image

The present invention generates an image of a character using a machine learning model. This image processing system comprises: an information acquisition unit that acquires character information for designating a character and a living organism image being an image of a living organism including a face; and an image generation unit that, by inputting input information to an image generation model using the character information and the living organism image, causes the image generation model to generate a decorative character image. The decorative character image is an image showing a decorative character being a character to which a decoration showing the face of the living organism shown in the living organism image is applied. The image generation model is trained so as to generate, on the basis of the inputted information, an image to which the decoration showing the face of the living organism shown in the living organism image is applied.
Owner:BROTHER KOGYO KK

Collaborative work-stealing scheduler

A method for use in a computing system having a central processing unit (CPU) and a graphics processing unit (GPU), the method comprising: assigning a first memory portion and a second memory portion to: a worker thread of a work-stealing scheduler and an execution unit that is part of the GPU; retrieving a task from a queue associated with the worker thread; having the worker thread detect whether a deadline condition for the task is met; if the deadline condition is not met, dividing the task into two or more additional tasks and adding the two or more additional tasks to the queue; if the deadline condition is met, storing a first data corresponding to the task in the second memory portion and issuing a memory fence fetch instruction; and storing a first value in the first memory portion.
Owner:RAYTHEON CO

Learning device, learning method, and program

To generate learning data suitable to improvement of the generalization performance of a neural network used for estimation.SOLUTION: A learning device comprises: a conversion unit which creates a second conversion image obtained by converting a first image into a second domain image in a pseudo manner by using a first neural network, converts the second conversion image into a first reconstructed image reconstructed to the first domain image by using a second neural network different from the first neural network, creates a first conversion image obtained by converting a second image into the first domain image in a pseudo manner by using the second neural network, and converts the first conversion image into a second reconstructed image reconstructed to the second domain image by using the first neural network; and an update unit which calculates a loss due to image conversion to update to parameters of the first neural network and the second neural network in which the loss becomes minimum.SELECTED DRAWING: Figure 3
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD