Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

800results about "Architecture with multiple processing units" patented technology

Implementing asymmetric processor cores to enable higher operating frequencies in processor-based devices

Implementing asymmetric processor cores to enable higher operating frequencies in processor-based devices is disclosed herein. In some aspects, a processor-based device provides a core cluster that comprises a plurality of processor cores and a corresponding phase-locked loop (PLL). Each processor core is based on a common instruction set architecture (ISA) and is configured to operate synchronously based on a same clock signal from the PLL of the core cluster. A first subset of processor cores within the core cluster is implemented with a different physical characteristic relative to a second subset of processor cores within the core cluster, wherein the different physical characteristic enables each processor core of the first subset of processor cores to operate at a higher operating frequency than each processor core of the second subset of processor cores.
Owner:QUALCOMM INC

Compute-in-memory chip, instruction scheduling method, and related apparatus

The present application discloses a compute-in-memory chip, an instruction scheduling method, and a related apparatus. The compute-in-memory chip comprises an instruction memory, an instruction scheduler, and at least one compute-in-memory memory; each compute-in-memory memory comprises at least one storage array; the instruction memory is used for acquiring a first tensor instruction to be executed; and the instruction scheduler is used for scheduling, on the basis of the association relationship between the first tensor instruction and a second tensor instruction and the state of a target storage array needing to be operated for executing the first tensor instruction, the first tensor instruction to the compute-in-memory memory to which the target storage array belongs so that the compute-in-memory memory executes the first tensor instruction. According to embodiments of the present application, diversified compute-in-memory computing can be supported, efficient out-of-order execution scheduling of a tensor instruction set is achieved on the basis of the compute-in-memory chip, and the requirements of compute-in-memory technology for high concurrency and high throughput rate are met.
Owner:HUAWEI TECH CO LTD

Information processing device, information processing method, and program

[Problem] To enable generation of more effective advertisements. [Solution] An information processing device comprising a control unit that edits a tag group obtained by converting advertisement content to text, extracts a derivable element from the advertisement content, and generates, by deriving the corresponding element according to the edited tag group, derived advertisement content based on the advertisement content.
Owner:SONY GROUP CORP

Data processing method and device, processor and electronic equipment

The invention relates to the technical field of artificial intelligence chips, and discloses a data processing method and device, a processor and electronic equipment. Determining a first jump step length required for rearranging the block data based on the block size of the input data in the first dimension; secondly, performing jump reading on the block data according to a first jump step length, rearranging and loading read input elements into a thread private register so as to realize transposition storage of the block data in the thread private register, and mapping a thread direction on a second dimension so as to obtain a data basis of convolution correlation calculation; convolution correlation calculation is executed based on the transposed data in the thread private register and the weight data in the thread private register, repeated data reading in the calculation process is reduced during sliding window calculation, and an operator is converted into calculation performance bottleneck operation from memory access performance bottleneck operation.
Owner:SHANGHAI BIREN TECH CO LTD

Programmable in-memory computing accelerator for low-precision deep neural network inference

A programmable in-memory computing (IMC) accelerator for low-precision deep neural network inference, also referred to as PIMCA, is provided. Embodiments of the PIMCA integrate a large number of capacitive-coupling-based IMC static random-access memory (SRAM) macros and demonstrate large-scale integration of IMC SRAM macros. For example, a 28 nm prototype integrates 108 capacitive-coupling-based IMC SRAM macros of a total size of 3.4 megabytes (Mb), demonstrating one of the largest IMC hardware to date. In addition, a custom instruction set architecture (ISA) is developed featuring IMC and single-instruction-multiple-data (SIMD) functional units with hardware loop to support a range of deep neural network (DNN) layer types. The 28 nm prototype chip achieves a peak throughput of 4.9 tera operations per second (TOPS) and system-level peak energy-efficiency of 437 TOPS per watt (TOPS / W) at 40 megahertz (MHz) with a 1 volt (V) supply.
Owner:THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK +1

Caching method and system for access unit of superscalar processor

The invention belongs to the field of integrated circuits and computer system structures, and provides a caching method and system for a memory access unit of a superscalar processor, and the method comprises the steps: receiving a plurality of memory access instructions in the same period, and determining a corresponding Bank in a to-be-accessed cache through the memory access instructions; after the memory access instruction obtains a cache access permission, if cache line missing occurs, generating a missing request, merging all the missing requests, and performing parallel prefetching training on the merged missing requests by utilizing a mode of fusing a constant step length prefetching mode and a complex step length prefetching mode to obtain a prefetching request and a prefetching cache address corresponding to the prefetching request; requesting a missing cache line from the first-level cache to the second-level cache based on the missing queue, and writing the missing cache line back to the cache line of the corresponding data cache in the first-level cache; and storing the bus consistency request by using the sniffing queue, judging whether the data in the multi-core cache are consistent or not by using the consistency request, and performing consistency modification according to a judgment result. The cache hit rate and the bandwidth utilization rate are improved.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Vector processor, high performance processor, and electronic device

The invention provides a vector processor, a high-performance processor and electronic equipment. The vector processor comprises a vector program control unit, a plurality of functional units, a matrix register file and a scalar register, wherein the vector program control unit is used for fetching instructions and transmitting the instructions; the vector program control unit interacts with the scalar register; the function unit is used for performing function processing according to the instruction; the matrix register file is used for returning data after receiving the read-write request; rearranging the data and then returning the data; performing read-write interaction with the functional unit; and configuring a configuration register of the vector program control unit through data in the matrix register file. The vector processor provided by the invention can efficiently process the vector data.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Discrete three-dimensional processor

A discrete three-dimensional (3-D) processor comprises stacked first and second dice. The first die comprises three-dimensional memory (3D-M) arrays, whereas the second die comprises at least a portion of a logic / processing circuit and an off-die peripheral-circuit component of the 3D-M array(s). The preferred 3-D processor can be used to compute non-arithmetic function / model. In other applications, the preferred 3-D processor may also be a 3-D configurable computing array, a 3-D pattern processor, or a 3-D neuro-processor.
Owner:HONG KONG HAICUN TECHNOLOGY CO LTD

Neural processing unit having direct data pathway to external memory

A neural network processing unit (NPU) includes a plurality of processing elements for performing the ANN model computations using weight parameters and input activation data; an NPU internal memory operatively coupled to the plurality of processing elements for storing at least one of the weight parameters or the activation data; and a dedicated external memory interface configured for a direct data pathway to an external main memory system storing ANN model data, the direct data pathway facilitating transfer of a portion of the ANN model data, wherein the NPU is configured to receive the ANN model data via the dedicated external memory interface for processing.
Owner:DEEPX CO LTD

Discrete Three-Dimensional Processor

A discrete three-dimensional (3-D) processor comprises vertically stacked and communicatively coupled first and second dice. The first die comprises memory arrays, while the second die comprises non-memory circuits and off-die peripheral-circuit components of the memory arrays. The total-number difference of the BEOL layers between the memory arrays and the off-die peripheral-circuit components is substantially larger than the total-number difference of the BEOL layers between the non-memory circuits and the off-die peripheral-circuit components.
Owner:HONG KONG HAICUN TECHNOLOGY CO LTD

Large language model reasoning system and method based on multi-chip parallel computing

The invention provides an inference system and method of a large language model based on multi-chip parallel computing, and relates to the technical field of artificial intelligence. The system comprises a pre-calculation module used for processing input instruction information to generate to-be-reasoned data, and the to-be-reasoned data is in a matrix form; the expert parallel module is used for sending the to-be-reasoned data to accelerator chips in the expert parallel module and determining sub-reasoning data processed by the activation expert units corresponding to the accelerator chips respectively, so that the activation expert units carry out calculation based on the corresponding sub-reasoning data and complete parallel calculation result data is determined. The input data is broadcasted to all the accelerator chips, each accelerator chip selects the corresponding input data for calculation according to the set activation expert unit, the same complete calculation result is obtained through global protocol operation among all the accelerator chips, and the overall operation performance and efficiency are improved.
Owner:SHENZHEN CORERAIN TECH CO LTD

SYSTEMS AND METHODS IN THE FIELD OF SELF-SUPERVISED DETECTION OF FACIAL FLAGSHIPS

SYSTEMS AND METHODS IN THE FIELD OF SELF-SUPERVISED DETECTION OF FACIAL BOUNDARIES. Systems and methods for self-supervised learning (SSL) in facial detection networks are proposed. In one embodiment, a facial detection network comprises encoder components configured to encode facial features, the encoder components including trained components of a masked image modeling (MIM) network configured to process non-overlapping patches determined from the input image, the MIM network trained with an SSL objective; and decoder components configured by training to determine local matches between features to determine estimates for facial landmarks. In one embodiment, the MIM network is an MAE network.In one embodiment, the decoder components are derived from those of a second trained network comprising the encoder components as trained but fixed, wherein the decoder components of the second network are trained using locality constraint repulsion loss (LCR). Methods are proposed for SSL training of the encoder and decoder components. Figure for abstract: none.
Owner:LOREAL SA

Methods and circuits for streaming data to processing elements in stacked processor-plus-memory architecture

A stacked processor-plus-memory device includes a processing die with an array of processing elements of an artificial neural network. Each processing element multiplies a first operand—e.g. a weight—by a second operand to produce a partial result to a subsequent processing element. To prepare for these computations, a sequencer loads the weights into the processing elements as a sequence of operands that step through the processing elements, each operand stored in the corresponding processing element. The operands can be sequenced directly from memory to the processing elements or can be stored first in cache. The processing elements include streaming logic that disregards interruptions in the stream of operands.
Owner:RAMBUS INC

Multi-core processor chip and storage access method and device of multi-core processor chip

The invention provides a multi-core processor chip and a storage access method and device of the multi-core processor chip, and relates to the technical field of computer processors, the multi-core processor chip comprises at least one processor core, the processor core comprises a general processor core, a first address space mapping module and a first core interconnection interface, the universal processor core supports high-speed cache consistency; at least one input / output core grain, wherein the input / output core grain comprises a memory interface, a directory, a second address space mapping module and a second core grain interconnection interface; both the first address space mapping module and the second address space mapping module store an address space mapping table, and the address space mapping table stores a mapping relationship between a memory address accessed by a processor core grain and a memory interface of an input / output core grain; the directory stores cache consistency information of the cache blocks corresponding to the memory interfaces of the input / output core grain and other input / output core grains corresponding to the directory.
Owner:BEIJING VCORE TECH CO LTD

Compiling an application having polynomial operations to produce directed acyclic graphs having commands to execute in a near memory processing device

Provided are a computer program product, system, and method for compiling an application having polynomial operations to produce directed acyclic graphs having commands to execute in a near memory processing device. An application is compiled including operations on a polynomial having coefficients, decomposed into a number of levels of coefficient elements, to generate hierarchical directed acyclic graphs (DAGs) having nodes indicating commands for execution by a hierarchy of hardware components in a near memory processing (NMP) device. The hierarchy of hardware components includes a plurality of enclaves of tiles. Each tile includes memory and a processing element to perform operations on the decomposed coefficients stored in the memory of the tile. Each of the hardware components includes a controller to process the commands in the DAG generated for the hardware components. The DAGs are provided to a hierarchical DAG tracker to generate commands for the NMP device.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION +1

Methods and apparatus to facilitate write miss caching in cache system

Methods, apparatus, systems and articles of manufacture to facilitate write miss caching in cache system are disclosed. An example apparatus includes a first cache storage; a second cache storage, wherein the second cache storage includes a first portion operable to store a first set of data evicted from the first cache storage and a second portion; a cache controller coupled to the first cache storage and the second cache storage and operable to: receive a write operation; determine that the write operation produces a miss in the first cache storage; and in response to the miss in the first cache storage, provide write miss information associated with the write operation to the second cache storage for storing in the second portion.
Owner:TEXAS INSTRUMENTS INC

Ultra-wide RISC-V long vector processor

The invention relates to the technical field of processor hardware design, and discloses an ultra-wide RISC-V long vector processor, which comprises a 64-bit scalar RISC-V core and a plurality of vector clusters, wherein each vector cluster comprises an instruction dispatcher and an instruction scheduler, the instruction dispatcher receives vector instructions sent from the scalar core and distributes the instructions to idle processing channels, and the instruction scheduler controls the execution sequence of the instructions in time; each channel is connected with the mask unit, the sliding unit and the vector read-write unit through the full-interconnection crossbar switch, and the full-interconnection crossbar switch and the mask unit carry out conditional execution on elements in a vector instruction based on the mask register. According to the method, the long vector can be quickly and efficiently calculated, the expandability problem of a full interconnection structure is solved by adopting a special layered pipeline interconnection structure, and then long vector support can be carried out on an extended vector processor architecture.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Vector mask buffers in a vector instruction execution pipeline

Systems and methods related to vector mask buffers in a vector instruction execution pipeline are disclosed herein. The vector instruction execution pipeline may include several lanes. Each lane may include a vector register file, a vector mask buffer, and a functional processing unit. The vector register file may store operand data and the vector mask buffer may store a vector mask associated with the operand data. In a lane, the operand data may be read from the register file into a functional processing unit, and the vector mask may be read from the vector mask buffer to the functional processing unit. The functional processing unit may process the operand data based on the vector mask. The lane-specific vector mask buffers improve the efficiency of the vector instruction execution pipeline by storing the vector masks proximate to where the vector masks will be used.
Owner:TENSTORRENT USA INC

Irregular Cadence Data Processing Units

Aspects of the disclosure are directed to an architecture including a dynamic serialization buffer and / or dynamic deserialization buffer coupled between a vector processing unit and a matrix multiplication unit. The dynamic serialization buffer and / or dynamic deserialization buffer allow for streaming any integer of vectors per cycle when performing acceleration of matrix multiplication operations. The matrix multiplication unit receives vectors equivalent to an amount of data from the vector processing unit at an arbitrary rate of vectors per cycle. The matrix multiplication unit processes the vectors to generate resulting vectors that are output at the arbitrary rate.
Owner:GOOGLE LLC

Processor, tensor processing method, and device

The present application relates to the technical field of data processing, and is used for realizing universal and efficient processing of a multi-dimensional tensor. Provided are a processor, a tensor processing method, and a device. A tensor processing unit in the processor comprises: a preprocessing circuit, which is used for adjusting first dimension information of a multi-dimensional tensor, so as to obtain second dimension information, wherein the product of the size of any dimension in the second dimension information and the capacity of a unit access storage space is smaller than or equal to the capacity of a data cache; an address generation circuit, which is used for generating, on the basis of the second dimension information, a plurality of source addresses corresponding to a plurality of pieces of data in the multi-dimensional tensor; an access control circuit, which is used for determining a plurality of pieces of cache mapping information corresponding to the plurality of source addresses, and sending a plurality of access requests on the basis of the plurality of source addresses, and is used for acquiring from a first memory the plurality of pieces of data of the multi-dimensional tensor; and the data cache, which is used for caching the plurality of pieces of data on the basis of the plurality of pieces of cache mapping information.
Owner:HUAWEI TECH CO LTD

Data processing method, electronic equipment and storage medium

The invention discloses a data processing method, electronic equipment and a storage medium. The data processing method comprises the following steps: receiving a first original tensor; performing dimension conversion processing on the first original tensor to obtain a second original tensor; carrying out size expansion on at least one dimension of the second original tensor, and storing a third original tensor obtained by size expansion into a memory; receiving shape information of the tensor to be processed and an initial coordinate in a preset coordinate system; in combination with the third original tensor, the shape information of the tensor to be processed and the initial coordinate of the tensor to be processed, determining a plurality of requests for loading the tensor to be processed to the cache region; and sending a plurality of requests in sequence, and writing the sub-data returned by each request into the cache region in sequence to load the tensor to be processed to the cache region. According to the method, dimension conversion is supported, the data loading complexity of dimension conversion is reduced, the bandwidth utilization rate is improved, the data bandwidth is effectively utilized, and the loading efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Virtual GPU system and application method therefor, and device and storage medium

PCT designated stage expiredWO2025134011A1Resource allocationProcessor architectures/configuration
Provided in the present disclosure are a virtual GPU system and an application method therefor, and a device and a storage medium. In the system, a server cluster composed of a plurality of CPU service nodes simulates a virtual GPU, load balancing between the CPU service nodes is completed by a scheduler, a vector driving assembly and an algorithm plug-in are deployed on each CPU service node, and the vector driving assembly can invoke the algorithm plug-in to execute a task called by the scheduler, so that the simulation of the server cluster on the GPU is realized. A plurality of CPU service nodes are used to simulate a single GPU, so that the number of GPU cores is greatly increased. On the basis of test data, the performance of each simulated CPU core is equivalent to 10% of the performance of a GPU. When the number of CPU service nodes in the cluster exceeds 10, the performance of the whole system starts to exceed a common GPU.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Method for controlling information processing apparatus, information processing apparatus, and program

To reduce, in a more appropriate way, the time and effort of a user related to a prompt instruction to process a plurality of images by using a model configured through machine learning.SOLUTION: A layout data DB 400 stores a first prompt including an instruction for processing of a first image, the first prompt having been received when a generative model 415 outputs a processed image obtained by processing the first image, the generative model being trained to process an input image based on an input prompt. A prompt reuse unit 404 generates a second prompt including at least an instruction for processing of a designated second image, based on the first prompt.SELECTED DRAWING: Figure 4
Owner:CANON KK

Performance optimization method for accelerator, and processor and storage medium

PCT designated stage expiredWO2025134026A2Resource allocationInterprogram communication
Provided in the embodiments of the present disclosure are a performance optimization method for an accelerator, and a processor and a storage medium. In the embodiments of the present disclosure, modification performed on a driver of an accelerator built in a processor is proposed. Before a working task sent in a user mode is delivered to a port provided by the accelerator, a process identifier and a thread identifier are pre-configured for the working task, and the two identifiers are used as labels of the working task; and then, the working task, which carries the labels, is delivered to the port provided by the accelerator. In this way, the isolation of working tasks between different processes can be ensured in an accelerator by means of labels, thereby allowing a port of the accelerator to be shared among threads under a plurality of processes. Accordingly, in the present embodiment, the accelerator built in the processor can allow acceleration services to be simultaneously provided for more user-mode processes, thereby providing higher acceleration performance in a high-concurrency scene of a user-mode process.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Tensor dimension recombination method and device for tensor processing unit, and chip

The invention relates to the field of tensor processing, and provides a tensor dimension recombination method and device for a tensor processing unit, and a chip. The method comprises the following steps: checking whether the total number of elements of an input tensor is equal to that of elements of a target output tensor; based on the shape parameters of the input tensor and the target output tensor, storage layout information and hardware architecture characteristics of a tensor processing unit, performing mode judgment on the current dimension recombination operation to judge whether the current dimension recombination operation is matched with a preset high-frequency special judgment mode or not; when the current dimension recombination operation is matched with the high-frequency special judgment mode, a hardware acceleration execution mode corresponding to the high-frequency special judgment mode is adopted, and tensor dimension recombination is completed in an on-chip storage range; when the current dimension recombination operation is not matched with the high-frequency special judgment mode, tensor dimension recombination is completed through parallel computing and cooperative processing in a general recombination execution mode oriented to a hierarchical storage structure. The respe execution efficiency can be improved, occupation of on-chip memory resources is reduced, and waste of bandwidth resources is avoided.
Owner:ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD

Convolution weight gradient calculation method and device, medium, equipment and product

The invention discloses a convolution weight gradient calculation method and device, a medium, equipment and a product, and the method comprises the steps: carrying out the role distribution of an input feature map and a corresponding output gradient map in a global memory, and determining a dynamic sliding map and a static fixed map; respectively extracting a sliding part and a corresponding fixed part from the dynamic sliding graph and the static fixed graph, and enabling each thread to hold a data row and a corresponding aligned data block in the sliding part; and in each thread, carrying out synchronous sliding and point multiplication operation on the held data row and the aligned data block to obtain a local calculation result of a current sliding position, accumulating the local calculation result to a cache position of the weight gradient matrix, and after synchronous traversal calculation of all sliding parts and corresponding fixed parts is completed, obtaining the weight gradient matrix of the convolution kernel. According to the method, data multiplexing in the sliding direction can be effectively realized, and the data volume read repeatedly is reduced, so that the calculation efficiency of the weight gradient is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Master control election method for multiple micro-control units and related device

The invention discloses a main control election method for multiple micro-control units and a related device, and relates to the technical field of information analysis, and the method comprises the steps: calculating a priority value of each micro-control unit based on a starting timestamp, an internal temperature, a central processing unit utilization rate and voltage stability; each micro-control unit broadcasts a current voting round and proposes a master control identity identification number and a priority value so as to perform priority evaluation on each micro-control unit by utilizing a triple comparison rule; performing proposal updating based on a priority evaluation result to obtain proposal updating information; determining a main control unit and a plurality of standby control units by using a preset agreement mechanism based on the proposal update information; and the main control unit broadcasts a heartbeat frame, each standby control unit judges whether the main control unit has a fault based on the heartbeat frame, and if the main control unit is judged to have the fault, the main control election is performed again. According to the method, the consistency and reliability of the master control election process are guaranteed, and meanwhile self-judgment and self-recovery after the master control fails are achieved.
Owner:SHANGHAI FUKUN AVIATION TECH CO LTD

Mainboard, processor board and computing system

The present application discloses a mainboard, a processor board card, and a computing system. The mainboard includes a plurality of interfaces, a first interface of the plurality of interfaces is configured to be connected to the processor board card having a processor circuit, a second interface of the plurality of interfaces is configured to be connected to a non-processor board card, and the first interface and the second interface are connected to each other via a communication circuit; for the mainboard, a connection-centric design idea is adopted, and the processor board card is regarded to have the same status as the non-processor board card; and because no processor circuit is provided, and only the first interface connected to the processor board card needs to be provided, compared with an original mainboard, an area of the board card in the present application is reduced, and the processor board card and other non-processor board cards may be connected to the mainboard in a stacked manner, thereby reducing the required space, reducing the requirements for the space in the device, and improving the running flexibility of the processor board card simultaneously.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Quantization prediction for block data

A scalar processor associated with a vector processor reduces the quantization error for blocked data with a relatively small register size by predicting adjustments for shared scalars used in runtime quantization. The scalar processor provides a recommended scale value to the vector processor for scaling a block of data from a wide data type format to a narrow data type format. The scalar processor and the vector processor share a register at which the scalar processor stores the recommended scale value and from which the vector processor accesses the recommended scale value. The vector processor performs an operation to quantize at least a portion of the block of data by applying a scale value that is based on the recommended scale value.
Owner:ADVANCED MICRO DEVICES INC +1

Heterogeneous sensing adaptive low-bit neural network deployment method

The invention provides a heterogeneous perception adaptive low-bit neural network deployment method, and relates to the technical field of artificial intelligence and heterogeneous computing, and the method comprises the steps: firstly obtaining the hierarchical computing feature information of each layer of a neural network and the dynamic feature parameter information of heterogeneous hardware, and forming a multi-level basic information set; and then a hierarchical efficiency association model is constructed to describe the association relationship among the calculation precision, the hardware dynamic characteristics and the layer calculation efficiency. During online operation, a hardware real-time load state and an energy efficiency constraint condition are tracked to generate a dynamic state monitoring result, and the dynamic state monitoring result and the dynamic state monitoring result are combined to generate a hierarchical deployment configuration scheme by adopting a multi-objective optimization algorithm. And finally, a lightweight runtime scheduler is called to allocate calculation tasks according to the scheme, and interlayer dependency data is loaded and executed, so that dynamic neural network deployment across hardware equipment is realized, and the deployment effect and the operation efficiency are improved.
Owner:XINGFAN XINGQI (CHENGDU) TECH CO LTD