Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

294results about "Architecture with multiple processing units" patented technology

Large language model reasoning system and method based on multi-chip parallel computing

The invention provides an inference system and method of a large language model based on multi-chip parallel computing, and relates to the technical field of artificial intelligence. The system comprises a pre-calculation module used for processing input instruction information to generate to-be-reasoned data, and the to-be-reasoned data is in a matrix form; the expert parallel module is used for sending the to-be-reasoned data to accelerator chips in the expert parallel module and determining sub-reasoning data processed by the activation expert units corresponding to the accelerator chips respectively, so that the activation expert units carry out calculation based on the corresponding sub-reasoning data and complete parallel calculation result data is determined. The input data is broadcasted to all the accelerator chips, each accelerator chip selects the corresponding input data for calculation according to the set activation expert unit, the same complete calculation result is obtained through global protocol operation among all the accelerator chips, and the overall operation performance and efficiency are improved.
Owner:SHENZHEN CORERAIN TECH CO LTD

Methods and circuits for streaming data to processing elements in stacked processor-plus-memory architecture

A stacked processor-plus-memory device includes a processing die with an array of processing elements of an artificial neural network. Each processing element multiplies a first operand—e.g. a weight—by a second operand to produce a partial result to a subsequent processing element. To prepare for these computations, a sequencer loads the weights into the processing elements as a sequence of operands that step through the processing elements, each operand stored in the corresponding processing element. The operands can be sequenced directly from memory to the processing elements or can be stored first in cache. The processing elements include streaming logic that disregards interruptions in the stream of operands.
Owner:RAMBUS INC

Tensor dimension recombination method and device for tensor processing unit, and chip

The invention relates to the field of tensor processing, and provides a tensor dimension recombination method and device for a tensor processing unit, and a chip. The method comprises the following steps: checking whether the total number of elements of an input tensor is equal to that of elements of a target output tensor; based on the shape parameters of the input tensor and the target output tensor, storage layout information and hardware architecture characteristics of a tensor processing unit, performing mode judgment on the current dimension recombination operation to judge whether the current dimension recombination operation is matched with a preset high-frequency special judgment mode or not; when the current dimension recombination operation is matched with the high-frequency special judgment mode, a hardware acceleration execution mode corresponding to the high-frequency special judgment mode is adopted, and tensor dimension recombination is completed in an on-chip storage range; when the current dimension recombination operation is not matched with the high-frequency special judgment mode, tensor dimension recombination is completed through parallel computing and cooperative processing in a general recombination execution mode oriented to a hierarchical storage structure. The respe execution efficiency can be improved, occupation of on-chip memory resources is reduced, and waste of bandwidth resources is avoided.
Owner:ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD

Quantization prediction for block data

A scalar processor associated with a vector processor reduces the quantization error for blocked data with a relatively small register size by predicting adjustments for shared scalars used in runtime quantization. The scalar processor provides a recommended scale value to the vector processor for scaling a block of data from a wide data type format to a narrow data type format. The scalar processor and the vector processor share a register at which the scalar processor stores the recommended scale value and from which the vector processor accesses the recommended scale value. The vector processor performs an operation to quantize at least a portion of the block of data by applying a scale value that is based on the recommended scale value.
Owner:ADVANCED MICRO DEVICES INC +1

Configurable processor element arrays for implementing convolutional neural networks

PendingUS20260154525A1Neural architecturesPhysical realisationData streamProcessor element
Example apparatus disclosed herein include an array of processor elements, the array including rows each having a first number of processor elements and columns each having a second number of processor elements. Disclosed example apparatus also include configuration registers to store descriptors to configure the array to implement a layer of a convolutional neural network based on a dataflow schedule corresponding to one of multiple tensor processing templates, ones of the processor elements to be configured based on the descriptors to implement the one of the tensor processing templates to operate on input activation data and filter data associated with the layer of the convolutional neural network to produce output activation data associated with the layer of the convolutional neural network. Disclosed example apparatus further include memory to store the input activation data, the filter data and the output activation data associated with the layer of the convolutional neural network.
Owner:INTEL CORP

System

A system is provided.SOLUTION: A system comprising: means for registering a face image of a user; means for extracting feature points from the registered face image; means for generating a 3D model of a face based on the feature points; means for automatically generating an angle, a pose, and an expression to be combined with a background picture; means for representing the face image by the 3D model according to the generated angle, pose, and expression and naturally combining the face image with the background picture; and means for presenting the combined image to the user and storing the combined image after checking.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Computing core particle, computing chip and computing system

A computing core particle, a computing chip, and a computing system are provided. Provided is a computing core particle, characterized in that the computing core particle comprises: a storage module, the storage module comprising a plurality of storage layers; the logic module comprises one or more logic layers, each logic layer in the one or more logic layers comprises a control core and a plurality of computing cores, and the control core of each logic layer is connected with the plurality of computing cores of the same logic layer; each computing core of each logic layer is connected with the adjacent computing core; wherein the plurality of storage layers and the one or more logic layers are vertically stacked, and each of the plurality of computing cores of each logic layer is connected to one or more storage layers of the plurality of storage layers.
Owner:张江国家实验室

Apparatus and method for dynamic core management

An apparatus and method are described for intelligently scheduling threads across a plurality of logical processors. For example, one embodiment of a processor comprises: a plurality cores and power management circuitry to associate a plurality of performance values and a plurality of efficiency values with the plurality of cores. In some implementations, each core is associated with at least one performance value and at least one efficiency value. The performance values and efficiency values are used by a scheduler for scheduling threads on the plurality of cores. Some implementations include dynamic core configuration hardware logic coupled to or integral to the power management circuitry to resolve a plurality of configuration hints into a consolidated hint for updating one or more performance values of the plurality of performance values and / or one or more efficiency values of the plurality of efficiency values.
Owner:INTEL CORP

Learning data generation device, learning data generation method, and learning data generation program

This enables the generation of effective image data to improve detection accuracy without human intervention. [Solution] The image conversion unit 21 receives image data, which is training data 31, as input to the image conversion model 41 that converts image data, and uses this training data to train an object detection model that detects objects with the attributes of the target object from the image data. The image conversion model 41 then converts the image data to obtain a converted image. The training data addition unit 22 adds the converted image obtained by the image conversion unit 21 to the training data 31.
Owner:MITSUBISHI ELECTRIC DIGITAL INNOVATION CORP

Image processing device, learning method, and program

To simultaneously train a feature extractor and a restorer which restores an ideal image from a feature quantity outputted from the feature extractor, so that they are totally optimized.SOLUTION: An image processing apparatus includes: acquisition means which acquires a first image including an object, a second image which includes the object and is different from the first image in condition relating to photographing, and identification information for identifying the object; extraction means which extracts a feature quantity from the second image; classification means which uses the feature quantity extracted by the extraction means to classify the object; restoration means which uses the feature quantity extracted by the extraction means to generate a restored image brought closer to the first image from the second image; and training means which trains the extraction means and the restoration means so that a value corresponding to a first difference relating to a relation between a classification result of the object obtained by the classification means and the identification information and a second difference relating to a relation between the restored image generated by the restoration means and the first image is reduced.SELECTED DRAWING: Figure 4
Owner:CANON KK

Image processing system and method for generating decorative character image

The present invention generates an image of a character using a machine learning model. This image processing system comprises: an information acquisition unit that acquires character information for designating a character and a living organism image being an image of a living organism including a face; and an image generation unit that, by inputting input information to an image generation model using the character information and the living organism image, causes the image generation model to generate a decorative character image. The decorative character image is an image showing a decorative character being a character to which a decoration showing the face of the living organism shown in the living organism image is applied. The image generation model is trained so as to generate, on the basis of the inputted information, an image to which the decoration showing the face of the living organism shown in the living organism image is applied.
Owner:BROTHER KOGYO KK

Learning device, learning method, and program

To generate learning data suitable to improvement of the generalization performance of a neural network used for estimation.SOLUTION: A learning device comprises: a conversion unit which creates a second conversion image obtained by converting a first image into a second domain image in a pseudo manner by using a first neural network, converts the second conversion image into a first reconstructed image reconstructed to the first domain image by using a second neural network different from the first neural network, creates a first conversion image obtained by converting a second image into the first domain image in a pseudo manner by using the second neural network, and converts the first conversion image into a second reconstructed image reconstructed to the second domain image by using the first neural network; and an update unit which calculates a loss due to image conversion to update to parameters of the first neural network and the second neural network in which the loss becomes minimum.SELECTED DRAWING: Figure 3
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Accelerated computation of direct memory access scatter context for get response

A system receives an instruction corresponding to a Get request packet of a message and indicating a pattern type associated with direct memory access (DMA) write operations for the Get response. The system determines a descriptor and starting context associated with the Get request packet if the type of pattern indicates nested loops associated with a multi-dimensional array structure. The system stores the starting context in a hardware table, providing access to the starting context in response to processing a Get response packet corresponding to the Get request packet. The system processes the instruction in cycles until a byte count of bytes hypothetically transferred is equal to or greater than a size of the Get request payload. The system obtains an ending context comprising updated loop counters and byte offset and stores the ending context in a cache as the starting context for a next instruction of a same message.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

A method for master election of multiple micro control units and related device

The application discloses a kind of master control election methods of multiple micro control units and related devices, it is related to information analysis technical field, the method includes: based on start time stamp, internal temperature, central processing unit utilization and voltage stability degree calculate the priority value of each micro control unit;Each micro control unit broadcasts current voting round, proposed master identity identification number and priority value to utilize triple comparison rule to carry out the priority evaluation of each micro control unit;Based on priority evaluation result carries out proposal update, obtains proposal update information;Based on proposal update information utilizes preset agreement mechanism to determine master unit and several spare control units;Master unit broadcasts heartbeat frame, each spare control unit judges whether master unit fails based on heartbeat frame, if it is judged that master unit fails, then re-performs master control election.The application guarantees the consistency and reliability of master control election process, simultaneously realizes self-judgment and self-recovery after master control failure.
Owner:SHANGHAI FUKUN AVIATION TECH CO LTD

Data processing method and device, equipment, storage medium and computer program product

The invention discloses a data processing method, the method is applied to data processing equipment, and a processor architecture of the data processing equipment at least comprises a universal central processing unit (CPU) and a plurality of matrix acceleration modules. The method comprises the following steps: reading a real part data set and an imaginary part data set of a first data type stored in a first register in the complex matrix acceleration module; wherein the real part data set and the imaginary part data set correspond to a preset number of matrix data blocks; performing multiply-accumulate operation on the real part data set and the imaginary part data set through a first calculation unit corresponding to the first data type in the complex matrix acceleration module to obtain a first calculation result; and storing the first calculation result to a first storage area corresponding to the complex matrix acceleration module. The invention further discloses a data processing device and equipment, a storage medium and a computer program product.
Owner:CHINA MOBILE COMM LTD RES INST +1

Image processing device and image processing method

To provide an image processing device which can easily reduce noise in an object image by performing teacherless learning after pre-learning with teacher to a CNN.SOLUTION: An image processing device 1 comprises an input image creation unit 10, a first calculation unit 20 and a second calculation unit 30, and creates a noise reduction image by reducing noise in an object image 52. The first calculation unit 20 includes a first CNN processing part 21 and a first CNN learning part 22 and performs processing of pre-learning with teacher. A first input image 40 is obtained by changing a pixel value of a partial region on the basis of a teacher image 42. The second calculation unit 30 includes a second CNN processing part 31 and a second CNN learning part 32 and performs processing of teacherless learning.SELECTED DRAWING: Figure 1
Owner:HAMAMATSU PHOTONICS KK

Medical image processing method, medical image processing apparatus, X-ray CT apparatus, and medical image processing program

To generate a medical image in which object visibility and image quality in the medical image are improved.SOLUTION: A medical image processing method according to an embodiment includes: obtaining a first set of projection data by performing, with a first CT apparatus including a detector with a first pixel size, a first CT scan of an object by using a first imaging region of the detector; obtaining a first CT image with a first resolution by performing reconstruction processing on the first set of projection data; obtaining a processed CT image with a resolution higher than the first resolution by applying a machine-learning model for resolution enhancement to the first CT image; and displaying the processed CT image or outputting the processed CT image for analysis processing. The machine-learning model is obtained by machine learning using a second CT image based on a second set of projection data acquired by executing a second CT scan of the object by using a second imaging region smaller than the first imaging region of the detector with a second CT apparatus including a detector with a second pixel size smaller than the first pixel size.SELECTED DRAWING: Figure 3
Owner:CANON MEDICAL SYST CORP

System for biometric identity enrollment

User enrollment in a biometric identification system begins with a pre-enrollment process on a selected generic input device (GID), such as a smartphone. The user enters identifying data, such as their name, and may use the GID's camera to capture first image data, such as of their hand. The first image data is processed to determine a first representation. Upon presenting the hand at the biometric input device, second image data is captured. The second image data is processed to determine a second representation. If the second representation is deemed associated with the first representation, the enrollment process may be completed by saving the second representation for later use.
Owner:AMAZON TECH INC

Dpu data read-write method, device and equipment supporting multi-hypervisor device

The present disclosure relates to a DPU data read-write method, device and equipment supporting multiple semi-virtualization devices. The DPU data read-write method supporting multiple semi-virtualization devices is applied to a DPU, comprising: in response to a value in a target register changing, initiating a direct memory data access operation, obtaining metadata corresponding to a target read-write request from a host side, the value of the target register being used to notify the DPU that there is a new data read-write request to be processed; determining the request type corresponding to the target read-write request, and determining the device type corresponding to the target read-write request based on the target register; determining the target processing unit corresponding to the target read-write request based on the device type, and downlinking the target read-write request and the metadata to the target processing unit, and executing the target read-write request based on the request type and the target processing unit, thereby the same hardware unit can be flexibly configured into different types of devices in the case of no change in the hardware configuration in the DPU.
Owner:YUSUR TECH CO LTD

Processor and memory access method

The invention relates to a processor and a memory access method. The processor comprises a UAV instruction sending unit, a UAV register management unit, a general register unit, a UAV data caching unit and a loading storage unit; the loading storage unit is configured to receive an SIMD instruction; reading resource information of the target UAV register from the UAV register management unit according to the address information of the target UAV register; under the condition that the SIMD instruction is determined to be a loading instruction according to the operation code, reading channel logic address information of each channel from the general register unit; under the condition that the memory addresses to be accessed by the multiple channels of the SIMD instruction are judged to be continuous, calculating the memory access address of the first channel, and calculating the memory access addresses of the other channels based on the memory access address of the first channel; and generating a memory access request based on the memory access address of each channel. According to the method, the throughput rate of the loading storage instruction of the continuous address is improved, the number of ALUs calculated by the memory address is reduced, and meanwhile, the power consumption is reduced.
Owner:GLENFLY TECH CO LTD

Device and method for performing addition operation between quantized tensors

Disclosed is an operation method for executing an addition operation on quantized matrices. This operation method comprises the steps in which: an NPU matrix multiplication unit executes a matrix multiplication operation on a first matrix and a second matrix, wherein the first matrix has a size of R0*2C0 and is obtained by concatenating a first input tensor and a second input tensor, each of which is composed of quantized values and has a size of R0*C0, and the second matrix is obtained by concatenating a first diagonal matrix having a size of C0*C0 and having all main diagonal element values equal to integer N1 and a second diagonal matrix having a size of C0*C0 and having all main diagonal element values equal to integer N2; and an NPU vector processing unit calculates elements of a non-quantized tensor obtainable by adding the second input tensor to the first input tensor, by adding real number B to a product of real number R and a third matrix calculated by the matrix multiplication operation.
Owner:OPENEDGES TECH INC

Processor synchronization system and electronic equipment

PendingCN122086639AEliminate instruction cycle penaltiesEliminate explicit state access operationsProgram synchronisationInterprogram communicationComputer architectureSynchronous control
The invention provides a processor synchronization system and electronic equipment, and relates to the field of digital integrated circuits, the processor synchronization system comprises a plurality of execution engines, each execution engine is configured to execute an instruction stream, and the instruction stream comprises a synchronization mark instruction used for indicating a dependency relationship across the execution engines; each synchronous control module is embedded in the corresponding execution engine, and each synchronous control module is configured to generate a synchronous request signal or a synchronous event signal according to the current state of the execution engine in response to the execution of the execution engine to the synchronous mark instruction; and the inter-engine synchronization module is coupled with the plurality of execution engines to receive the synchronization request signals and the synchronization event signals and return synchronization completion signals to the corresponding execution engines when the dependency relationship is detected to be met, so that all explicit state access operations can be eliminated, and the performance of the execution engines is improved. And conversion of a synchronization mechanism from software explicit scheduling to hardware implicit response is realized.
Owner:MOFFETT AI TECHNOLOGY SHENZHEN CO LTD

Hardware accelerator

A hardware accelerator (4) includes an array (20) of direct memory access (DMA) systems (7, 8) and processing elements (PEs). Each PE (20a) includes two data inputs (40, 41) and two data outputs (42, 43) and can perform selectable logical or arithmetic operations. The array (20) includes configurable interconnects (23) for selectively connecting the outputs of the PEs to the inputs of the PEs. A first data buffer (21) includes two or more first edge circular registers (21a) for connecting the DMA systems (7, 8) to selected data inputs at the first edge of the PE array (20). A second data buffer (22) includes two or more second edge linear or circular shift registers for connecting selected data outputs at the second edge of the PE array (20) to the DMA systems.
Owner:NORDIC SEMICONDUCTOR

Systems, methods and apparatus for implementing tracked data communications on a chip

An electronic chip, chip assembly, device, system, and method enabling tracked data communications. The electronic chip comprises a plurality of processing cores and at least one hardware interface coupled to at least one of the one or more processing cores. At least one processing core implements a game and / or simulation engine, at least one processing core implements a position engine, and at least one processing core implements a gyroscope and, optionally, an IMU. The at least one position engine obtains pose data from an external positioning system comprising GNSS augmented by millimeter-wave cellular networks and / or Wi-Fi; and internal pose data from the gyroscope, optional IMU, and game and / or simulation engine, the data comprising inertial, 3D structure, and simulation data, thereby computing a 6 DOF pose of the client device, driving processing of 3D applications by the one or more game and / or simulation engine.
Owner:THE CALANY HLDG S.À.R L

In-memory computing macro devices and electronic devices

This application discloses an in-memory computing macro device and an electronic device. The in-memory computing macro device includes an array of in-memory computing units comprising a plurality of in-memory computing units. First data is divided into at least two bit groups, the at least two bit groups including a first bit group and a second bit group, the first bit group being the most significant bit of the first data, the second bit group being the least significant bit of the first data, and the bit groups are respectively loaded into in-memory computing units in different columns of the in-memory computing unit array. The electronic device includes at least one in-memory computing macro and at least one processing circuit. The processing circuit is configured to receive parallel outputs corresponding to columns of the in-memory computing unit array and perform operations on the parallel outputs, wherein the parallel outputs include a plurality of corresponding groups, and each of the corresponding groups includes the most significant bit and the least significant bit of the output activation.
Owner:NOVATEK MICROELECTRONICS CORP

Controllers in data processing engine columns

Embodiments herein describe a hardware accelerator with an array of data processing engines (DPEs) which includes a controller (e.g., a microcontroller) for multiple columns of the array. The controllers can be hardened circuitry that executes software code (or firmware) that controls the hardware accelerator. In one embodiment, the task of the controller is to control and orchestrate the functions performed by the hardware accelerator.
Owner:XILINX INC

Image super-resolution processing method and system based on WebGPU technology

The invention provides an image super-division processing method and system based on a WebGPU technology, and the method comprises the following steps: S1, designing an image super-division model: designing the image super-division model through the combination of Transform and a convolutional neural network; s2, image super-division model training: offline training an image super-division model which is suitable for a WebGPU computing architecture and based on a lightweight neural network, and improving the reasoning performance and the memory utilization rate through model pruning, weight quantization and operator fusion for a resource-limited scene of a browser end; s3, constructing a front-end platform: constructing a front-end website by adopting a mainstream Web development framework; s4, model loading and reasoning process implementation; and S5, displaying and post-processing an output result. The method solves the problems that an existing image super-division model is low in computing resource utilization rate, poor in real-time performance, insufficient in platform compatibility and the like in the Web end deployment process.
Owner:GUANGDONG BOHUA UHD INNOVATION CENT CO LTD

Efficient processing on a dual core processing system

A system may include: a plurality of processing cores including at least a first processing core and a second processing core; a shared memory communicatively coupled to each of the plurality of processing cores, and accessible by each of the plurality of processing cores; a global monitor communicatively coupled to each of the plurality of processing cores and configured to control exclusive access to the memory by each of the plurality of processing cores; and a software architecture embodied in a non-transitory computer readable medium and configured to divide a plurality of processing tasks between the first processing core and the second processing core when read and executed by the multi-core processor.
Owner:CIRRUS LOGIC INT SEMICON LTD

Data processing array

Apparatuses, computer programs and methods are disclosed, relating 2D arrays of data elements and a 2D array of processing elements. A first 2D array of data elements provides data values for processing by each processing element and a second 2D array of data elements provides control values controlling the processing. The 2D array of processing elements has a data flow direction across the 2D array of processing elements, and the data flow direction proceeds from a starting set of processing elements of the 2D array of processing elements. For each processing element not in the starting set of processing elements, the data processing operation preformed takes as an operand a respective data value provided by a neighbouring data element of the first 2D array of data elements and selection of the neighbouring processing element is controlled by a corresponding source control value in the second 2D array of data elements.
Owner:ARM LTD