Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Processor array" patented technology

A processor array is like a storage array but contains and manages processing elements instead of storage elements.

Pipeline architecture for bitwise multiplier-accumulator (MAC)

A unit for accumulating multiplied bit values includes an array of bit-line processors. The unit is implemented in an in-memory associative processor, and each bit-line processor includes multiple memory cells coupled to a bit-line. The array of processors is arranged in rows and columns. The array passes bits of a first multiplicand vertically down a column and provides bits of a second multiplicand horizontally across a row. The array generates carry bits and passes them vertically to a subsequent processor in the same column. The array also generates sum bits and passes them diagonally to a subsequent processor in an adjacent column. The array includes multiplying processors, summing processors, and accumulator processors. Multiplying processors perform an XOR operation by simultaneously activating two memory cells and then perform a full adder operation. Summing processors perform a full adder operation. Accumulator processors perform a full adder operation that includes a feedback sum bit from a previous cycle.
Owner:GSI TECHNOLOGY INC

Compute time point processor array for solving partial differential equations

Embodiments relate to a system for solving partial differential equations. The system receives a problem to be solved comprising a partial differential equation and a domain. A solver stores a plurality of nodes of the domain corresponding to a first time-step, and processes the nodes over a plurality of time-steps using an array of point processors. Each point processor comprises an ALU and a register file, and is configured to receive data corresponding to a respective node of a domain and generate a value for the node for a next time step, based upon instructions received over time via an instruction stream.
Owner:VORTICITY INC

RTL code simulation method and device, storage medium, computer program product, server and system

The invention discloses an RTL code simulation method and device, a storage medium, a computer program product, a server and a system.The RTL code simulation method comprises the steps that a to-be-simulated RTL code is obtained; in response to a simulation acceleration mode selected as a periodic accurate simulation acceleration mode, decomposing the RTL code into a plurality of sub-modules, and performing feature extraction on each sub-module; configuring an FPGA (Field Programmable Gate Array) according to the respective characteristics of the plurality of sub-modules so as to configure hardware resources on the FPGA into an initial processor array matched with the plurality of sub-modules; and mapping the RTL code into an instruction stream, and downloading the instruction stream to the initial processor array for simulation, wherein the instruction stream is generated according to an instruction set used by the initial processor array. Therefore, the processor array can be reconstructed, free switching of multiple simulation modes is realized, and the simulation efficiency and the simulation flexibility are improved.
Owner:ZHEJIANG YIFANG HANGCHUANG TECHNOLOGY CO LTD

Method and system for unloading configuration data in a reconfigurable processor array

A method and system for unloading configuration data in a reconfigurable processor array comprises a bus system, and array of processor units connected to bus system, the processor units in the array including configuration data stores to store unit files comprising plurality of subfiles of configuration data particular to corresponding processor units. A configuration unload controller is connected to the bus system, including logic to execute an array configuration unload process, including distributing a command to plurality of the processor units in array to unload the unit files particular to corresponding processor units, the unit files each comprising plurality of ordered sub-files, receiving sub-files via bus system from the array of process units, and assembling an unload configuration file by arranging the received subfiles in memory according to the process unit of the unit file of which the subfile is a part, and order of the subfile in unit file.
Owner:SAMBANOVA SYSTEMS INC

Hybrid reconfigurable computing chip based on encryption and decryption computing and computer equipment

The invention relates to a hybrid reconfigurable computing chip based on encryption and decryption computing and computer equipment. The method comprises the following steps: acquiring a first prime number and a second prime number, and determining a public key and an Euler function value according to the first prime number and the second prime number; the public key comprises a public key index and a modulus; performing linear operation in a private key calculation process by using the public key index and the Euler function value to obtain a first calculation result; outputting a first calculation result to a Boolean processor calculation core of a fine-grained reconfigurable architecture through a processing unit array data transmission unit and a Boolean processor array data transmission unit in sequence, so that the Boolean processor calculation core performs structured operation in a private key calculation process by using the first calculation result and an Euler function value, and a private key is obtained and output to the coarse-grained calculation core through the Boolean processor array data transmission unit and the processing unit array data transmission unit in sequence. By adopting the method, the key calculation efficiency can be improved.
Owner:TSINGHUA UNIVERSITY

HYBRID DATA MODEL PARALLELITY FOR EFFICIENT DEEP LEARNING

Method (500) which exhibits: Selecting a hybrid parallelism technique (140) for distributing the operational load of a layer of a neural network across an array (200, 600) of processors (110, 155), wherein each processor in the array of processors can transfer data to neighboring processors in a first direction (610) and in a second direction (605); and Assigning tasks corresponding to the layer of the neural network to the array of processors using the selected hybrid parallelism technique (145), wherein the hybrid parallelism technique includes the use of a first parallelism technique when data is transferred between processors in the array of processors in the first direction (610), and the use of a second, different parallelism technique when data is transferred between processors in the array of processors in the second direction (605).
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Modular design process

A method and apparatus for constructing a canonical data structure for a module of a multiprocessor array (MPA) chip. The normative data structure includes: parameters for combining a plurality of register transfer language (RTL) templates for sub-modules of the module into RTL description of the module; the parameter combination module is used for combining a plurality of test bench templates for corresponding sub-modules into parameters of the test bench for the modules; the combination module is used for combining a plurality of physical design script templates for corresponding sub-modules into parameters of physical design scripts for the modules; and / or for constructing parameters of an API for a module based on a set of functional criteria for module operation. The RTL description, the test bench, the physical design script, and / or the API will be built and stored in memory for designing and manufacturing the module.
Owner:HYPEX HOLDINGS LLC

Large model reasoning asynchronous scheduling method and device

The invention provides a large model reasoning asynchronous scheduling method and device, and the method comprises the steps: receiving an input data load, and dividing the input data load into a first data subset and a second data subset; dividing a plurality of physical cores in the spatial processor array into a calculation partition and a communication partition; inputting the first data subset into a calculation partition, executing calculation processing and outputting a first intermediate result; performing converged communication and normalization processing on the first intermediate result in the communication partition to generate a first final result; when the communication partition processes the first intermediate result, inputting the second data subset into the calculation partition, executing calculation processing and outputting a second intermediate result; converged communication and normalization processing are continued to determine a second final result. According to the method, the tasks can be cooperatively scheduled in space and time, the communication overhead is remarkably reduced, and the reasoning throughput and efficiency of the data stream processor are improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

3D semiconductor device and structure with logic circuits, memory cells, and processor array

PendingUS20260129877A1Memory cellDevice material
An integrated semiconductor device including: a first level including single crystal silicon and logic circuits each include first transistors; a second level, disposed above the first level and includes arrays of first memory cells, where the second level includes a plurality of second transistors, where each of the first memory cells includes at least one of the second transistors, where the first level is bonded to the second level; an array of processors; a plurality of SerDes circuits; and a third level, where the third level includes a plurality of third transistors, where the third level is disposed above the second level and includes a plurality of arrays of second memory cells, where each of the second memory cells includes at least one of the third transistors, where the device includes a substrate having an area greater than 1,000 mm2, and where the substrate includes at least one interconnect.
Owner:MONOLITHIC 3D INC

Method and system for parallel processing of neural network model on processor array

The invention discloses a method for parallel processing of a neural network model on a processor array, and the method comprises the steps: employing a mixed weight division strategy for the to-be-processed Transform neural network model, distributing the to-be-processed Transform neural network model to the processor array for parallel calculation, and carrying out the parallel calculation of the to-be-processed Transform neural network model in the processor array, dividing weight matrixes of query Wq, key Wk and Wv values in an attention module of the neural network model into different processor groups along a column direction, and dividing an output projection matrix Wo into the same processor group along a row direction; in the processor array, different expert networks in the feedforward network module are completely and uniformly distributed to each processor in the processor array, and a router weight matrix is copied to all processors. According to the method and the device, the neural network model calculation is efficiently mapped to the multi-chip system, and the method and the device have high parallel calculation efficiency.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Rapid rooting method and device for double error correction BCH and RS code decoding

The invention provides a rapid rooting method and device for double error correction BCH and RS code decoding, and the method comprises the steps: S1, pre-calculating elements of a finite field, and obtaining a binary matrix; step S2, converting the matrix into a row simplest shape, recording a corresponding row transformation matrix, and if yes; s3, inputting a row transformation matrix obtained through pre-calculation and all-zero row position information of the matrix through an input component, and converting elements into binary column vectors of the elements; s4, a judgment vector is calculated through the parallel processor array; s5, comparing the calculation result of the judgment vector with the position of the all-zero line of the matrix, and if the calculation result is the same as the position of the all-zero line of the matrix, outputting a root; if not, outputting no root. According to the optimized technical scheme, the calculation complexity of the rooting process and the occupation of resources and energy consumption are effectively reduced, the hardware design is simplified, and the rooting efficiency and adaptability are improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

Digital twin safety system of key infrastructure

The utility model relates to the technical field of safety protection of information technology and physical infrastructure fusion, in particular to a digital twin safety system of a key infrastructure. The system is composed of a multi-source heterogeneous data acquisition module, a secure transmission module, a digital twin modeling engine, a security protection module and a visual interaction module. The multi-source heterogeneous data acquisition module is in butt joint with a target facility through a wired / wireless dual-mode communication interface; the secure transmission module is configured with a physically isolated data channel; the digital twin modeling engine comprises a parallel processor array and a storage medium; the safety protection module is provided with a bidirectional data bus which is respectively connected with the digital twin modeling engine and the visual interaction module; the visual interaction module comprises a display terminal and an alarm unit. According to the utility model, full-chain safety guarantee from physical perception to virtual protection is realized through cooperation of multiple modules.
Owner:GP CAPITAL GROUP LTD

Reducing overhead in processor array searching

A processor with instruction storage configured to store processor instructions, data storage configured to store processor data representing an array, the array including plural data elements, a controller, and an instruction pipeline. The instruction pipeline includes: a load stage circuit configured to load an array element from the data storage, a compare stage circuit configured to compare the array element to a reference value, a store stage circuit configured to store a set of results that includes a result of the comparison of the array element to the reference value, and a loop hit detect stage circuit configured to determine whether any of the set of results is associated with a hit on the reference value.
Owner:TEXAS INSTRUMENTS INC

Graph streaming neural network processing system and method thereof

Disclosed herein is a graph streaming neural network processing system comprising a first processor array, a second processor, and a thread scheduler. The thread scheduler dispatches a thread of a first node to the first processor array or the second processor, wherein the thread is executed to generate output data comprising a data unit stored in a private data buffer of the second processor. The thread scheduler determines that the data unit is sufficient for executing a thread of a second node. The second node is dependent on the output data generated by execution of a plurality of threads of the first node. Upon determining that the data unit is sufficient, the thread scheduler dispatches the thread of the second node. The thread scheduler determines to dispatch a subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer.
Owner:BLAIZE INC

Mixed precision controller and method for the same

A Mixed Precision controller configured to operate in a High-Performance Computing, HPC, system, where the HPC system includes a high-precision processor array and a low-precision processor array. The Mixed Precision controller is configured to receive a system of partial differential equations and utilize the high-precision processor array and the low-precision processor array to solve the system of partial differential equations utilizing a fixed-point iterative scheme, g. The Mixed Precision controller is characterized in that the Mixed Precision controller is configured to utilize a defect correction term which is defined as the difference between a high-precision residual, RH, and a low-precision residual, RL, where the defect correction term is added to an independent constants vector of the system of partial differential equations. The use of the MP controller in the HPC system enables high-precision accuracy of final solution while reducing computational cost and overall execution time of non-linear partial differential equations.
Owner:HUAWEI TECH CO LTD +1

Method and device for controlling a vehicle's alert system in the presence of an object on the vehicle's roof

The present invention relates to a method and control system for a vehicle warning system belonging to a system comprising an array of external cameras connected in communication with the vehicle. To this end, the method is implemented by an array of processors in the system and comprises receiving (51) image data representing an image of the vehicle's roof from at least one camera in the array of external cameras, and detecting (52) an object on the roof by an object detection model based on the image data. Warning data is transmitted (53) to the vehicle when an object is present on the vehicle's roof, and the vehicle warning system is controlled (54) based on the warning data received by the vehicle. Figure 5
Owner:STELLANTIS AUTO SAS +1

Point processor array for solving partial differential equations

Embodiments relate to a system for solving partial differential equations. The system receives problem packages corresponding to problems to be solved, each comprising at least a partial differential equation and a domain. A solver stores a plurality of nodes of the domain corresponding to a first time-step, and processes the nodes over a plurality of time-steps using an array of point processors. Each point processor comprises a series of tiles, each having a computational element and a router, and are configured and connected based on a discretized form of the partial differential equation, to allow each point processor to receive a node of the domain and generate a value for the node for a next time step. Because all the data and computational requirements of the point processors are determined at compile time, no dynamic scheduling needs to be performed, allowing for more efficient usage of computational resources.
Owner:VORTICITY INC

A three-dimensional processor array reconstruction method for removing bottlenecks and compensation

The application discloses a three-dimensional processor array reconstruction method for removing a bottleneck and compensation, which removes a bottleneck surface limiting the size of a logic array to break the limitation of the size of array reconstruction, and uses a mechanism of compensating for a faulty unit of a neighbor by using a faultless processor unit on the bottleneck surface, so as to effectively improve the utilization rate of the faultless processor unit in the reconstruction process, break the limitation of the size of processor array reconstruction, increase the reconstruction size of the three-dimensional processor array, and further set a maximum size replacement time to realize time limitation, so that invalid iteration can be terminated in time, and the timeliness of the algorithm is ensured.
Owner:GUANGXI NORMAL UNIV

Deep displacement monitoring system

The utility model relates to a deep displacement monitoring system, which comprises an acquisition instrument main body, an array displacement meter and a cable assembly, the acquisition instrument main body is electrically connected with the array displacement meter through the cable assembly; the acquisition instrument main body comprises a shell, a sealing cover, a bracket, a power supply device and a processor; the support is installed in a cavity of the shell, the processor is fixedly arranged on the support, and the power supply device is installed in the cavity of the shell. The power supply device comprises a battery and a voltage converter. The output end of the battery is electrically connected with the input end of the voltage converter, and the output end of the voltage converter is electrically connected with the processor and the array displacement meter. A low-energy-consumption control module is arranged on the processor; the input end of the low-energy-consumption control module is electrically connected with the output end of the voltage converter; the output end of the low-energy-consumption control module is electrically connected with the input end of the processor and the input end of the array type displacement meter. The low-energy-consumption control module is suitable for controlling the acquisition instrument body and the array type displacement meter to be switched between the standby state and the awakening state according to the preset switching frequency.
Owner:BEIJING ZHONGHONG TAIKE TECH CO LTD

Reconstruction method for maximizing three-dimensional processor array

The invention discloses a reconstruction method for maximizing a three-dimensional processor array, which adopts a preprocessing mechanism to mark unavailable units in advance, so that a logic plane construction process does not need to be backtracked, and the reconstruction efficiency is greatly improved; a low-efficiency strategy of carrying out full-amount scanning on a physical plane through a bottleneck identification strategy based on a unit cluster is adopted, so that the calculation redundancy and the time overhead are remarkably reduced; a preprocessing mechanism and a fast identification bottleneck strategy improve the utilization rate of a fault-free processing unit in the processor to a certain extent, the reconstruction efficiency of the three-dimensional reconfigurable processor array is improved, the reconstruction time is shortened under the condition that harvest is not lost, the scale of the reconfigurable processor array is increased, and the reconstruction efficiency of the reconfigurable processor array is improved. Therefore, the size of the target array is obviously increased. And meanwhile, in the iterative logic array reconstruction process, an iterative method and a multi-direction construction strategy are combined, and the optimization is selected by comparing reconstruction results of different dimensions so as to adapt to different fault distributions, so that larger-scale subarray construction is realized.
Owner:GUANGXI NORMAL UNIV

Graph streaming neural network processing system and method thereof

Disclosed herein is a graph streaming neural network processing system comprising a first processor array, a second processor, and a thread scheduler. The thread scheduler dispatches a thread of a first node to the first processor array or the second processor, wherein the thread is executed to generate output data comprising a data unit stored in a private data buffer of the second processor. The thread scheduler determines that the data unit is sufficient for executing a thread of a second node. The second node is dependent on the output data generated by execution of a plurality of threads of the first node. Upon determining that the data unit is sufficient, the thread scheduler dispatches the thread of the second node. The thread scheduler determines to dispatch a subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer.
Owner:BLAIZE INC

Reconstruction method of maximized three-dimensional processor array containing switch fault

The invention discloses a reconstruction method of a maximized three-dimensional processor array containing switch faults, which considers the switch faults and processing unit faults at the same time, utilizes the idea of preprocessing before reconstruction when reconstructing a three-dimensional physical array, firstly preprocesses fault switches and related processing units and planes, and then reconstructs the fault switches and the related processing units and planes. The preprocessing realizes dimension reduction fault tolerance by eliminating connectivity limited planes, so that the existing reconstruction algorithm can be flexibly embedded. And on the basis of preprocessing, performing logic array reconstruction by using an existing GPR algorithm. According to the method, the utilization rate of the processing unit is effectively improved in a complex fault scene, and the maximum connected logic array can be constructed under constraint conditions.
Owner:GUANGXI NORMAL UNIV

3D semiconductor device and structure with logic circuits, memory cells, and processor array

An integrated semiconductor device including: a first level; a second level, where the first level includes single crystal silicon and a plurality of logic circuits, where the plurality of logic circuits each include first transistors, where the second level is disposed above the first level and includes a plurality of arrays of first memory cells, where the second level includes second transistors, where each of the first memory cells includes at least one of the second transistors, where the first level is bonded to the second level; an array of processors; and a third level, where the third level includes third transistors, where the third level is disposed above the second level and includes a plurality of arrays of second memory cells, where each of the second memory cells includes at least one of the third transistors, where the device includes a substrate area greater than 1,000 mm2.
Owner:MONOLITHIC 3D INC

Unmanned aerial vehicle thrust line measuring device and measuring method based on inertial measurement unit

The application provides a kind of unmanned aerial vehicle thrust line measuring device and measuring method based on inertial measurement unit, comprising: measuring cylinder installation connecting seat, it includes circular bottom plate and annular side wall, also includes tubular connecting piece, connecting piece includes first section and second section connected with each other, circular bottom plate is opened through hole, tubular connecting piece is fixed at through hole;PCB circuit board is installed in connecting seat, processor, IMU array and external communication interface are arranged on PCB circuit board, IMU array is used to collect the attitude angle of measuring cylinder, processor is used to data analysis processing, and is exported through external communication interface;Upper cover is fixedly installed on connecting seat;Measuring cylinder and boom, the top of measuring cylinder is fixed with the second section of connecting piece, the bottom is fixed on unmanned aerial vehicle, boom is successively penetrated measuring cylinder, connecting piece and upper cover, and the bottom is suspended unmanned aerial vehicle.The application realizes the automation of measuring process, and the measuring data can be traced back, improves the measuring efficiency and measuring accuracy.
Owner:INST OF ENGINEERING THERMOPHYSICS - CHINESE ACAD OF SCI

Advanced state inspection and dynamic execution control for reconfigurable processors

A data processing system includes compile time logic configured to generate configuration files for applications executing on reconfigurable processors, and execution flow logic configured to manage execution with conditional breakpoint capabilities. The system provides state inspection interfaces for examining processor state, memory contents, and application variables during stopped execution. Runtime logic executes configuration files with support for conditional breakpoints triggered by runtime data values, performance metrics, or error conditions. The system enables multi-processor synchronization with coordinated breakpoint handling across processor arrays, dynamic reconfiguration of processor configurations during stopped execution, and performance profiling at breakpoint locations. Breakpoint conditions can be defined at multiple granularity levels for computational graphs and dataflow applications. The system supports interactive debugging with step-through execution control and real-time modification of execution parameters.
Owner:SAMBANOVA SYSTEMS INC

Rtl code simulation method and device, storage medium, computer program product, server and system

An RTL code simulation method and device, a storage medium, a computer program product, a server and a system, wherein the RTL code simulation method comprises: obtaining RTL code to be simulated; in response to a simulation acceleration mode being selected as a cycle-accurate simulation acceleration mode, decomposing the RTL code into a plurality of sub-modules, and performing feature extraction on each of the sub-modules respectively; configuring an FPGA according to the features of the plurality of sub-modules respectively, to configure hardware resources on the FPGA as an initial processor array matched with the plurality of sub-modules; and mapping the RTL code to an instruction stream and downloading the instruction stream onto the initial processor array for simulation, the instruction stream being generated according to an instruction set used by the initial processor array. Thus, the processor array can be reconstructed, free switching of multiple simulation modes can be achieved, and simulation efficiency and simulation flexibility are improved.
Owner:ZHEJIANG YIFANG HANGCHUANG TECHNOLOGY CO LTD