Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

263 results about "Multi processor" patented technology

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

High performance code parallelization compiler with loop level parallelization

PendingCN121464428ACode compilationComputer architectureLoop level parallelism
A symmetric auto-compiler system (1) and method for high performance hardware optimized auto-parallelization of program code (3) executed by a multi-core or multi-processor parallel processing system (2) having a plurality of processing units (21) that simultaneously process instructions for data in the parallel processing system (2) by executing the program code (3). The automatic compiler system (1) converts serial source code (31) of program code (3) into parallel processing machine code (32) comprising a plurality of instructions executable by a plurality of processing units (21) of the parallel processing system (2) or controlling operation of the plurality of processing units (21).
Owner:MINATIX INC

Method, apparatus and system for monitoring ultra-high frequency partial discharge of hydro-generator

The present disclosure provides a method, apparatus and system for monitoring ultra-high frequency partial discharge of a hydro-generator, and belongs to the technical field of hydro-generator partial discharge monitoring. The method includes: cleaning a partial discharge pulse sequence using a cleaning threshold to obtain a valid pulse sequence; performing redundant data filtering on each data unit divided from the valid pulse sequence to obtain a first target pulse sequence for short-period partial discharge monitoring; determining sub-sequences that are partial discharge events from the valid pulse sequence, forming a second target pulse sequence after associating an amplitude statistical feature of the sub-sequences, and storing the second target pulse sequence. The aforementioned method combines data cleaning, redundant data filtering, and partial discharge event identification, which enhances the real-time performance of partial discharge monitoring and records the long-period partial discharge change trend. The corresponding system adopts a multi-buffer zone and multi-processor architecture, thus further improving the real-time performance of partial discharge monitoring.
Owner:GUODIAN SCI & TECH RES INST

Communication architecture for multicore system

A device can provide a unified and scalable multiprocessor communication framework that enables communication between multiple processor cores in a multicore device. For example, the device may be configured to perform data payload management using a Smart Message Queue (SMQ) and / or shared memory. Additionally or alternatively, the communication framework may enable the processors to communicate with each other and / or peripherals of the device while abstracting details of various communication protocols, hardware interfaces, and / or the like. For example, the communication framework may provide a common interface for applications, enabling a first application associated with a processor to establish a connection to any processor and / or peripheral without knowing details of a communication protocol for each connection.
Owner:AMAZON TECH INC

Telemetry assisted hybrid load balancing on switch fabric paths

Techniques described herein can use a hybrid load balancing approach to balance loads on paths in a switch fabric. The switch fabric can deliver synchronization data between processors in a multi-processor cluster, and the synchronization data can load the paths on which it is sent. First paths can be identified in the switch fabric, and first synchronization data can be distributed to the first paths using a first load balancing approach, such as a telemetry assisted load balancing approach. Second paths can be identified in the switch fabric, and second synchronization data can be distributed to the second paths using a second load balancing approach, such as a packet spraying load balancing approach.
Owner:CISCO TECHNOLOGY INC

Fully cache coherent virtual partitions in multitenant configurations in a multiprocessor system

Various embodiments include techniques for processing memory operations in a computing system. The computing system includes a central processing unit (CPU) and an auxiliary processor, such as a parallel processing unit (PPU). The PPU can be divided into multiple partitions. Although the partitions are included in a single PPU, the CPU can track the partitions as if the partitions are independent devices rather than different portions of a single device. When two different partitions generate memory operations that access the same memory address in CPU memory address space, the two partitions employ two different data paths. The CPU can use path information for the two different paths to identify which partition generated each memory operation. As a result, the CPU can maintain data consistency and memory coherency in a system where a PPU is divided into multiple partitions.
Owner:NVIDIA CORP

High performance code parallelization compiler with loop level parallelization

PendingCN121399576ACode compilationHandling CodeComputer architecture
A system and method for universal static multi-transmit CPU design with a static pipeline for automatically parallelizing code is presented. A multi-core and / or multi-processor integrated circuit (2) has a plurality of processing units (21) and / or processing pipelines (53) that simultaneously process instructions for data by executing parallel processing machine code (32). The execution of the parallelized processing code (32) by the parallel processing multi-core and / or multi-processor integrated circuit (2) comprises the occurrence of a delay time (26), wherein the delay time is given by an idle time between the processing unit (21) returning the data after processing a specific instruction block of the processing code (32) for the data and receiving the data required by the processing unit (21) to execute a consecutive instruction block of the processing code (32). The parallel pipeline (53) comprises means for: (i) forwarding by providing a data forwarding from a MEM stage as an EX / MEM register to an EX stage as an ID / EX-stage register; (ii) exchanging by making the results of the Ex-ME-phase registers accessible by the Ex phase of a parallel pipeline (53) to provide a result exchange between the pipelines; and (iii) implementing branch pipeline refresh by providing control conflicts by refreshing only those pipelines (53) dependent on one pipeline (53) based on branch address computation conditions.
Owner:MINATIX INC

A Typed Task Co-scheduling System and Method Based on Heterogeneous Multi-core Architecture

This invention relates to the field of multi-processor multi-task joint scheduling technology, specifically disclosing a typed task joint scheduling system and method based on a heterogeneous multi-core architecture. It introduces a joint scheduling mechanism to divide and sort tasks according to task size and priority. Based on this, for the time-limited characteristics of real-time tasks, an inertial weight coefficient particle swarm optimization method is used to iteratively update the load balancing strategy, minimizing the maximum response time of real-time tasks while meeting their schedulability requirements. Simultaneously, for the low-priority characteristics of non-real-time tasks, the problem-solving is simplified using a Lagrange-based convex optimization approach, and an energy-constrained binary search algorithm is employed to effectively reduce the average response time of non-real-time tasks, achieving optimal load distribution for the system.
Owner:CHONGQING UNIV +2

Multi-processor communication framework using TCP / IP architecture

A multi-processor communication framework (MCF) based on Transmission Control Protocol / Internet Protocol (TCP / IP) architecture is defined that can be used for inter-processor communication between heterogenous processing nodes. The MCF includes an application layer, a transport layer, a network layer, and an inter-processor communication (IPC) layer for each processing node. The IPC layer can be used to facilitate data transfer between two processing nodes via a physical channel within an electronic device, or through shared memory. The MCF can support multicast, broadcast, unicast, and zero-copy features like the TCP / IP stack.
Owner:AMAZON TECH INC

GEMM load-oriented GPU modeling method

A GEMM load-oriented GPU modeling method is characterized in that through a multi-stage collaborative modeling mechanism, cache behaviors, instruction overhead and calculation intensity are deeply coupled, accurate performance prediction of GPU execution GEMM operators is realized, the method can be widely applied to scheduling optimization of GPU intensive scenes such as AI training and scientific calculation, firstly, a three-stage cache weight distribution mechanism is established, and then, a three-stage cache weight distribution mechanism is established; quantifying the contribution of the L1 / L2 cache hit rate and the DRAM bandwidth degradation factor to the effective bandwidth; secondly, an instruction-level memory access overhead correction mechanism is introduced, and the mixing precision and the real calculation strength of a sparse calculation scene are captured through dynamic parameter adjustment and optimization; then combining the calculation force peak value and the bandwidth upper limit to construct a double-boundary constraint model, and generating a theoretical performance critical value; further predicting a stream multiprocessor utilization rate based on a neural network, and quantifying efficiency loss caused by hardware resource contention through a multi-layer perceptron structure; and finally, the integration module outputs task execution time to realize end-to-end performance prediction.
Owner:BEIHANG UNIV

Communication calculation parallel optimization method, multiprocessor system, medium and program product

The invention discloses a communication computing parallel optimization method, a multiprocessor system, a medium and a program product, and the method comprises the steps: configuring a first computing core and a second computing core on a single task flow for a general matrix multiplication subtask allocated to a single processor; wherein the first calculation core is used for executing a general matrix multiplication subtask, and the second calculation core is used for executing a set communication task; the set communication task comprises a full accumulation operator or a protocol dispersion operator; then, starting scheduling is conducted on the first calculation core and the second calculation core according to a preset dependency mechanism, so that the general matrix multiplication subtask and the set communication task are executed asynchronously in an overlapped mode; according to the method, the problem of poor reusability of the original Kernel caused by intrusive modification of the GEMM or re-implementation of the Kernel can be effectively avoided through parallel optimization of GEMM calculation and ensemble communication operation, and the performance overhead of the processor is reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Hardware-aware attention mechanism with dynamic workload distribution for transformer models

A technique for optimizing attention mechanism computations in transformer-based language models improves computational efficiency during both prefill and decode phases. The approach unequally partitions attention operations across multiple streaming multiprocessors of a hardware processing unit (e.g., such as a graphics processing unit, or GPU) to maximize hardware utilization. By leveraging the associative property of online softmax calculation as a reduction operation and employing stream-K style decomposition, the technique enables parallelization across all modes of the attention matrix, including the context length dimension. This allows for efficient distribution of computational workload across available GPU resources while ensuring equal total work allocation. The approach delivers significant speedup over existing methods, particularly for long context lengths, by maintaining near 100% GPU occupancy through optimal workload distribution and single-kernel execution.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Systems, methods, apparatus, and core particles for interrupt handling

Disclosed are a system, method, device and core particle for interrupt handling, the system comprising: a manager, a repeater and an interrupter wherein the manager is used for sorting the response sequence of interrupt signals to send the sorted interrupt signals to the repeater; the repeater is used for filtering the sorted interrupt signals based on a filtering condition so as to send the filtered interrupt signals to a target core particle; and the interrupter is used for distributing the filtered interrupt signal received by the target core particle to the processor in the target core particle for processing according to the response sequence. Through the scheme of the invention, the interrupt distribution and processing flow among the core particles can be optimized by utilizing the characteristics of a multi-processor architecture.
Owner:SHANGHAI PROCESSOR TECH INNOVATION CENT

Heterogeneous file generation method and device, equipment, medium and product

The invention provides a heterogeneous file generation method and device, equipment, a medium and a product. The method comprises the steps that before a binary code of any processor architecture is written into a heterogeneous file, whether a binary code of a first processor architecture corresponding to the processor architecture exists in the heterogeneous file or not is determined; if yes, metadata of binary codes of the processor architecture is stored in a storage area of the heterogeneous file, the binary codes of the processor architecture do not execute write-in operation, and the metadata comprises identification information of the processor architecture, offset and the size of a memory occupied by the binary codes. The offset and the occupied memory size of the binary code of the processor architecture are the same as the offset and the occupied memory size of the binary code of the first processor architecture. According to the method and the device, repeated binary codes are not stored repeatedly, and multiplexing of a multiprocessor architecture can be realized while the file volume is reduced.
Owner:JIANGSU DAWN INFORMATION TECH CO LTD

Systems and methods of preconfiguring coherency protocol for computing systems

A multi-processor computing system (e.g., a system-on-chip) can store, in a shared memory, (i) a reservation table that is accessible by the one or more workload processors, and (ii) a scheduling program. The system can further execute the scheduling program to schedule execution of a set of workloads by one or more workload processors in accordance with an optimized compute graph, an optimized data positioning graph, and a coherence protocol that is precomputed based on the optimized compute graph and the optimized data positioning graph.
Owner:MERCEDES BENZ GROUP AG

A GPU warp scheduling method and device based on a chessboard scheduling strategy

The application provides a GPU Warp scheduling method and device based on a chessboard scheduling strategy, and the method comprises the following steps: a stream multiprocessor assigns a Warp number to each task of different task types received through a facing Warp numbering mechanism, the priority of each Warp is determined according to the task type and the Warp number, the Warp with high priority is preferentially executed, if the current Warp execution is overdue, the priority of the Warp is dynamically adjusted, and the Warp with low priority is executed in turn. The facing Warp numbering mechanism is used to realize the priority ordering of the Warps of different task types, and the dynamic priority adjustment mechanism based on the execution time threshold can guarantee that the high-priority task is preferentially executed, also take into account that the low-priority task can be executed in a hidden manner, ensure fairness, and improve resource utilization.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Server

PCT designated stageWO2026174750A1Uniprocessor systemMulti processor
The embodiments of the present application relate to the technical field of servers, and specifically relate to a server. The embodiments of the present application aim to solve the problem of difficulty in switching between a single-processor system and a multi-processor system. The server provided in the embodiments of the present application comprises: a housing, which comprises a panel; a first processor and a second processor, which are arranged in the housing; and an expansion circuit board, which is pluggably connected to the panel. The expansion circuit board comprises a first connector and a second connector, wherein the first connector is configured to couple to the first processor, and the second connector is configured to couple to the second processor; and the expansion circuit board is provided with a baseboard management controller, wherein the first connector and the second connector are each coupled to the baseboard management controller; alternatively, the first connector and the second connector are coupled to each other, and the first connector and the second connector are coupled to the baseboard management controller. The server can switch between a single-processor system and a multi-processor system, and the switching operation is relatively simple.
Owner:HUAWEI TECH CO LTD

Access address configuration method, processor, multiprocessor system, medium and product

The invention discloses an access address configuration method, a processor, a multiprocessor system, a medium and a product. The method comprises the following steps: configuring an independent register block for each thread bundle in the processor; wherein the register group is used for storing coordinate information of tensor data used in the execution process of a corresponding thread bundle program; reading the coordinate information of the tensor data to be processed from the register group, and sending the read coordinate information to a coprocessor to carry out access address calculation of the corresponding tensor data; according to the method, each thread bundle is provided with a group of entity registers for storing the coordinate information of the tensor data, multiple multiplexing of the coordinate information can be supported, and a large number of instructions for specifying coordinates are saved, so that the number of assembly instructions is reduced, the instruction overhead is effectively reduced, the execution time of a processor is saved, and the processing efficiency of the processor is improved.
Owner:SHANGHAI BIREN TECH CO LTD

GPGPU thread block scheduling method and system based on data space locality

The invention belongs to the field of GPGPU chip design, and particularly relates to a GPGPU thread block scheduling method and system based on data space locality, and the method comprises the steps: carrying out the statistics of Bank feature information of access data of all thread blocks, and classifying the thread blocks accessing the same Bank into the same Bank group according to the Bank feature information; preferentially distributing the thread blocks of the same Bank group to the same programmable multiprocessor until the resources of the programmable multiprocessor are saturated; setting a private row cache space for each Bank in each programmable multiprocessor; when the data in the private line cache space is updated, synchronously updating the corresponding data in the L1 cache, the L2 cache and the DRAM; counting the hit rate of the private line cache in real time, and closing the private line cache function when the hit rate is lower than a preset threshold value. High-delay external storage access is reduced, the data reuse rate is improved, and the execution efficiency is remarkably improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Shared queue for data exchange between stacks

A device can provide a unified and scalable multiprocessor communication framework that enables communication between multiple processor cores in a multicore device. For example, the device may be configured to perform data payload management using a Smart Message Queue (SMQ) and / or shared memory. Additionally or alternatively, the communication framework may enable the processors to communicate with each other and / or peripherals of the device while abstracting details of various communication protocols, hardware interfaces, and / or the like. For example, the communication framework may provide a common interface for applications, enabling a first application associated with a processor to establish a connection to any processor and / or peripheral without knowing details of a communication protocol for each connection.
Owner:AMAZON TECH INC

DRAM cache with stacked, heterogenous tag and data dies

A high-capacity cache memory is implemented by multiple heterogenous DRAM dies, including a dedicated tag-storage DRAM die architected for low-latency tag-address retrieval and thus rapid hit / miss determination, and one or more capacity-optimized cache-line DRAM dies that render a net cache-line storage capacity orders of magnitude beyond that of state-of-the art SRAM cache implementations. The tag-storage die serves double-duty in some implementations, yielding rapid tag hit / miss determination for cache-line read / write requests while also serving as a high-capacity snoop-filter in a memory-sharing multiprocessor environment.
Owner:RAMBUS INC

Sharing memory among multiple processors

Apparatuses, systems, and techniques to facilitate memory management. In at least one embodiment, data from one or more first shared physical memory locations is accessed based, at least in part, on one or more virtual addresses corresponding to one or more second shared physical memory locations.
Owner:NVIDIA CORP

Graphics processor resource management methods, apparatus, electronic devices, and readable media

This invention provides a graphics processing unit (GPU) resource management method, apparatus, electronic device, and readable medium. The method includes: detecting whether there is a stagnant streaming multiprocessor in the GPU; if there is a stagnant streaming multiprocessor, merging the remaining memory resources of the stagnant streaming multiprocessor with the remaining memory resources of at least one other streaming multiprocessor; determining at least one thread bundle to be executed in the stagnant streaming multiprocessor as a target thread bundle; and allocating memory resources to the target thread bundle based on the merged remaining memory resources using a preset thread scheduler in the GPU, so that the target thread bundle executes preset instructions based on the reallocated memory resources. This method allows for the rational allocation of remaining memory resources in the streaming multiprocessor, enabling the thread bundles in the streaming multiprocessor to run efficiently and improving the resource utilization efficiency of the streaming multiprocessor in the GPU.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

INDEPENDENT ON-CHIP MULTI-CHECKING

System, comprehensive: a plurality of processor clusters, wherein a given processor cluster comprises one or more processors; a large number of graphics processing units; a variety of storage controllers configured to control access to storage devices; a multitude of agents; and a multitude of network switches coupled to the multitude of processor clusters, the multitude of graphics processing units, the multitude of memory controllers, and the multitude of agents, wherein: a first subset of the multitude of network switches is interconnected to form a network of a central processing unit (CPU) between the multitude of processor clusters and the multitude of memory controllers, a second subset of the multitude of network switches is interconnected to form an input / output (I / O) network between the multitude of processor clusters, the multitude of agents, and the multitude of memory controllers, a third subset of the multitude of network switches is interconnected to form a loosely sorted network between the multitude of graphics processing units, selected agents of the multitude of agents, and the multitude of memory controllers, The CPU network, the I / O network, and the loosely sorted network are independent of each other. the CPU network and the I / O network are coherent and The network with loosened sorting is incoherent and has reduced sorting constraints compared to the CPU network and the I / O network.
Owner:APPLE INC

Switching equipment, on-network computing method thereof and multiprocessor system

The invention relates to switching equipment, an online computing method thereof and a multiprocessor system, and belongs to the technical field of communication. The switching equipment comprises N ports and N INC agent modules, the N ports are used for connecting M processors, N and M are integers greater than or equal to 1, and N is greater than or equal to M; the N INC agent modules are connected with the N ports one by one, and the N INC agent modules are sequentially connected to form a ring topology; and each INC agent module comprises an ALU unit. According to the invention, the INC agent module is added to the switching equipment side, so that the switching equipment side has the online computing capability, all intermediate data uploaded to the switching equipment side by the processors can be transmitted and computed in the switching equipment, finally the computing result is returned to the processors, the intermediate data does not need to be transmitted between the processors, and the computing efficiency is improved. Therefore, the data transmission quantity is reduced, and the utilization rate of the processor and the network bandwidth can be improved.
Owner:海光信息技术(成都)有限公司

Bpf-based on-chip heterogeneous multiprocessor system task scheduling system and method

A task scheduling system and method of a BPF-based on-chip heterogeneous multiprocessor system, comprising: a scheduling module and a sending module located at a general operating system Linux end, a receiving module located at a real-time operating system end, a BPF bytecode analysis and execution module and a calculation module, the application uses BPF virtual machine technology, uses BPF bytecode as the transmission carrier between the general operating system Linux and the real-time operating system of the on-chip heterogeneous multiprocessor, and schedules the to-be-processed application program in the part of the general operating system Linux to the real-time operating system in real time, and executes in the BPF instruction analysis and BPF virtual machine. The application can achieve certain load balancing and maximize the use of on-chip heterogeneous multiprocessor resources of the system without increasing the hardware design cost.
Owner:SHANGHAI JIAOTONG UNIV

Hardware-optimized symmetric high-performance automated parallelization system with loop-level parallelization and method thereof

PendingJP2026516776ACode compilationSource to sourceComputer architectureLoop level parallelism
A symmetric automatic compiler system (1) and method for high-performance, hardware-optimized automatic parallelization of program code (3) for execution by a multicore or multiprocessor parallel processing system (2) having multiple processing units (21) that simultaneously process instructions for data in a parallel processing system (2) by executing program code (3). The automatic compiler system (1) converts the sequential source code (31) of program code (3) into parallel processing machine code (32) which includes several instructions that can be executed by multiple processing units (21) of the parallel processing system (2) or that control the operation of multiple processing units (21).
Owner:マイナティックス アーゲー

Deployment method and system of multi-client operating system for preventing resource contention

The invention discloses a method and system for deploying a multi-client operating system for preventing resource contention, and the method comprises the steps: generating a grouping information table and a resource distribution table based on the business demands and the total number of hardware resources of a multi-core processor system, the number of the processing groups is configured to be smaller than or equal to the total number of hardware resources which can be independently divided in the multi-core processor system, and the resource allocation table is used for storing a plurality of pieces of hardware resource allocation information corresponding to the processing groups; dividing the plurality of processing cores into each processing group based on the service requirements of the multi-core processor system and the grouping information table; and based on the resource allocation table, performing hardware resource configuration on each processing group so as to obtain a plurality of hardware resource isolated client operating systems. Based on the deployment method, isolation of hardware resources among multi-client operating systems can be realized, and the performance and resource occupation of a multi-processor system are optimized.
Owner:SHENZHEN YUXIAN MICROELECTRONICS COMPUTING CO LTD

Detecting user inactivity in a multiprocessor communication device

Examples include a communication device including a user interface (UI), a first electronic processor configured to detect and process UI input events, and a second electronic processor communicatively connected to the first electronic processor and configured to detect and process UI input events. The first electronic processor executes an application having a user inactivity timeout feature by initializing a countdown timer for a first time period, and, in response to expiration of the countdown timer, determining a user inactivity time for the communication device. The user inactivity time is a lesser of a time since a UI input event was last detected by the first electronic processor and a time since a UI input event was last detected by the second electronic processor. In response to the user inactivity time being greater than or equal to the first time period, the first electronic processor performs an application timeout function.
Owner:MOTOROLA SOLUTIONS INC