Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

378 results about "Multi processor" patented technology

Multi-processor core communication method and device, electronic equipment and storage medium

The invention relates to the field of satellite navigation anti-interference and the technical field of computers, and discloses a multi-processor core communication method and device, electronic equipment and a storage medium, and the method comprises the steps that a second processor core writes target information into a shared memory, and obtains a first idle channel number; the second processor obtains a first interrupt signal based on an interrupt trigger register corresponding to the first idle channel number; the first processor core reads target information in a shared memory according to the first interrupt signal; the first processor core processes the target information to obtain a processing result, and writes the processing result into the shared memory to obtain a second idle channel number; the first processor core obtains a second interrupt signal based on an interrupt trigger register corresponding to the second idle channel number; and the second processor core reads the processing result in the shared memory according to the second interrupt signal, and writes the processing result into the cache region. The multi-processor core communication method is suitable for different operating systems and is high in transportability.
Owner:CHANGSHA HAIGE BEIDOU INFORMATION TECH CO LTD

Fault-tolerant multi-processor systems and methods for an aircraft

Aspects of the present disclosure generally relate to systems and methods for flight control of aircrafts driven by electric propulsion systems and in other types of vehicles. In some embodiments, a computer-implemented method for controlling an aircraft is disclosed, comprising: receiving, from a source processor, a first copy of a signal corresponding to an input device; sending a second copy of the signal to all other processors; receiving a number of second copies of the signal from all other processors, the number of second copies being equal to the number of all other processors excluding the source processor; determining a consensus signal based on the first copy and the second copies of the signal; and determining a command signal for an effector of the aircraft based on the consensus signal, and wherein no two processors are configured to receive signals from a same input device.
Owner:ARCHER AVIATION INC

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Reservation policies for real-time processing tasks in multi-processor systems

Approaches presented herein provide systems and methods for allocating streaming multiprocessors (SMs) to execute one or more tasks. A utilization for a given task may be determined by one or more parameters, such as a task execution time or a period. The utilization may then be used to assign a proportionate number of SMs associated with a processing unit, such as a graphics processing unit (GPU) or other type of processing unit, executing the task. The SMs may then be identified, allocated, and reserved until execution of the task is complete.
Owner:NVIDIA CORP

Thread block scheduling module, general purpose computing graphics processing unit, device and product

The invention provides a thread block scheduling module, a general-purpose computing graphic processing unit, equipment and a product, and relates to the technical field of data process.The method comprises the steps that an operation data analysis unit conducts data analysis on all first thread blocks of a computing task issued by a host, and second thread blocks sharing data with all the first thread blocks are determined; the mapping relation between the first thread block and the second thread block is written into a thread block data mapping table of the queue management unit; the resource management unit is used for recording available resources and used resources of each streaming multiprocessor and information of a thread block currently allocated to each streaming multiprocessor; and the thread block allocation unit is used for allocating the thread blocks of the first target thread block to the corresponding stream multiprocessors according to the thread block data mapping table of the first target thread block under the condition of detecting that any stream multiprocessor recorded by the resource management unit has idle available resources.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Multi-processor management method and device based on parallel computing chip, and medium

The invention relates to the technical field of parallel computing chips, in particular to a multiprocessor management method and device based on a parallel computing chip and a medium. The method comprises the steps that a plurality of stream processors are organized according to a hierarchical arrangement mode, a hierarchical SM cluster is formed, and each stage of SM core only keeps cache consistency with a lower stage of SM core directly controlled by the SM core; data of an external storage module enters a scheduling core through an L2 cache, and the scheduling core splits the data into thread bundles and dynamically allocates the thread bundles to a target SM core; when the thread processing demand of the target SM core exceeds the current capacity, starting a lower-level SM core according to a preset hierarchy rule, and expanding the processing capacity step by step; and skipping an unprocessed thread bundle over the current SM core through a thread bundle scheduler, splitting the unprocessed thread bundle, transmitting the split unprocessed thread bundle to a lower-level SM core, and merging calculation results of the lower-level SM core step by step through a shared memory. According to the invention, extra resources allocated by the system can be reduced, and the whole system is more efficient and ordered on the basis of ensuring consistency.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

High performance code parallelization compiler with loop level parallelization

PendingCN121464428ACode compilationComputer architectureLoop level parallelism
A symmetric auto-compiler system (1) and method for high performance hardware optimized auto-parallelization of program code (3) executed by a multi-core or multi-processor parallel processing system (2) having a plurality of processing units (21) that simultaneously process instructions for data in the parallel processing system (2) by executing the program code (3). The automatic compiler system (1) converts serial source code (31) of program code (3) into parallel processing machine code (32) comprising a plurality of instructions executable by a plurality of processing units (21) of the parallel processing system (2) or controlling operation of the plurality of processing units (21).
Owner:MINATIX INC

Task processing method and system, electronic equipment, storage medium and program product

The invention discloses a task processing method and system, electronic equipment, a storage medium and a program product, and relates to the technical field of computers. The method comprises the following steps: when multiple processors cooperatively execute a to-be-processed task, determining that the to-be-processed task belongs to a communication dominant task according to memory access behavior characteristics of the to-be-processed task, and synchronizing far-end data of each processor to a local far-end data cache space. Reading target read data from the far-end data cache space when the far-end read operation is carried out; when far-end write operation is carried out, target write data written into the target local cache space is written into the local far-end data cache space, and meanwhile corresponding data in the target local cache space is invalid and synchronized to the target local cache space. According to the invention, the problem of low calculation efficiency of multiprocessor cooperative processing tasks in the prior art can be solved, the communication delay of the multiprocessor cooperative processing tasks can be reduced, the efficiency of the multiprocessor cooperative processing tasks is effectively improved, and the task processing performance of the multiprocessor is improved.
Owner:INSPUR (BEIJING) ELECTRONICS INFORMATION IND CO LTD

Thread block scheduling method and system for universal graphics processor

The invention provides a thread block scheduling method and system for a universal graphics processor, and belongs to the technical field of computational graphics process.The method comprises the steps that a computational task from a host is received, and the computational task comprises at least one thread block; according to the workload of each thread block, all the thread blocks are distributed to at least one thread block queue, and each thread block queue corresponds to one capacity; obtaining the free resource quantity of each streaming multiprocessor in the universal graphics processor; determining a matched thread block queue according to the free resource quantity of any available streaming multiprocessor, extracting thread blocks from the matched thread block queue, and distributing the thread blocks to the available streaming multiprocessors; the available streaming multiprocessors are streaming multiprocessors of which the free resource quantity meets a preset thread block allocation standard. According to the method, a plurality of thread block queues are arranged, and the thread blocks are dynamically selected according to the real-time idle resources of the streaming multiprocessor for fine distribution, so that the overall calculation performance and the operation efficiency of the GPGPU are improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Server

The invention discloses a server, and relates to the technical field of servers, the server comprises at least one mainboard and at least one backboard, each mainboard comprises a central processing unit, the central processing units are mutually independently arranged, each backboard is provided with at least one hard disk, each hard disk is based on a disk sequence of the hard disks, and each hard disk is provided with at least one hard disk. According to the technical scheme, the central processing unit is connected with the nearest central processing unit, so that the central processing units can independently operate and execute the functions of data storage and data operation, the problem of load imbalance existing in the multiple central processing units in the server can be avoided, and the operation efficiency in a multi-processor mode is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Software-defined tensor streaming multiprocessor for large-scale machine learning

A system contains a network of processors arranged in a plurality of nodes. Each node comprises a respective plurality of processors connected via local links, and different nodes are connected via global links. The processors of the network communicate with each other to establish a global counter for the network, enabling deterministic communication between the processors of the network. A compiler is configured to explicitly schedule communication traffic across the global and local links of the network of processors based upon the deterministic links between the processors, which enable software-scheduled networking with explicit send or receive instructions executed by functional units of the processors at specific times, to establish a specific ordering of operations performed by the network of processors. In some embodiments, the processors of the network of processors are tensor streaming processors (TSPs).
Owner:GROQ INC

Methods and apparatus for deep learning network execution pipeline on multi-processor platform

Methods and systems are disclosed using an execution pipeline on a multi-processor platform for deep learning network execution. In one example, a network workload analyzer receives a workload, analyzes a computation distribution of the workload, and groups the network nodes into groups. A network executor assigns each group to a processing core of the multi-core platform so that the respective processing core handle computation tasks of the received workload for the respective group.
Owner:INTEL CORP

Method, apparatus and system for monitoring ultra-high frequency partial discharge of hydro-generator

The present disclosure provides a method, apparatus and system for monitoring ultra-high frequency partial discharge of a hydro-generator, and belongs to the technical field of hydro-generator partial discharge monitoring. The method includes: cleaning a partial discharge pulse sequence using a cleaning threshold to obtain a valid pulse sequence; performing redundant data filtering on each data unit divided from the valid pulse sequence to obtain a first target pulse sequence for short-period partial discharge monitoring; determining sub-sequences that are partial discharge events from the valid pulse sequence, forming a second target pulse sequence after associating an amplitude statistical feature of the sub-sequences, and storing the second target pulse sequence. The aforementioned method combines data cleaning, redundant data filtering, and partial discharge event identification, which enhances the real-time performance of partial discharge monitoring and records the long-period partial discharge change trend. The corresponding system adopts a multi-buffer zone and multi-processor architecture, thus further improving the real-time performance of partial discharge monitoring.
Owner:GUODIAN SCI & TECH RES INST

Communication architecture for multicore system

A device can provide a unified and scalable multiprocessor communication framework that enables communication between multiple processor cores in a multicore device. For example, the device may be configured to perform data payload management using a Smart Message Queue (SMQ) and / or shared memory. Additionally or alternatively, the communication framework may enable the processors to communicate with each other and / or peripherals of the device while abstracting details of various communication protocols, hardware interfaces, and / or the like. For example, the communication framework may provide a common interface for applications, enabling a first application associated with a processor to establish a connection to any processor and / or peripheral without knowing details of a communication protocol for each connection.
Owner:AMAZON TECH INC

Telemetry assisted hybrid load balancing on switch fabric paths

Techniques described herein can use a hybrid load balancing approach to balance loads on paths in a switch fabric. The switch fabric can deliver synchronization data between processors in a multi-processor cluster, and the synchronization data can load the paths on which it is sent. First paths can be identified in the switch fabric, and first synchronization data can be distributed to the first paths using a first load balancing approach, such as a telemetry assisted load balancing approach. Second paths can be identified in the switch fabric, and second synchronization data can be distributed to the second paths using a second load balancing approach, such as a packet spraying load balancing approach.
Owner:CISCO TECHNOLOGY INC

Systolic arithmetic on sparse data

Embodiments described herein provided for an instruction and associated logic to enable a processing resource including a tensor accelerator to perform optimized computation of sparse submatrix operations. One embodiment provides a parallel processor comprising a processing cluster coupled with the cache memory. The processing cluster includes a plurality of multiprocessors coupled with a data interconnect, where a multiprocessor of the plurality of multiprocessors includes a tensor core configured to load tensor data and metadata associated with the tensor data from the cache memory, wherein the metadata indicates a first numerical transform applied to the tensor data, perform an inverse transform of the first numerical transform, perform a tensor operation on the tensor data after the inverse transform is performed, and write output of the tensor operation to a memory coupled with the processing cluster.
Owner:INTEL CORP

Fully cache coherent virtual partitions in multitenant configurations in a multiprocessor system

Various embodiments include techniques for processing memory operations in a computing system. The computing system includes a central processing unit (CPU) and an auxiliary processor, such as a parallel processing unit (PPU). The PPU can be divided into multiple partitions. Although the partitions are included in a single PPU, the CPU can track the partitions as if the partitions are independent devices rather than different portions of a single device. When two different partitions generate memory operations that access the same memory address in CPU memory address space, the two partitions employ two different data paths. The CPU can use path information for the two different paths to identify which partition generated each memory operation. As a result, the CPU can maintain data consistency and memory coherency in a system where a PPU is divided into multiple partitions.
Owner:NVIDIA CORP

Systems and methods of preconfiguring coherency protocol for computing systems

A multi-processor computing system (e.g., a system-on-chip) can store, in a shared memory, (i) a reservation table that is accessible by the one or more workload processors, and (ii) a scheduling program. The system can further execute the scheduling program to schedule execution of a set of workloads by one or more workload processors in accordance with an optimized compute graph, an optimized data positioning graph, and a coherence protocol that is precomputed based on the optimized compute graph and the optimized data positioning graph.
Owner:MERCEDES BENZ GROUP AG

High performance code parallelization compiler with loop level parallelization

A system and method for universal static multi-transmit CPU design with a static pipeline for automatically parallelizing code is presented. A multi-core and / or multi-processor integrated circuit (2) has a plurality of processing units (21) and / or processing pipelines (53) that simultaneously process instructions for data by executing parallel processing machine code (32). The execution of the parallelized processing code (32) by the parallel processing multi-core and / or multi-processor integrated circuit (2) comprises the occurrence of a delay time (26), wherein the delay time is given by an idle time between the processing unit (21) returning the data after processing a specific instruction block of the processing code (32) for the data and receiving the data required by the processing unit (21) to execute a consecutive instruction block of the processing code (32). The parallel pipeline (53) comprises means for: (i) forwarding by providing a data forwarding from a MEM stage as an EX / MEM register to an EX stage as an ID / EX-stage register; (ii) exchanging by making the results of the Ex-ME-phase registers accessible by the Ex phase of a parallel pipeline (53) to provide a result exchange between the pipelines; and (iii) implementing branch pipeline refresh by providing control conflicts by refreshing only those pipelines (53) dependent on one pipeline (53) based on branch address computation conditions.
Owner:MINATIX INC

Cross-database SQL (Structured Query Language) dynamic conversion adaptation method based on chain of responsibility

The invention provides a cross-database SQL (structured query language) dynamic conversion adaptation method based on a responsibility chain, which comprises the following steps of: loading adaptation processors into the responsibility chain, and sequencing the processors to finish system initialization operation; according to the configuration file parameters, whether SQL adaptive conversion configuration is started or not is judged, and if yes, whether SQL conversion cache is started or not continues to be judged; when the SQL conversion cache is started, SQL conversion is executed, cache query is carried out, and available SQL statements of a target database are output. According to the method, one-time development and multi-database adaptation are achieved, the migration maintenance cost is greatly reduced, the opening and closing principle is achieved through a plug-in type, a responsibility chain and a multi-processor mode, future expansion, enhancement or customization are facilitated, and concurrency efficiency and adaptation flexibility are guaranteed.
Owner:CHINA UNICOM XIONGAN IND INTERNET CO LTD

A Typed Task Co-scheduling System and Method Based on Heterogeneous Multi-core Architecture

This invention relates to the field of multi-processor multi-task joint scheduling technology, specifically disclosing a typed task joint scheduling system and method based on a heterogeneous multi-core architecture. It introduces a joint scheduling mechanism to divide and sort tasks according to task size and priority. Based on this, for the time-limited characteristics of real-time tasks, an inertial weight coefficient particle swarm optimization method is used to iteratively update the load balancing strategy, minimizing the maximum response time of real-time tasks while meeting their schedulability requirements. Simultaneously, for the low-priority characteristics of non-real-time tasks, the problem-solving is simplified using a Lagrange-based convex optimization approach, and an energy-constrained binary search algorithm is employed to effectively reduce the average response time of non-real-time tasks, achieving optimal load distribution for the system.
Owner:CHONGQING UNIV +2

Deadlock prevention utilizing distributed resource reservations

In multi-threaded or multi-processor computing systems, a deadlock may occur when two or more processes or threads are unable to proceed because they are each waiting for a resource that the other holds. As a result, progress is halted because conflicting entities are stuck in a circular dependency, and none can release the resources they hold to let the others continue. Systems and methods are provided wherein a resource reservation is carried out in two steps. The first step causes query nodes to add an identifier to a queue and, upon a request and the identifier being in a first position, a non-sharable resource is reserved. As a result, non-sharable resources are reserved in order and when needed, thereby preventing deadlocks.
Owner:ROCKET SOFTWARE

Application programming interface for reading from data structure

The invention discloses an application programming interface for reading from a data structure. Apparatuses, systems, and techniques for performing computing operations are disclosed. In at least one embodiment, a processor executes an application programming interface such that one or more numbers of one or more indicators of one or more streaming multiprocessors in one or more processors are read from one or more data structures that store the one or more indicators.
Owner:NVIDIA CORP

Multi-processor communication framework using TCP / IP architecture

A multi-processor communication framework (MCF) based on Transmission Control Protocol / Internet Protocol (TCP / IP) architecture is defined that can be used for inter-processor communication between heterogenous processing nodes. The MCF includes an application layer, a transport layer, a network layer, and an inter-processor communication (IPC) layer for each processing node. The IPC layer can be used to facilitate data transfer between two processing nodes via a physical channel within an electronic device, or through shared memory. The MCF can support multicast, broadcast, unicast, and zero-copy features like the TCP / IP stack.
Owner:AMAZON TECH INC

Self-adaptive task scheduling method based on multiprocessor cooperative system

The invention relates to a self-adaptive task scheduling method based on a multiprocessor cooperative system, and belongs to the technical field of electronic information. Comprising the following steps: executing CPU-GPU task scheduling; calculating to obtain respective data size distribution of the CPU and the GPU, and storing the data size distribution into a text; performing parallel computing by the CPU and the GPU according to the obtained respective data volume distribution of the CPU and the GPU; summarizing the CPU data and the GPU data, outputting a result, and ending; reading historical data from the text to an array; reading array data, and obtaining respective data size distribution of the CPU and the GPU; and constructing a linear regression equation for array data by using a least square method to obtain a prediction result. According to the invention, a new task scheduling method is designed, so that the balance point of the CPU and the GPU for executing the tasks can be efficiently found, and the tasks can be reasonably distributed. According to the method, the array and the linear regression model are combined, and the extra time consumption of task scheduling can be better reduced by utilizing historical data and the prediction model.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Application programming interface to read from a data structure

Apparatuses, systems, and techniques to perform computing operations. In at least one embodiment, a processor performs an application programming interface to cause one or more indicators of one or more numbers of one or more streaming multiprocessors of one or more processors to be read from one or more data structures storing the one or more indicators.
Owner:NVIDIA CORP

Method and device for accessing memory

The invention discloses a method and equipment for accessing a memory, relates to the technical field of terminals, and can solve the problem that a terminal device with a multiprocessor architecture cannot realize scientific and ordered memory access. In the application, the plurality of processors of the terminal equipment share the same memory, so that any processor can be ensured to normally store data, modify or read stored data and the like in any use scene. Moreover, any one of the plurality of processors can realize mutual exclusion access of the plurality of processors to the memory based on the acquired condition that whether other processors are accessing the memory or not, so that more scientific and ordered memory sharing is realized. According to the scheme that the processors of the terminal equipment share the same memory, the production cost of the terminal equipment can be reduced, and support is provided for miniaturization of the terminal equipment.
Owner:HUAWEI TECH CO LTD

GEMM load-oriented GPU modeling method

A GEMM load-oriented GPU modeling method is characterized in that through a multi-stage collaborative modeling mechanism, cache behaviors, instruction overhead and calculation intensity are deeply coupled, accurate performance prediction of GPU execution GEMM operators is realized, the method can be widely applied to scheduling optimization of GPU intensive scenes such as AI training and scientific calculation, firstly, a three-stage cache weight distribution mechanism is established, and then, a three-stage cache weight distribution mechanism is established; quantifying the contribution of the L1 / L2 cache hit rate and the DRAM bandwidth degradation factor to the effective bandwidth; secondly, an instruction-level memory access overhead correction mechanism is introduced, and the mixing precision and the real calculation strength of a sparse calculation scene are captured through dynamic parameter adjustment and optimization; then combining the calculation force peak value and the bandwidth upper limit to construct a double-boundary constraint model, and generating a theoretical performance critical value; further predicting a stream multiprocessor utilization rate based on a neural network, and quantifying efficiency loss caused by hardware resource contention through a multi-layer perceptron structure; and finally, the integration module outputs task execution time to realize end-to-end performance prediction.
Owner:BEIHANG UNIV

Multi-processor architecture intelligent adaptive bottom plate of nuclear power DCS (Distributed Control System) controller

The invention discloses a multiprocessor architecture intelligent adaptive base plate of a nuclear power DCS controller, and relates to the field of nuclear power control, and the base plate comprises a signal processing module which is used for collecting and preprocessing each path of signal and then realizing bidirectional data communication with a core board card main processor; the PCIE network controller expansion module is used for realizing gigabit network interface expansion; the storage module is used for storing data; the redundancy synchronization module is used for realizing redundancy switching and network synchronization of the main controller and the standby controller; the bottom plate power supply system is used for providing voltage; the bottom plate clock system is used for providing time information; the connector is used for connecting the core board card and the backboard; and each peripheral module. Each peripheral module comprises a USB module, a serial port module, a network module, an output driving module, an isolation communication module and a B code circuit module. According to the method, the compatibility can be improved, the complexity and cost of system upgrading are reduced, and the flexibility of technical iteration is improved.
Owner:CHINA NUCLEAR CONTROL SYST ENG

Asynchronous release operation in multiprocessor system

Embodiments of the present disclosure relate to asynchronous release operations in a multiprocessor system. Various embodiments include techniques for performing memory synchronization operations between processors in a multi-processor computing system. The first processor transfers the data by issuing a memory operation to store the data to the shared memory. And the first processor sends an asynchronous release operation to the loading storage unit. In response, the load memory unit issues a memory synchronization operation to ensure that data associated with the memory operation is visible in the shared memory. When the asynchronous release operation is suspended, the first processor can issue further instructions and perform other operations. When data associated with the memory operation is visible in the shared memory, the memory synchronization operation is completed and the load memory unit writes a flag to a separate memory location. Once it is detected that the flag has been written, the second thread and / or other thread may reliably read data stored in the shared memory.
Owner:NVIDIA CORP