Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8 results about "Automatic parallelization" patented technology

Automatic parallelization, also auto parallelization, autoparallelization, or parallelization, the last one of which implies automation when used in context, refers to converting sequential code into multi-threaded or vectorized (or even both) code in order to utilize multiple processors simultaneously in a shared-memory multiprocessor (SMP) machine. The goal of automatic parallelization is to relieve programmers from the hectic and error-prone manual parallelization process. Though the quality of automatic parallelization has improved in the past several decades, fully automatic parallelization of sequential programs by compilers remains a grand challenge due to its need for complex program analysis and the unknown factors (such as input data range) during compilation.

High performance code parallelization compiler with loop level parallelization

PendingCN121464428ACode compilationComputer architectureLoop level parallelism
A symmetric auto-compiler system (1) and method for high performance hardware optimized auto-parallelization of program code (3) executed by a multi-core or multi-processor parallel processing system (2) having a plurality of processing units (21) that simultaneously process instructions for data in the parallel processing system (2) by executing the program code (3). The automatic compiler system (1) converts serial source code (31) of program code (3) into parallel processing machine code (32) comprising a plurality of instructions executable by a plurality of processing units (21) of the parallel processing system (2) or controlling operation of the plurality of processing units (21).
Owner:MINATIX INC

High performance code parallelization compiler with loop level parallelization

PendingCN121399576ACode compilationHandling CodeComputer architecture
A system and method for universal static multi-transmit CPU design with a static pipeline for automatically parallelizing code is presented. A multi-core and / or multi-processor integrated circuit (2) has a plurality of processing units (21) and / or processing pipelines (53) that simultaneously process instructions for data by executing parallel processing machine code (32). The execution of the parallelized processing code (32) by the parallel processing multi-core and / or multi-processor integrated circuit (2) comprises the occurrence of a delay time (26), wherein the delay time is given by an idle time between the processing unit (21) returning the data after processing a specific instruction block of the processing code (32) for the data and receiving the data required by the processing unit (21) to execute a consecutive instruction block of the processing code (32). The parallel pipeline (53) comprises means for: (i) forwarding by providing a data forwarding from a MEM stage as an EX / MEM register to an EX stage as an ID / EX-stage register; (ii) exchanging by making the results of the Ex-ME-phase registers accessible by the Ex phase of a parallel pipeline (53) to provide a result exchange between the pipelines; and (iii) implementing branch pipeline refresh by providing control conflicts by refreshing only those pipelines (53) dependent on one pipeline (53) based on branch address computation conditions.
Owner:MINATIX INC

Hardware-optimized symmetric high-performance automated parallelization system with loop-level parallelization and method thereof

PendingJP2026516776ACode compilationSource to sourceComputer architectureLoop level parallelism
A symmetric automatic compiler system (1) and method for high-performance, hardware-optimized automatic parallelization of program code (3) for execution by a multicore or multiprocessor parallel processing system (2) having multiple processing units (21) that simultaneously process instructions for data in a parallel processing system (2) by executing program code (3). The automatic compiler system (1) converts the sequential source code (31) of program code (3) into parallel processing machine code (32) which includes several instructions that can be executed by multiple processing units (21) of the parallel processing system (2) or that control the operation of multiple processing units (21).
Owner:マイナティックス アーゲー

System and method for designing general-purpose static multi-issue integrated circuits using static automatic parallelization code

A system and method for designing a general-purpose static multi-issue CPU using static pipelining of automatically parallelized code is proposed. A multicore and / or multiprocessor integrated circuit has multiple processing units and / or processing pipelines that simultaneously process instructions for data by executing parallelized processing machine code. The execution of parallelized processing code by a multicore and / or multiprocessor integrated circuit that performs parallel processing involves the occurrence of latency, which is given by the idle time of the processing unit between the time the processing unit processes a particular block of instructions in the processing code for data and sends the data back, and the time the processing unit receives the data necessary to execute a successive block of instructions in the processing code. The parallel pipeline includes means for (i) forwarding by providing data forwarding from the MEM stage as an EX / MEM register to the EX stage as an ID / EX stage register, (ii) exchange by providing result exchange between pipelines by making the Ex-MEM stage register results accessible to the EX stage of the parallel pipeline, and (iii) means for branch pipeline flushing with control hazards by flushing only pipelines that depend on one pipeline that calculates conditions based on branch addresses.
Owner:マイナティックス アーゲー

Resource Parallel Scheduling and Optimization Method and System for Large-Scale Difference Operators

The present invention discloses a resource parallel scheduling and optimization method and system for large-scale differential operators. Guided by distributed computing, computer systems and architectures, and sharding technology, the present invention uses a distributed framework to connect task execution units to a host to form a cluster, and uses middleware in the cluster environment, which has the functions of automatic parallelization translation and resource scheduling. For large-scale operators, analyze the operator structure, find commonalities, extract their differences as parameters to be passed, and translate their parallelized code. The present invention can not only achieve efficient computing of operators on the cluster, but also help with the topic of converting other serial programs into parallel programs. After flexible changes, in addition to enabling the tasks to be processed to run on the cluster, if the task volume is small, CPU\GPU co-computation can also be achieved on a single host.
Owner:WUHAN UNIV

Automatic parallelization method and apparatus for large model, and storage medium

PCT designated stage expiredWO2025123655A1Neural learning methodsComputational scienceParallel computing
The present invention relates to an automatic parallelization method and apparatus for a large model, and a storage medium. The method comprises: acquiring a model structure of a large model that needs to be trained, and cluster information used for training; generating a search space on the basis of the product of the number of nodes and the number of compute cards included in each node; on the basis of the model structure of the large model and the cluster information combined with a pre-configured rule, pruning the search space to remove some candidate parallelization schemes; calculating the computing time and communication time of each of the remaining candidate parallelization schemes in the search space, and obtaining the total time consumption of each candidate parallelization scheme on the basis of the computing time and communication time of each candidate parallelization scheme; and selecting a candidate parallelization scheme having the minimum total time consumption to perform parallel training on the large model. Compared with the prior art, the present invention has the advantages of improving the search efficiency, etc.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Two-stage strategy gradient optimization automatic parallelization method based on self-adaptive Fimanban

The invention discloses a two-stage strategy gradient optimization automatic parallel method based on adaptive graph Mangbar, which comprises the following steps of: firstly, acquiring a public AI model data set, performing operator fusion on each computational graph in the data set, and acquiring a feature matrix X of the computational graph; then X is used as an initial coding matrix and is input into a Pitman-Ba neural network to obtain a feature code, after a feature # imgabs0 # is generated through a multi-layer perceptron, # imgabs1 # is input into a disturbance two-stage strategy gradient algorithm, j equipment placement strategies are generated through disturbance noise in each training period and internal stage, and parameters involved in the disturbance noise generation process are updated; and in the external stage, carrying out equipment placement strategy evaluation, selecting a minimum training time strategy to carry out global parameter updating, and outputting an optimal equipment placement strategy. According to the method, the problem of limited receptive field of a traditional graph neural network is solved by effectively extracting node features and capturing a long-term dependency relationship, and an optimal equipment parallel strategy is obtained.
Owner:SHANXI AGRI UNIV

Operator parallel partitioning method and system for general tensor processor

This invention proposes a method and system for parallel operator partitioning for a general-purpose tensor processor, including: calculating the output shape of the operator; obtaining a factor set for each dimension of the operator output shape based on the operator output shape; obtaining the partitioning scheme with the shortest operator runtime in the TPU, where the partitioning scheme is a combination of elements of the factor set, and the total number of partitioning schemes is the product of the size of the factor set. The runtime is calculated as: runtime = input read time + weight read time + computation time + output write time. The input read time is calculated using the sub-operator input size as a variable, which is obtained by analyzing the dependencies of the neural network; and presetting a solution space threshold. If the total number of partitioning schemes is not greater than the solution space threshold, an enumeration algorithm is used to find the optimal partitioning scheme; if it is greater, a heuristic algorithm is used to find the optimal partitioning scheme. This invention can implement an automatic parallelization scheme for neural network algorithms based on TPUs.
Owner:XIAMEN YIPU SMART TECH CO LTD