Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Loop unrolling" patented technology

Loop unrolling, also known as loop unwinding, is a loop transformation technique that attempts to optimize a program's execution speed at the expense of its binary size, which is an approach known as space–time tradeoff. The transformation can be undertaken manually by the programmer or by an optimizing compiler.

Unrolling an infinite loop during ray query traversal

This disclosure provides systems, devices, apparatus, and methods, including computer programs encoded on storage media, for unrolling an infinite loop during ray query traversal. A processor obtains, during a compile time, a number of loops associated with a BVH traversal based on a number of ray triangle intersections and a number of ray box intersections and / or a set of features associated with the BVH traversal and code generation associated with a shader. The processor determines, during the compile time, a loop unroll factor based on at least one of the obtained number of loops or the obtained set of features. The processor adjusts, during the compile time, a number of iterations of a loop associated with the BVH traversal based on the loop unroll factor. The processor outputs an indication of the adjusted number of iterations.
Owner:QUALCOMM INC

Compiling method and apparatus, device, storage medium, and program product

PCT designated stageWO2025180252A1Program controlCode compilationTime conditionObject code
A compiling method and apparatus, a device, a storage medium, and a program product. The compiling method comprises: unrolling an instruction sequence in a loop BB to generate multiple BBs and multiple DAGs corresponding to the multiple BBs, wherein the multiple BBs comprise an instruction sequence unrolled at least three times, and the multiple DAGs are used for indicating a data dependency relationship between instructions in the multiple BBs (810); performing window scheduling on the basis of the multiple BBs, and determining a target scheduling window and a scheduling result corresponding to the target scheduling window, wherein the window scheduling is used for dividing the instruction sequence in the loop BB into two parts on the basis of the scheduling window, such that the instruction sequence of the second part in a current loop and the instruction sequence of the first part in the next loop are synchronously executed, the target scheduling window is a scheduling window of which the execution time satisfies a time condition, and the execution time is the time required for executing instructions in the window and calculated on the basis of the data dependency relationship in the multiple DAGs (820); and on the basis of the scheduling result corresponding to the target scheduling window, performing compiling to obtain a target code of the loop BB in a target program (830).
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Microprocessor that builds inconsistent loop that iteration count unrolled loop multi-fetch block macro-op cache entries

A microprocessor includes a prediction unit (PRU) that predicts a sequence of fetch blocks (FBlks) in a program instruction stream, a macro-op (MOP) cache (MOC) that comprises MOC entries (MEs), a fusion engine. An ME holds MOPs into which architectural instructions of one or more FBlks are decoded. The PRU detects a loop body ME within the program instruction stream, accumulates loop iteration count information about a series of instances of a loop on the loop body ME in the program instruction stream, updates a consistency counter of the loop body ME while accumulating the loop iteration count information, and in response to detecting that the consistency counter has reached a threshold, instructs the fusion engine to use F copies of the MOPs of the loop body ME to build in the MOC an unrolled loop multi-FBlk ME; F is a loop unroll factor that is at least two.
Owner:VENTANA MICRO SYSTEMS INC

Prediction unit that predicts successor fetch block start address of multi-fetch block macro-op cache entry

A microprocessor includes a prediction unit (PRU) comprising a fetch block (FBlk) predictor (FBP) that predicts a sequence of FBlks, each FBlk having a corresponding fetch block start address (FBSA), and branch predictors; a macro-op (MOP) cache (MOC) includes MOC entries (MEs) including multi-FBlk MOC entries (MF-MEs) for holding MOPs decoded from instructions of multiple FBlks. The PRU detects a hit of a current FBSA on an MF-ME; performs a set of actions K times: looking up the current FBSA in the FBP and branch predictors to obtain outputs, using the outputs to predict a successor FBSA of a successor FBlk; and making the current FBSA the successor FBSA; and predicts that an FBSA of a successor FBlk to the MF-ME is the current FBSA resulting from performing K times the set of actions. K is a number of FBlks built into the MF-ME (alternatively times a loop unroll factor).
Owner:VENTANA MICRO SYSTEMS INC

Static analysis method and device for program problems, electronic equipment, readable storage medium and program product

The embodiment of the invention provides a program problem static analysis method and device, electronic equipment, a readable storage medium and a program product, and relates to the technical field of program static analysis. The method comprises the steps that a target program is analyzed based on a rule constraint set of target syntax, and an abstract syntax tree is generated; constructing an annotation control flow graph according to the abstract syntax tree; according to the annotation control flow diagram, identifying a periodic task in the target program; and on the basis of performing loop expansion on the execution process of the periodic task, performing symbolic execution analysis on the annotation control flow diagram, and identifying the cross-period conflict problem of the periodic task. By identifying the periodic task and circularly expanding the periodic task, the analysis limitation on the program problem in the related technology is solved, the cross-period conflict problem in the program is efficiently identified, and the static guarantee capability on the program quality is improved.
Owner:SHANGHAI FORMAL TECH INFORMATION TECH CO LTD

Integrating loop unrolling and loop splitting to reduce control overheads

ActiveUS12632233B2Code compilationControl flowLoop splitting
Described are techniques for reducing overhead controls. A loop tree is constructed from a program, such as a structured control flow program. Structured control flow refers to a programming concept where the flow of control to a block or region is based on single entry and single-exist methodology (SESE). A loop tree refers to a tree-like data structure that graphically represents loop(s) and / or an if-condition(s) in a program, such as a structured control flow program. A loop splitting operation or a loop unrolling operation may then be performed in connection with the node of the loop tree that is identified as having the highest benefit (ratio of execution cycles gained to the increase in code size) representing an if-condition or a loop, respectively, provided that the resultant code fits in the instruction buffer.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Integrating loop unrolling and loop splitting to reduce control overheads

ActiveUS20250251920A1Code compilationControl flowLoop splitting
Described are techniques for reducing overhead controls. A loop tree is constructed from a program, such as a structured control flow program. Structured control flow refers to a programming concept where the flow of control to a block or region is based on single entry and single-exist methodology (SESE). A loop tree refers to a tree-like data structure that graphically represents loop(s) and / or an if-condition(s) in a program, such as a structured control flow program. A loop splitting operation or a loop unrolling operation may then be performed in connection with the node of the loop tree that is identified as having the highest benefit (ratio of execution cycles gained to the increase in code size) representing an if-condition or a loop, respectively, provided that the resultant code fits in the instruction buffer.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

A method for optimizing cross-layer memory access bandwidth to support deep neural network acceleration

The present invention relates to the field of neural network optimization technology, and in particular to a method for optimizing cross-layer memory access bandwidth to support deep neural networks. The method comprises the following steps: 1. designing a bandwidth model for single-layer block optimization and loop sequence optimization; 2. performing single-layer independent block optimization and loop sequence optimization; 3. performing cross-layer joint block optimization and loop sequence optimization; and 4. performing composite block optimization and loop sequence optimization. The method can implement a data block strategy for minimum DRAM access, a loop unrolling control strategy, and a data storage mapping update strategy, achieving a fine-grained data access and data flow solution that maximizes data reuse, and effectively supporting automatic code generation and computational mapping optimization for optimizing computational and memory access efficiency.
Owner:HANGZHOU DIANZI UNIV

An Optimization Method and System for Heterogeneous Parallel Cholesky Decomposition Based on ShenWei Architecture

The present invention proposes a Cholesky decomposition heterogeneous parallel optimization method and system based on the Shenwei architecture, which relates to the field of high-performance computing technology. The method comprises: dividing a symmetric positive definite matrix into sub-blocks based on a distributed parallel allocation scheme and iteratively completing the matrix decomposition; allocating each sub-block to a different process through the MPI programming model, exchanging data between processes through asynchronous communication, and performing coarse-grained task-level parallel acceleration; utilizing the master-slave core acceleration parallel characteristics of the Shenwei architecture to perform two-level parallel acceleration on the four operations in the Cholesky decomposition; wherein, for GEMM and SYRK operations, the column vectors of the matrix are mapped to the slave core array, and the columns are divided according to the number of slave cores; the calculation process is optimized through a double buffering mechanism, vectorized operations and loop unrolling technology to improve parallel efficiency; for the TRSM operation, it is decomposed into multiple TRSV operations, which are allocated to the slave cores for parallel execution, and the data dependency is reduced by loop reading and data broadcasting to achieve efficient parallel computing.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Dynamic cross-layer data reuse method for efficient deep convolutional computation

The present application relates to the technical field of neural network optimization, and more particularly to a dynamic cross-layer data multiplexing method for efficient deep convolution calculation, comprising: step 1, designing a single-layer block optimization and loop sequence optimization model; step 2, performing separate block optimization and loop sequence optimization; step 3, performing multi-layer joint block optimization and loop sequence optimization; and step 4, performing dynamic composite block optimization and loop sequence optimization of single-layer and multi-layer. The present application can finally realize a multi-layer joint data block strategy with minimum DRAM access, a loop unrolling control strategy, and a data storage mapping update strategy, and realize a data access and data flow scheme with maximized data multiplexing of fine granularity, which can effectively support code automatic generation and calculation mapping optimization for optimizing calculation and memory efficiency.
Owner:HANGZHOU DIANZI UNIV

Binary program size optimizer based on loop folding

The present invention discloses a binary program volume optimizer based on loop folding, comprising the following steps: Step 1, running the program and collecting performance analysis data; Step 2, inputting the binary program and performance analysis data into the optimizer; Step 3, outputting the optimized program. Compared with existing code size optimization methods, the present invention optimizes at the binary level, lowering the optimization threshold, and at the same time utilizes the performance analysis data of the program to design and implement a more finely controlled optimization strategy in response to the shortcomings of existing loop expansion strategies. Compared with existing binary optimizers, the present invention focuses on the optimization of program volume, optimizing program volume without affecting the original performance.
Owner:SUN YAT SEN UNIV

A falcon signature FFT acceleration method for RISC-V architecture

The application provides a Falcon signature FFT acceleration method for a RISC-V architecture. The method comprises the following steps: S1, adopting a Union data processing mechanism to complete data type conversion, directly calling RISC-V native double-precision hardware floating point instructions for FFT complex number operation, and accurately controlling register allocation; S2, introducing a cache prefetch mechanism, storing batch of discrete distributed complex rotation factors in a continuous intermediate cache array in advance before executing the FFT core calculation loop, changing the discrete real-time data fetching mode, and sequentially addressing to complete rotation factor reading; S3, adopting multi-loop loop unrolling transformation for the innermost calculation loop of FFT, performing parallel processing on multiple groups of calculation units in single loop iteration, and reducing the proportion of jump branch instructions; S4, after completing the FFT operation, writing the operation result into a target storage area through an optimized instruction sequence. The method provided by the application solves the problems of low efficiency of instruction mapping, low memory access efficiency and insufficient hardware resource scheduling of the existing Falcon algorithm when running on a RISC-V platform.
Owner:NANJING UNIV OF POSTS & TELECOMM

Computer-implemented method, computer program, and system (integrating loop unrolling and loop splitting to reduce control overhead)

To describe techniques for reducing overhead controls.SOLUTION: A loop tree is constructed from a program, such as a structured control flow program. Structured control flow refers to a programming concept where the flow of control to a block or region is based on single entry and single-exist methodology (SESE). A loop tree refers to a tree-like data structure that graphically represents loop(s) and / or an if-condition(s) in a program, such as a structured control flow program. A loop splitting operation or a loop unrolling operation may then be performed in connection with the node of the loop tree that is identified as having the highest benefit (ratio of execution cycles gained to the increase in code size) representing an if-condition or a loop, respectively, provided that the resultant code fits in the instruction buffer.SELECTED DRAWING: Figure 14
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Compiler infrastructure based on parameterized control flow diagram and implementation method thereof

A traditional compiler widely uses a static single assignment (SSA) form in intermediate representation (IR), and the problem of variable value conflict during control flow merging is solved through a Phi function. However, the Phi function results in deep coupling of the control flow with the data flow, resulting in IR expansion, limited optimization and difficult maintenance. In order to solve the problem, the invention provides a compiler infrastructure based on a parameterized control flow diagram and an implementation method thereof, and relates to a compiler design and code optimization technology. The explicit parameterization basic block is used for replacing a Phi function, control flow and data flow analysis is thoroughly decoupled, data flow analysis is simplified, the complexity of intermediate representation is remarkably reduced, optimization efficiency is greatly improved, and good maintainability and friendly debugging performance are achieved. The method is suitable for scenes of loop expansion, redundancy elimination and the like, and a modularized and efficient solution is provided for a modern compiler.
Owner:朱征赜

Vectorization method for homomorphic encryption compiler

The invention discloses a vectorization method for a homomorphic encryption compiler, and the method comprises the following steps: S1, obtaining a cloud homomorphic encryption calculation-oriented source program written by a developer; s2, converting a source program into an intermediate representation, and executing loop expansion optimization on the intermediate representation to obtain a scalar intermediate representation without control flow; s3, converting the scalar intermediate representation into a directed acyclic graph G (V, E); s4, calculating the ASAP and ALAP attribute of each node in the directed acyclic graph, and determining the maximum lane width required by the directed acyclic graph; s5, constructing an integer linear programming model based on the parameters; and S6, solving the integer linear programming model, and generating a vectorized homomorphic encryption program according to a solving result and a homomorphic encryption library at the rear end of the target hardware. According to the method, manual intervention is not needed, regular and irregular programs are supported, the speed is averagely increased by 5.95 times at the CPU end and 3.79 times at the GPU end compared with a scalar program, and the homomorphic encryption technology is assisted.
Owner:SHANGHAI JIAOTONG UNIV

Multi-heterogeneous main body system optimization method and system oriented to real-time processing scene

The invention discloses a multi-heterogeneous main body system optimization method and system oriented to a real-time processing scene, and the method comprises the steps: a CPU employs a bare core BMP mode for parallel division of labor, the CPU comprises a plurality of large cores and a plurality of small cores, a control task and a protection task are operated on the large cores, a master station command is received, and a background task is operated on the small cores; the FPGA writes back and reads data between a DDR (Double Data Rate) memory of the CPU and an internal cache of the FPGA according to a set time slot or period, and data preprocessing carried out in the data writing back and reading process is migrated from the CPU to an FPGA layer to be executed; three-buffer independent access and atomic state switching are combined, and data interaction among multiple cores is carried out; the compiler automatically inserts a prefetch instruction in a loop expansion and pipeline scheduling stage; and the compiler performs linear address arrangement on the data members in a physical address space. The instruction execution efficiency and the data access efficiency are improved.
Owner:BEIJING SIFANG JIBAO ENG TECH +1

A General Neural Network Acceleration Method Based on 3D Loop Unrolling

The present invention belongs to the technical field of neural network acceleration and processing unit design, and proposes a general neural network acceleration method based on three-dimensional loop unrolling. The method includes: receiving instructions and sending decoded instruction information to configure each control module and the core computing unit; enabling after receiving the start signal; reading input data and weight data; reading the input data and weights from the zero-level buffer and writing them into the core computing unit; the core computing unit continuously generates pre-reading enable signals to respectively control the read-write control modules for partial sum data, quantization weights, and output data; the output data read-write control module reads the output data from the core computing unit and writes it into the first-level buffer, and after completing all the writing work, sends a calculation end signal. The method performs parallel computing in three dimensions of the input channels, output channels, and output feature maps, has higher computing performance and data reuse rate, and reduces the power consumption caused by memory access.
Owner:BEIJING INST OF TECH

Cyclic vectorization method and apparatus

PendingCN122285004Areduce in quantityReduce risk of spillageAlgorithmControl theory
A loop vectorization method and apparatus are disclosed. The loop vectorization method includes: in response to determining that the current loop can be rolled back, creating a first data structure corresponding to the current loop after the loop rollback; for the current loop and the first data structure, determining the maximum vectorization factor and candidate vectorization factors of the current loop that will not cause register overflow, respectively; determining the optimal vectorization factor from the candidate vectorization factors; in response that the optimal vectorization factor is a candidate vectorization factor determined for the first data structure, performing a loop rollback on the current loop to reduce the number of registers used in each iteration of the current loop after the loop rollback; selectively performing loop unrolling on the current loop after the loop rollback to obtain a first target loop; and performing a vectorization transformation on the first target loop based on the optimal vectorization factor.
Owner:SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1

Method for accelerating AI operator library based on data stream architecture

A method for accelerating an AI operator library based on a data stream architecture comprises the steps that S1, the operation type in operators is recognized, if the operation type is calculation-intensive operation, the step S2 is executed, and if the operation type is data-intensive operation, the step S3 is executed; s2, the step S2 comprises the following steps of S21 to S25: S21, generating a configuration file according to the matrix dimension and the hardware parameters; s22, generating a calculation task scheduling table according to the configuration file and the matrix dimension information; s23, generating an optimization assembly instruction; s24, constructing a data flow diagram and a PE communication relation diagram; s25, generating a master control code, a scheduling strategy and a DMA transmission strategy; s3, the step S3 comprises the following steps of S31 to S34: S31, defining a mapping relation between the PE and the loop unroll; s32, establishing and compiling a task folder; s33, establishing mapping among graph nodes, csv numbers and PE numbers; and S34, establishing an edge between the nodes.
Owner:SUZHOU RICORE IC TECH LTD

Information processing apparatus and information processing method

An information processing apparatus that generates an operation configuration including functional units to be used, a connection path between the functional units, and an output path of an operation result, for a reduction operation, based on a plurality of pieces of operation input data including sets of two pieces of data, the information processing apparatus comprising, a memory, and a processor coupled to the memory and configured to, generate an initial operation configuration based on a loop unrolling factor that is a number of the functional units that perform the predetermined operation with the operation input data as a direct input, and generate a first operation configuration by adding an output path of an operation result for a different degree of parallelism in a case where the predetermined operation is repeated with a predetermined degree of parallelism to the generated initial operation configuration.
Owner:FUJITSU LTD

Compilation method, device, equipment and storage medium

The present application discloses a compilation method, apparatus, device, and storage medium, belonging to the field of compilers. The method includes: loop unrolling an instruction sequence in a loop block (BB) to generate a triple BB and a triple DAG corresponding to the triple BB; the triple BB includes an instruction sequence that loops three times; the triple DAG is used to indicate the data dependencies between the instructions in the triple BB; window scheduling is performed based on the triple BB to determine a target scheduling window; the window scheduling is used to split the instruction sequence in the loop block into two parts based on the scheduling window; the target scheduling window is a scheduling window whose execution time meets a time condition, and the execution time is the time required to execute instructions within the window calculated based on the data dependencies in the triple DAG; and based on the scheduling result corresponding to the target scheduling window, the target code of the loop block in the target program is compiled. The above method can improve the compilation and execution efficiency of the loop block.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Automatic tuning method for improving heterogeneous parallel computing performance

The invention relates to the technical field of automatic tuning of Triton compilers, and discloses an automatic tuning method for improving heterogeneous parallel computing performance, which comprises the following steps of: constructing a parameter space containing thread quantity, circular pipeline depth, data block size, register use quantity limitation and circular expansion factors; a variational automatic encoder is utilized to encode historical optimization data, the dependency relationship among parameters is learned, and joint probability distribution of a parameter space is generated. The combination of a Triton compiler and machine learning is realized for the first time, an optimal parameter combination is selected from rich configurations, the compiling overhead is reduced, the reasoning efficiency of a front-end model is improved, the energy consumption of a parallel system is saved, and the prediction precision, the balance performance and the energy efficiency are continuously improved through iterative optimization. The problem that parallel parameters cannot be combined with application programs and hardware feature information to make optimal solutions is solved.
Owner:ZHENGZHOU UNIV