Automatic parallelization method and system for enhancing dependency analysis

Through flow-sensitive pointer analysis and alias analysis technology, the limitations of traditional automatic parallelization methods in the face of complex program logic and new programming language characteristics are solved, more accurate data dependency analysis and parallelization decision-making are achieved, and the operation efficiency of programs on multi-core hardware is improved.

CN120066520APending Publication Date: 2025-05-30SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171083.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional automatic parallelization method has limitations in the face of complex program logic and new programming language characteristics, and it is difficult to accurately analyze the dynamic changes of pointers, the dependencies between complex data structures and functions, resulting in inaccuracy of wrong judgments and parallelized decisions.

Method used

Flow-sensitive pointer analysis and alias analysis technology are used to accurately track the dynamic changes of pointers under different program structures and operation statements, identify the alias relationship between pointers and other variables or pointers, perform data-dependent classification marking, and then guide parallel decision-making.

Benefits of technology

It significantly improves the accuracy of data-dependent analysis, reduces error judgments, enhances the accuracy and effectiveness of parallel decisions, improves the operation efficiency of programs on multi-core hardware, reduces execution time and improves system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066520A_ABST
    Figure CN120066520A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic parallelization method and system for enhancing dependency analysis, and belongs to the technical field of computer software, and the method comprises the following steps: tracking the dynamic change of a pointer under different program structures and operation statements according to a program execution sequence, and accurately mastering the pointing direction and the value range of the pointer; when the pointer variable is identified, starting to check variable information in the program from the declaration and the definition source of the variable, and comprehensively judging the alias relationship between the pointer and other variables or the pointer according to the assignment condition and the function calling scene, so as to accurately identify the data dependency relationship in the program; according to results obtained by stream sensitive pointer analysis and alias analysis, performing classification marking on data dependence in the internal representation of the program; and the compiler implements parallelization processing on the part without data dependence or capable of parallel dependence according to the mark. According to the method, misjudgment easily occurring in a traditional method can be effectively reduced, and the accuracy and effectiveness of parallelization decision making are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer software, and more specifically, to an automatic parallelization method and system for enhancing dependency analysis. Background Art

[0002] In the current rapid development of computer technology, multi-core processors and parallel computing architectures dominate, and the complexity of software is increasing continuously. Automatic parallelization technology has emerged to improve the running efficiency of programs on multi-core hardware.

[0003] Although traditional automatic parallelization methods can identify some parallel structures, they have many limitations in the face of complex program logic. In terms of dependency analysis, based on simple static analysis, they have insufficient analysis capabilities for dynamic changes of pointers, complex data structures, and dependencies between functions. For example, they are prone to errors in scenarios such as pointer linked lists and nested structure arrays, and it is also difficult to exploit parallelism in recursive functions.

[0004] With the emergence of new programming language features, traditional technologies are even more difficult to adapt, and they are in trouble in terms of dependency analysis and parallelization decision-making, unable to meet the requirements of efficient parallel execution of modern complex software. There is an urgent need for new automatic parallelization methods and systems for enhancing dependency analysis to break through the bottleneck. Summary of the Invention

[0005] The technical task of the present invention is to address the above deficiencies and provide an automatic parallelization method and system for enhancing dependency analysis, which can effectively reduce the error judgments prone to traditional methods, enhance the accuracy and effectiveness of parallelization decisions, reduce execution time, improve system throughput, and improve software development efficiency and quality, while reducing time and labor costs.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] An automatic parallelization method for enhancing dependency analysis, the implementation of this method includes the following steps:

[0008] 1) Flow-sensitive pointer analysis: Trace the dynamic changes of pointers in different program structures (such as loops, branches, function calls, etc.) and operation statements (such as pointer assignments, arithmetic operations, indirect accesses, etc.) in detail according to the program execution order, and accurately master the pointer pointing and value range.

[0009] 2) Alias analysis: After identifying pointer variables, check the variable information in the program from the source of variable declarations and definitions, and comprehensively judge the alias relationship between pointers and other variables or pointers according to the assignment situation and function call scenarios, so as to accurately identify the data dependency relationships in the program, overcoming the inaccuracy problems in dependency analysis of traditional technologies.

[0010] 3) Dependency marking: Based on the results obtained from flow-sensitive pointer analysis and alias analysis, classify and mark data dependencies in the internal representation of the program, including true dependencies, anti-dependencies, and output dependencies, providing a clear and accurate basis for subsequent parallelization decisions;

[0011] 4) Parallelization decision: Based on the above markings, the compiler performs parallelization processing on parts without data dependencies or with parallelizable dependencies, making full use of the hardware's parallel computing capabilities; for parts with complex data dependencies, combined with hardware characteristics (such as the number of processor cores, cache size, memory bandwidth, etc.), formulate flexible and diverse parallelization strategies, including: data chunking, task splitting, reasonably introducing synchronization mechanisms, or maintaining serial execution, etc., to ensure that the program balances performance and correctness during the parallelization process, effectively improving the compiler's ability and effect of automatic parallelization, solving the problems in parallelization decision-making in traditional technologies, which is also the core advantage of this method compared with previous technologies.

[0012] Furthermore, for the flow-sensitive pointer analysis, after the compiler receives the code of the program to be compiled, it checks each variable in it. If it is a pointer variable, it starts the analysis process; according to the control flow graph and execution order of the program, it traverses the basic blocks and internal statements of the program in sequence;

[0013] When encountering a pointer assignment statement, immediately update the pointing of the pointer;

[0014] For pointer arithmetic statements, accurately calculate the new pointing position;

[0015] In a loop structure, for each iteration and different loop conditions, carefully record the state changes of the pointer, including key information such as its pointing position and possible value range;

[0016] For conditional branch structures (such as "if-else" structures), according to different branch conditions, respectively record the state evolution of the pointer under each branch and properly store this information, providing accurate pointer data for subsequent dependency analysis to effectively handle complex program structures.

[0017] Furthermore, when encountering a pointer assignment statement of "p = q", immediately update the pointing of pointer "p" to the pointing of "q";

[0018] For a pointer arithmetic statement of "p = p + offset", based on the current pointing of "p" and the value of "offset", accurately calculate the new pointing position of "p";

[0019] The loop structure includes "for" loops or "while" loops;

[0020] The conditional branch structure includes "if-else" structures.

[0021] Furthermore, for the alias analysis, the assignment cases include direct assignment and indirect assignment; the function call scenarios include parameter passing and return value handling;

[0022] For the direct assignment statement "p = q", quickly identify and mark "p" and "q" as aliases;

[0023] For the indirect assignment statement "*r = p", by means of static analysis including analyzing the scope and storage location of variables, carefully analyze the multiple variables that "r" may point to to determine whether alias situations will occur;

[0024] In the function call scenario, fully consider the alias possibilities of parameter passing and return values; for example, when the function parameter is a pointer, deeply analyze the alias situations of the pointer during the function call process, covering the pointer passed to the function and the pointer returned by the function, to ensure comprehensive alias analysis.

[0025] Furthermore, for the dependency marking,

[0026] If the calculation of a variable depends on the current value of another variable, mark it as a true dependency;

[0027] If the read operation of a variable affects the subsequent update of another variable, mark it as an anti-dependency;

[0028] For the situation of updating the same variable multiple times, mark it as an output dependency;

[0029] These marked information will become an important basis for subsequent parallelization decisions, effectively ensuring the accuracy and scientific nature of the decisions.

[0030] Furthermore, for the dependency marking, the internal representation of the program includes an abstract syntax tree or intermediate representation.

[0031] Furthermore, for the cumulative summation of array elements, the process of enhancing the automatic parallelization method of dependency analysis is as follows:

[0032] Step 1: First, the compiler reads and parses the C language program containing the cumulative summation function array_sum, and constructs the abstract syntax tree (AST) and control flow graph (CFG) of the program; for the array_sum function, the control flow graph contains a simple for loop that traverses the array and updates the value of sum;

[0033] Step 2: Analyze the loop using the flow-sensitive pointer analysis method. In the for loop of the array_sum function, the compiler checks the statement sum += arr[i]; For the sum variable, its state change is recorded at each iteration. Initially, the value of sum is 0. In each iteration, the compiler tracks the update operation of sum and finds that the new value of sum is its old value plus the value of arr[i]. For example, in the first iteration, sum changes from 0 to 0 + arr[0]; in the second iteration, it changes from 0 + arr[0] to (0 + arr[0]) + arr[1], and so on. The compiler records this state change precisely, which is similar to the pointer state change recording in flow-sensitive pointer analysis, except that here we focus on the value change of the variable sum.

[0034] Step 3: Check the variables using the alias analysis method. For the arr array, the compiler checks whether it will be aliased in other parts of the program. If arr is passed to other functions or there are other operations that may cause aliasing, the compiler will analyze these situations.

[0035] Step 4: Mark data dependencies. For the operation sum += arr[i]; the compiler will perform dependency marking based on the results of the flow-sensitive pointer analysis. Since the update operation of sum depends on its own old value, the compiler will mark it as an output dependency. This means that the current update of sum depends on the previous value of sum and cannot be updated simultaneously with other threads to avoid data races. For the access to arr[i], the compiler will mark it as a read operation on the arr array and will determine its dependency on the elements of the arr array based on the value of i.

[0036] Step 5: Based on the precise dependency marking, the compiler will recognize that there is an output dependency in the update operation of sum. Therefore, the entire for loop cannot be simply parallelized because simultaneous updates of sum by multiple threads will lead to data races. For example, if the entire for loop is wrongly parallelized, two threads may simultaneously read the old value of sum, calculate the new value, and update it, resulting in incorrect results. This method will divide the array into multiple parts and allocate the summation tasks of different parts to different threads according to the number of processor cores or other performance metrics. Specifically, it includes:

[0037] First, use omp_get_num_threads() to obtain the number of threads and divide the array arr into multiple chunks.

[0038] Each thread obtains its own thread ID tid through omp_get_thread_num() and calculates the starting start and ending end indices of the array part it is responsible for.

[0039] Each thread uses local_sum to calculate the sum of the elements in the array block it is responsible for, avoiding concurrent modification of sum;

[0040] Each thread stores its partial sum in partial_sums[tid];

[0041] Step 6, Result aggregation: After all threads complete their respective summation tasks, add the partial sums of each thread to obtain the final result.

[0042] The present invention also claims a system for automatic parallelization that enhances dependency analysis, including:

[0043] A flow-sensitive pointer analysis module, which is used to carefully track the dynamic changes of pointers under different program structures (such as loops, branches, function calls, etc.) and operation statements (such as pointer assignments, arithmetic operations, indirect accesses, etc.) according to the program execution order, and accurately master the pointer pointing and value range;

[0044] An alias analysis module, which is used to check the variable information in the program starting from the source of variable declarations and definitions after identifying pointer variables, and comprehensively judge the alias relationship between pointers and other variables or pointers according to the assignment situation (direct assignment, indirect assignment) and function call scenarios (parameter passing, return value processing), so as to accurately identify the data dependency relationship in the program;

[0045] A dependency marking module, which is used to classify and mark data dependencies in the internal representation of the program according to the results obtained from flow-sensitive pointer analysis and alias analysis, including true dependencies, anti-dependencies, and output dependencies, providing a clear and accurate basis for subsequent parallelization decisions;

[0046] A parallelization decision module. The compiler performs parallelization processing on the parts without data dependencies or parallelizable dependencies according to the markings, making full use of the hardware parallel computing ability; for the parts with complex data dependencies, flexible and diverse parallelization strategies are formulated in combination with hardware characteristics (such as the number of processor cores, cache size, memory bandwidth, etc.), including: data chunking, task splitting, reasonably introducing synchronization mechanisms or maintaining serial execution, etc., to ensure that the program takes into account both performance and correctness during the parallelization process;

[0047] This system implements the above-mentioned method for automatic parallelization that enhances dependency analysis.

[0048] The present invention also claims an implementation device for automatic parallelization that enhances dependency analysis, including: at least one memory and at least one processor;

[0049] The at least one memory is used to store machine-readable programs;

[0050] The at least one processor is used to call the machine-readable program to implement the above method.

[0051] The present invention also claims protection for a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the above-mentioned method is implemented.

[0052] Compared with the prior art, the automatic parallelization method and system for enhancing dependence analysis of the present invention have the following beneficial effects:

[0053] The present invention innovatively integrates flow-sensitive, context-sensitive pointer analysis, and advanced alias analysis technologies, deeply analyzes the program control flow and data flow, which greatly improves the accuracy of data dependence analysis, accurately tracks pointer dynamics, accurately identifies alias relationships, and effectively reduces the misjudgments prone to traditional methods. Based on accurate dependence analysis, the compiler can accurately identify the parallelizable parts, efficiently utilize the parallel computing of multi-core processors, and at the same time formulate strategies for complex dependence parts in combination with hardware characteristics to avoid blind parallelization, greatly enhancing the accuracy and effectiveness of parallelization decisions. These improvements comprehensively improve the running efficiency of the program on multi-core hardware, greatly reduce the execution time, increase the system throughput, and have particularly prominent advantages in scenarios with high performance requirements such as large-scale data processing, fully exploiting the parallel potential of the program and greatly improving the user experience.

[0054] In software development, it reduces the debugging and repair work caused by parallelization errors, shortens the development cycle, enables developers to focus on business logic, improves the software development efficiency and quality, and reduces time and labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a flowchart of an automatic parallelization method for enhancing dependence analysis provided by an embodiment of the present invention;

[0056] Figure 2 is a diagram showing the cumulative summation process of array elements based on the automatic parallelization method for enhancing dependence analysis provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The present invention will be further described below in conjunction with specific embodiments.

[0058] An embodiment of the present invention provides an automatic parallelization method for enhancing dependence analysis, including:

[0059] 1. Precise Dependency Analysis: Through a comprehensive analysis of the program's control flow and data flow, in flow-sensitive pointer analysis, according to the program execution order, carefully track the dynamic changes of pointers under different program structures (such as loops, branches, function calls, etc.) and operation statements (such as pointer assignments, arithmetic operations, indirect accesses, etc.), and accurately master the pointer pointing and value range; in alias analysis, starting from the source of variable declarations, comprehensively consider various assignment situations (direct assignment, indirect assignment) and function call scenarios (parameter passing, return value handling), and comprehensively judge the alias relationship between pointers and other variables or pointers, so as to accurately identify the data dependency relationships in the program, overcoming the inaccuracy of traditional technologies in dependency analysis.

[0060] 2. Systematic Parallelization Process: Based on the results of precise dependency analysis, classify and mark data dependencies (true dependencies, anti-dependencies, output dependencies) in the internal representation of the program, providing a clear and accurate basis for subsequent parallelization decisions. The compiler, based on these marks, performs parallelization processing on parts without data dependencies or parallelizable dependencies, making full use of the hardware parallel computing power; for parts with complex data dependencies, combined with hardware characteristics (number of processor cores, cache size, memory bandwidth, etc.), formulate flexible parallelization strategies (such as data chunking, task splitting, reasonably introducing synchronization mechanisms or maintaining serial execution, etc.), ensuring that the program takes into account both performance and correctness during the parallelization process, effectively improving the ability and effect of the compiler's automatic parallelization, solving the problems in parallelization decisions of traditional technologies, which is also the core advantage of this method compared with previous technologies.

[0061] In traditional compiler automatic parallelization technologies, parallelism mainly relies on relatively simple static analysis and basic program structure recognition. These technologies usually only conduct a rough analysis of the program's surface structure, and are not precise enough in data dependency judgment, especially when dealing with pointer operations and complex program structures, there are many limitations. They are difficult to effectively track the changes of pointers at different program execution stages, nor can they accurately identify the alias relationships between pointers, resulting in frequent errors in the parallelization decision-making process, seriously affecting the program performance and the correctness of the results. The improved compiler automatic parallelization technology proposed by this method is a breakthrough in traditional technologies. It deeply integrates flow-sensitive and context-sensitive pointer analysis technologies with advanced alias analysis technologies to construct a complete automatic parallelization process. The overall process is as Figure 1 shown, and the specific steps are as follows:

[0062] 1. Flow-Sensitive Pointer Analysis:

[0063] After the compiler receives the program code to be compiled, it checks each variable in it. If it is a pointer variable, it starts the analysis process. According to the program's control flow graph and execution order, it traverses the basic blocks and internal statements of the program in turn.

[0064] When encountering a pointer assignment statement "p = q", immediately update the pointing of pointer "p" to the pointing of "q".

[0065] For a pointer arithmetic operation statement like "p = p + offset", accurately calculate the new pointing position of "p" based on the current pointing of "p" and the value of "offset".

[0066] In a loop structure (such as a "for" loop or a "while" loop), for each iteration and different loop conditions, carefully record the state changes of the pointer, including key information such as its pointing position and possible value range.

[0067] For a conditional branch structure (such as an "if-else" structure), according to different branch conditions, respectively record the state evolution of the pointer under each branch, and properly store this information to provide accurate pointer data for subsequent dependence analysis to effectively handle complex program structures.

[0068] 2. Alias analysis:

[0069] When a pointer variable is identified, start checking the variable information in the program from the source of variable declaration and definition.

[0070] For a direct assignment statement "p = q", quickly identify and mark "p" and "q" as aliases.

[0071] For an indirect assignment statement "*r = p", by means of static analysis such as analyzing the scope and storage location of variables, carefully analyze the multiple variables that "r" may point to to determine whether an alias situation will occur.

[0072] In the scenario of function calls, fully consider the possibility of aliases in parameter passing and return values. For example, when a function parameter is a pointer, deeply analyze the alias situation of this pointer during the function call, covering the pointer passed to the function and the pointer returned by the function, to ensure comprehensive alias analysis.

[0073] 3. Dependence marking:

[0074] According to the results obtained from the previous flow-sensitive pointer analysis and alias analysis, carry out the classification marking work of data dependence relationships in the internal representation of the program (such as an abstract syntax tree or an intermediate representation), including true dependence, anti-dependence, and output dependence, to provide a clear and accurate basis for subsequent parallelization decisions.

[0075] If the calculation of a variable depends on the current value of another variable, mark it as a true dependence;

[0076] If the read operation of a variable affects the subsequent update of another variable, mark it as an anti-dependence;

[0077] For the case of updating the same variable multiple times, it is marked as output dependence;

[0078] These marked information will become an important basis for subsequent parallelization decisions, effectively ensuring the accuracy and scientific nature of the decisions.

[0079] 4. Parallelization decision:

[0080] Based on the marked data dependence relationships, the compiler starts to make parallelization decisions. For the program parts that are identified as having no data dependence or having parallelizable data dependence relationships, they are determined as parallelizable parts, and parallelization processing means will be adopted to allocate them to different processor cores or threads to achieve parallel execution, so as to make full use of the parallel computing power of the hardware. For the parts with complex data dependence, the tightness of the dependence relationships and the characteristics of the hardware, including the number of processor cores, cache size, memory bandwidth, etc., will be comprehensively considered.

[0081] The following takes the cumulative summation of array elements as an example to further illustrate the specific application of this method:

[0082]

[0083]

[0084] In the array_sum function, traditional automatic parallelization techniques may not be able to accurately identify the dependence relationship of the operation sum += arr[i]; This operation actually has output dependence because the value of sum depends on its previous value and is updated in each iteration. However, traditional techniques may wrongly consider each iteration as independent and not take into account that sum is a shared variable, resulting in potential parallelization errors.

[0085] The processing method of this method for the above example is as follows:

[0086] Step 1. First, the compiler reads and parses the C language program containing the cumulative summation function array_sum, and constructs the abstract syntax tree (AST) and control flow graph (CFG) of the program; For the array_sum function, the control flow graph contains a simple for loop that traverses the array and updates the value of sum.

[0087] Step 2: Analyze the loop using the flow-sensitive pointer analysis method: In the for loop of the array_sum function, the compiler checks the statement sum += arr[i]; For the sum variable, its state change is recorded at each iteration. Initially, the value of sum is 0. At each iteration, the compiler tracks the update operation of sum and finds that the new value of sum is the sum of its old value and the value of arr[i]. For example, in the first iteration, sum changes from 0 to 0 + arr[0]; in the second iteration, it changes from 0 + arr[0] to (0 + arr[0]) + arr[1], and so on. The compiler records this state change precisely, which is similar to the pointer state change record in flow-sensitive pointer analysis, except that here we focus on the value change of the variable sum.

[0088] Step 3: Check the variables using the alias analysis method: In this example, although there are no complex pointer alias problems, for the arr array, the compiler checks whether it will be aliased in other parts of the program. If arr is passed to other functions or there are other operations that may cause aliasing, the compiler will analyze these situations.

[0089] Step 4: Mark data dependencies: For the operation sum += arr[i]; the compiler will perform dependency marking based on the results of the flow-sensitive pointer analysis. Since the update operation of sum depends on its own old value, the compiler will mark it as an output dependency. This means that the current update of sum depends on the previous value of sum and cannot be updated simultaneously with other threads to avoid data races. For the access to arr[i], the compiler will mark it as a read operation on the arr array and will determine its dependency on the elements of the arr array based on the value of i.

[0090] Step 5: Based on the precise dependency marking, the compiler will identify that there is an output dependency in the update operation of sum. Therefore, the entire for loop cannot be simply parallelized because multiple threads updating sum simultaneously will lead to data races. For example, if the entire for loop is wrongly parallelized, two threads may simultaneously read the old value of sum, calculate the new value, and update it, resulting in incorrect results; this method will divide the array into multiple parts and allocate the summation tasks of different parts to different threads according to the number of processor cores or other performance metrics. The following are the specific steps:

[0091] / / Parallelized code

[0092]

[0093] First, use omp_get_num_threads() to obtain the number of threads and divide the array arr into multiple chunks;

[0094] Each thread obtains its own thread ID tid through omp_get_thread_num(), and calculates the starting start and ending end indices of the array part it is responsible for;

[0095] Each thread uses local_sum to calculate the sum of elements in the array block it is responsible for, avoiding concurrent modification of sum;

[0096] Each thread stores its partial sum in partial_sums[tid].

[0097] Step 6, Result aggregation: After all threads complete their respective summation tasks, add the partial sums of each thread to obtain the final result:

[0098]

[0099] Here, the local_sum of each thread is private, avoiding data competition, and the operation of finally adding these partial sums is sequential, ensuring the correctness of the result.

[0100] The embodiment of the present invention also provides an automatic parallelization system for enhancing dependence analysis, including:

[0101] A flow-sensitive pointer analysis module, which is used to carefully track the dynamic changes of pointers under different program structures (such as loops, branches, function calls, etc.) and operation statements (such as pointer assignments, arithmetic operations, indirect accesses, etc.) according to the program execution order, and accurately master the pointer pointing and value range;

[0102] An alias analysis module, which is used to check the variable information in the program from the source of variable declarations and definitions after identifying pointer variables, and comprehensively judge the alias relationship between pointers and other variables or pointers according to the assignment situation (direct assignment, indirect assignment) and function call scenarios (parameter passing, return value processing), so as to accurately identify the data dependence relationship in the program;

[0103] A dependence marking module, which is used to classify and mark data dependencies in the internal program representation according to the results obtained from flow-sensitive pointer analysis and alias analysis, including true dependencies, anti-dependencies, output dependencies, providing a clear and accurate basis for subsequent parallelization decisions;

[0104] A parallelization decision module. The compiler performs parallelization processing on the parts without data dependencies or parallelizable dependencies according to the markings, making full use of the hardware parallel computing power; for complex data dependency parts, flexible and diverse parallelization strategies are formulated in combination with hardware characteristics (number of processor cores, cache size, memory bandwidth, etc.), including: data chunking, task splitting, reasonably introducing synchronization mechanisms or maintaining serial execution, etc., to ensure that the program takes into account both performance and correctness during the parallelization process;

[0105] This system implements the automatic parallelization method for enhanced dependency analysis described in the above embodiments. Specifically as follows:

[0106] 1. Flow-sensitive pointer analysis:

[0107] After the compiler receives the program code to be compiled, it checks each variable therein. If it is a pointer variable, the analysis process is started. According to the control flow graph and execution order of the program, the basic blocks and internal statements of the program are traversed in sequence.

[0108] When encountering a pointer assignment statement "p = q", immediately update the pointing of pointer "p" to the pointing of "q";

[0109] For a pointer arithmetic operation statement like "p = p + offset", accurately calculate the new pointing position of "p" based on the current pointing of "p" and the value of "offset";

[0110] In a loop structure (such as a "for" loop or a "while" loop), for each iteration and different loop conditions, carefully record the state changes of the pointer, including key information such as its pointing position and possible value range;

[0111] For a conditional branch structure (such as an "if-else" structure), according to different branch conditions, respectively record the state evolution of the pointer under each branch, and properly store this information to provide accurate pointer data for subsequent dependency analysis to effectively handle complex program structures.

[0112] 2. Alias analysis:

[0113] When a pointer variable is identified, check the variable information in the program starting from the source of variable declaration and definition.

[0114] For a direct assignment statement "p = q", quickly identify and mark "p" and "q" as aliases;

[0115] For an indirect assignment statement "*r = p", by means of static analysis such as analyzing the scope and storage location of the variable, carefully analyze the multiple variables that "r" may point to to determine whether an alias situation will occur.

[0116] In the scenario of function calls, fully consider the possibility of aliases in parameter passing and return values. For example, when a function parameter is a pointer, deeply analyze the alias situation of this pointer during the function call, covering the pointer passed to the function and the pointer returned by the function to ensure comprehensive alias analysis.

[0117] 3. Dependency marking:

[0118] Based on the results obtained from the previous flow-sensitive pointer analysis and alias analysis, classify and mark the data dependencies in the internal representation of the program (such as the abstract syntax tree or intermediate representation), including true dependencies, anti-dependencies, and output dependencies, providing a clear and accurate basis for subsequent parallelization decisions.

[0119] If the calculation of a variable depends on the current value of another variable, mark it as a true dependency;

[0120] If the read operation of a variable affects the subsequent update of another variable, mark it as an anti-dependency;

[0121] For the case of updating the same variable multiple times, mark it as an output dependency;

[0122] These marked information will become an important basis for subsequent parallelization decisions, effectively ensuring the accuracy and scientific nature of the decisions.

[0123] 4. Parallelization Decision:

[0124] Based on the marked data dependencies, the compiler starts making parallelization decisions. For the parts of the program that are determined to have no data dependencies or have parallelizable data dependencies, they are identified as parallelizable parts, and parallelization processing means will be adopted to allocate them to different processor cores or threads for parallel execution to fully utilize the parallel computing power of the hardware. For the parts with complex data dependencies, the tightness of the dependencies and the characteristics of the hardware, including the number of processor cores, cache size, memory bandwidth, etc., will be comprehensively considered.

[0125] The embodiment of the present invention also provides an automatic parallelization implementation device for enhancing dependency analysis, including: at least one memory and at least one processor;

[0126] The at least one memory is used to store machine-readable programs;

[0127] The at least one processor is used to call the machine-readable program to implement the automatic parallelization method for enhancing dependency analysis described in the above embodiment.

[0128] The embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the automatic parallelization method for enhancing dependency analysis described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is made to read and execute the program codes stored in the storage medium.

[0129] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0130] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0131] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by making an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.

[0132] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is made to execute part or all of the actual operations, thereby realizing the functions of any one of the above embodiments.

[0133] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.

Claims

1. An automatic parallelization method for enhancing dependency analysis, characterized in that: The implementation of this method includes the following steps: 1) Flow-sensitive pointer analysis: Track the dynamic changes of pointers under different program structures and operation statements according to the program execution order, and accurately grasp the pointer pointing and value range; 2) Alias ​​analysis: After identifying a pointer variable, check the variable information in the program from the source of the variable declaration and definition. According to the assignment and function call scenario, comprehensively determine the alias relationship between the pointer and other variables or pointers, so as to accurately identify the data dependency in the program. 3) Dependency marking: Based on the results of flow-sensitive pointer analysis and alias analysis, data dependencies are classified and marked in the internal representation of the program, including true dependency, anti-dependency, and output dependency, providing a basis for subsequent parallelization decisions; 4) Parallelization decision: The compiler implements parallel processing for parts with no data dependency or parallel dependency based on the above tags; for parts with complex data dependency, various parallelization strategies are formulated in combination with hardware characteristics, including: data segmentation, task division, reasonable introduction of synchronization mechanism or maintaining serial execution, to ensure that the program takes into account both performance and correctness during the parallelization process.

2. The automatic parallelization method for enhancing dependency analysis according to claim 1, characterized in that: In the flow-sensitive pointer analysis, after receiving the program code to be compiled, the compiler checks each variable therein, and starts the analysis process if it is a pointer variable; according to the control flow graph and execution order of the program, the basic blocks and internal statements of the program are traversed in sequence; When a pointer assignment statement is encountered, the pointer is immediately updated; For pointer arithmetic statements, accurately calculate the new pointing position; In the loop structure, for each iteration and different loop conditions, the pointer state changes are recorded in detail, including the position it points to and the possible value range; For conditional branch structures, the state evolution of the pointer under each branch is recorded according to different branch conditions, and this information is stored to provide accurate pointer data for subsequent dependency analysis.

3. The automatic parallelization method for enhancing dependency analysis according to claim 2, characterized in that: When encountering a pointer assignment statement of "p = q", the pointer "p" is immediately updated to point to "q"; For the pointer arithmetic statement "p = p + offset", the new position of "p" is accurately calculated based on the current position of "p" and the value of "offset"; The loop structure includes a "for" loop or a "while" loop; The conditional branch structure includes an "if-else" structure.

4. The automatic parallelization method for enhancing dependency analysis according to claim 1 or 3, characterized in that: The alias analysis, the assignment conditions include direct assignment and indirect assignment; the function call scenarios include parameter passing and return value processing; For direct assignment statements "p=q", quickly identify and mark "p" and "q" as aliases; For the indirect assignment statement "*r=p", static analysis methods are used to analyze the scope and storage location of the variable, and to analyze the multiple variables that "r" may point to, so as to determine whether aliasing will occur; In the function call scenario, the alias possibility of parameter passing and return value is analyzed; when the function parameter is a pointer, the alias situation of the pointer during the function call process is analyzed, covering the pointer passed to the function and the pointer returned by the function to ensure comprehensive alias analysis.

5. The automatic parallelization method for enhancing dependency analysis according to claim 1, characterized in that: The dependency tag, If the computation of a variable depends on the current value of another variable, mark it as true dependency; If the read operation of one variable affects the subsequent update of another variable, it is marked as anti-dependency; In the case of updating the same variable multiple times, it is marked as output dependency.

6. The automatic parallelization method for enhancing dependency analysis according to claim 1 or 5, characterized in that: The dependency marker and the internal representation of the program include an abstract syntax tree or an intermediate representation.

7. The automatic parallelization method for enhancing dependency analysis according to claim 1, characterized in that: For the cumulative summation of array elements, the automatic parallelization method of enhanced dependency analysis is as follows: Step 1: First, the compiler reads and parses the C language program containing the cumulative sum function array_sum, and builds the program's abstract syntax tree and control flow graph; for the array_sum function, the control flow graph contains a simple for loop that traverses the array and updates the value of sum; Step 2: Use flow-sensitive pointer analysis to analyze the loop: In the for loop of the array_sum function, the compiler checks the statement sum + = arr[i]; for the sum variable, it records its state change at each iteration; initially, the value of sum is 0; at each iteration, the compiler tracks the update operation of sum and finds that the new value of sum is its old value plus the value of arr[i]; the compiler records this state change accurately; Step 3: Use alias analysis to check variables: For the arr array, the compiler checks whether it is aliased in other parts of the program; if arr is passed to other functions or there are other operations that may cause aliasing, the compiler will analyze these situations; Step 4: Mark data dependency: For the sum += arr[i] operation, the compiler will mark the dependency based on the result of flow-sensitive pointer analysis. Since the update operation of sum depends on its old value, the compiler will mark it as an output dependency. For the access of arr[i], the compiler will mark it as a read operation on the arr array, and will determine its dependency on the arr array elements based on the value of i. Step 5: Based on the precise dependency marking, the compiler will recognize that the update operation of sum has output dependency, divide the array into multiple parts, and assign the summing tasks of different parts to different threads according to the number of processor cores or other performance indicators; the details are as follows: First, use omp_get_num_threads() to get the number of threads and divide the array arr into multiple blocks; Each thread gets its own thread number tid through omp_get_thread_num() and calculates the start and end indexes of the array part it is responsible for; Each thread uses local_sum to calculate the sum of the elements of the array block it is responsible for, avoiding concurrent modifications to sum; Each thread stores its partial sum in partial_sums[tid]; Step 6: Result aggregation: After all threads complete their respective summation tasks, add up the partial sums of each thread to get the final result.

8. An automatic parallelization system for enhanced dependency analysis, characterized in that: include: Flow-sensitive pointer analysis module, which is used to track the dynamic changes of pointers under different program structures and operation statements according to the program execution order, and accurately grasp the pointer pointing and value range; Alias ​​analysis module, which is used to check the variable information in the program from the source of variable declaration and definition after identifying the pointer variable, and comprehensively judge the alias relationship between the pointer and other variables or pointers according to the assignment situation and function call scenario, so as to accurately identify the data dependency relationship in the program; The dependency marking module is used to classify and mark data dependencies in the program internal representation according to the results of flow-sensitive pointer analysis and alias analysis, including true dependency, anti-dependency, and output dependency, to provide a basis for subsequent parallelization decisions; Parallelization decision module: The compiler implements parallel processing for parts without data dependency or parallel dependency according to the above mark; for parts with complex data dependency, various parallelization strategies are formulated in combination with hardware characteristics, including: data segmentation, task division, reasonable introduction of synchronization mechanism or maintaining serial execution, to ensure that the program takes into account both performance and correctness during the parallelization process; The system implements the method described in any one of claims 1 to 7.

9. An automatic parallelization implementation device for enhanced dependency analysis, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.