A Code Automatic Vectorization Optimization Method, Device and Medium
By performing standardized conversion, grouping processing and vector length matching processing on the scalar instruction graph of the calculation task, automatic vectorization of code is realized, and the problem of inconvenience in automatic vectorization of code in the prior art is solved, which significantly improves computing performance and reduces energy consumption.
Patent Information
- Application Number
- CN202510104501.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The automatic vectorization of existing codes is relatively inconvenient, and manual operations account for a large proportion, making it difficult to support vectorization of irregular or load-dependent codes, which is not conducive to improving computing performance and reducing energy consumption.
By obtaining the scalar instruction diagram of the calculation task, performing standard conversion processing, grouping according to the high-level organizational structure of the instructions, and splitting and matching the instruction vector based on the hardware's limited vector length, and finally generating a vector operation diagram to realize automatic vectorization of the code.
This method can significantly improve data processing speed, reduce memory access times, reduce energy consumption, and improve cache utilization. It can automatically optimize code without the need for programmers to manually perform complex optimization work, improving code portability and parallel processing capabilities.
Smart Images

Figure CN119536744B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of high-performance computing, and in particular, to a method, device, and medium for automatically optimizing code vectorization. Background Art
[0002] In the field of high-performance computing, parallel computing was initially widely used in scientific computing. In recent years, with the development of artificial intelligence and machine learning, the demand for using parallel computing to widely and effectively utilize computing resources has become more urgent. Vectorization is an important feature of parallel computing in modern computer architectures, which allows a single instruction to be executed on multiple data, that is, single instruction multiple data.
[0003] To utilize this function, the program must explicitly use vector instructions. Currently, this requirement mostly relies on professional programmers engaged in high-performance computing to manually vectorize the code in assembly language or using built-in functions. To reduce the professional knowledge required for code vectorization and improve the portability of the code, the backend compiler needs to vectorize the program when generating machine code or at runtime.
[0004] Therefore, developing more intelligent compilers and tools to more widely support automatic vectorization is an important research direction in the current field of high-performance computing. There is an urgent need for a code automatic vectorization optimization system that supports vectorizing irregular or load-dependent code, thereby improving computing performance and reducing energy consumption. Summary of the Invention
[0005] Embodiments of this application provide a method, device, and medium for automatically optimizing code vectorization to solve the following technical problems: Existing code automatic vectorization is relatively inconvenient, with a large proportion of manual operations, and it is difficult to support vectorizing irregular or load-dependent code, which is not conducive to improving computing performance and reducing energy consumption.
[0006] Embodiments of this application adopt the following technical solutions:
[0007] On the one hand, embodiments of this application provide a method for automatically optimizing code vectorization, including: obtaining a scalar instruction graph of a computing task according to the computing kernel of parallel computing; performing a reduction transformation process on the scalar instruction graph to obtain a reduced instruction graph; performing grouping processing related to mapping features on the reduced instruction graph according to the high-level organizational structure of the instructions to obtain a grouped graph; performing split matching processing on the instruction vectors in the grouped graph based on the limited vector length of the hardware to obtain a vector matching grouped graph; performing execution configuration on the elements of each group in the vector matching grouped graph to obtain a vector operation graph.
[0008] In the embodiments of the present application, by converting scalar instructions into vector instructions, this method can significantly improve the speed of data processing, especially when dealing with a large amount of data. Vectorized operations can generally reduce the number of memory accesses because they can process multiple data elements in one operation, thus reducing the pressure on memory bandwidth. Since vectorized operations can make more effective use of processor resources, it is possible to reduce energy consumption while increasing the computing speed. Vectorized instructions tend to make better use of the processor cache because they can process multiple data items at once, reducing cache misses. Through reduction transformation and grouped processing, this method can automatically optimize the code without the programmer having to manually perform complex optimization work. Vectorization techniques are usually closely related to specific hardware architectures. This method can improve the portability of the code across different hardware platforms because it can automatically apply vectorization on multiple hardware. Vectorization is a form of parallel processing, and this method can enhance the parallel processing ability of the application, especially when dealing with vector or matrix operations. Automatic vectorization can simplify the programming model because it reduces the complexity that the programmer needs to optimize manually. Through automated processing, vectorization optimization becomes easier to implement, especially for developers who are not familiar with vectorized programming.
[0009] In a feasible implementation manner, according to the computing cores of parallel computing, a scalar instruction graph of a computing task is obtained, specifically including: through the computing cores, performing assignment configuration of registers on the computing task to determine the assignment nodes between scalar instructions in the computing task; through specifying address data, performing storage configuration of the registers in the computing task to determine the load nodes between the scalar instructions; based on the operation logic corresponding to the input and output in the registers, determining the operation nodes between the scalar instructions; performing external memory processing of the operation results in the registers to a specified storage address and generating the storage nodes between the scalar instructions; according to the data dependency order and data execution order between the scalar instructions, and based on the assignment nodes, the load nodes, the operation nodes, and the storage nodes, performing tree-like analysis processing on each scalar instruction in the computing task under the logical operation relationship to obtain the scalar instruction graph.
[0010] In a feasible implementation manner, the scalar instruction graph is subjected to a reduction conversion process to obtain a reduced instruction graph, which specifically includes: performing parallel loading processing on each loading array in the scalar instruction graph; through a preset reduction operation, performing reduction processing on the element attributes in the loading array to obtain loading array element graph nodes; performing reduction processing on the loading execution algorithms of each element in the loading array to obtain element execution algorithm graph nodes; performing reduction processing on the storage results of each element in the loading array to obtain element storage graph nodes; performing corresponding conversion on the loading array element graph nodes, the element execution algorithm graph nodes, and the element storage graph nodes and integrating them into the scalar instruction graph to generate the reduced instruction graph.
[0011] In a feasible implementation manner, the reduction operation is a single operation jointly combined by arithmetic operations and logical operations on multiple elements.
[0012] In a feasible implementation manner, according to the high-level organizational structure of the instruction, the reduced instruction graph is subjected to grouping processing related to mapping characteristics to obtain a grouped graph, which specifically includes: according to the type characteristics and dependency characteristics of the element attributes, grouping and dividing the element attributes in the same loading array in the reduced instruction graph to obtain element attribute groups; according to the loading characteristics of the input array, grouping and dividing the input arrays under the same loading instruction in the reduced instruction graph to obtain loading instruction groups; according to the storage characteristics of the output array, grouping and dividing the output arrays under the same storage instruction in the reduced instruction graph to obtain storage instruction groups; according to each node in the scalar instruction graph, grouping the corresponding elements in the same processing node in each scalar instruction array to obtain node structure groups; through the high-level organizational structure, performing an integration process on the element attribute groups, the loading instruction groups, the storage instruction groups, and the node structure groups to obtain the grouped graph.
[0013] In a feasible implementation, based on the hardware-limited vector length, the instruction vectors in the grouped graph are split and matched to obtain a vector matching grouped graph, which specifically includes: comparing and determining the instruction vector length of the instruction vectors in the grouped graph with the limited vector length; if the length of each instruction vector is less than the limited vector length, then using the limited vector length as a matching template, and performing split matching processing on the length of the next instruction vector to obtain a supplementary instruction vector length; performing combined matching processing on the supplementary instruction vector length and the current instruction vector length to obtain a combined instruction vector length; wherein, the combined instruction vector length is equal to the limited vector length; if the length of each instruction vector is greater than the limited vector length, then based on the limited vector length as a matching template, performing split matching processing on the current instruction vector length to obtain a split instruction vector length; wherein, the split instruction vector length is equal to the limited vector length; based on the split instruction vector length and the combined instruction vector length, determining the instruction vector length that is less than the limited vector length as the remaining instruction vector length; wherein, the remaining instruction vector length is the value of the residual instruction vector after split and / or combined processing; based on the split instruction vector length, the combined instruction vector length, and the remaining instruction vector length, obtaining the vector matching grouped graph after completing vector length matching; wherein, the vector matching grouped graph is the grouped graph after array instruction division and matching.
[0014] In a feasible implementation, the elements of each group in the vector matching grouped graph are configured for execution to obtain a vector operation graph, which specifically includes: identifying the elements of each array in the vector matching grouped graph, and sorting the elements according to the data execution format and the data loading order to obtain the element order in each array; through the element order, performing corresponding permutation or extraction configuration on the elements in each array to obtain the vector operation graph; wherein, the vector operation graph is used to guide the vector instructions to be loaded into the register.
[0015] In a feasible implementation, after the elements of each group in the vector matching grouped graph are configured for execution to obtain a vector operation graph, the method further includes: converting the vector operation graph into a vector instruction list through a parallel computing backend; wherein, the vector instruction table contains the vector instructions to be executed on the target hardware; through the vector instruction list, performing vectorization processing on the code in the computing task to complete the compilation processing of the parallel computing data.
[0016] Second aspect, an embodiment of the present application further provides a code automatic vectorization optimization device, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute a code automatic vectorization optimization method according to any of the above embodiments.
[0017] Third aspect, an embodiment of the present application further provides a non-volatile computer storage medium, the storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, each program includes instructions, and when the instructions are executed by a terminal, the terminal executes a code automatic vectorization optimization method according to any of the above embodiments.
[0018] The present application provides a code automatic vectorization optimization method, device and medium. Compared with the prior art, the embodiments of the present application have the following beneficial technical effects:
[0019] 1. Improve computing efficiency: By converting scalar instructions into vector instructions, this method can significantly improve the speed of data processing, especially when processing a large amount of data.
[0020] 2. Reduce memory access: Vectorized operations can usually reduce the number of memory accesses, because they can process multiple data elements in one operation, thus reducing the pressure on memory bandwidth.
[0021] 3. Reduce energy consumption: Since vectorized operations can utilize processor resources more effectively, it is possible to reduce energy consumption while increasing the computing speed.
[0022] 4. Improve cache utilization: Vectorized instructions can often make better use of the processor cache, because they can process multiple data items at once, reducing cache misses.
[0023] 5. Code optimization: Through reduction transformation and grouped processing, this method can automatically optimize the code without the need for programmers to manually perform complex optimization work.
[0024] 6. Improve portability: Vectorization techniques are usually closely related to specific hardware architectures. This method can improve the portability of the code across different hardware platforms, because it can automatically apply vectorization on multiple hardware.
[0025] 7. Enhance parallel processing ability: Vectorization is a form of parallel processing, and this method can enhance the parallel processing ability of the application, especially when dealing with vector or matrix operations.
[0026] 8. Simplified programming model: Auto-vectorization can simplify the programming model as it reduces the complexity that programmers need to optimize manually.
[0027] 9. Compatibility: The method may be designed to be compatible with existing programming languages and compilers, enabling existing codebases to achieve performance improvements through optimization.
[0028] 10. Easy to implement: Through automated processing, vectorization optimization becomes easier to implement, especially for developers who are not familiar with vectorized programming. Description of the Drawings
[0029] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0030] Figure 1 is a flowchart of a method for automatically optimizing code vectorization provided by an embodiment of the present application;
[0031] Figure 2 is a schematic structural diagram of a scalar instruction graph obtained for a given computing task provided by an embodiment of the present application;
[0032] Figure 3 is a schematic structural diagram of a mapping vector grouping graph provided by an embodiment of the present application;
[0033] Figure 4 is a schematic structural diagram of a device for automatically optimizing code vectorization provided by an embodiment of the present application. Detailed Embodiments
[0034] To enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0035] An embodiment of the present application provides a method for automatically optimizing code vectorization. As Figure 1 shown, the method for automatically optimizing code vectorization specifically includes steps S101 - S105:
[0036] S101. Obtain the scalar instruction graph of the computing task according to the computing kernels for parallel computing.
[0037] Specifically, it is necessary to first perform register assignment configuration on the computing task through the computing kernels to determine the assignment nodes between the scalar instructions in the computing task.
[0038] Furthermore, perform storage configuration on the registers in the computing task through the specified address data to determine the load nodes between the scalar instructions.
[0039] Furthermore, determine the operation nodes between the scalar instructions based on the operation logic corresponding to the input and output in the registers.
[0040] Furthermore, perform external memory processing of the operation results in the registers to a specified storage address and generate storage nodes between the scalar instructions.
[0041] Furthermore, according to the data dependence order and data execution order between the scalar instructions, and based on the assignment nodes, load nodes, operation nodes, and storage nodes, perform tree-shaped analysis processing on each scalar instruction in the computing task under the logical operation relationship to obtain the scalar instruction graph.
[0042] As a feasible implementation manner, Figure 2 FIG. is a schematic structural diagram of the scalar instruction graph obtained for a given computing task provided by an embodiment of the present application. As Figure 2 shown, the scalar instruction graph represents the dependence order and execution order between all scalar instructions in the kernel. The scalar instruction graph consists of four types of nodes: assignment, load, operation, and storage. Assignment means assigning a specified register to a given value. Load means taking out the data at the specified address from the external memory and storing it in the register. Operation means using the values in one or more registers as inputs and obtaining the corresponding output values through certain arithmetic or logical operations. Storage means writing the value in the register back to the specified address in the external memory.
[0043] In one embodiment, as Figure 2 shown, assume the computing task is as follows: obtain the sum of all elements of arrays a and b each having 6 elements and store the calculation result in array c. Among them, the hardware limit vector size for executing the computing task is 4. First, obtain its scalar instruction graph during the compilation process of optimizing the system according to the given computing kernels.
[0044] S102. Perform reduction conversion processing on the scalar instruction graph to obtain a reduced instruction graph.
[0045] Specifically, perform parallel loading processing on each loaded array in the scalar instruction graph.
[0046] Further, through a preset reduction operation, the element attributes in the loaded array are processed by graph node reduction to obtain the loaded array element graph nodes. Among them, the reduction operation is a single operation jointly merged by arithmetic operations and logical operations on multiple elements.
[0047] Further, the loading execution algorithms of the respective elements in the loaded array are processed by graph node reduction to obtain the element execution algorithm graph nodes.
[0048] Further, the storage results of the respective elements in the loaded array are processed by graph node reduction to obtain the element storage graph nodes.
[0049] Further, the loaded array element graph nodes, the element execution algorithm graph nodes, and the element storage graph nodes are all correspondingly transformed and fused into the scalar instruction graph to generate a reduced instruction graph.
[0050] As a feasible implementation, the scalar instruction graph is transformed by reduction to obtain a reduced instruction graph. Reduction is a single operation jointly merged by arithmetic and logical operations on multiple elements. If there are arithmetic, logical operation modes, etc. in the scalar instruction graph, the scalar instruction graph is transformed so as to more effectively utilize these reduction operations in the subsequent vectorization process.
[0051] In one embodiment, as Figure 2 shown, in a computing task, the respective elements of array a and array b can be loaded in parallel because reduction can reduce the element attributes of the loaded array to a graph node, convert the arithmetic addition of the corresponding elements to a graph node, and convert the storage result to a graph node.
[0052] S103. According to the high-level organizational structure of the instruction, the reduced instruction graph is grouped according to relevant mapping features to obtain a grouped graph.
[0053] Specifically, according to the type features and dependency features of the element attributes, the element attributes in the reduced instruction graph that are in the same loaded array are grouped and divided to obtain an element attribute group.
[0054] Further, according to the loading features of the input array, the input arrays in the reduced instruction graph that are under the same loading instruction are grouped and divided to obtain a loading instruction group.
[0055] Further, according to the storage features of the output array, the output arrays in the reduced instruction graph that are under the same storage instruction are grouped and divided to obtain a storage instruction group.
[0056] Further, according to each node in the scalar instruction graph, the corresponding elements in each scalar instruction array that are at the same processing node are grouped to obtain a node structure group.
[0057] Furthermore, through a high-level organizational structure, the element attribute group, the loading instruction group, the storage instruction group and the node structure group are integrated and processed to obtain a grouping graph.
[0058] As a feasible implementation method, Figure 2 As shown, the instructions of the reduced instruction graph need to be grouped according to the mapping features to obtain a grouping graph. This grouping graph represents the higher-level organizational structure of the instructions, so that the instructions can be mapped to vector instructions more easily in the future. When grouping, the dependencies and data types between instructions are considered to ensure that the grouped instructions can be executed in parallel. For example, the load instructions for the same input array are grouped into one group, and the storage instructions for the same output array are grouped into another group.
[0059] In one embodiment, Figure 2 In the calculation task, the 6 elements of the loaded array a are first divided into one group, the 6 elements of the loaded array b are divided into one group, the arithmetic addition operations of the corresponding elements of the two arrays are divided into one group, the accumulation operations of the elements of the new array obtained after the arithmetic addition calculation of the two arrays are completed are divided into one group, and the storage of the final result is divided into one group.
[0060] S104. Based on the limited vector length of the hardware, split and match the instruction vectors in the grouping graph to obtain a vector matching grouping graph.
[0061] Specifically, it is also necessary to compare the instruction vector length of the instruction vector in the group diagram with the limited vector length:
[0062] If the length of each instruction vector is less than the limited vector length, the limited vector length is used as a matching template, and the next instruction vector length is split and matched to obtain a supplementary instruction vector length.
[0063] Furthermore, the supplementary instruction vector length is merged and matched with the current instruction vector length to obtain a merged instruction vector length, wherein the merged instruction vector length is equal to the limited vector length.
[0064] If the length of each instruction vector is greater than the limited vector length, the current instruction vector length is split and matched based on the limited vector length as a matching template to obtain a split instruction vector length, wherein the split instruction vector length is equal to the limited vector length.
[0065] Further, based on the split instruction vector length and the merge instruction vector length, the instruction vector length that is less than the limited vector length is determined as the remaining instruction vector length, wherein the remaining instruction vector length is the value of the remaining instruction vector after the split and / or merge process is completed.
[0066] Further, based on the split instruction vector length, the merge instruction vector length, and the remaining instruction vector length, a vector matching grouping graph after completing vector length matching is obtained. Among them, the vector matching grouping graph is the grouping graph after array instruction division matching.
[0067] As a feasible implementation manner, it is also necessary to split the instruction vectors in the grouping graph to match the given vector size. This vector size (limited vector length) is determined by the vector register size of the target hardware. For example: Since the number of instructions within a group (instruction vector length) grouped in the previous step may exceed the given vector size (limited vector length) and is not an integer multiple of the vector size, it is necessary to split the instructions within the group into multiple columns according to the vector size, that is, perform split matching processing on the current instruction vector length, with the length of each column being the limited vector size, and ensuring that the instruction column with a length less than the vector size (remaining instruction vector length) does not exceed one.
[0068] In one embodiment, since the number of array elements is 6 and the limited vector size is 4, the optimization system splits the instructions within the group according to the limited vector size. First, use the instructions within the group to fill up the instruction column with the given limited vector size, and preferably form an instruction column with a length less than the limited vector size with the remaining instructions in the array, that is, the remaining instruction vector length.
[0069] S105. Perform execution configuration on the elements of each group in the vector matching grouping graph to obtain a vector operation graph.
[0070] Specifically, first identify the elements of each array in the vector matching grouping graph, and sort the elements according to the data execution format and the data loading order to obtain the element order in each array.
[0071] Further, through the element order, perform corresponding permutation or extraction configuration on the elements in each array to obtain a vector operation graph. Among them, the vector operation graph is used to guide the vector instructions to be loaded into the register.
[0072] Further, through the parallel computing backend, convert the vector operation graph into a vector instruction list. Among them, the vector instruction table contains the vector instructions that need to be executed on the target hardware; through the vector instruction list, vectorize the code in the computing task to complete the compilation processing of the parallel computing data.
[0073] In one embodiment, Figure 3 This is a schematic structural diagram of a mapping vector grouping graph provided by an embodiment of the present application, as Figure 3As shown, first, set the execution order for the elements of each array in the vector matching grouped graph, and perform necessary permutation or extraction operations to ensure that the data is loaded into the vector register in the correct format and order. Then, use the backend to convert the finally generated vector operation graph into a list of vector instructions. This list of vector instructions contains vector instructions that can be executed on the target hardware and can efficiently process data in parallel.
[0074] In addition, the embodiment of the present application also provides a code automatic vectorization optimization device, such as Figure 4 As shown, the code automatic vectorization optimization device 400 specifically includes:
[0075] At least one processor 401. And, a memory 402 communicatively connected to the at least one processor 401. Among them, the memory 402 stores instructions that can be executed by the at least one processor 401, so that the at least one processor 401 can execute:
[0076] Obtain a scalar instruction graph of a computing task according to the computing kernel of parallel computing;
[0077] Perform a reduction conversion process on the scalar instruction graph to obtain a reduced instruction graph;
[0078] According to the high-level organizational structure of the instructions, perform grouping processing of the reduced instruction graph with respect to mapping characteristics to obtain a grouped graph;
[0079] Based on the hardware-limited vector length, perform split matching processing on the instruction vectors in the grouped graph to obtain a vector matching grouped graph;
[0080] Perform execution configuration on the elements of each group in the vector matching grouped graph to obtain a vector operation graph.
[0081] The embodiments of the present application can support the vectorization of irregular or load-dependent code, thereby improving computing performance and reducing energy consumption. At the same time, by converting scalar instructions into vector instructions, this method can significantly increase the speed of data processing, especially when dealing with a large amount of data. Vectorization operations can usually reduce the number of memory accesses because they can process multiple data elements in one operation, thus reducing the pressure on the memory bandwidth. Since vectorization operations can utilize processor resources more effectively, they can reduce energy consumption while increasing the computing speed. Vector instructions tend to make better use of the processor cache because they can process multiple data items at once, reducing cache misses. Through reduction transformation and grouped processing, this method can automatically optimize the code without the need for programmers to manually perform complex optimization work. Vectorization techniques are usually closely related to specific hardware architectures. This method can improve the portability of the code across different hardware platforms because it can automatically apply vectorization on multiple hardware. Vectorization is a form of parallel processing, and this method can enhance the parallel processing ability of applications, especially when dealing with vector or matrix operations. Automatic vectorization can simplify the programming model because it reduces the complexity that programmers need to optimize manually.
[0082] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0083] The devices and media provided by the embodiments of the present application correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0084] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0085] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the specified functions in one or more blocks or more blocks.
[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the specified functions in one or more blocks or more blocks.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the specified functions in one or more blocks or more blocks.
[0088] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0089] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0090] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0091] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0092] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the specification of the present application.
Claims
1. A method for automatic code vectorization optimization, characterized in that: The method comprises: According to the computing kernel of parallel computing, the scalar instruction graph of the computing task is obtained, including: By means of the computing kernel, the computing task is assigned a value to a relevant register, and an assignment node between scalar instructions in the computing task is determined; By specifying address data, the registers in the computing task are configured for storage, and a load node between the scalar instructions is determined; Determining the operation nodes between the scalar instructions based on the operation logic corresponding to the input and output in the register; Performing external memory processing of the operation result in the register at a designated storage address, and generating storage nodes between the scalar instructions; According to the data dependency order and data execution order between the scalar instructions, and based on the assignment node, the load node, the operation node and the storage node, each scalar instruction in the computing task is subjected to tree analysis processing under a logical operation relationship to obtain the scalar instruction graph; Performing protocol conversion processing on the scalar instruction graph to obtain a protocol instruction graph specifically includes: Performing parallel loading processing on each load array in the scalar instruction graph; By means of a preset reduction operation, the element attributes in the loading array are subjected to a graph node reduction process to obtain a loading array element graph node; The loading execution algorithm of each element in the loading array is subjected to graph node specification processing to obtain an element execution algorithm graph node; Performing graph node reduction processing on the storage results of each element in the loading array to obtain an element storage graph node; The load array element graph node, the element execution algorithm graph node and the element storage graph node are all converted accordingly and merged into the scalar instruction graph to generate the reduced instruction graph; According to the high-level organizational structure of the instructions, the protocol instruction graph is grouped according to the mapping features to obtain a grouping graph; Based on the limited vector length of the hardware, the instruction vectors in the grouping graph are split and matched to obtain a vector matching grouping graph; The elements of each group in the vector matching grouping graph are configured for execution to obtain a vector operation graph.
2. A method for automatic code vectorization optimization according to claim 1, characterized in that: The reduction operation is a single operation that combines arithmetic operations and logical operations on multiple elements.
3. A method for automatic code vectorization optimization according to claim 1, characterized in that: According to the high-level organizational structure of the instructions, the protocol instruction graph is grouped according to the mapping features to obtain a grouping graph, which specifically includes: According to the type characteristics and dependency characteristics of the element attributes, the element attributes in the same load array in the protocol instruction graph are grouped and divided to obtain element attribute groups; According to the loading characteristics of the input array, the input arrays under the same loading instruction in the protocol instruction graph are divided into groups to obtain loading instruction groups; According to the storage characteristics of the output arrays, the output arrays under the same storage instruction in the protocol instruction graph are grouped and divided to obtain storage instruction groups; According to each node in the scalar instruction graph, corresponding elements in the same processing node in each scalar instruction array are grouped and processed to obtain a node structure group; The element attribute group, the loading instruction group, the storage instruction group and the node structure group are integrated and processed through the high-level organizational structure to obtain the grouping graph.
4. The method for automatic code vectorization optimization according to claim 1, characterized in that: Based on the limited vector length of the hardware, the instruction vectors in the grouping graph are split and matched to obtain a vector matching grouping graph, which specifically includes: Compare and judge the instruction vector length of the instruction vector in the group diagram with the limited vector length; If the length of each instruction vector is less than the limited vector length, the limited vector length is used as a matching template, and the next instruction vector length is split and matched to obtain a supplementary instruction vector length; The supplementary instruction vector length is merged and matched with the current instruction vector length to obtain a merged instruction vector length; wherein the merged instruction vector length is equal to the limited vector length; If the length of each instruction vector is greater than the limited vector length, then based on the limited vector length as a matching template, split matching processing is performed on the current instruction vector length to obtain a split instruction vector length; wherein the split instruction vector length is equal to the limited vector length; Based on the split instruction vector length and the merge instruction vector length, the instruction vector length that is less than the limited vector length is determined as the remaining instruction vector length; wherein the remaining instruction vector length is the value of the residual instruction vector after the splitting and / or merging process is completed; Based on the split instruction vector length, the merged instruction vector length and the remaining instruction vector length, the vector matching grouping graph after completing the vector length matching is obtained; wherein the vector matching grouping graph is a grouping graph after the array instruction partition matching.
5. The method for automatic code vectorization optimization according to claim 1, characterized in that: The elements of each group in the vector matching grouping graph are executed and configured to obtain a vector operation graph, which specifically includes: Identify the elements of each array in the vector matching grouping diagram, and sort the elements according to the data execution format and the data loading order to obtain the order of elements in each array; The elements in each array are correspondingly replaced or extracted and configured according to the element sequence to obtain the vector operation graph; wherein the vector operation graph is used to guide the vector instructions to be loaded into the register.
6. A method for automatic code vectorization optimization according to claim 5, characterized in that: After executing and configuring the elements of each group in the vector matching grouping graph to obtain a vector operation graph, the method further includes: The vector operation graph is converted into a vector instruction list through a parallel computing backend; wherein the vector instruction list includes vector instructions to be executed on the target hardware; The code in the computing task is vectorized through the vector instruction list to complete the compilation process of the parallel computing data.
7. A code automatic vectorization optimization device, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute the code automatic vectorization optimization method according to any one of claims 1-6.
8. A non-volatile computer storage medium, characterized in that: The storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, each of which includes instructions, and when the instructions are executed by a terminal, the terminal executes a code automatic vectorization optimization method according to any one of claims 1-6.