Control flow quantization method and device based on conditional selection instruction
Through the control flow vector quantization method based on conditional selection instructions, the problem of difficult to vectorize the control flow structure on the processor platform without native mask fetching instructions is solved, efficient vector code generation is achieved, and the computing efficiency of the processor is improved.
Patent Information
- Application Number
- CN202510311471.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
On processor platforms that lack native mask fetch instructions, it is difficult for existing compilers to effectively vectorize the control flow structure, resulting in the failure to fully explore vectorization opportunities.
The control flow vector quantization method based on conditional selection instructions is adopted, and the control flow is flattened, the cost evaluation is performed, and vector code is generated based on the planarization code.
It realizes effective vectorization processing of control flow programs, generates efficient object code, fully utilizes the computing efficiency of the processor, and improves the vectorization ability of the compiler.
Smart Images

Figure CN120196366A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vectorization technology, and in particular to a control flow vectorization method based on conditional selection instructions and a control flow vectorization device based on the ShenWei platform compiler SWGCC. Background Art
[0002] Single Instruction Multiple Data (SIMD) technology transforms serial scalar operations on multiple data elements into parallel vector operations, which can effectively improve data processing efficiency and has become a basic component of modern high-performance processors. Automatic vectorization is an important means to explore SIMD parallelism in programs. The compiler automatically performs dependence analysis and program transformation on program code through its built-in vectorization engine, converting the originally serially executed scalar code into vector code on the target platform, which can effectively improve the usage efficiency of SIMD instructions. Control flow code, as common data processing logic in programs, has a crucial impact on the running efficiency of programs.
[0003] Currently, mainstream production compilers, such as GCC and LLVM, both use masked memory access instructions as the intermediate representation of control flow vectorization. For processor platforms lacking native masked instruction support, the compiler cannot fully exploit potential vectorization opportunities in programs. Summary of the Invention
[0004] The purpose of the present invention is to provide a control flow vectorization method based on conditional selection instructions to at least solve one of the above technical problems.
[0005] One aspect of the present invention provides a control flow vectorization method based on conditional selection instructions, and the control flow vectorization method based on conditional selection instructions includes: Obtain the statement to be processed in the source program; Perform control flow flattening on the statement to be processed in the source program to obtain flattened code; Evaluate the cost of the flattened code to determine whether to generate code. If so, Generate vector code according to the flattened code.
[0006] Optionally, before obtaining the statement to be processed in the source program, the control flow vectorization method based on conditional selection instructions further includes: Design a masked memory access instruction template suitable for the target platform at the compiler backend to support flattening the control flow with masked instructions at the if conversion stage.
[0007] Optionally, after performing control flow flattening on the statement to be processed in the source program to obtain flattened code, the control flow vectorization method based on conditional selection instructions further includes: Obtain the data type to be converted for the statement to be processed in the source program. According to the length of the vector type corresponding to the scalar data, set the vector mask type so that the mask type is consistent with the condition type in the conditional selection instruction.
[0008] Optionally, the cost evaluation of the flattened code to determine whether to perform code generation includes: Obtain a cost model, where the cost of the mask load instruction and the cost of the mask store instruction are added to the cost model; Perform cost evaluation on the flattened code through the code model.
[0009] Optionally, the cost of the mask instruction includes: The cost of the mask load instruction is equal to the sum of the cost of the regular vector load instruction and the vector selection instruction; The cost of the mask store instruction is equal to the sum of the cost of the regular vector load instruction, the vector selection instruction, and the regular vector store instruction.
[0010] Optionally, the generation of vector code according to the flattened code includes: Establish an instruction template; Generate vector code according to the flattened code through the instruction template.
[0011] Optionally, the generation of vector code according to the flattened code through the instruction template includes: Generate mask load vector code; Generate mask store vector code.
[0012] Optionally, the generation of mask load vector code includes: Transfer the data to a temporary register vtmp, and then use the vector conditional selection instruction to selectively write the elements in the temporary register vtmp into the target register according to the values in the mask vector register.
[0013] Optionally, the generation of mask store vector code includes: Send the original data at the storage address to the temporary register vtmp, and then generate a vector conditional selection instruction. The data in the temporary register vtmp is selectively concatenated with the data to be stored according to the values in the mask register and then written into the storage location.
[0014] This application also provides a control flow quantization device based on conditional selection instructions. The control flow quantization device based on conditional selection instructions includes: A source program statement to be processed acquisition module, which is used to acquire the source program statement to be processed; A flattening module, which is used to perform control flow flattening on the statements to be processed in the source program, so as to obtain flattened code; An evaluation module, which is used to evaluate the cost of the flattened code, so as to determine whether to generate code; A vector generation module, which is used to generate vector code according to the flattened code after the evaluation module evaluates to yes.
[0015] Beneficial effects: Aiming at the problem that the control flow structure is difficult to be effectively vectorized due to the lack of native mask memory access instructions in the processor instruction set, the control flow quantization method based on conditional selection instructions of the present application proposes a control flow quantization method based on conditional selection instructions, aiming to realize the vectorization processing of control flow programs, generate correct and efficient target code, and give full play to the computing efficiency of the processor. Description of the drawings
[0016] Figure 1 is a schematic flowchart of a control flow quantization method based on conditional selection instructions according to an embodiment of the present application.
[0017] Figure 2 is a schematic diagram of an electronic device for implementing a control flow quantization method based on conditional selection instructions according to an embodiment of the present application.
[0018] Figure 3 is a schematic diagram of an instruction template according to an embodiment of the present application.
[0019] Figure 4 is a schematic diagram of the specific operation logic in an embodiment of the present application. Detailed implementation manners
[0020] To make the purpose, technical solutions and advantages of the implementation of the present application clearer, the technical solutions in the embodiments of the present application will be described in more detail below with reference to the accompanying drawings in the embodiments of the present application. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, of the embodiments of the present application. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0021] As Figure 1 shown, the control flow quantization method based on conditional selection instructions includes: Obtain the statements to be processed in the source program; Perform control flow flattening on the statements to be processed in the source program to obtain flattened code; Perform cost evaluation on the flattened code to determine whether to perform code generation. If so, Generate vector code according to the flattened code.
[0022] A control flow vectorization method based on conditional selection instructions in this application aims at the problem that the lack of native masked memory access instructions in the processor instruction set makes it difficult to effectively vectorize the control flow structure. It proposes a control flow vectorization method based on conditional selection instructions, aiming to realize the vectorization of control flow programs, generate correct and efficient target code, and give full play to the computing efficiency of the processor.
[0023] In this embodiment, before obtaining the statements to be processed in the source program, the control flow vectorization method based on conditional selection instructions further includes: Design a masked memory access instruction template suitable for the target platform at the compiler backend to support flattening the control flow with masked instructions at the if-conversion stage.
[0024] Based on the characteristics of the target platform instruction set, design the operation type and data type of the masked instruction. Add different types of masked instruction templates at the compiler backend according to different combinations of operation types and data types. Detect the load / store statements in the control flow structure at the if-conversion stage, match the access type and data type of the statements with the masked instruction templates. If the match is successful, the corresponding masked load / store instruction can be used to replace the original control flow structure to achieve the flattening of the control flow code.
[0025] For example, the operation types include load and store; The data types include: octaword integer (8 32-bit integer data), quadword integer (4 64-bit integer data), quad single-precision floating point (4 single-precision floating point data), quad double-precision floating point (4 double-precision floating point data); forming the following eight masked operation instructions (octaword integer masked load, quadword integer masked load, quad single-precision floating point masked load, quad double-precision floating point masked load, octaword integer masked store, quadword integer masked store, quad single-precision floating point masked store, quad double-precision floating point masked store).
[0026] Taking the octaword integer masked load as an example: Load 8 32-bit integers from the starting address according to the value in the mask register into the target register.
[0027] Taking the octaword integer masked store as an example: Store 8 32-bit integers at the target address according to the value in the mask register.
[0028] In this embodiment, after performing control flow flattening on the statements to be processed in the source program to obtain flattened code, the control flow quantization method based on conditional selection instructions further includes: Set the vector mask type according to the vector type length corresponding to the scalar data, so that the mask type is consistent with the condition type in the conditional selection instruction.
[0029] For example, when the data type is a single-precision floating-point mask instruction, the mask type corresponding to quadruple single-precision floating-point data is quadruple long word integer. In this embodiment, the cost evaluation of the flattened code to determine whether to perform code generation includes: Obtain a cost model, where the cost of mask load instructions and the cost of mask store instructions are added to the cost model; in this embodiment, the instruction costs used in the cost model are encoded inside the compiler. Here, the cost encoding of mask instructions is newly added, and its calculation method is as follows: mask load = vector load + conditional selection; mask store = vector load + conditional selection + vector store.
[0030] Perform cost evaluation on the flattened code through a code model.
[0031] In this embodiment, the cost of the mask instruction includes: The cost of the mask load instruction is equal to the sum of the conventional vector load instruction and the vector selection instruction; The cost of the mask store instruction is equal to the sum of the conventional vector load instruction, the vector selection instruction, and the conventional vector store instruction.
[0032] In this embodiment, generating vector code according to the flattened code includes: Establish an instruction template; Generate vector code according to the flattened code through the instruction template.
[0033] In this embodiment, extension functions for generating vector code are built into the instruction template. When the compiler backend generates code, it will match the corresponding mask instruction template according to the operation type and data type of the mask instruction, and call the extension function to generate specific vector code.
[0034] In this embodiment, generating vector code according to the flattened code through the instruction template includes: Generate mask load vector code; Generate mask store vector code.
[0035] In one embodiment, see Figure 4 , the generating of the mask load vector code includes: Transfer the data to a temporary register vtmp, and then use a vector conditional selection instruction to selectively write the elements in the temporary register vtmp into the destination register according to the values in the mask vector register.
[0036] In this embodiment, the generating of the mask storage vector code includes: Transfer the data to a temporary register vtmp through a single conventional vector load from the memory address of array b, and then use a vector conditional selection instruction to selectively write the elements in the temporary register vtmp into the destination register vb according to the values in the mask vector register vp.
[0037] This application has the following advantages over the prior art: (1) Design a mask memory access instruction template suitable for the target platform at the compiler backend to support flattening the control flow with mask instructions at the if-conversion stage.
[0038] (2) At the vectorization analysis stage, adjust the data type of the mask instruction according to the mask instruction template to ensure that the generated vector intermediate representation meets the requirements of the compiler backend. Here, it refers to analyzing the data type. If none of the data types are supported, it means there is no corresponding vector instruction support, and the vectorization cost analysis cannot be carried out. The mask instruction template stipulates the data type of the mask operation and its mask type. During the compilation process, the data type and the mask type need to strictly comply with the design convention.
[0039] (3) For the code generation strategy of mask instructions, design a special cost model to analyze the vectorization benefit and avoid the situation of negative acceleration.
[0040] (4) Design an extended function corresponding to the mask instruction. At the code generation stage, automatically call the extended function in the instruction template according to the type of the mask instruction to expand the vectorized mask instruction into a native instruction sequence on the target platform.
[0041] In view of the problem that the lack of native mask memory access instructions in the processor instruction set makes it difficult to effectively vectorize the control flow structure, the present invention proposes a control flow vectorization method based on conditional selection instructions. Through the above method, the present invention effectively improves the vectorization ability of the production compiler on the target platform, can generate more efficient executable code for the target platform, and improves the execution efficiency of the program. In addition, it also further completes the vectorization function of the processor platform and provides a reference for the development of future new-generation processors.
[0042] Embodiment: A control flow vectorization method based on conditional selection instructions includes the following steps: (1) Add the required mask instruction template to the machine description file simd.md of SWGCC.
[0043] In the vector type analysis stage, mask instructions with the data type of single-precision floating point are processed so that the mask type corresponding to the quadruple single-precision floating-point data is quadruple long word integer.
[0044] In the cost model, add the cost of mask load / store instructions, and the calculation method is as follows: the cost of the mask load instruction is equal to the sum of the cost of the regular vector load instruction and the vector select instruction; the mask store instruction needs to add the cost of the regular vector store on top of this.
[0045] Define the extended function corresponding to the mask instruction in sw64.c. According to the operation type and data type of the mask instruction, expand the vectorized mask instruction into a native instruction sequence on the ShenWei platform.
[0046] In the control flow flattening stage, the compiler scans the instruction sequence in the region. When a load / store type statement is found, call the built-in function to check whether the target machine supports the corresponding type of mask load / store instruction. By adding the instruction template of mask load / store in the machine description file, the SWGCC compiler can provide mask instruction support here.
[0047] In the vectorized data type analysis, according to the characteristics of the ShenWei SIMD instruction set, special processing needs to be performed on the mask instruction to ensure that the obtained vector type matches the compiler backend and maintain the correctness of subsequent analysis. Take the process of obtaining the mask type with the single-precision floating-point vector type as an example. When it is detected that the vector data type is V4SF (quadruple single-precision floating point) and the scalar mask type is SI (word integer), switch the data type to V4DF (quadruple double-precision floating point) and the mask type to DI (long word integer). Then send it as a parameter to the compiler internal subroutine get_vectype, and the correct vector mask type V4DI (quadruple long word integer) can be automatically obtained. Before returning the vector mask type, the input parameters need to be restored to avoid errors in subsequent analysis.
[0048] On the ShenWei platform, mask memory access can be completed by adding additional vector conditional select instructions and vector load instructions in addition to the regular memory access. The additional overhead brought by the auxiliary instructions needs to be considered in the vectorized cost model. The cost of the mask load instruction includes one regular vector load and vector conditional select; the mask store instruction needs to add the cost of the regular vector store on top of this. When the cost of the vector code is less than the cost of the scalar code, the compiler converts the code into SIMD vector instructions.
[0049] During the code generation phase, the SWGCC compiler finds the corresponding instruction template according to the operation type and data type of the mask instruction, and calls the expansion function in the instruction template to expand the vectorized mask instruction into a native instruction sequence on the Shenwei platform. Taking the mask store instruction with the data type of quadruple single-precision floating-point as an example, its template definition is as Figure 3 shown. The keyword define_expand defines the name of the instruction template, maskstore represents the mask store operation, and v4sfv4di represents the corresponding data type and mask type. The content within the curly braces is the expansion function for generating the target instruction. The mask instruction expansion function in the instruction template defines a set of instruction sequences specific to the Shenwei processor, which can achieve the same operation effect as the native mask vector instruction. Its specific operation logic is as Figure 4 shown.
[0050] Taking the mask load instruction as an example, its specific method is as follows: (1) "First, through a regular vector load from the memory address of array b, transfer the data to a temporary register vtmp", which means generating a vector load instruction from the load address to send the data to the temporary register.
[0051] (2) "Then, use the vector conditional selection instruction to selectively write the elements in the temporary register vtmp into the target register vb according to the values in the mask vector register vp", which means generating a vector conditional selection instruction to selectively write the data in the temporary register into the target register according to the values in the mask register.
[0052] Code generation of the mask store instruction: (1) Generate a vector load instruction from the storage address to send the data to the temporary register.
[0053] (2) Generate a vector conditional selection instruction to selectively splice the data in the temporary register with the data to be stored according to the values in the mask register, and then write it into the target register.
[0054] (3) Generate a vector store instruction to write the content in the target register to the storage location.
[0055] In the mask storage stage, in addition to the data to be written back in vb, there are redundant elements. Therefore, it is necessary to perform a splicing operation with the original data that may be damaged at the write-back address. Here, the vector conditional selection instruction is used to selectively extract elements from va and vb according to the values in the mask vector register vp, so as to obtain the correct write-back data, and the result is stored in the target register vtmp. Finally, a conventional vector storage is performed to write the content of vtmp into the memory address of array a. Through the above two stages, the operation function of the mask instruction is completed.
[0056] The present application also provides a control flow quantization device based on the ShenWei platform compiler SWGCC, and the control flow quantization device based on the ShenWei platform compiler SWGCC includes: A source program statement to be processed acquisition module, which is used to acquire the source program statement to be processed; A flattening module, which is used to perform control flow flattening processing on the source program statement to be processed, so as to obtain flattened code; An evaluation module, which is used to perform cost evaluation on the flattened code, so as to determine whether to perform code generation; A vector generation module, which is used to generate vector code according to the flattened code after the evaluation module evaluates it as yes.
[0057] It should be noted that the foregoing explanations of the method embodiments are also applicable to the system of this embodiment, and will not be repeated here.
[0058] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the control flow quantization method based on conditional selection instructions as described above.
[0059] The present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the control flow quantization method based on conditional selection instructions as described above.
[0060] Figure 2 It is an exemplary structural diagram of an electronic device capable of implementing the control flow quantization method based on conditional selection instructions provided by an embodiment of the present application.
[0061] As Figure 2As shown, the electronic device includes an input device 501, an input interface 502, a central processing unit 503, a memory 504, an output interface 505, and an output device 506. Among them, the input interface 502, the central processing unit 503, the memory 504, and the output interface 505 are interconnected via a bus 507. The input device 501 and the output device 506 are respectively connected to the bus 507 through the input interface 502 and the output interface 505, and then connected to other components of the electronic device. Specifically, the input device 501 receives input information from the outside and transmits the input information to the central processing unit 503 through the input interface 502; the central processing unit 503 processes the input information based on the computer-executable instructions stored in the memory 504 to generate output information, stores the output information temporarily or permanently in the memory 504, and then transmits the output information to the output device 506 through the output interface 505; the output device 506 outputs the output information to the outside of the electronic device for user use.
[0062] That is to say, Figure 2 The electronic device shown can also be implemented to include: a memory storing computer-executable instructions; and one or more processors that, when executing the computer-executable instructions, can implement the control flow quantization method based on conditional selection instructions in combination with Figure 1 the description.
[0063] In one embodiment, Figure 2 The electronic device shown can be implemented to include: a memory 504 configured to store executable program code; one or more processors configured to run the executable program code stored in the memory 504 to execute the control flow quantization method based on conditional selection instructions in the above embodiment.
[0064] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0065] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0066] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0067] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the figures. For example, two consecutive blocks marked may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or overall flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0068] In this embodiment, the so-called processor may be a central processing unit (CPU), or it may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0069] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by invoking the data stored in the memory, the processor can implement various functions of the system / terminal device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0070] In this embodiment, if the modules / units integrated in the system / terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or system that can carry the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. Although this application is disclosed above with preferred embodiments, it is not actually used to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the protection scope of this application should be subject to the scope defined by the claims of this application.
[0071] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0072] In addition, it is obvious that the term "including" does not exclude other units or steps. The multiple units, modules, or systems stated in the system claims can also be implemented by one unit or a general system through software or hardware.
[0073] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A control flow vectorization method based on conditional selection instructions, characterized in that: The control flow vectorization method based on conditional selection instructions includes: Get the source program statements to be processed; Perform control flow flattening on statements to be processed in the source program to obtain flattened code; Evaluate the cost of the flattened code to determine whether to generate code. If so, A vector code is generated according to the flattened code.
2. The control flow vectorization method based on conditional selection instructions according to claim 1, characterized in that: Before obtaining the statements to be processed in the source program, the control flow vectorization method based on conditional selection instructions further includes: A masked memory access instruction template suitable for the target platform is designed in the compiler backend, which supports flattening the control flow with masked instructions in the if conversion stage.
3. The control flow vectorization method based on conditional selection instructions according to claim 2, characterized in that: After the control flow flattening process is performed on the statements to be processed in the source program to obtain the flattened code, the control flow vectorization method based on the conditional selection instruction further includes: Get the data type to be converted for the statement to be processed in the source program, set the vector mask type according to the vector type length corresponding to the scalar data, and make the mask type consistent with the condition type in the conditional selection instruction.
4. The control flow vectorization method based on conditional selection instructions according to claim 3, characterized in that: The cost evaluation of the flattened code to determine whether to generate the code includes: Obtaining a cost model, wherein the cost model includes a cost of a mask load instruction and a cost of a mask store instruction; The cost of the flattened code is evaluated through a code model.
5. The control flow vectorization method based on conditional selection instructions according to claim 4, characterized in that: The cost of the mask instruction includes: The cost of a masked load instruction is equal to the sum of the regular vector load instruction and the vector select instruction; The cost of a masked store instruction is equal to the sum of a regular vector load instruction, a vector select instruction, and a regular vector store instruction.
6. The control flow vectorization method based on conditional selection instructions according to claim 5, characterized in that: Generating a vector code according to the flattened code comprises: Create instruction templates; Generate vector code from the flattened code using instruction templates.
7. The control flow vectorization method based on conditional selection instructions according to claim 6, characterized in that: Generating vector code according to the flattened code by using the instruction template includes: Generate mask load vector code; Generates code for mask storage vectors.
8. The control flow vectorization method based on conditional selection instructions according to claim 7, characterized in that: The generating mask loading vector code comprises: The data is transferred to a temporary register vtmp, and then the vector conditional selection instruction is used to selectively write the elements in the temporary register vtmp into the target register according to the value in the mask vector register.
9. The control flow vectorization method based on conditional selection instructions according to claim 8, characterized in that: The generating mask storage vector code comprises: The original data at the storage address is sent to the temporary register vtmp, and then a vector conditional selection instruction is generated to select and splice the data in the temporary register vtmp with the data to be stored according to the value in the mask register, and then write it into the storage location.
10. A control flow vectorization device based on conditional selection instructions, characterized in that: The control flow vectorization device based on conditional selection instruction includes: A source program pending statement acquisition module, wherein the source program pending statement acquisition module is used to acquire source program pending statement; A flattening module, the flattening module is used to perform control flow flattening processing on the statements to be processed in the source program, so as to obtain flattened code; An evaluation module, the evaluation module is used to evaluate the cost of the flattened code, so as to determine whether to perform code generation; A vector generation module is used to generate a vector code according to the flattened code after the evaluation module evaluates to be yes.
Citation Information
Patent Citations
Flattening conditional statements
US20140013090A1