Program vectorization method and device
By building a mapping table of scalar instructions to vector instructions and using machine learning models to generate vector packages, the compiler's insufficient vector adaptability on the new instruction set is solved, and a wider vector support is achieved.
Patent Information
- Application Number
- CN202510164623.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, compilers have low adaptability when vectorizing scalar statements and cannot effectively perform vectorization when facing new instruction sets in the architecture.
By constructing a mapping table of scalar instructions to vector instructions, and using machine learning models to combine the mapping table to generate vector packets, the scalar instructions to be converted are converted into vector instructions, and data processing instructions are inserted to update registers.
Improve the support and scope of the compiler's vectorization function, so that the compiler can effectively vectorize program when facing a new instruction set.
Smart Images

Figure CN120067044A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a program vectorization method and device. Background Art
[0002] Vector instructions improve the execution rate of a program by processing multiple data at a time, and are the main method for current computer architectures to improve the single-core execution rate. The methods of using vector instructions include writing vector programs manually and automatically generating vector programs by compilers. The former requires developers to be familiar with the vector instructions of the architecture, and the latter requires the compiler to provide automatic vectorization support for the target architecture.
[0003] During the vector optimization process, the main difficulty of vector optimization is to keep the semantics of the optimized vector statements the same as those of the scalar statements before optimization. As a general optimization means, automatic compiler vectorization focuses more on the stability of the program after optimization, and will lose some vectorization opportunities. At the same time, the processor instruction set architecture is also an important factor affecting vectorization. The vectorization support for a certain instruction set is often specific, and the vector support and optimization methods for a certain instruction set often cannot be used in other instruction sets, resulting in weak adaptation of the compiler to new instruction sets.
[0004] Therefore, how to improve the vectorization function support and vectorization range of the compiler is a technical problem to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to solve the technical problem that in the prior art, the adaptability of the compiler during scalar statement vectorization is low and it cannot effectively perform vectorization when facing new instruction sets in the architecture.
[0006] To achieve the above technical purpose, on the one hand, the present invention provides a program vectorization method, which includes: Constructing a mapping table from scalar instructions to vector instructions based on the scalar instructions and vector instructions in the target instruction set; Processing the training program set by the compiler to obtain a training assembly file, and inputting the training assembly file into a pre-established machine learning model, and generating a vector package by the machine learning model in combination with the mapping table; Processing the program to be converted by the compiler to obtain a to-be-converted assembly file, and converting each scalar instruction in the to-be-converted assembly file into a corresponding vector instruction according to the mapping relationship in the vector package, and updating the register after inserting data processing instructions.
[0007] Further, before generating the vector package by the machine learning model in combination with the mapping table, the method further includes: Encode the basic blocks in the training assembly file into a population, and encode the mapping relationship from scalar instructions to vector instructions in the same basic block in the training assembly file into an individual based on the mapping table. The length of the individual is the number of vector instructions corresponding to the basic block to which the individual belongs, and the encoded bits of the individual represent the vector instructions after mapping at that position.
[0008] Further, the generating of the vector packet by the machine learning model in combination with the mapping table specifically includes: Based on the individual, move the consecutive scalar instructions and non-consecutive scalar instructions in the population to consecutive positions and then map them to the first vector instruction; Take all the first vector instructions and data processing instructions belonging to the same basic block as the second individual, iterate the second individual, and then after the iteration ends, pack all the mapping relationships in the second individual into sub-packets and make labels for the sub-packets; Pack all the sub-packets into a vector packet.
[0009] Further, after obtaining the first vector instruction, the method further includes reloading the register as a vector register, determining the legality of the individuals in the population at this time through the process chain, and deleting the individuals that do not meet the legality. The process chain is the process chain of data from initialization to use in the data structure, and the legality specifically means that the scalar instruction is not overwritten within the moving range after being moved.
[0010] Further, after packing all the sub-packets into a vector packet, the method further includes deleting the individuals in the vector packet whose fitness is less than the preset fitness. The fitness of the individual is specifically determined by the following formula: ; In the formula, is the fitness of the th individual, is the original running time of the program, is the running time of the th individual.
[0011] Further, the mapping table includes a first mapping and a second mapping. The first mapping is specifically the mapping between the scalar instructions in the target instruction set and their directly corresponding vector instructions, and the second mapping is specifically the mapping of the scalar instructions that do not have directly corresponding vector instructions in the target instruction set.
[0012] Further, after determining the mapping table, the method further includes: Determine the performance value of each mapping relationship in the mapping table before optimization for the vectorization of the training assembly file through the machine learning model; Delete the mapping relationships with performance values lower than the preset threshold to obtain the optimized mapping table.
[0013] On the other hand, the present invention also provides a program vectorization device, which includes: A construction module for constructing a mapping table from scalar instructions to vector instructions based on scalar instructions and vector instructions in a target instruction set; A compiler for processing a training program set to obtain a training assembly file, inputting the training assembly file into a pre-established machine learning model, generating a vector packet through the machine learning model in combination with the mapping table, and the compiler is also used to process a program to be converted to obtain a to-be-converted assembly file; A conversion module for converting each scalar instruction in the to-be-converted assembly file into a corresponding vector instruction according to the mapping relationship in the vector packet, and updating registers after inserting data processing instructions.
[0014] Compared with the prior art, the program vectorization method and device provided by the present invention first construct a mapping table from scalar instructions to vector instructions based on scalar instructions and vector instructions in a target instruction set; then, the compiler processes a training program set to obtain a training assembly file, inputs the training assembly file into a pre-established machine learning model, and generates a vector packet through the machine learning model in combination with the mapping table; finally, the compiler processes a program to be converted to obtain a to-be-converted assembly file, converts each scalar instruction in the to-be-converted assembly file into a corresponding vector instruction according to the mapping relationship in the vector packet, and updates registers after inserting data processing instructions, which can improve the adaptability of compiler vectorization, so that the compiler can effectively vectorize a program when facing a new target instruction set. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It shows a schematic flowchart of the program vectorization method provided by the embodiments of the present specification; Figure 2 It shows a schematic structural diagram of the program vectorization device provided by the embodiments of the present specification. Detailed Embodiments
[0017] To enable those of ordinary skill in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0018] As Figure 1 shown is a schematic flowchart of the program vectorization method provided in the embodiments of this specification. Although this specification provides the method operation steps or device structures shown in the following embodiments or drawings, based on routine or without creative efforts, more or fewer operation steps or module units may be included in the method or device. In steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments of this specification or drawings. When the described method or module structure is applied to actual devices, servers or terminal products, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or drawings (for example, in an environment with parallel processors or multi-threaded processing, or even including an implementation environment of distributed processing and server clusters).
[0019] The program vectorization method provided in the embodiments of this specification can be applied to terminal devices such as clients and servers. As Figure 1 shown, the method specifically includes the following steps: Step S101, construct a mapping table from scalar instructions to vector instructions based on the scalar instructions and vector instructions in the target instruction set; Specifically, the mapping table includes a first mapping and a second mapping. The first mapping is specifically the mapping between scalar instructions in the target instruction set and their directly corresponding vector instructions. The second mapping is specifically the mapping of scalar instructions in the target instruction set that do not have directly corresponding vector instructions. That is to say, in this step, for any target instruction set that supports vector instructions, this step constructs the corresponding instruction mapping relationship according to the correspondence between scalar instructions and vector instructions supported by the architecture. The mapping relationship is the correspondence from scalar instructions to vector instructions. This mapping relationship is divided into a simple mapping, that is, the first mapping, and a complex mapping, that is, the second mapping. The simple mapping directly corresponds to the relative vector instruction by the scalar instruction. The simple mapping is only affected by the width of the machine-supported register. Different numbers corresponding to the same scalar instruction may correspond to different vector instructions. The complex mapping is the mapping of scalar instructions that do not have corresponding vector instructions. This mapping method uses multiple vector instructions to be spliced to generate the corresponding scalar instruction. At the same time, the complex mapping also includes the mapping from multiple scalar instructions with different operations to vector instructions. The mapping relationship from scalar instructions to vector instructions in the complex mapping can be incremental or decremental. The following Table 1 is a partial schematic diagram of the exemplary first mapping: Table 1 ; It should be noted that the creation of the mapping table can be manually created based on the instruction manual of the target instruction set. The instruction manual does not include the mapping table. The instruction manual only includes the instructions supported by the architecture. The creation process of this table is constructed according to the vector and scalar instructions in the instruction table. For example, an addition instruction corresponds to a vector addition, but the registers used by the vector may be different. Therefore, a scalar addition may correspond to a 4-bit, 8-bit, or 16-bit addition. It can also be used as training data for the machine learning model through the instruction manual and the mapping table created manually, so as to generate the mapping table according to the instruction manual through the trained machine learning module.
[0020] In addition, after determining the mapping table, the method further includes determining, by the machine learning model, the performance value of each mapping relationship in the mapping table before optimization for the vectorization of the training assembly file; and then deleting the mapping relationships with performance values lower than the preset threshold to obtain the optimized mapping table.
[0021] Step S102: Process the training program set through a compiler to obtain a training assembly file, and input the training assembly file into a pre-established machine learning model, and generate a vector packet through the machine learning model in combination with the mapping table.
[0022] Specifically, the training program set is a set of source programs, that is, the original program. The corresponding training assembly file is generated through the compiler, and the machine learning model is trained through the training assembly file so that the machine learning model generates a vector packet.
[0023] In an embodiment of the present application, before generating a vector packet by combining the machine learning model with the mapping table, the method further includes: Encoding basic blocks in the training assembly file into a population, and encoding the mapping relationship from scalar instructions to vector instructions in the same basic block in the training assembly file into individuals, where the length of an individual is the number of vector instructions corresponding to the basic block to which the individual belongs, and the encoded bits of the individual represent the vector instructions after mapping at that position.
[0024] Specifically, this step encodes basic blocks in the training assembly file into a population, encodes the mapping relationship from scalar instructions to vector instructions into individuals, and at the same time constructs individuals for scalar instructions in the same basic block according to the mapping mode in the mapping table. After encoding, the individuals are independent of each other. The length of an individual is the number of vector instructions corresponding to the basic block to which the individual belongs, that is, the number of instructions that can be vectorized in the belonging basic block. An assembly program, that is, an assembly file, includes multiple basic blocks. A basic block contains multiple assembly instructions. The lengths of different individuals may be different due to different matching methods. The encoded bits of an individual represent the vector instructions after mapping at that position. Each mapping from scalar to vector corresponds to a different individual. Because different vectorization methods are used, different numbers of vectorized instructions may be generated in the same basic block.
[0025] In an embodiment of the present application, the generating of the vector packet by combining the machine learning model with the mapping table specifically includes: Moving continuous scalar instructions and discontinuous scalar instructions in the population to continuous positions based on the individuals and mapping them to first vector instructions; Taking all the first vector instructions and data processing instructions belonging to the same basic block as a second individual, iterating the second individual, and then after the iteration ends, packing all the mapping relationships in the second individual into sub-packets and making labels for the sub-packets; Packing all the sub-packets into a vector packet.
[0026] Specifically, this step moves discontinuous scalar instructions in the population to continuous positions according to the mapping method determined in the individuals, then takes all the first vector instructions and data processing instructions belonging to the same basic block as a second individual, iterates the second individual, and then after the iteration ends, packs all the mapping relationships in the second individual into sub-packets and makes labels for the sub-packets. After obtaining the second individual, the program containing the second individual is also run to obtain a fitness value according to a formula, and it is iterated based on the fitness value. The setting of the iteration can be flexibly set by those skilled in the art according to the actual situation.
[0027] Among them, the movement of scalar instructions follows the exchange legality of the program's definition-use chain; after scalar instructions are consecutive, the scalar instructions are mapped to vector instructions, and the registers are overloaded as vector registers; the overloading follows the definition-use chain legality in the population, and individuals that do not meet the replacement conditions are deleted. The definition-use chain (def-use chain) is also the process chain, and the process chain is the process chain of data from initialization to use in the data structure. In a process chain, data has a dependency relationship. For vectorization, vectorization requires that the process chain of the data to be processed by the vector is complete, and there is no overwriting of this data in the middle of the process chain. The movement here is to match the mapping in the mapping table, and the legality of the movement is determined by the legality of the process chain of the data being moved. The data being moved has no overwriting within the movement range, that is, the def in the process chain, which is the initialization of the data, is not overwritten. The specific legality is that the scalar instruction is not overwritten within the movement range after being moved.
[0028] This step inserts additional data processing instructions in the basic block to support vector instructions.
[0029] As mentioned above, the label includes the location of the instruction, the starting scalar instruction, the scalar instruction to be exchanged, the vector instruction width, and the vector instruction after matching. This label can be used for the iteration of the machine learning model. Each sub-packet is mutually exclusive and irrelevant, and there are no two sub-packets that contain the same instructions.
[0030] In the embodiment of the present application, after all sub-packets are packed into a vector packet, the method further includes deleting individuals in the vector packet whose fitness is less than the preset fitness. The fitness of the individual is specifically determined by the following formula: ; In the formula, is the fitness of the th individual, is the original running time of the program, is the th individual running time.
[0031] Step S103: Process the program to be converted through the compiler to obtain an assembly file to be converted, and convert each scalar instruction in the assembly file to be converted into a corresponding vector instruction according to the mapping relationship in the vector packet, and update the register after inserting the data processing instruction.
[0032] Specifically, scalar instructions in the assembly file to be converted are matched according to the mapping relationships in the vector packet. When there are multiple sets of corresponding relationships, the mapping relationship with the highest fitness is selected. Fitness is the ratio of the execution time of the program after being mapped using this mapping relationship to the execution time of the original program. A value less than 1 indicates acceleration, and a value greater than 1 indicates deceleration. Then, the matched scalar instructions are transformed into vector instructions; and additional data processing instructions are inserted in the basic block to support the vector instructions, and the registers are updated.
[0033] Based on the above program vectorization method, one or more embodiments of this specification also provide a program vectorization platform and terminal. The platform or terminal may include devices, software, modules, plugins, servers, clients, etc. that use the method described in the embodiments of this specification, combined with the necessary implementation hardware devices. Based on the same innovative concept, the systems in one or more embodiments provided by the embodiments of this specification are as described in the following embodiments. Since the implementation solutions of the systems for solving problems are similar to the methods, the implementation of the specific systems in the embodiments of this specification can refer to the implementation of the foregoing methods, and the repeated parts will not be elaborated. The terms "unit" or "module" used hereinafter may be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware and software combined implementations are also possible and contemplated.
[0034] Specifically, Figure 2 is a schematic diagram of the module structure of an embodiment of the program vectorization device provided in this specification. As Figure 2 shown, the program vectorization device provided in this specification includes: A construction module 201, configured to construct a mapping table from scalar instructions to vector instructions based on scalar instructions and vector instructions in a target instruction set; A compiler 202, configured to process a training program set to obtain a training assembly file, input the training assembly file into a pre-established machine learning model, generate a vector packet through the machine learning model in combination with the mapping table, and the compiler is further configured to process a program to be converted to obtain an assembly file to be converted; A conversion module 203, configured to convert each scalar instruction in the assembly file to be converted into a corresponding vector instruction according to the mapping relationship in the vector packet, and update the register after inserting data processing instructions.
[0035] It should be noted that the above system may also include other implementation manners according to the description of the corresponding method embodiments. The specific implementation manners may refer to the description of the corresponding method embodiments above, and will not be elaborated here one by one.
[0036] An embodiment of the present application also provides an electronic device, including: A processor; A memory for storing the executable instructions of the processor; The processor is configured to execute the method provided in the above embodiment.
[0037] The electronic device provided by the embodiment of the present application stores the executable instructions of the processor through a memory. When the processor executes the executable instructions, it can first construct a mapping table from scalar instructions to vector instructions based on the scalar instructions and vector instructions in the target instruction set; then process the training program set through a compiler to obtain a training assembly file, and input the training assembly file into a pre-established machine learning model. The machine learning model generates a vector packet in combination with the mapping table; finally, the compiler processes the program to be converted to obtain a converted assembly file, and converts each scalar instruction in the converted assembly file into a corresponding vector instruction according to the mapping relationship in the vector packet, and updates the register after inserting data processing instructions, which can improve the adaptability of the compiler vectorization, so that the compiler can effectively vectorize the program when facing a new target instruction set.
[0038] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0039] The method or device described in the above embodiments provided by this specification can implement business logic through a computer program and record it on a storage medium. The storage medium can be read and executed by a computer to achieve the effects of the solutions described in the embodiments of this specification, such as: Construct a mapping table from scalar instructions to vector instructions based on the scalar instructions and vector instructions in the target instruction set; Process the training program set through a compiler to obtain a training assembly file, and input the training assembly file into a pre-established machine learning model. The machine learning model generates a vector packet in combination with the mapping table; Process the program to be converted through the compiler to obtain a converted assembly file, and convert each scalar instruction in the converted assembly file into a corresponding vector instruction according to the mapping relationship in the vector packet, and update the register after inserting data processing instructions.
[0040] The storage medium may include a physical device for storing information, usually by digitizing the information and then storing it in a medium using electrical, magnetic, or optical means. The storage medium may include: devices that store information using electrical energy, such as various memories, such as RAM, ROM, etc.; devices that store information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memories, magnetic bubble memories, USB flash drives; devices that store information using optical means, such as CDs or DVDs. Of course, there are also other types of readable storage media, such as quantum memories, graphene memories, and so on.
[0041] The embodiments of this specification are not limited to those that must conform to industry communication standards, standard computer resource data update and data storage rules, or the situations described in one or more embodiments of this specification. Some industry standards or implementation methods described using custom methods or embodiments can also achieve the same, equivalent, or similar implementation effects as the above embodiments, or implementation effects that can be predicted after deformation. The embodiments obtained by applying these modified or deformed data acquisition, storage, judgment, processing methods, etc. still fall within the scope of the optional implementation methods of the embodiments of this specification.
[0042] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium that stores computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the controller can be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0043] The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or plugins can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0044] These computer program instructions can also be loaded onto a computer or other programmable resource data update device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps for the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 the steps of the functions specified in one block or multiple blocks.
[0045] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. The relevant parts can refer to the description of the method embodiment. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0046] Those of ordinary skill in the art will realize that the embodiments described here are to help the reader understand the principles of the present invention. It should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A program vectorization method, characterized in that: The method comprises: Building a mapping table from scalar instructions to vector instructions based on scalar instructions and vector instructions in the target instruction set; Processing the training program set by a compiler to obtain a training assembly file, and inputting the training assembly file into a pre-established machine learning model, and generating a vector packet by combining the machine learning model with a mapping table; The compiler processes the program to be converted to obtain the assembly file to be converted, and converts each scalar instruction in the assembly file to be converted into a corresponding vector instruction according to the mapping relationship in the vector package, and updates the register after inserting the data processing instruction.
2. The program vectorization method according to claim 1, characterized in that: Before generating the vector packet by combining the machine learning model with the mapping table, the method further includes: The basic blocks in the training assembly file are encoded as populations, and based on the mapping table, the mapping relationship from scalar instructions to vector instructions in the same basic block in the training assembly file is encoded as individuals, the length of the individual is the number of vector instructions corresponding to the basic block to which the individual belongs, and the encoding bits of the individual represent the vector instructions after mapping at that position.
3. The program vectorization method according to claim 2, characterized in that: The generating of the vector package by combining the machine learning model with the mapping table specifically includes: Based on the individuals, continuous scalar instructions and non-continuous scalar instructions in the population are moved to continuous positions and then mapped into first vector instructions; All first vector instructions and data processing instructions belonging to the same basic block are taken as a second entity, and the second entity is iterated, and then after the iteration, all mapping relationships in the second entity are packaged into sub-packages, and labels are generated for the sub-packages; Package all subpackages into a vector package.
4. The program vectorization method according to claim 3, characterized in that: The method also includes, after obtaining the first vector instruction, overloading the register as a vector register, and determining the legitimacy of the individuals in the population at this time through a process chain, and deleting the individuals that do not meet the legitimacy. The process chain is a process chain from initialization to use of data in a data structure, and the legitimacy specifically refers to that the scalar instruction is not overwritten within the moving range after being moved.
5. The program vectorization method according to claim 3, characterized in that: After all sub-packets are packaged into a vector pack, the method further includes deleting individuals in the vector pack whose fitness is less than a preset fitness, and the fitness of the individuals is specifically determined by the following formula: ; In the formula, For the The fitness of an individual, is the original running time of the program, For the Individual running time.
6. The program vectorization method according to claim 1, characterized in that: The mapping table includes a first mapping and a second mapping, wherein the first mapping is specifically a mapping between scalar instructions in a target instruction set and vector instructions directly corresponding thereto, and the second mapping is specifically a mapping of scalar instructions that do not have a directly corresponding vector instruction in the target instruction set.
7. The program vectorization method according to claim 6, characterized in that: After determining the mapping table, the method further includes: Determine the performance value of each mapping relationship in the mapping table before optimization for vectorization of the training assembly file through a machine learning model; The mapping relationships whose performance values are lower than a preset threshold are deleted to obtain the optimized mapping table.
8. A program vectorization device, characterized in that: The device comprises: A construction module, for constructing a mapping table of scalar instructions to vector instructions based on scalar instructions and vector instructions in a target instruction set; A compiler, used to process the training program set to obtain a training assembly file, and input the training assembly file into a pre-established machine learning model, and generate a vector packet through the machine learning model combined with a mapping table, and the compiler is also used to process the program to be converted to obtain an assembly file to be converted; The conversion module is used to convert each scalar instruction in the assembly file to be converted into a corresponding vector instruction according to the mapping relationship in the vector packet, and to update the register after inserting the data processing instruction.