Method, apparatus, device and storage medium for processing real number vector operations
Through hybrid programming, the high-level language is used for smaller real vectors and the low-level language is used for operation for larger real vectors, which solves the problem that high-bit register storage space is not fully utilized in the existing technology and achieves the improvement of processor performance.
Patent Information
- Application Number
- CN202510361136.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In the prior art, the four-character operation of real vectors and constants cannot fully utilize the storage space of high-bit registers, resulting in wasted processor resources.
Adopting hybrid programming methods, small real vectors are used to operate in high-level languages, and large real vectors are used to operate in low-level languages. The instructions in low-level languages are more streamlined, and hardware resources are directly accessed to realize parallel calculations of multiple real elements.
Through mixed programming methods, the computing performance of real vectors and real constants of different scales is improved, and the advantages of registers processing data in parallel are fully utilized to achieve a significant improvement in processor performance.
Smart Images

Figure CN119883373B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly, to a method, apparatus, device, and storage medium for processing real number vector operations. Background Art
[0002] The four arithmetic operations (i.e., addition, subtraction, multiplication, and division) of real number vectors and constants are the most commonly used vector operations in various fields. After performing the four arithmetic operations on each element of the real number vector and the constant, the operation results of each element can be used as the operation result of the real number vector.
[0003] In the prior art, when performing the four arithmetic operations on a real number vector and a constant, it is necessary to use the general registers of the Central Processing Unit (CPU) for serialized operations. For a processor with high-level registers, this method cannot fully utilize the storage space of the high-level registers, resulting in a waste of processor resources. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and storage medium for processing real number vector operations to solve the problem that the prior art cannot fully utilize the storage space of high-level registers, resulting in a waste of processor resources.
[0005] To achieve the above object, the technical solution adopted in this application is as follows:
[0006] In a first aspect, this application provides a method for processing real number vector operations for performing operations on a real number vector and a real number constant. The method includes:
[0007] Obtain the vector length of the real number vector to be operated and a preset length threshold;
[0008] If the vector length is greater than the preset length threshold, read the real number vector and the real number constant from the memory based on the first target instruction set of the first programming language, and perform operations on the real number vector and the real number constant to obtain an operation result;
[0009] If the vector length is less than or equal to the preset length threshold, read the real number vector and the real number constant from the memory based on the second target instruction set of the second programming language, and perform operations on the real number vector and the real number constant to obtain an operation result, where the level of the first programming language is lower than the level of the second programming language.
[0010] In a second aspect, this application provides a device for processing real number vector operations. The device includes:
[0011] An acquisition module, configured to acquire the vector length of a real number vector to be operated and a preset length threshold;
[0012] A first operation module, configured to, if the vector length is greater than the preset length threshold, read a real number vector and a real number constant from a memory based on a first target instruction set of a first programming language, and perform an operation on the real number vector and the real number constant to obtain an operation result;
[0013] A second operation module, configured to, if the vector length is less than or equal to the preset length threshold, read a real number vector and a real number constant from a memory based on a second target instruction set of a second programming language, and perform an operation on the real number vector and the real number constant to obtain an operation result, where the level of the first programming language is lower than the level of the second programming language.
[0014] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor, a storage medium, and a bus, where the storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of a real number vector operation processing method according to any one of the first aspect.
[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it performs the steps of a real number vector operation processing method according to any one of the first aspect.
[0016] The beneficial effects of the present application are as follows: By adopting a hybrid programming method, the operation of real number vectors with a smaller scale is implemented using a high-level language, and the operation of real number vectors with a larger scale is implemented using a low-level language, so that the operation performance of real number vectors and real number constants with different scales can be maximally improved. Among them, the low-level language can directly access and process hardware resources, the instructions are more concise, and the generated binary code occupies less memory resources. Therefore, the storage resources of the computer system can be better utilized, thereby effectively improving the calculation performance of the algorithm in the scenario of large-scale operations. And, the method of the present application can implement parallel calculation of multiple real number elements, thereby making full use of the advantage of parallel processing data by existing registers and achieving a significant improvement in the performance of the processor.
[0017] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and cooperates with the attached drawings to make detailed descriptions as follows. Description of the Drawings
[0018] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0019] Figure 1 It shows a flowchart of a real number vector operation processing method provided by an embodiment of the present application;
[0020] Figure 2 It shows a flowchart of performing operations on real number vectors and real number constants based on a first target instruction set of a first programming language;
[0021] Figure 3 It shows a flowchart of reading real number elements provided by an embodiment of the present application;
[0022] Figure 4 It shows another flowchart of reading real number elements provided by an embodiment of the present application;
[0023] Figure 5 It shows a flowchart of storing operation results provided by an embodiment of the present application;
[0024] Figure 6 It shows another flowchart of storing operation results provided by an embodiment of the present application;
[0025] Figure 7 It shows an overall flowchart of performing real number vector operations based on a first target instruction set provided by an embodiment of the present application;
[0026] Figure 8 It shows a flowchart of performing operations on real number vectors and real number constants based on a second target instruction set of a second programming language provided by an embodiment of the present application;
[0027] Figure 9 It shows an overall flowchart of performing real number vector operations based on a second target instruction set provided by an embodiment of the present application;
[0028] Figure 10 It shows a schematic structural diagram of a real number vector operation processing device provided by an embodiment of the present application;
[0029] Figure 11 It shows a schematic structural diagram of an electronic device 110 provided by an embodiment of the present application. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Usually, the components of the embodiments of this application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents the selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of this application.
[0031] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude adding other features.
[0032] In the prior art, when performing the four arithmetic operations on a real number vector and a constant, it is necessary to use the general registers of the CPU for serialized operations. For example, for a real number vector of length N performing the four arithmetic operations of addition, subtraction, multiplication, and division with a constant a, specifically including the following steps:
[0033] Step S1: Fetch the first element of the vector from memory and place it in register D0;
[0034]
[0035] Step S2: Store the constant alpha in register D1;
[0036] Step S3: Perform an addition (subtraction / multiplication / division) operation on registers D0 and D1, and save the operation result to register D2, and store the calculation result in D2 to the memory associated with the first element of the vector ;
[0037] Step S4: Repeat Step S1 and Step S3 until all N elements of vector A that need to perform the four arithmetic operations are completed.
[0038] In the above calculation process, when performing the four arithmetic operations on a real number vector and a constant, multiple loops are required. In each loop process, only one element of the input vector is subjected to the four arithmetic operations and the operation result is stored in the address memory of the output vector.
[0039] However, the current register capacity is generally large. Taking the registers provided by the ARM (Advanced RISC Machine) Cortex-A series processors as an example, the length of its registers can reach up to 128 bits at most, that is, it can perform calculations on 128-bit data in parallel each time. Then, for the complex convolution four arithmetic operations of only one element calculated each time, there will be a large number of idle positions remaining in the register. Therefore, for a processor with high-order registers, this method cannot fully utilize the storage space of high-order registers, and thus will also cause waste of processor resources.
[0040] Based on the above problems, this application proposes a real number vector operation processing method. By performing four arithmetic operations on real number vectors of different dimensions using different programming methods, parallel calculation of elements can be achieved, fully utilizing the advantage of parallel processing of data with a 128-bit register width, and improving the performance of the processor.
[0041] Figure 1 It is a schematic flowchart of the real number vector operation processing method provided by the embodiment of this application. The execution subject of this method can be a computer device with computing and processing capabilities, and a processor with a high-order memory is deployed on this computer device. Exemplarily, this processor can be an ARM Cortex-A series processor. As Figure 1 shown, this method includes:
[0042] S101. Obtain the vector length of the real number vector to be operated and a preset length threshold.
[0043] Among them, the real number vector includes multiple elements, and the vector length can be the number of elements in the real number vector. Exemplarily, a real number vector with a length of N can be expressed as .
[0044] The preset length threshold can be an integer greater than 1, and the specific value is not limited here. If the number of elements of the real number vector is greater than the preset length threshold, it means that the real number vector is of a larger dimension. At this time, the instruction set corresponding to a lower-level programming language can be used for operation, such as based on assembly language. If the number of elements of the real number vector is less than the preset length threshold, it means that the real number vector is of a smaller dimension. At this time, the instruction set corresponding to a higher-level programming language can be used for operation, such as based on C language.
[0045] S102. If the vector length is greater than the preset length threshold, based on the first target instruction set of the first programming language, read the real number vector and the real number constant from the memory, and perform operations on the real number vector and the real number constant to obtain an operation result.
[0046] Optionally, when the vector length of the real number vector is greater than a preset length threshold, it indicates that the dimension of the real number vector is large. At this time, in order to improve the operation performance of the processor, the first target instruction set of the first programming language can be used for operation processing.
[0047] Among them, the first programming language can be a low-level language, such as assembly language, machine language, or symbolic language, etc. Taking assembly language as an example, assembly code for the operation of the real number vector is written in assembly language, and then the assembly code is converted into machine code and an executable file is generated through a compiler. The low-level language can directly access and process hardware resources, the instructions are more concise, and the generated binary code occupies less memory. Therefore, it can make better use of the storage resources of the computer system and effectively improve the computing performance of the algorithm.
[0048] The first target instruction set includes multiple instructions. Data can be read from memory and operations can be performed by calling the instructions in the first target instruction set of the first programming language. Taking the first programming language as assembly language as an example, the first target instruction set includes multiple assembly instructions. The C language functions provided by the NEON intrinsics instruction set can be directly mapped to NEON instructions, and the compiler will convert them into corresponding assembly instructions. The assembly instructions in the first target instruction set in this application can be the assembly instructions after NEON instruction conversion, such as the fadd instruction for adding the data in two registers, the prfm pldllkeep instruction for prefetching the data in memory to the cache, the ldp instruction for storing the data in memory into a register, and the stp instruction for storing the data in the register into memory, etc.
[0049] In this application, different instructions in the first target instruction set can be used to perform operations on the real number vector and the real number constant, including four arithmetic operations such as addition, subtraction, multiplication, and division. Taking the addition operation of the real number vector and the real number constant as an example, it can be to add each element of the real number vector to the real number constant respectively, and the added result is used as the operation result.
[0050] S103: If the vector length is less than or equal to the preset length threshold, then based on the second target instruction set of the second programming language, read the real number vector and the real number constant from memory, and perform operations on the real number vector and the real number constant to obtain an operation result, where the level of the first programming language is lower than the level of the second programming language.
[0051] Among them, the second programming language can be a high-level language, such as C language, C++ language, Java language, and Python language, etc.
[0052] Optionally, when the vector length of the real number vector is less than the preset length threshold, it indicates that the dimension of the real number vector is small. At this time, in order to improve the operation performance of the processor, the second target instruction set of the second programming language can be used for operation processing.
[0053] It should be noted that when running an executable file, both high-level languages and low-level languages will call the compiler for optimization. For large-scale data, writing code in a low-level language and then optimizing it through the compiler has higher performance. For small-scale data, optimizing it on its own basis and then through the compiler has higher performance. Taking the C language as an example, when the scale of the data is small, the code written in the C language can effectively improve the computational performance of the algorithm after optimizing itself and then through the compiler again. Therefore, in this application, adopting a hybrid programming method for large-scale data and small-scale data can effectively improve the overall computational performance of the algorithm.
[0054] The second target instruction set includes multiple instructions. By calling the instructions in the second target instruction set of the second programming language, data can be read from memory and operations can be performed. Taking the second programming language as the assembly language as an example, the first target instruction set includes instructions in the NEON intrinsics instruction set, such as the vaddq_f32 instruction for adding data in two registers, the vld1q_f32 instruction for storing data in memory into a register, and the vst1q_f32 instruction for storing data in a register into memory, etc.
[0055] In this application, different instructions in the second target instruction set can be used to perform operations on real number vectors and real number constants, including four arithmetic operations such as addition, subtraction, multiplication, and division. Taking the addition operation of a real number vector and a real number constant as an example, it can be to add each element of the real number vector to the real number constant respectively, and take the added result as the operation result.
[0056] In the embodiment of this application, the vector length of the real number vector to be operated and a preset length threshold are obtained. If the vector length is greater than the preset length threshold, based on the first target instruction set of the first programming language, the real number vector and the real number constant are read from memory, and operations are performed on the real number vector and the real number constant to obtain an operation result. If the vector length is less than or equal to the preset length threshold, based on the second target instruction set of the second programming language, the real number vector and the real number constant are read from memory, and operations are performed on the real number vector and the real number constant to obtain an operation result, where the level of the first programming language is lower than the level of the second programming language.
[0057] By adopting a mixed programming approach, the operation of real number vectors with a relatively small scale is implemented in a high-level language, and the operation of real number vectors with a relatively large scale is implemented in a low-level language, so that the operation performance of real number vectors and real number constants with different scales can be maximally improved. Among them, the low-level language can directly access and process hardware resources, the instructions are more concise, and the memory resources occupied by the generated binary code are smaller. Therefore, the storage resources of the computer system can be better utilized, thereby effectively improving the computing performance of the algorithm in the scenario of large-scale operations. Moreover, the method of the present application can implement parallel computing of multiple real number elements, so as to fully utilize the advantage of parallel data processing of existing registers and achieve a significant improvement in the performance of the processor.
[0058] The following is a further description of reading a real number vector and a real number constant from memory and performing operations on the real number vector and the real number constant to obtain an operation result based on the first target instruction set of the first programming language, as Figure 2 shown, the above step S102 includes:
[0059] S201. Read a real number constant from memory, copy the real number constant based on the first target instruction set to obtain multiple real number constants, and store the multiple copied real number constants in a constant register.
[0060] Optionally, instructions in the first target instruction set can be used to read a real number constant from memory, and instructions in the first target instruction set can be used to copy the real number constant multiple times and store it in a constant register.
[0061] S202. Read 2M real number elements in the real number vector based on the first target instruction set, and store each real number element in the first register, the second register, the third register, and the fourth register in sequence, where M is an integer greater than 0.
[0062] In the first implementation manner, instructions in the first target instruction set can be used to directly read 2M real number elements from memory, and the read real number elements are stored in the first register, the second register, the third register, and the fourth register in the order of reading, and the number of real number elements stored in each register is the same.
[0063] In the second implementation manner, instructions in the first target instruction set can be used to read 2M real number elements from memory, and instructions in the first target instruction set can be used to prefetch N real number elements from memory, and store the prefetched real number elements in a cache, so that the real number elements can be directly read from the cache when reading real number elements next time.
[0064] In the third implementation method, when reading real number elements each time, instructions in the first target instruction set can be used to read M real number elements from the memory, and instructions in the first target instruction set can be used to prefetch M real number elements from the memory into the cache, that is, 2M real number elements are read from the memory each time, and the first M elements are sequentially stored in the first register, the second register, the third register, and the fourth register in the order of reading.
[0065] Exemplarily, assuming that the real number vector is of the single-precision float type, taking the 128-bit register processing process of the ARM Cortex-A series processor as an example, 16 real number elements can be read each time, and 16 real number elements can be prefetched. The first 16 real number elements read are sequentially stored in the registers in the order of reading, and the prefetched real number elements are stored in the cache. 4 elements are stored in each register, and the real number constant can be copied 4 times and stored in register D0.
[0066] S203. Respectively perform operations on the real number constant in the constant register and the real number elements in the first register, the second register, the third register, and the fourth register based on the first target instruction set, and store the operation results in the associated memory of the real number vector.
[0067] Repeat steps S202 - S203 until the operations on all real number elements in the real number vector are completed, and use the operation results in the associated memory as the operation results of the real number vector and the real number constant.
[0068] Optionally, using the instructions in the first target instruction set to perform operations on the real number constant in the constant register and the real number elements in the first register, the second register, the third register, and the fourth register respectively, the operation results of the real number elements and the real number constant in each register can be obtained. Next, the instructions in the first target instruction set can be used to store the operation results in the associated memory of the real number vector.
[0069] Among them, when storing the operation results in the associated memory of the real number vector, the operation results of each real number element and the real number constant can be stored in the associated memory of the real number vector in the order of each real number element.
[0070] It should be noted that when the number of elements in the real number vector is large, the above steps S202 - S203 may need to be executed multiple times. After storing the operation result of the last real number element of the real number vector in the associated memory of the real number vector, the operation results in the associated memory can be used as the operation results of the real number vector and the real number constant.
[0071] It should be understood that, in addition to the above steps S201 - S203, the operation on the real - number vector and the real - number constant based on the first target instruction set of the first programming language may also include other implementation manners. In one possible implementation manner, the real - number constant can be read from the memory based on the first target instruction set, and the real - number elements in the real - number vector are read sequentially. The operation is performed on the real - number constant and the currently read real - number element to obtain an operation result, and the operation result is sequentially stored into the associated memory of the real - number vector in the order of the read real - number elements.
[0072] The following is a further description of the second implementation manner of reading 2M real - number elements in the real - number vector based on the first target instruction set, as Figure 3 shown, the above step S202 includes:
[0073] S301. If it is the first - round loop currently, then based on the first target instruction set, read 2M real - number elements in the real - number vector from the memory, and pre - fetch the first N real - number elements among the unread elements into the cache, where N is an integer greater than 0.
[0074] Optionally, when reading real - number elements from the memory for the first time, the instructions of the first target instruction set can be used to sequentially read 2M real - number elements of the real - number vector from the memory from front to back, and starting from the (2M + 1) - th element, use the instructions in the first target instruction set to pre - fetch N consecutive real - number elements into the cache.
[0075] Among them, N is an integer greater than 0, and N can also be set as an integer multiple of 2M, which can reduce the number of times of reading real - number elements from the memory. Subsequently, directly read real - number elements from the cache, further improving the performance during the algorithm operation.
[0076] S302. If it is a non - first - round loop currently and the number of unread elements in the cache is greater than or equal to 2M, then based on the first target instruction set, read the first 2M real - number elements among the unread elements from the cache.
[0077] Optionally, after reading real - number elements and pre - fetching real - number elements into the cache in the first round, real - number elements can be directly read from the cache during each subsequent loop process. Before reading real - number elements in the cache each time, the number of unread real - number elements in the cache can be compared with the number of elements to be read. If the number of elements to be read 2M is less than the number of unread real - number elements in the cache, then the instructions in the first target instruction set can be used to read the first 2M real - number elements among the unread elements from the cache.
[0078] S303. If it is a non - first - round loop currently and the number of unread elements in the cache is less than 2M, then based on the first target instruction set, read the first 2M real - number elements among the unread elements from the memory, and pre - fetch the first N real - number elements among the unread elements into the cache.
[0079] If the number of real elements read in the current loop is greater than the number of unread real elements in the cache, that is, the number of unread elements in the cache is less than 2M, then the first 2M real elements of the unread elements can be read from the memory using the instructions in the first target instruction set.
[0080] As a possible implementation, the number of pre-fetched real elements can be set to an integer multiple of 2M, the number of real elements read each time. In this way, it can be ensured that after an integer number of loops, the number of elements in the cache is exactly 0. When the number of elements in the cache is equal to 0, the first 2M elements of the unread elements can be read from the memory again, and the first N real elements of the unread elements can be pre-fetched into the cache.
[0081] In another possible implementation, the position of the currently read real element in the real vector can also be recorded. For example, the array subscript of the real element in the real vector is recorded. When the number of real elements in the cache is less than 2M, the unread elements corresponding to the array subscript can be determined in the memory based on the subscript of the next real element to be read, and 2M real elements can be read from the unread elements using the instructions in the first target instruction set. Starting from the next unread real element after the last read real element, N real elements can be pre-fetched from the memory into the cache using the instructions in the first target instruction set.
[0082] In the embodiments of the present application, by pre-fetching real elements from the memory into the cache after reading real elements from the memory, for the scenario of large-scale data operations, the number of times of reading data from the memory can be reduced, and the performance of the processor can be improved.
[0083] The above steps S301 - S303 are an implementation of reading real elements. In another implementation, M real elements of the real vector can also be read from the memory based on the first target instruction set, and M real elements of the unread elements can be pre-fetched from the memory at the same time. After the operations on the M real elements read from the memory are completed, M real elements can be read from the cache for operations. After the operations on the M real elements in the cache are completed, the steps of reading M real elements of the unread elements from the memory and pre-fetching M real elements of the unread elements from the memory are re-executed until all operations on real elements and real constants are completed.
[0084] The following is a further description of reading 2M real elements of the real vector from the memory based on the first target instruction set and storing each real element in the first register, the second register, the third register, and the fourth register in sequence, as Figure 4 shown, step S202 above includes:
[0085] S401. Call the first instruction in the first target instruction set to read the first M real number elements from 2M real number vectors, and store each read real number element in the first register and the second register.
[0086] Among them, the first target instruction set can be the assembly instruction set corresponding to the NEON intrinsics instruction set, and the first instruction can be the assembly instruction corresponding to the vld1q_f32 instruction in the NEON intrinsics instruction set, that is, the ldp instruction. The vld1q_f32 instruction is used to fetch data from memory and store the data in a register.
[0087] Optionally, when reading real number elements, the first M real number elements can be read sequentially from front to back according to the order of each real number element in the real number vector, and each real number element is stored in the first register and the second register according to the read order.
[0088] S402. Call the first instruction in the first target instruction set to read the last M real number elements from 2M real number vectors, and store each real number element in the third register and the fourth register.
[0089] After storing the first M real number elements in the register, the ldp instruction can be called again to sequentially read the last M real number elements from memory, and each real number element is stored in the third register and the fourth register according to the read order.
[0090] It should be noted that the above is an implementation method of reading and storing real number elements in a register. In the second implementation method, the first instruction can also be called to read 2M real number elements from memory, and the real number elements are stored in the first register, the second register, the third register, and the fourth register according to the read order.
[0091] In the third implementation method, the first instruction can be called to read the first M real number elements from memory, and at the same time, the prefetch instruction in the first target instruction set can be called to prefetch the last M real number elements from memory into the cache. At this time, in the above S402 step, the first instruction can be to read the last M real number elements from the cache and store each real number element in the third register and the fourth register. Among them, the prefetch instruction can be the prfm pldllkeep instruction in the NEON intrinsics instruction set.
[0092] Next, the steps of reading real number constants from memory, copying multiple real number constants based on the first target instruction set, and storing the multiple copied real number constants in the constant register will be described. The above S201 step includes:
[0093] Call the second instruction in the first target instruction set to copy the real constants, and sequentially store the multiple copied real constants into the constant registers.
[0094] Among them, the second instruction can be the assembly instruction corresponding to the vdupq_n_f32 instruction in the NEON intrinsics instruction set. The vdupq_n_f32 instruction is used to copy the input constants and store the copied constants into the constant registers.
[0095] Optionally, the number of copied real constants can be determined based on the number of registers storing real elements. If the number of registers storing real elements is 4, then the real constants can be copied 4 times and stored in the constant registers.
[0096] The following is a further description of performing operations on the real constants in the constant register and the real elements in the first register, the second register, the third register, and the fourth register respectively based on the first target instruction set, and storing the operation results in the associated memory of the real vector, as Figure 5 shown. The above step S203 includes:
[0097] S501. Call the third instruction in the first target instruction set to read the real elements from the first register, the second register, the third register, and the fourth register respectively, read the real constants from the constant register, perform operations on the real elements and the real constants, and store the operation results of each real element in the fifth register corresponding to the first register, the sixth register corresponding to the second register, the seventh register corresponding to the third register, and the eighth register corresponding to the fourth register respectively.
[0098] Among them, the third instruction can be the assembly instruction corresponding to the arithmetic instruction in the NEON intrinsics instruction set, such as the fadd instruction corresponding to the vaddq_f32 addition instruction.
[0099] Optionally, the third instruction can be called four times to calculate the real elements in the first register, the second register, the third register, and the fourth register respectively with the real constants in the constant register, obtain the operation results of the real elements and the real constants in each register, save the operation result of the first register in the fifth register, save the operation result of the second register in the sixth register, save the operation result of the third register in the seventh register, and save the operation result of the fourth register in the eighth register.
[0100] Exemplarily, assume that it is necessary to calculate the operation result of adding a real vector and a real constant. Then, the fadd instruction can be called four times to calculate the addition results of the real elements in the first register, the second register, the third register, and the fourth register with the real constant, and store the addition results of the first register, the second register, the third register, and the fourth register into the fifth register, the sixth register, the seventh register, and the eighth register respectively.
[0101] S502. Call the fourth instruction in the first target instruction set to store the operation results in the fifth register, the sixth register, the seventh register, and the eighth register into the associated memory of the real vector.
[0102] Among them, the fourth instruction can be the assembly instruction corresponding to the vst1q_f32 instruction in the NEON intrinsics instruction set, that is, the stp instruction. Among them, the vst1q_f32 instruction is used to store the data in the register into the memory.
[0103] When calling the fourth instruction to store the operation results into the associated memory of the real vector, the operation results can be sequentially stored into the associated memory of the real vector in the order of the real elements corresponding to the operation results in the real vector.
[0104] The following is a further description of the above-mentioned calling the fourth instruction in the first target instruction set to store the operation results in the fifth register, the sixth register, the seventh register, and the eighth register into the associated memory of the real vector, as Figure 6 shown. The above S502 step includes:
[0105] S601. Call the fourth instruction in the first target instruction set to store the operation results in the fifth register and the sixth register into the position indicated by the current pointing address of the associated memory, and offset the current pointing address backward by a preset number of bytes to obtain a new current pointing address.
[0106] The current pointing address is used to indicate the starting position of data storage. By calling the fourth instruction, the operation results in the fifth register and the sixth register can be stored into the position indicated by the current pointing address, and the current pointing address is offset backward by a preset number of bytes to obtain a new current pointing address.
[0107] S602. Call the fourth instruction in the first target instruction set to store the operation results in the seventh register and the eighth register into the position indicated by the new current pointing address, and offset the new current pointing address backward by a preset number of bytes to obtain the current pointing address for the next loop.
[0108] Among them, the preset number of bytes can be the total number of bytes of the operation results of the fifth register and the sixth register, and the new current pointing address can point to the next position in memory of the last operation result of the sixth register. Call the fourth instruction again, store the operation results in the seventh register and the eighth register into the associated memory, and continue to offset the current pointing address backward by the preset number of bytes. The new current pointing address can point to the last operation result of the eighth register.
[0109] Next, in combination with Figure 7 The process of operating on a real number vector and a real number constant based on the first programming language will be described. Among them, the data type of the real number vector is single-precision type.
[0110] Refer to Figure 7 , real number vector , since the vector length K of the real number vector is greater than the preset length threshold M, the real number vector and the real number constant can be operated on based on the first programming language.
[0111] First, call the ldp instruction to read the first 16 float real number elements a0, a1... a15 from the memory associated with the vector and store them into registers q0, q1, q2, and q3 respectively. Each register stores 4 elements. Among them, "ldp q0,q1,[%[ap]],0x20" means to read 32 bytes from the address pointed to by ap, that is, read 8 float data at a time and store them in registers q0 and q1. q0 stores the first four a0 - a3, q1 stores a4 - a7, and the address pointed to by ap is offset backward by 32 bytes. Call the prfmpldllkeep prefetch instruction to prefetch 1024 bytes of data from the address pointed to by ap into the cache so that the real number elements can be directly read from the cache when reading data next time.
[0112] Then, call the fadd instruction 4 times to add registers q0 - q3 to register q8 respectively. Among them, 4 constants b are stored in register q8, and the addition results are stored in registers q4 - q7. Each register stores 16 added float elements / data in sequence.
[0113] Next, call the stp instruction to store the 16 float elements on registers q4 - q7 into the memory addresses associated with the first 16 elements of vector R in sequence: among them, "stp q4,q5,[%[rp]],0x20" means to store 32 bytes to the memory pointed address, that is, store 8 float data to registers q4 and q5 at a time, and the memory pointed address is offset backward by 32 bytes.
[0114] Call the vaddq_f32 instruction 8 times to add registers D1 - D8 to register D0 respectively, and store the addition results in registers D1 - D8. Among them, each register stores 32 added elements in sequence. Call the vst1q_f32 instruction to sequentially store 32 float elements on registers D1 - D8 into the first 32 items of vector R. Among them, vector R is the associated memory of the real vector.
[0115] Repeat the above steps until all real elements of the real vector and the addition results of the real constants are stored in the associated memory. If there are still remainders, perform operations on the remaining real elements using the conventional real element and real constant addition operation method.
[0116] The following is a further description of the above - mentioned second - target instruction set based on the second programming language, which reads real vectors and real constants from memory and performs operations on real vectors and real constants to obtain operation results. As Figure 8 shown, the above step S203 includes:
[0117] S801: Read a real constant from memory, and based on the second - target instruction set, copy multiple real constants and store them in constant registers.
[0118] Optionally, instructions in the second - target instruction set can be used to read real constants from memory, and instructions in the second - target instruction set can be used to copy multiple real constants and store them in constant registers.
[0119] S802: Based on the second - target instruction set, starting from the memory - pointed address in memory, read nQ real elements of the real vector, sequentially store each real element in Q ninth registers, and modify the memory - pointed address to the address of the next real element after the last read real element. n is an integer greater than 0, and Q is an integer greater than 0.
[0120] In the first implementation, instructions in the second - target instruction set can be used to sequentially read n real elements from memory in the storage order of real elements, store the read n real elements in the first register, and then loop through the step of reading n real elements until all real elements of the real vector are stored in Q ninth registers. After each read of a real element, the current memory - pointed address can be offset to the next memory address of the last read real element.
[0121] In the second implementation, instructions in the second - target instruction set can be used to read nQ real elements from memory in the storage order of real elements, store the real elements in Q registers in the read order, and offset the current memory - pointed address to the next memory address of the last read real element.
[0122] S803. Perform operations on the real constants in the constant register and the real elements in each ninth register respectively based on the second target instruction set, and store the operation results in each ninth register.
[0123] Optionally, instructions in the second target instruction set can be used to copy multiple real constants and store them in the constant register, and instructions in the second target instruction set can be used to read real elements from the ninth register, read real constants from the constant register, and perform operations on the read real constants and real elements.
[0124] When storing the operation results, the operation results can be stored at the positions of the real elements in each ninth register in the order of the real elements.
[0125] S804. Store the operation results in each ninth register into the associated memory of the real vector based on the second target instruction set.
[0126] Repeat steps S802 - S804 until the operations on all real elements in the real vector are completed, and use the operation results in the associated memory as the operation results of the real vector and the real constant.
[0127] Optionally, the associated memory is used to store the operation results. When the number of elements in the real vector is relatively large, the operations on the real vector and the real constant may require multiple rounds. After each round ends, the operation results in the ninth register can be stored in the associated memory of the real vector. After the operations on all real elements in the real vector are completed and all the operation results are stored in the associated memory, the operation results in the associated memory can be used as the operation results of the real vector and the real constant.
[0128] Among them, when storing the operation results in the associated memory, the operation results can be sequentially stored in the associated memory in the order of the operation results in the ninth register, and the first operation result of the next round can be stored at the next memory position after the last operation result in the associated memory.
[0129] Optionally, the process of reading real constants from the memory, copying multiple real constants based on the second target instruction set, and storing them in the constant register includes:
[0130] Call the fifth instruction in the second target instruction set to copy the real constant, and sequentially store the multiple copied real constants in the constant register.
[0131] Among them, the second target instruction set may be the NEON intrinsics instruction set, and the fifth instruction may be the vdupq_n_f32 instruction in the NEON intrinsics instruction set. The vdupq_n_f32 instruction is used to copy multiple input real constants and store the multiple copied real constants in the constant register.
[0132] Exemplarily, the vdupq_n_f32 instruction can be used to copy the input constant b four times and store the four constants b in the constant register in sequence.
[0133] It should be noted that the number of copied real constants can be determined based on the number of ninth registers. In one possible implementation, the number of copied real constants can be the same as the number of ninth registers, and each real constant is respectively operated on with the real elements in a ninth register.
[0134] The following is a further description of the above process of reading nQ real elements of the real vector from the memory pointed address in the memory based on the second target instruction set and storing each real element in the Q ninth registers in sequence. The above S802 step includes:
[0135] Call the sixth instruction in the second target instruction set to read n real elements of the real vector from the memory and store the read real elements in the first ninth register.
[0136] Repeat the step of calling the sixth instruction in the second target instruction set to read n real elements of the real vector from the memory and store the real elements in the next ninth register until the nQ real elements of the real vector are respectively stored in the Q ninth registers.
[0137] Among them, the sixth instruction may be the vld1q_f32 instruction in the NEON intrinsics instruction set. Using the vld1q_f32 instruction, n real elements can be read from the memory and the read real elements are stored in the first ninth register in the order of reading.
[0138] Continue to call the vld1q_f32 instruction until nQ real elements are read from the real vector and the read nQ real elements are respectively stored in the Q ninth registers. When the real vector is of single-precision type, the value of n can be 4 and the value of Q can be 8.
[0139] The following is a further description of the above process of respectively operating on the real constants in the constant register and the real elements in each ninth register based on the second target instruction set and storing the operation results in each ninth register. The above S803 step includes:
[0140] Invoke the seventh instruction in the second target instruction set, read real number elements from each ninth register, read real number constants from the constant register, perform four arithmetic operations on the real number elements and the real number constants, and store the operation results of each real number element at the positions corresponding to the real number elements in each ninth register.
[0141] Optionally, the seventh instruction can be an arithmetic instruction in the NEON intrinsics instruction set. Taking the addition operation of a real number vector and a real number constant as an example, the seventh instruction can be the vaddq_f32 instruction.
[0142] Invoke the seventh instruction to read real number constants from the constant register, read real number elements from each ninth register, perform four arithmetic operations on the real number constants and the real number elements, and store the operation results at the positions where the real number elements are located in the ninth register. After completing the operations on all real number elements in the current ninth register, the operation results in the ninth register can be sequentially saved in the associated memory.
[0143] The process of storing the operation results in each ninth register into the associated memory of the real number vector based on the second target instruction set includes:
[0144] Invoke the eighth instruction in the second target instruction set to store the operation results in each ninth register into the associated memory of the real number vector.
[0145] Among them, the eighth instruction can be the vst1q_f32 instruction in the NEON intrinsics instruction set, which is used to take out the operation results in the ninth register and store the taken-out operation results into the associated memory.
[0146] Figure 9 It is a flowchart for performing operations on a real number vector and a real number constant based on a second programming language. Among them, the data type of the real number vector is single-precision type.
[0147] Refer to Figure 9 , the real number vector , since the vector length K of the real number vector is less than the preset length threshold M, operations on the real number vector and the real number constant can be performed based on the second programming language.
[0148] First, invoke the vld1q_f32 instruction to read from the vector Read the first four floating-point real elements a0, a1, a2, a3 from the associated memory and store them in register D1. Repeatedly call the vld1q_f32 instruction to store 28 elements from a4 - a31 into registers D2 - D8 respectively. Then use the vdupq_n_f32 instruction to copy the real constant b four times and store it in register D0. Call the vaddq_f32 instruction 8 times to add registers D1 - D8 to register D0 respectively, and store the addition results in registers D1 - D8. Among them, each register stores 32 added elements in sequence. Call the vst1q_f32 instruction to store 32 floating-point elements on registers D1 - D8 into the first 32 items of vector R in sequence, where vector R is used to indicate the position of the real vector in the associated memory.
[0149] After repeating the above steps m times, the addition results of n elements of vector and constant b can be stored in the associated memory. If there are remaining items, the remaining real elements are operated using the conventional real element and real constant addition operation method to obtain the operation result of the real vector and the real constant.
[0150] Based on the same inventive concept, an apparatus for processing real vector operations corresponding to the real vector operation processing method is further provided in an embodiment of the present application. Since the principle of solving problems by the apparatus in the embodiment of the present application is similar to the above real vector operation processing method in the embodiment of the present application, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be described again.
[0151] Figure 10 It is a module structure diagram of an apparatus for processing real vector operations provided in an embodiment of the present application. As Figure 10 shown, the apparatus includes:
[0152] An acquisition module 1001, configured to acquire the vector length of the real vector to be operated and a preset length threshold;
[0153] A first operation module 1002, configured to, if the vector length is greater than the preset length threshold, read the real vector and the real constant from the memory based on a first target instruction set of a first programming language, and perform an operation on the real vector and the real constant to obtain an operation result;
[0154] A second operation module 1003, configured to, if the vector length is less than or equal to the preset length threshold, read the real vector and the real constant from the memory based on a second target instruction set of a second programming language, and perform an operation on the real vector and the real constant to obtain an operation result, where the level of the first programming language is lower than the level of the second programming language.
[0155] As a possible implementation manner, the first operation module 1002 is specifically configured to:
[0156] A. Read real constants from memory, copy the real constants based on the first target instruction set to obtain multiple real constants, and store the multiple copied real constants in a constant register;
[0157] B. Read 2M real elements in a real vector based on the first target instruction set, and sequentially store each real element in a first register, a second register, a third register, and a fourth register, where M is an integer greater than 0;
[0158] C. Perform operations on the real constants in the constant register and the real elements in the first register, the second register, the third register, and the fourth register respectively based on the first target instruction set, and store the operation results in the associated memory of the real vector;
[0159] D. Repeat steps B - C until all real elements in the real vector are operated on, and use the operation results in the associated memory as the operation results of the real vector and the real constants.
[0160] As a possible implementation, the first operation module 1002 is specifically configured to:
[0161] If it is currently in the first round of the loop, read 2M real elements in the real vector from memory based on the first target instruction set, and prefetch the first N real elements among the unread elements into the cache, where N is an integer greater than 0;
[0162] If it is currently in a non - first - round loop and the number of unread elements in the cache is greater than or equal to 2M, read the first 2M unread elements from the cache based on the first target instruction set;
[0163] If it is currently in a non - first - round loop and the number of unread elements in the cache is less than 2M, read the first 2M unread elements from memory based on the first target instruction set, and prefetch the first N real elements among the unread elements into the cache.
[0164] As a possible implementation, the first operation module 1002 is specifically configured to:
[0165] Call the first instruction in the first target instruction set to read the first M real elements in 2M real vectors, and store each read real element in the first register and the second register;
[0166] Call the first instruction in the first target instruction set to read the last M real elements in 2M real vectors, and store each real element in the third register and the fourth register.
[0167] As a possible implementation, the first operation module 1002 is specifically configured to:
[0168] Invoke the second instruction in the first target instruction set to copy real constants and sequentially store the multiple copied real constants in the constant register.
[0169] As a possible implementation, the first operation module 1002 is specifically configured to:
[0170] Invoke the third instruction in the first target instruction set to separately read real elements from the first register, the second register, the third register, and the fourth register, read real constants from the constant register, perform operations on the real elements and the real constants, and respectively store the operation results of each real element in the fifth register corresponding to the first register, the sixth register corresponding to the second register, the seventh register corresponding to the third register, and the eighth register corresponding to the fourth register;
[0171] Invoke the fourth instruction in the first target instruction set to store the operation results in the fifth register, the sixth register, the seventh register, and the eighth register into the associated memory of the real vector.
[0172] As a possible implementation, the first operation module 1002 is specifically configured to:
[0173] Invoke the fourth instruction in the first target instruction set to store the operation results in the fifth register and the sixth register into the position indicated by the current pointing address of the associated memory, and offset the current pointing address backward by a preset number of bytes to obtain a new current pointing address;
[0174] Invoke the fourth instruction in the first target instruction set to store the operation results in the seventh register and the eighth register into the position indicated by the new current pointing address, and offset the new current pointing address backward by a preset number of bytes to obtain the current pointing address for the next loop.
[0175] As a possible implementation, the second operation module 1003 is specifically configured to:
[0176] A. Read real constants from the memory, copy multiple real constants based on the second target instruction set, and store them in the constant register;
[0177] B. Based on the second target instruction set, start from the memory pointing address in the memory, read nQ real elements of the real vector, sequentially store each real element in Q ninth registers, and modify the memory pointing address to the address of the next real element after the last read real element, where n is an integer greater than 0 and Q is an integer greater than 0;
[0178] C. Based on the second target instruction set, perform operations on the real constants in the constant register and the real elements in each ninth register respectively, and store the operation results in each ninth register;
[0179] D. Based on the second target instruction set, store the operation results in each ninth register into the associated memory of the real vector;
[0180] E. Repeat steps B - D until the operations on all real elements in the real vector are completed, and use the operation results in the associated memory as the operation results of the real vector and the real constant.
[0181] As a possible implementation, the second operation module 1003 is specifically configured to:
[0182] Call the fifth instruction in the second target instruction set to copy the real constant, and sequentially store the multiple copied real constants into the constant register.
[0183] As a possible implementation, the second operation module 1003 is specifically configured to:
[0184] Call the sixth instruction in the second target instruction set to read n real elements of the real vector from the memory, and store the read real elements into the first ninth register;
[0185] Repeat the step of calling the sixth instruction in the second target instruction set, read n real elements of the real vector from the memory, and store the real elements into the next ninth register until the nQ real elements of the real vector are respectively stored into Q ninth registers.
[0186] As a possible implementation, the second operation module 1003 is specifically configured to:
[0187] Call the seventh instruction in the second target instruction set to read the real elements from each ninth register, read the real constant from the constant register, perform four arithmetic operations on the real elements and the real constant, and store the operation results of each real element at the position corresponding to the real element in each ninth register.
[0188] As a possible implementation, storing the operation results in each ninth register into the associated memory of the real vector based on the second target instruction set includes:
[0189] Call the eighth instruction in the second target instruction set to store the operation results in each ninth register into the associated memory of the real vector.
[0190] The embodiment of the present application also provides a computer device 110, such as Figure 11As shown in the figure, it is a schematic structural diagram of a computer device 110 provided by an embodiment of the present application, including: a processor 1101, a memory 1102, and optionally, a bus 1103 may also be included. The memory 1102 stores machine-readable instructions executable by the processor 1101 (for example, Figure 10 execution instructions corresponding to the acquisition module 1001, the first operation module 1002, and the second operation module 1003 in the device in
[0191] When the computer device 110 runs, communication between the processor 1101 and the memory 1102 is carried out through the bus 1103. When the machine-readable instructions are executed by the processor 1101, the steps of the real number vector operation processing method in the above method embodiment are executed.
[0192] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the method embodiments, which will not be elaborated in this application. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other forms.
[0193] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0194] The above are only specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.
Claims
1. A method for processing real number vector operations, characterized in that, For performing operations on a real vector and a real constant, the method includes: Obtaining the vector length of the real vector to be operated on and a preset length threshold; If the vector length is greater than the preset length threshold, based on the first target instruction set of the first programming language, read the real constant from memory, copy and store the real constant in a constant register, and read the real vector and batch-store multiple real elements in the real vector into multiple registers, and perform operations on the real constant stored in the constant register and the real elements stored in the multiple registers respectively based on the first target instruction set to obtain an operation result, where the first target instruction set includes multiple instructions written in the first programming language; If the vector length is less than or equal to the preset length threshold, based on the second target instruction set of the second programming language, read the real constant from memory, copy and store the real constant in a constant register, and read the real vector and store multiple real elements in the real vector into multiple registers, and perform operations on the real constant stored in the constant register and the real elements stored in the multiple registers respectively based on the second target instruction set to obtain an operation result, where the level of the first programming language is lower than the level of the second programming language, and the second target instruction set includes multiple instructions written in the second programming language.
2. The method according to claim 1, wherein The step of reading the real vector and the real constant from memory and performing operations on the real vector and the real constant based on the first target instruction set of the first programming language to obtain an operation result includes: A. Read the real constant from memory, copy the real constant based on the first target instruction set to obtain multiple real constants, and store the multiple copied real constants in a constant register; B. Based on the first target instruction set, read 2M real elements in the real vector, and sequentially store each of the real elements in a first register, a second register, a third register, and a fourth register, where M is an integer greater than 0, and the number of real elements stored in the first register, the second register, the third register, and the fourth register is the same; C. Based on the first target instruction set, perform operations on the real constant in the constant register and the real elements in the first register, the second register, the third register, and the fourth register respectively, and store the operation result in the associated memory of the real vector; D. Repeat steps B - C until the operations on all real elements in the real vector are completed, and use the operation result in the associated memory as the operation result of the real vector and the real constant.
3. The method according to claim 2, characterized in that, The step of reading 2M real elements in the real vector based on the first target instruction set includes: If it is the first round of loop currently, based on the first target instruction set, read 2M real elements in the real vector from memory, and prefetch the first N real elements among the unread elements into the cache, where N is an integer greater than 0; If it is not the first round of loop currently and the number of unread elements in the cache is greater than or equal to 2M, then read the first 2M real elements among the unread elements from the cache based on the first target instruction set; If it is not the first round of loop currently and the number of unread elements in the cache is less than 2M, then read the first 2M real elements among the unread elements from the memory based on the first target instruction set, and prefetch the first N real elements among the unread elements into the cache in the memory.
4. The method according to claim 2, characterized in that, The reading of 2M real elements in the real number vector based on the first target instruction set and sequentially storing each of the real elements in the first register, the second register, the third register, and the fourth register includes: Call the first instruction in the first target instruction set to read the first M real elements in the 2M real number vector, and store each of the read real elements in the first register and the second register; Call the first instruction in the first target instruction set to read the last M real elements in the 2M real number vector, and store each of the real elements in the third register and the fourth register.
5. The method according to claim 2, wherein The reading of the real constant from the memory, copying the real constant based on the first target instruction set to obtain a plurality of real constants, and storing the plurality of copied real constants in the constant register includes: Call the second instruction in the first target instruction set to copy the real constant, and sequentially store the plurality of copied real constants into the constant register.
6. The method according to claim 2, wherein The performing operations on the real constants in the constant register and the real elements in the first register, the second register, the third register, and the fourth register respectively based on the first target instruction set, and storing the operation results in the associated memory of the real number vector includes: Call the third instruction in the first target instruction set to respectively read the real elements from the first register, the second register, the third register, and the fourth register, read the real constant from the constant register, perform operations on the real elements and the real constant, and store the operation results of each real element in the fifth register corresponding to the first register, the sixth register corresponding to the second register, the seventh register corresponding to the third register, and the eighth register corresponding to the fourth register respectively; Call the fourth instruction in the first target instruction set to store the operation results in the fifth register, the sixth register, the seventh register, and the eighth register into the associated memory of the real number vector.
7. The method according to claim 6, characterized in that, The calling of the fourth instruction in the first target instruction set to store the operation results in the fifth register, the sixth register, the seventh register, and the eighth register into the associated memory of the real number vector includes: Call the fourth instruction in the first target instruction set to store the operation results in the fifth register and the sixth register into the position indicated by the current pointing address of the associated memory, and offset the current pointing address backward by a preset number of bytes to obtain a new current pointing address; Invoke the fourth instruction in the first target instruction set, store the operation results in the seventh register and the eighth register at the position indicated by the new current pointing address, and offset the new current pointing address backward by a preset number of bytes to obtain the current pointing address for the next loop.
8. The method according to claim 1, characterized in that The second target instruction set based on the second programming language reads a real number vector and a real number constant from the memory, and performs operations on the real number vector and the real number constant to obtain an operation result, including: A. Read a real number constant from the memory, and based on the second target instruction set, copy the real number constant multiple times and store them in the constant register; B. Based on the second target instruction set, starting from the memory pointing address in the memory, read nQ real number elements of the real number vector, store each of the real number elements in Q ninth registers in sequence, and modify the memory pointing address to the address of the next real number element after the last real number element read, where n is an integer greater than 0, Q is an integer greater than 0, and the number of real number elements stored in each ninth register is the same; C. Based on the second target instruction set, perform operations on the real number constant in the constant register and the real number elements in each of the ninth registers respectively, and store the operation results in each of the ninth registers; D. Based on the second target instruction set, store the operation results in each of the ninth registers into the associated memory of the real number vector; E. Repeat steps B - D until the operations on all real number elements in the real number vector are completed, and use the operation result in the associated memory as the operation result of the real number vector and the real number constant.
9. The method according to claim 8, wherein The step of reading a real number constant from the memory, and based on the second target instruction set, copying the real number constant multiple times and storing them in the constant register, includes: Invoke the fifth instruction in the second target instruction set to copy the real number constant, and store the multiple copied real number constants in the constant register in sequence.
10. The method according to claim 8, wherein The step of based on the second target instruction set, starting from the memory pointing address in the memory, reading nQ real number elements of the real number vector, and storing each of the real number elements in Q ninth registers in sequence, includes: Invoke the sixth instruction in the second target instruction set to read n real number elements of the real number vector from the memory, and store the read real number elements in the first ninth register; Repeat the step of invoking the sixth instruction in the second target instruction set, read n real number elements of the real number vector from the memory, and store the real number elements in the next ninth register until the nQ real number elements of the real number vector are respectively stored in Q ninth registers.
11. The method according to claim 8, characterized in that The step of based on the second target instruction set, performing operations on the real number constant in the constant register and the real number elements in each of the ninth registers respectively, and storing the operation results in each of the ninth registers, includes: Invoke the seventh instruction in the second target instruction set to read the real number elements from each of the ninth registers, read the real number constant from the constant register, perform four arithmetic operations on the real number elements and the real number constant, and store the operation results of each real number element in the position corresponding to the real number element in each of the ninth registers.
12. The method according to claim 8, wherein Storing the operation results in each of the ninth registers into the associated memory of the real vector based on the second target instruction set includes: Invoking an eighth instruction in the second target instruction set to store the operation results in each of the ninth registers into the associated memory of the real vector.
13. A real number vector operation processing device, characterized in that, Including: An acquisition module, configured to acquire the vector length of the real vector to be operated and a preset length threshold; A first operation module, configured to, if the vector length is greater than the preset length threshold, based on a first target instruction set of a first programming language, read a real constant from a memory, copy and store the real constant into a constant register, read the real vector, and batch-store a plurality of real elements in the real vector into a plurality of registers, and perform operations on the real constant stored in the constant register and the real elements stored in the plurality of registers respectively based on the first target instruction set to obtain operation results, where the first target instruction set includes a plurality of instructions written in the first programming language; A second operation module, configured to, if the vector length is less than or equal to the preset length threshold, based on a second target instruction set of a second programming language, read a real constant from a memory, copy and store the real constant into a constant register, read the real vector, and store a plurality of real elements in the real vector into a plurality of registers, and perform operations on the real constant stored in the constant register and the real elements stored in the plurality of registers respectively based on the second target instruction set to obtain operation results, where the level of the first programming language is lower than the level of the second programming language, and the second target instruction set includes a plurality of instructions written in the second programming language.
14. An electronic device, characterized in that, Including: A processor, a storage medium, and a bus, where the storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of a real vector operation processing method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of a real vector operation processing method according to any one of claims 1 to 12 are performed.
Citation Information
Patent Citations
Systems, apparatuses and methods for performing absolute difference calculation between corresponding packed data elements of two vector registers
CN104126169A
Method for super-high-speed interpretive execution of assembly instructions of TMS320C25 chip in X86 computer
CN107402799A