A Cache-based Multi-core DSP Parallel Programming Optimization Method
By building and updating the database record parallel algorithm operations and cycle times, the problem of inaccurate Cache size setting is solved, fast and efficient Cache size optimization is achieved, and the performance of multi-core DSP parallel programming is improved.
Patent Information
- Application Number
- CN202211493303.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-11-25
AI Technical Summary
In the prior art, in the multi-core DSP parallel programming optimization method, the setting of the cache size usually depends on experience values or multiple trials, resulting in the inability to quickly and accurately obtain a reasonable cache size, which takes a long time.
By building a database, record the number of operations, cycle times and running time of the parallel algorithm, generate the original database, use the database to quickly locate the cache size, update the database to improve accuracy, and optimize the cache size settings.
It realizes the rapid and accurate setting of the cache size, improves the efficiency of multi-core DSP parallel programming, and reduces the test time. With the update of the database, the cache size setting is becoming more and more accurate.
Abstract
Description
Technical Field
[0001] The invention belongs to the field of multi-core processors, and in particular relates to a multi-core DSP parallel programming optimization method based on Cache. Background Art
[0002] The current methods for parallel programming optimization of multi-core DSP parallel programs mainly include the following categories:
[0003] (1) Use the optimization options provided by the compiler: Some compilers provide some compilation options that directly affect or control program optimization. By properly configuring these optimization options, the goal of improving program running efficiency can be achieved.
[0004] (2) Change the code layout or calling method: The compiler provides many inline functions. The execution efficiency of inline functions is similar to that of assembly instructions, which can quickly optimize C code. In addition, SIMD (Single Instruction Multiple Data) is also one of the optimization methods to improve data processing efficiency. For example, when implementing single-precision multiplication operations, the operands can be stored in 32-bit form in the high or low part of the 64-bit register group, and a double-precision multiplication instruction can be used to simultaneously complete the multiplication of two pairs of single-precision data. This data packaging processing method can save instruction cycles when accessing and calculating data, thereby increasing the execution speed of the code.
[0005] (3) Improving C language loop programs: The purpose of software pipelining is to maximize the number of parallel instructions at the same time, thereby improving program execution efficiency. Loop unrolling is to expand the iterations of small loops to increase the number of possible parallel instructions, thus forming a better software pipelining arrangement. In addition, in nested multiple loops, only the inner loop can be software pipelining. Therefore, the inner loop of a multiple loop is generally unrolled.
[0006] (4) Use Cache to improve program performance: The access speed of Cache is higher than that of memory. Therefore, opening Cache in the program and putting frequently accessed programs and data into the cache will improve the running efficiency of the program. Generally, programs use two levels of cache. The first level cache is fixed and the second level cache can be configured. If the data that the CPU wants to access can be obtained directly from the cache every time, the cache hit rate is high and the program execution efficiency is high. Therefore, in order to improve the hit rate, it is necessary to reasonably configure the size of the second level cache. The L2 space can be configured as cache in its entirety, but the cache capacity is not the larger the better. It should be determined according to the calculation scale in the program. Sometimes, if the cache is opened too large, the execution efficiency of the program will not be significantly improved. Instead, it will affect the configuration of the SRAM capacity of the L2 space.
[0007] Currently, the cache size is generally set according to empirical values or through multiple trials. Setting according to empirical values may not result in a reasonable cache size, and multiple trials will take a long time. Summary of the Invention
[0008] (1) Technical issues to be resolved
[0009] The technical problem to be solved by the present invention is how to provide a cache-based multi-core DSP parallel programming optimization method to solve the problem that the cache size is generally set according to empirical values or set through multiple experiments. Setting according to empirical values may not result in a reasonable cache size, and multiple experiments will take a long time.
[0010] (2) Technical solution
[0011] In order to solve the above technical problems, the present invention proposes a cache-based multi-core DSP parallel programming optimization method, which includes the following steps:
[0012] S11. Generating an original database: Analyze the number of operations and loops included in various parallel algorithms, try setting different cache sizes, run a given parallel algorithm program, and obtain the running time; record the number of operations and loops of various parallel algorithms, as well as the five cache size settings with the shortest running time, use the running time, number of operations, and number of loops as parameters, and store the cache size setting as a query result as a data item;
[0013] S12. Use the database to set the cache size of the new algorithm: calculate the number of operations and loops of the algorithm after parallelization, and perform corresponding searches in the database;
[0014] If there is a similar parameter setting, perform a predetermined number of operations based on the cache value in the database and select the cache size setting with the shortest time from the results;
[0015] If there are similar results for both parameters but they are not in the same entry, the optimal value will be calculated separately;
[0016] If the number of operations and the number of loops are both close and are for the same entry, search and filter out a preset number of cache sizes with the shortest time, set the cache size according to the search results, obtain the running time, and set the cache size with the shortest time from the preset number of search results;
[0017] S13, updating the database: If the number of operations and the number of cycles of the new algorithm do not have similar values in the database, the operation results of the new algorithm are added to the database according to the method of S11.
[0018] Furthermore, the number of cycles includes the number of single-layer cycles or the number of inner cycles in multi-layer cycles.
[0019] Furthermore, in step S12, the query conditions are the number of operations and the number of loops, but only one of the two conditions may be hit. According to the query result, the cache size is set according to the cache size of the entry hit in the database, and parallel calculation of the parallel algorithm is performed, and the running completion time is recorded.
[0020] Furthermore, in step S12, the optimal value means that both the number of operations and the number of loops are hit, but they are not the same entries. In this case, the cache size needs to be set according to the cache size of all hit entries to see which cache size setting takes the shortest time to run the established parallel algorithm.
[0021] Furthermore, the preset number is 3.
[0022] Furthermore, in step S12, the predetermined number of times is 5 times.
[0023] A multi-core DSP parallel programming optimization method based on cache, the method comprising the following steps:
[0024] S21. Database construction: Calculate the number of operations and loops of relatively mature parallel algorithms, and try to set different cache values. Take the five data with the shortest time and record them. The database entry is set as one data with the number of operations, loops, running time, and cache size.
[0025] S22, calculating the number of runs and the number of loops of the cache size algorithm to be set, and performing a search using these two parameters;
[0026] S23. If only one similar entry is found for a parameter, perform a predetermined number of operations according to the recorded cache size setting to obtain the optimal cache size value, set it, and record the result in the database;
[0027] S24. If two parameters have similar values but are not in the same entry, calculate the two parameters separately according to S23, take the optimal value, and record the result in the database;
[0028] S25. If similar values are found for the two parameters in the same entry, a preset number of cache sizes with the shortest time are selected for calculation to obtain the running time. The cache size with the shortest time is selected from the preset number of results for setting, and the result is recorded in the database.
[0029] S26. If no similar value of the corresponding parameter is found, calculate and record and update the database according to S21.
[0030] Furthermore, the predetermined number of times is 5 times.
[0031] Furthermore, the preset number is 3.
[0032] Furthermore, in step S24, the optimal value means that both the number of operations and the number of loops are hit, but they are not the same entries. In this case, the cache size needs to be set according to the cache size of all hit entries to see which cache size setting takes the shortest time to run the established parallel algorithm.
[0033] (3) Beneficial effects
[0034] This paper proposes a cache-based multi-core DSP parallel programming optimization method. It also proposes a database-based cache sizing method to quickly and accurately determine the cache size. This method generates an initial database through initial training of basic algorithms. As new algorithms are used, the database is updated based on the new data. As the database data continues to grow, the optimal cache size setting becomes increasingly accurate. DETAILED DESCRIPTION
[0035] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below with reference to the embodiments.
[0036] With the rapid development of information technology, the demand for data information processing is also growing. Simply improving the performance of single-core processors can no longer meet the real-time requirements of the system. Therefore, multi-core processors have become the development direction of subsequent processors. Parallel programming and optimization methods for multi-core processors have also become a hot topic of subsequent research and are the key to improving the actual application performance of multi-core processors.
[0037] The multi-core DSP parallel programming optimization method based on Cache of the present invention comprises the following steps:
[0038] S11. Generating an original database: Analyze the number of operations and the number of single-layer loops or inner loops in a multi-layer loop for various parallel algorithms, try setting different cache sizes, run a given parallel algorithm program, and obtain the runtime. Record the number of operations and loops for the various parallel algorithms, as well as the five cache size settings with the shortest runtime. The runtime, number of operations, and number of loops are used as parameters, and the cache size setting is stored as a data item as a query result. In one embodiment, the data is stored in a database, such as a MySQL database.
[0039] S12. Use the database to set the cache size of the new algorithm: calculate the number of operations of the parallelized algorithm and the number of inner loops in a single-layer loop or a multi-layer loop, and perform corresponding searches in the database;
[0040] If there is a similar parameter setting, perform five operations based on the cache value in the database and select the cache size setting with the shortest time from the results;
[0041] If there are similar results for both parameters but they are not in the same entry, the optimal value will be calculated separately;
[0042] If the number of operations and the number of loops are both close and are for the same entry, search and select the three cache sizes with the shortest time. Set the cache size based on the search results, and then determine the running time. Then, select the cache size with the shortest time from the three search results and set it.
[0043] The best results of the above three situations are recorded in the database.
[0044] In a certain embodiment, the query conditions are the number of operations and the number of loops, but only one of the two conditions may be hit. Based on the query results, the cache size is set according to the cache size of the hit entry in the database, and parallel calculations of the parallel algorithm are performed, and the running completion time is recorded.
[0045] In one embodiment, the optimal value means that both the number of operations and the number of loops are hit, but they are not the same entry. In this case, the cache size needs to be set based on the cache size of all hit entries to see which cache size setting takes the shortest time to run the established parallel algorithm.
[0046] S13. Update the database: If the number of operations of the new algorithm and the number of inner loops in a single-layer loop or a multi-layer loop have no similar values in the database, the operation results of the new algorithm are added to the database in the same way as the original database is constructed.
[0047] The multi-core DSP parallel programming optimization method based on Cache of the present invention comprises the following steps:
[0048] S21. Database construction: Calculate the number of operations and loops of relatively mature parallel algorithms, and try to set different cache values. Take the five data with the shortest time and record them. The database entry is set as one data with the number of operations, loops, running time, and cache size.
[0049] S22, calculating the number of runs and the number of loops of the cache size algorithm to be set, and performing a search using these two parameters;
[0050] S23. If only one similar entry is found for a parameter, perform five operations according to the setting of the recorded cache size to obtain the optimal cache size value, set it, and record the result in the database;
[0051] S24. If two parameters have similar values but are not in the same entry, calculate the two parameters separately according to S23, take the optimal value, and record the result in the database;
[0052] S25. If similar values are found for the two parameters in the same entry, the three cache sizes with the shortest time are selected for setting calculation to obtain the running time. The cache size with the shortest time is selected from the three results for setting, and the result is recorded in the database.
[0053] S26. If no similar value of the corresponding parameter is found, calculate and record and update the database according to S21.
[0054] This paper proposes a database-based cache sizing method to quickly and accurately determine the cache size. This method generates an initial database through initial training of basic algorithms. As new algorithms are used, the database is updated based on the new data. As the database data continues to grow, the optimal cache size setting becomes increasingly accurate.
[0055] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A cache-based multi-core DSP parallel programming optimization method, characterized in that: The method comprises the following steps: S11. Generating an original database: Analyze the number of operations and loops included in various parallel algorithms, try setting different cache sizes, run a given parallel algorithm program, and obtain the running time; record the number of operations and loops of various parallel algorithms, as well as the five cache size settings with the shortest running time, use the running time, number of operations, and number of loops as parameters, and store the cache size setting as a query result as a data item; S12. Use the database to set the cache size of the new algorithm: calculate the number of operations and loops of the algorithm after parallelization, and perform corresponding searches in the database; If there is a similar parameter setting, perform a predetermined number of operations based on the cache value in the database and select the cache size setting with the shortest time from the results; If there are similar results for both parameters but they are not in the same entry, the optimal value will be calculated separately; If the number of operations and the number of loops are both close and are for the same entry, search and filter out a preset number of cache sizes with the shortest time, set the cache size according to the search results, obtain the running time, and set the cache size with the shortest time from the preset number of search results; S13, update the database: if the number of operations and the number of cycles of the new algorithm do not have similar values in the database, then add the operation results of the new algorithm to the database according to the method of S11; in, In step S12, the query conditions are the number of operations and the number of loops. Based on the query results, the cache size is set according to the cache size of the entry hit in the database, and the parallel calculation of the parallel algorithm is performed, and the completion time of the operation is recorded; In step S12, the optimal value means that both the number of operations and the number of loops are hit, but they are not the same entries. In this case, the cache size needs to be set according to the cache size of all hit entries to see which cache size setting takes the shortest time to run the established parallel algorithm.
2. The cache-based multi-core DSP parallel programming optimization method according to claim 1, characterized in that: The number of cycles includes the number of single-layer cycles or the number of inner cycles in multi-layer cycles.
3. The cache-based multi-core DSP parallel programming optimization method according to claim 1, characterized in that: The default number is 3.
4. The cache-based multi-core DSP parallel programming optimization method according to claim 1, characterized in that: In step S12, the predetermined number of times is 5 times.
5. A cache-based multi-core DSP parallel programming optimization method, characterized in that: The method comprises the following steps: S21. Database construction: Calculate the number of operations and loops of relatively mature parallel algorithms, and try to set different cache values. Take the five data with the shortest time and record them. The database entry is set as one data with the number of operations, loops, running time, and cache size. S22, calculating the number of runs and the number of loops of the cache size algorithm to be set, and performing a search using these two parameters; S23. If only one similar entry is found for a parameter, perform a predetermined number of operations according to the recorded cache size setting to obtain the optimal cache size value, set it, and record the result in the database; S24. If two parameters have similar values but are not in the same entry, calculate the two parameters separately according to S23, take the optimal value, and record the result in the database; S25. If similar values are found for the two parameters in the same entry, a preset number of cache sizes with the shortest time are selected for calculation to obtain the running time. The cache size with the shortest time is selected from the preset number of results for setting, and the result is recorded in the database. S26. If no similar value of the corresponding parameter is found, calculate and record and update the database according to S21; in, In step S24, the optimal value means that both the number of operations and the number of loops are hit, but they are not the same entry. In this case, the cache size needs to be set according to the cache size of all hit entries to see which cache size setting takes the shortest time to run the established parallel algorithm.
6. The cache-based multi-core DSP parallel programming optimization method according to claim 5, characterized in that: The scheduled number of times is 5 times.
7. The cache-based multi-core DSP parallel programming optimization method according to claim 5, characterized in that: The default number is 3.
Citation Information
Patent Citations
Method of read-optimized memory database T-tree index structure
CN103902693A
Multi-core parallel hash partitioning optimizing method based on column storage
CN104133661A